Fault tree generation method for cloud computing service and computer device

By acquiring heterogeneous spatiotemporal data to construct logical fault trees and annotating their attributes, a target fault tree is generated. This solves the problem that traditional fault trees are difficult to adapt to dynamic changes in cloud computing services, and achieves efficient and reliable reliability analysis.

CN122395080APending Publication Date: 2026-07-14LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LENOVO (BEIJING) LTD
Filing Date
2026-04-22
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Traditional fault trees are insufficient to meet the reliability analysis needs of cloud computing business architectures in complex and dynamically changing time-varying environments.

Method used

By acquiring heterogeneous spatiotemporal data, constructing logical fault trees, and combining the heterogeneous spatiotemporal data for attribute labeling, a target fault tree is generated, enabling reliability analysis of cloud computing services.

Benefits of technology

It improves the efficiency and reliability of fault tree construction, can adapt to the dynamic changes in the cloud computing environment, and meets the needs of reliability analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122395080A_ABST
    Figure CN122395080A_ABST
Patent Text Reader

Abstract

The application provides a fault tree generation method and computer equipment for cloud computing services, acquires heterogeneous spatio-temporal data of a first cloud platform, which includes spatial data and monitoring time series data from different data sources, then determines each fault event existing in the first cloud platform and the causal logical relationship between each fault event based on the heterogeneous spatio-temporal data, carries out fault tree modeling on each fault event and the causal logical relationship, obtains at least one logical fault tree, and at least based on the heterogeneous spatio-temporal data, carries out attribute labeling on each event object contained in the at least one logical fault tree, generates a target fault tree for the first cloud platform, and is used for reliability analysis of cloud computing services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud computing, and in particular to a fault tree generation method and computer device for cloud computing services. Background Technology

[0002] Fault Tree Analysis (FTA) is a top-down, deductive fault analysis method used to identify all possible causes of system failure. It is a widely used technique in fields such as automotive and energy to achieve vulnerability analysis or fault diagnosis.

[0003] However, due to the complex and dynamic nature of cloud computing services, traditional fault trees, which are built in one go, are difficult to meet the reliability analysis needs of cloud computing in time-varying environments. Summary of the Invention

[0004] In view of the above problems, this application provides the following solution:

[0005] The first aspect of this application provides a fault tree generation method for cloud computing services, the method comprising:

[0006] Acquire heterogeneous spatiotemporal data from the first cloud platform; the heterogeneous spatiotemporal data includes spatial data and monitoring time-series data from different data sources;

[0007] Based on the heterogeneous spatiotemporal data, the various fault events existing in the first cloud platform and the causal logical relationships between the various fault events are determined.

[0008] Fault tree modeling is performed on each of the fault events and the causal logical relationships to obtain at least one logical fault tree;

[0009] Based at least on the heterogeneous spatiotemporal data, attribute annotations are performed on each event object contained in the at least one logical fault tree to generate a target fault tree for the first cloud platform; the target fault tree is used for reliability analysis of the cloud computing service.

[0010] A second aspect of this application provides a computer device, the computer device comprising: at least one communication element, at least one memory, and at least one processor, wherein:

[0011] The communication element is used to acquire heterogeneous spatiotemporal data from the first cloud platform, the heterogeneous spatiotemporal data including spatial data and monitoring time series data from different data sources;

[0012] The memory is used to store multiple computer instructions;

[0013] The processor is configured to load and execute the computer instructions to implement the various steps of the fault tree generation method for cloud computing services provided in the first aspect of this application. Attached Figure Description

[0014] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0015] Figure 1 This is a flowchart illustrating the fault tree generation method for cloud computing services proposed in Embodiment 1 of this application.

[0016] Figure 2 A schematic diagram of a local logical fault tree is provided for the fault tree generation method for cloud computing services proposed in the embodiments of this application.

[0017] Figure 3 This is a flowchart illustrating the fault tree generation method for cloud computing services proposed in Embodiment 2 of this application.

[0018] Figure 4 This is a flowchart illustrating the fault tree generation method for cloud computing services proposed in Embodiment 3 of this application.

[0019] Figure 5 This is a flowchart illustrating the fault tree generation method for cloud computing services proposed in Embodiment 4 of this application.

[0020] Figure 6 The flowchart shown is a schematic diagram of the fault tree generation method for cloud computing services proposed in Embodiment 5 of this application;

[0021] Figure 7 This is a flowchart illustrating the fault tree generation method for cloud computing services proposed in Embodiment Six of this application;

[0022] Figure 8 This is a flowchart illustrating the fault tree generation method for cloud computing services proposed in Embodiment 7 of this application.

[0023] Figure 9 This is a flowchart illustrating the fault tree generation method for cloud computing services proposed in Embodiment 8 of this application.

[0024] Figure 10 A schematic diagram of a fault tree generation device for cloud computing services provided in an embodiment of this application;

[0025] Figure 11 This is a schematic diagram of the hardware structure of a computer device for a fault tree generation method applicable to cloud computing services. Detailed Implementation

[0026] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. The embodiments of this application are described below with reference to the accompanying drawings. It will be understood by those skilled in the art that, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems. The terms "first," "second," etc., used throughout this application and in the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be interchanged where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or device that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to these processes, methods, products, or devices.

[0027] To address the technical problems described in the background section, and considering the complex and dynamically changing architecture of cloud computing services, this application proposes an automatic fault tree construction method for dynamic cloud computing architectures, rather than the traditional one-time construction of fault trees. This method parses multi-source data streams, reuses fault logic modules, and dynamically generates a fault tree model that adapts to resource binding changes and topology updates—essentially constructing a dynamic fault tree to reflect the dynamic state changes of the analyzed object. This application provides a fault tree generation method for cloud computing services. In this application, for time-varying cloud computing environments, heterogeneous spatiotemporal data from any cloud platform (such as the first cloud platform executing the business) is acquired in real-time or periodically. This includes spatial data from different data sources and monitoring time-series data. Based on this data, the relationships between various cloud computing resources and different cloud services provided by the first cloud platform, as well as the changes in the system state and performance indicators of the first cloud platform, are determined. Through analysis, various fault events existing on the first cloud platform and the causal logical relationships between different fault events can be automatically determined, thereby constructing at least one logical fault tree representing each fault event and its causal logical relationship. This significantly reduces the workload of engineers developing logical fault trees. Since the logical fault tree does not include heterogeneous spatiotemporal data from the first cloud platform, the data volume of the logical fault tree is greatly reduced. However, it reflects the time-varying characteristics of cloud computing. By combining actual heterogeneous spatiotemporal data, it achieves adaptive modeling of dynamic changes in resource binding and node topology updates. The resulting target fault tree can reliably meet the reliability analysis requirements of cloud computing in a time-varying environment, improving its construction efficiency and reliability. The fault tree generation method for cloud computing services in this application embodiment will be described in detail below with reference to the accompanying drawings.

[0028] Reference Figure 1 This is a flowchart illustrating the fault tree generation method for cloud computing services proposed in Embodiment 1 of this application. This embodiment can be applied to computer devices, such as cloud servers or resource-rich terminal devices. This application uses a server as an example for illustration. Figure 1 As shown, the fault tree generation method for cloud computing services proposed in this embodiment may include, but is not limited to:

[0029] Step S11: Obtain heterogeneous spatiotemporal data from the first cloud platform; the heterogeneous spatiotemporal data includes spatial data and monitoring time series data from different data sources.

[0030] In this embodiment, spatial data of the first cloud platform in the cloud computing environment can be obtained from different types of data sources such as external databases or datasets through tools such as ETL (Extraction Transformation Loading, which describes the process of extracting, transforming and loading data from the source to the destination), query programs, data engines, dedicated data acquisition systems, visualization and analysis platforms, etc., and integrated with the monitoring time series data (time series data) of the first cloud platform. Together with the spatial data, they constitute heterogeneous spatiotemporal data describing the cloud computing environment of the first cloud platform. This application does not limit the types of data content, acquisition methods and recording methods.

[0031] The spatial data of the first cloud platform within the cloud computing environment describes the relationships between various cloud computing resources and different cloud services provided by the first cloud platform, such as call relationships, connection relationships, and dependency relationships. It typically includes the hardware topology and spatial distribution information of the first cloud platform. For example, a container or virtual machine running on a host machine; a container connected to a certain type of GPU card inside the host machine (connection relationship); a server connected to an edge switch or core switch (connection relationship); an application running in an HA (High Availability) architecture or a certain cluster architecture; and the location information of the currently active service running under a certain HA architecture. Therefore, the data source for this spatial data can include, but is not limited to, configuration management databases, cloud platform databases, etc., so that computer devices can directly read from them. Optionally, this application can also access configuration or service status through a remote server and parse the spatial data accordingly. This application does not limit the method and source of obtaining the spatial data of the first cloud platform.

[0032] It should be noted that the spatial data in this application is not static. The resource call relationships, connection relationships, or dependencies in cloud computing may change at different times. To address this, the spatial data of the first cloud platform can be incrementally updated to obtain the latest spatial data. This can also be used to correct the logical fault tree, thereby improving the accuracy of the constructed or updated target fault tree. Furthermore, while the methods for obtaining spatial data are similar for different cloud platform architectures in a cloud computing environment, differences in architecture and other factors lead to variations in the spatial data between different cloud platforms, and their data sources may also differ. This application will not provide detailed examples of these variations.

[0033] The monitoring time-series data for the first cloud platform can describe the changes in the system status and performance indicators of the first cloud platform. This data can be obtained by capturing system status and indicator time-series data through the monitoring system, and also by parsing system-stored log data. Therefore, the data source for the monitoring time-series data in this application can include various databases such as the monitoring system database and the log database. It should be noted that many monitoring or operation and maintenance systems typically configure corresponding thresholds to implement monitoring alarms or operation and maintenance reminders. For example, an alarm is triggered when the remaining disk capacity is less than a certain threshold. Therefore, this application can also associate and store these thresholds and alarm / reminder information with the data at the corresponding time points in the monitoring time-series data.

[0034] Based on this, the monitoring time-series data obtained from the monitoring system in this application may include, but is not limited to: server performance indicators (such as CPU utilization, memory usage, disk I / O, etc.), network traffic, storage system I / O performance, and other related data. This data is obtained at a high sampling frequency (i.e., high frequency) according to actual needs, so that the monitoring time-series data can reflect the changes in the system over time in detail. Simultaneously, it can also be combined with IT service management system data, monitoring alarm documents, and log data to analyze abnormal states in heterogeneous spatiotemporal data. For example, analyzing abnormal events in IT service management fault report records, monitoring alarm documents, and error information in log files can identify abnormal system state patterns, thereby forming heterogeneous spatiotemporal data describing the first cloud platform in a cloud computing environment. Alternatively, it can be stored as supplementary content to heterogeneous spatiotemporal data for subsequent query verification of whether fault events in the constructed logical fault tree can be accurately triggered, ensuring that the logical fault count accurately reflects the abnormal state of the system.

[0035] In some embodiments, spatial data and monitoring time-series data of the first cloud platform are often correlated. To facilitate subsequent fault tree construction, this correlation can be established on the common content between the two to constitute heterogeneous spatiotemporal data of the first cloud platform. Optionally, unique identifiers describing computing resources or nodes (such as nodes through cloud services) in the spatial data can be determined, such as server names, IP addresses, UUIDs (Universally Unique Identifiers), etc., and corresponding identifiers in the monitoring time-series data can be determined. By using the common identifiers between the spatial data and the monitoring time-series data, a correlation can be established between the spatial data and the monitoring time-series data that have the specified identifier.

[0036] Furthermore, during the above analysis, the state / state abstraction of computing resources, cloud services, etc., within the same time slice can be obtained from both spatial data and monitoring time-series data, according to the time slices used to divide different time series in the monitoring time-series data (which can be a fixed, short time window, serving as a time node). This allows for the association of spatial and temporal data within the same time slice, forming heterogeneous spatiotemporal data. Optionally, this heterogeneous spatiotemporal data can be stored in a graph structure, such as a graph database or a memory-based data structure, to facilitate subsequent analysis and fault tree construction. This application does not restrict the content or storage method of the heterogeneous spatiotemporal data; it can be determined as appropriate.

[0037] In different types of reliability analysis, such as vulnerability analysis in System Reliability Engineering (e.g., inherent weaknesses that make a cloud platform more prone to failure or service degradation due to fixed architectural characteristics, operation and maintenance modes, and technology choices, i.e., places where cloud platforms may fail or have reliability problems) and fault diagnosis (i.e., the process of identifying the location, nature, and cause of failures in a system or device through monitoring, detection, and analysis technologies, with the aim of quickly restoring the system to normal operation), the content of heterogeneous spatiotemporal data required to construct their respective target fault trees will differ because the former focuses on prevention and the latter on recovery. This can be determined by combining their respective characteristics and implementation methods.

[0038] In some embodiments, during the acquisition of spatial data from the first cloud platform, the data primarily contains key information within the platform's architecture, and the level of detail varies depending on the analytical purpose. For example, in vulnerability analysis, spatial data focuses on aspects that significantly impact system reliability to identify overall system weaknesses and potential risks, providing foundational data support for vulnerability analysis. Taking the OpenStack architecture-based first cloud platform as an example, the acquired spatial data may include overall information about compute nodes (host machines), the running status of virtual machine instances, network IP allocation, and basic information about storage volumes. This can be obtained by querying the compute node table and virtual machine instance table in the OpenStack database, which may provide spatial data such as machine creation time, update time, deletion time (which may not exist), host machine super-resolution ratio, and number of CPU cores, but is not limited to these. In fault diagnosis, the spatial data acquired needs to be more detailed and specific, such as cloud computing resource dependencies, call relationships, and connection relationships, in order to accurately locate the specific node causing the fault. In this regard, in addition to the data required for vulnerability analysis as mentioned above, the spatial data used for fault diagnosis can also include device identifiers such as the universally unique identifier (UUID) of each machine to ensure its uniqueness in the entire system, so as to accurately identify each instance and ensure the reliability and accuracy of subsequent indexing.

[0039] Furthermore, to facilitate fault location and detection, spatial data can also be used to provide query interfaces for fault location and detection, enabling rapid narrowing of the fault scope and pinpointing the specific fault node when a fault occurs. Therefore, spatial data can also include data that serves as a query interface or is associated with a query interface. Taking the first cloud platform of the OpenStack architecture as an example, it can include querying detailed information such as the mount location and mount node of a storage volume through Cinder (OpenStack Block Storage Service, which provides persistent block storage devices for virtual machine instances in the OpenStack cloud), and querying data such as the running status and device ID of a network port through Neutron (OpenStack Networking Service, which provides network connectivity as a service to all components in the OpenStack cloud). However, it is not limited to the data content described in this application and can be flexibly adjusted as needed.

[0040] Step S12: Based on heterogeneous spatiotemporal data, determine the various fault events existing in the first cloud platform and the causal logical relationships between the various fault events;

[0041] Step S13: Perform fault tree modeling on each fault event and causal logical relationship to obtain at least one logical fault tree;

[0042] Following the analysis of heterogeneous spatiotemporal data above, this application can extract the states of each time slice. For example, if node x experiences a CPU spike at time t1, this can be abstracted as a basic event occurring on the first cloud platform. If the extracted states are such as service Y exhibiting a memory leak trend or a 5-fold increase in database latency, these can be abstracted as intermediate events occurring on the first cloud platform. If the extracted states are such as an increased error rate highly correlated with deployment time, they can be used to associate the basic fault events with the top event (service unavailability), thus determining the causal logical relationship between the two corresponding fault events. Therefore, this application can parse heterogeneous spatiotemporal data to extract states indicating fault events and / or the causal logical relationships (such as call / connection / dependency relationships) between different fault events, thereby obtaining all fault events currently existing on the first cloud platform, as well as the actual causal logical relationships between different fault events. This application does not limit the implementation method.

[0043] Subsequently, this application can treat each identified fault event as an event node, and represent the corresponding causal logical relationship using different types of logic gates such as AND gates, OR gates, XOR gates, and K / N gates to construct a logical fault tree. That is, each fault event is identified as a different event node, and a logical fault tree for the first cloud platform is constructed using logic gates of corresponding types for each causal logical relationship. The fault events existing in the cloud computing environment are events that may cause abnormal operation of cloud computing services on the first cloud platform, and can generally be divided into basic events (bottom events), intermediate events, and top events. Basic events are indivisible, low-level fault events, such as each specific inspection activity, like server CPU utilization levels. Intermediate events can be composite events composed of logic gates, such as database service unavailability caused by master-slave synchronization failure, or API gateway timeout caused by load balancer failure. Top events can be the final event state that triggers fault detection, such as server crashes, slow application operation, or the average API response time of web applications consistently exceeding 500 milliseconds at the P95 percentile. The top event is the target of fault tree analysis; it represents the final manifestation of a system failure. These events are interconnected. Figure 2 The logical relationship shown.

[0044] The fault propagation logic represented by the AND gate is as follows: the corresponding output event (parent node) will only occur if all input events (child nodes) occur simultaneously. The fault propagation logic represented by the OR gate is as follows: the occurrence of any input event (child node) will lead to the occurrence of the corresponding output event (parent node). The fault propagation logic represented by the XOR gate is as follows: the output event (parent node) will occur when only one input event (child node) occurs. The fault propagation logic represented by the K / N gate is as follows: the output event (parent node) will occur if at least K of the N input events (child nodes) occur. Based on this, this application can connect the various fault events (event nodes) that are input events to the fault events that are output events through AND gates (connection lines between corresponding event nodes), thereby forming at least one logical fault tree. For example, after identifying the top event, the direct causes of the top event can be analyzed layer by layer from top to bottom. These causes are connected by logic gates to form the logical fault tree (subtree) of the top event, thus completing the construction of the logical fault tree of the first cloud platform.

[0045] Therefore, each logical fault tree constructed in this application, in the form of event nodes and logic gates, initially describes the possible fault events of the first cloud platform and their causal logical relationships. Preferably, this application can focus the logical fault tree on the core business processes provided by the first cloud platform and identify the key fault modes (which can be represented by corresponding logic gates) and potential fault points to construct the basic framework as the logical fault tree. Therefore, in the implementation process of the above steps, after determining the core business processes provided by the first cloud platform, the various fault events existing in the core business processes and their causal logical relationships can be determined according to the analysis method described above to construct the logical fault tree, but it is not limited to this.

[0046] When constructing a fault tree in a cluster environment (redundant structure), a logical fault tree can be directly built as described above. This is typically a highly repetitive, "fat tree," which often results in high time and space complexity. Furthermore, as cloud computing services change, the size of redundant event nodes also changes, causing the fault tree structure to alter accordingly. To avoid the resource waste caused by frequently rebuilding the fat tree, this application constructs multiple logical fault trees (which can be considered subtrees of the fat tree) during the logical fault tree construction process. Specifically, a logical fault tree (subtree) is constructed for highly repetitive event nodes, requiring only the corresponding subtree to be updated. However, this approach is not limited to this. Subsequently, if a complete logical fault tree needs to be constructed, this application can progressively generate the complete logical fault tree through intermediate events during the construction of each subtree, reducing the time and space complexity of generating a complete logical fault tree all at once and improving construction efficiency.

[0047] In the aforementioned logical fault tree, information with redundancy characteristics can be indicated by logical gates, such as using K / N gates to represent HA structures or cluster structures; or using XOR gates to represent HA structures, etc. Figure 2 As shown, this application can use OR gates to represent a chain of calls, such as service A calling service B, and service B calling database C. Any basic event exception will cause the system to malfunction. Therefore, this application can combine the analyzed business states to model different logical gate structures, obtaining corresponding logical fault trees, i.e., subtrees. Furthermore, this application can use AND gates to represent fault modes triggered by multiple conditions. For example, in a distributed system, multiple nodes must fail simultaneously for the system to malfunction; AND gates can be used to connect the fault events of these nodes, accurately representing fault modes triggered by multiple conditions. To clarify the scenarios corresponding to the inputs and outputs of each logical gate, it can be marked with the resource type and necessary attribute information it accesses; the specific content is not limited in this application.

[0048] Step S14: Based at least on heterogeneous spatiotemporal data, attribute annotation is performed on each event object contained in at least one logical fault tree to generate a target fault tree for the first cloud platform; the target fault tree is used for reliability analysis of cloud computing services.

[0049] Based on the above analysis, in the modeling of logic fault trees, logic gates typically need to connect to two or more subtrees or basic events, but do not include heterogeneous spatiotemporal data. To facilitate the subsequent fusion of heterogeneous spatiotemporal data, this application can configure key attributes for each event object involved in the event nodes of the logic fault tree, especially the event objects in critical fault events, to define or explain the event object, providing detailed information for subsequent analysis. Therefore, for different types of reliability analysis needs, at least based on heterogeneous spatiotemporal data, attribute annotation can be performed on each event object contained in at least one logic fault tree (which can be multiple subtrees or a complete logic fault tree formed by them). For example, determining each event object in the logic fault tree that needs attribute annotation, or specifying each event object in each fault event, etc., and then extracting the attribute information that needs to be annotated from the heterogeneous spatiotemporal data accordingly, so as to annotate it to the corresponding event node or its associated corresponding event object for subsequent viewing or analysis. This application does not restrict the tools used for attribute annotation or the data format of the annotated attribute information.

[0050] The event object can be a component included in the first cloud platform, such as a device providing cloud services / computing resources. The attribute information annotated for it can include, but is not limited to, object type (i.e., component type, such as server, storage system, network device, virtual machine, and container) and basic attributes (such as device model, storage capacity, and data center location). For the event objects in the logical fault tree for fault diagnosis, the attribute information to be annotated can also include operational information, such as the device's IP address, port number, maintenance records, historical fault frequency, alarm threshold, and real-time performance indicators. Therefore, in one possible implementation, during the attribute annotation process for constructing the logical fault tree, the object type (i.e., component category) can be considered. After classifying the component categories according to, but not limited to, the methods described above, the attribute information corresponding to each category of components in the first cloud platform can be determined from heterogeneous spatiotemporal data. For example, basic attributes such as the server's device model, CPU frequency, memory capacity, storage expansion capability, and data center location; and detailed parameters such as the network device's IP address, MAC address, port number, supported network protocols, and connection relationships. This attribute information can come from at least one authoritative data source, such as the configuration management database of the corresponding component, the manufacturer's manual, and the deployment configuration file, to ensure the accuracy and reliability of the attribute information.

[0051] In some embodiments, when attribute information comes from multiple data sources, to ensure the accuracy and consistency of the attribute information, the attribute information from different sources can be fused. At least based on the fused attribute information, attribute annotations can be performed on the corresponding event objects (such as components at different levels) contained in the logical fault tree. Optionally, this application can adopt a multi-source data fusion strategy to integrate attribute information from different data sources and perform deduplication. For example, when the IP address information of a server is obtained simultaneously from both the configuration management database and the device maintenance log, data cleaning and comparison algorithms can be used to ensure that the IP address ultimately annotated on the logical fault tree is accurate and unique.

[0052] Preferably, this application can also automatically verify the attribute information after annotation through a pre-configured data verification mechanism. Upon discovering any non-compliant attribute information, a warning will be immediately issued, prompting for correction of the corresponding attribute information. Based on this, this application can perform compliance verification on various attribute information annotated in the logical fault tree; correct non-compliant attribute information; and based on the compliant attribute information and the corrected attribute information, perform attribute annotation on the corresponding event objects contained in the logical fault tree, such as updating previously incorrectly annotated attribute information, supplementing missing attribute information, or deleting incorrectly annotated attribute information. For example, it can check whether the IP address conforms to the standard IP address format, whether the port number is within the valid range, etc., which can be implemented according to known or preset standards / rules. Upon determining the existence of non-compliant attribute information, it can output alarm information containing the attribute information and its annotation location, such as sending it to the relevant engineer or the attribute annotation tool, to correct the annotated attributes in a timely manner and lay a reliable foundation for the subsequent generation of the target fault tree.

[0053] Based on the above analysis, this application can accurately map event nodes and their causal relationships in the logical fault tree to heterogeneous spatiotemporal data using labeled attribute information. In this way, this application can fuse heterogeneous spatiotemporal data with logical fault data according to this mapping relationship to generate a target fault tree for reliability analysis. It should be understood that for different reliability analyses of different cloud computing services, such as the aforementioned vulnerability analysis and fault diagnosis (analysis), the corresponding heterogeneous spatiotemporal data, logical fault trees, and their labeled attribute information will differ to some extent. Following the target fault tree generation method described above, a first target fault tree can be generated for vulnerability analysis to identify weak points and potential risks; a second target fault tree can be generated for fault analysis (i.e., the aforementioned fault diagnosis) to locate fault points and assess the scope of impact. Therefore, in the process of generating the target fault tree, in response to the vulnerability analysis request for cloud computing services, based on the inherent attributes and historical operating data of different types of components in the first cloud platform contained in the heterogeneous spatiotemporal data, attribute annotations are performed on each event object contained in at least one logical fault tree (e.g., annotations of key indicators such as device failure frequency, distribution, MTTR (Mean Time To Repair), MTBF (Mean Time Between Failures), and lifespan). The first spatiotemporal data contained in the heterogeneous spatiotemporal data is then fused with the attribute-annotated logical fault tree to generate the first target fault tree for vulnerability analysis of cloud computing services. This allows it to integrate information such as historical fault data, device reliability parameters, and system design documents. It is evident that the first spatiotemporal data includes multidimensional spatiotemporal data for vulnerability scoring.

[0054] Similarly, in response to requests for fault analysis of cloud computing services, based on heterogeneous spatiotemporal data and business operation and maintenance information (which may include device IP addresses, port numbers, real-time performance indicators, alarm thresholds, and maintenance records, etc.), attribute annotations are performed on each event object contained in the corresponding logical fault tree. The second spatiotemporal data contained in the heterogeneous spatiotemporal data is then fused with the attribute-annotated logical fault tree to generate a second target fault tree for fault analysis of cloud computing services. This enables the rapid location of the root cause of the fault when it occurs, and the precise location of the specific node in the second target fault tree that caused the fault, such as the basic events (leaf nodes) or minimal cut sets (the smallest set of basic events that cause the top event) in the second target fault tree, thereby providing strong support for rapid fault diagnosis and effective fault elimination. The second spatiotemporal data includes first spatiotemporal data and dynamic monitoring data, such as real-time monitoring data, alarm information, and log data. This application does not limit the content of various types of dynamic monitoring data.

[0055] In summary, this application utilizes a target fault model generated by combining a logical fault tree with heterogeneous spatiotemporal data. This model provides a more comprehensive, detailed, and accurate description of the logical relationships of system faults, laying a solid foundation for subsequent vulnerability analysis and fault diagnosis. In practical applications, cloud computing services based on target fault trees for reliability analysis can cover the entire technology stack, from basic resources to high-level services. These cloud computing services may include, but are not limited to: Infrastructure as a Service (IaaS) such as elastic computing services, storage services, and network services; Platform as a Service (PaaS) such as database services, middleware services, and big data platforms; Software as a Service (SaaS) such as enterprise applications and industry solutions such as healthcare, finance, and gaming; high-level cloud-native services such as serverless computing, AI platform services, and quantum computing cloud; services such as hybrid cloud connectivity and cloud management platforms; and emerging cutting-edge services such as web services, metaverse cloud, and sky cloud computing. While the target fault trees constructed according to the method of this application may differ depending on different business needs, the construction process is similar, and will not be detailed in this embodiment.

[0056] Reference Figure 3 This is a flowchart illustrating the fault tree generation method for cloud computing services proposed in Embodiment 2 of this application. This embodiment describes an optional implementation method for constructing a logical fault tree based on heterogeneous time data to determine fault events and causal logical relationships. Figure 3 As shown, this implementation method may include, but is not limited to:

[0057] Step S31: Based on the heterogeneous spatiotemporal data of the first cloud platform, determine the components at each level used by the target cloud computing service supported by the first cloud platform and the interaction relationships between different components, as well as the dependency relationships between the target business links in the target cloud computing service and the basic resources of the first cloud platform.

[0058] In this application embodiment, the target cloud computing service can be the core service supported by the first cloud platform. It can be determined by the main service mode and deployment mode. It can be providing users with on-demand, elastically scalable, and metered IT resources and services to be delivered through the network (such as the Internet). The aim is to reduce the complexity and cost of enterprise IT operation and maintenance, improve resource utilization and application innovation efficiency, etc. This application does not limit the type of core service and its execution process, and can be determined as appropriate.

[0059] During the construction of the logical fault tree for the first cloud platform, a logical fault tree focusing on the core business execution process, fault modes, and potential fault points can be built. Therefore, in the process of acquiring the various event nodes and relationships used to construct the logical fault tree, after determining the target cloud computing business supported by the first cloud platform, the system hierarchical architecture and target computing business processes can be analyzed based on the heterogeneous spatiotemporal data of the first cloud platform. This allows for the sorting out of the components involved in the target computing business processes at each level, as well as the interaction relationships between different components. For example, after determining the various layers of the first cloud platform architecture (such as the presentation layer, application layer, domain layer, data layer, and infrastructure layer) based on spatial data, for each core business process, all components involved in the core business process and their responsibilities / functions can be determined by combining heterogeneous spatiotemporal data. Then, a sequence diagram can be drawn for the core business process to show how different components call, pass messages and data. By sorting out these data flows, the interaction relationships between different components can be determined, such as calling relationships or connection / transmission relationships, but this is not limited to this.

[0060] In the above analysis, this application can also identify the key business links (target business links) in the core business process and their dependent basic resources, that is, determine the dependency relationship between the key business link and the basic resources. Key business links refer to those indispensable steps that directly determine the successful operation of the cloud computing business, directly affect user experience and company revenue, and whose failure or performance bottleneck will have an immediate negative impact on the cloud computing business. It is evident that, compared to regular business links, key business links are value realization points, single-point bottlenecks, resource-intensive, highly dependent on external factors, and have high data consistency requirements within the core business process. This application can identify the key business links in each business link included in the core business process based on these characteristics; the implementation process is not detailed here. For example, based on the above analysis, the key business links in the core business process of a user placing an order in an e-commerce platform may include: inventory verification and deduction, order creation, and payment processing; the key business links in the core business process of user registration / login may include: identity verification and user profile creation.

[0061] Optionally, in the process of identifying the components and their interactions involved in the core business processes, it is also possible to determine the components involved in the key business processes and their directly or indirectly dependent basic resources, such as computing resources, storage resources, network resources, and security resources—the basic components for building applications on the first cloud platform. This application does not limit the content or form of these resources. This application can record the direct or indirect dependencies between key business processes and these basic resources as causal logical relationships. These relationships can be represented by a dependency matrix between construction levels and can also generate a visual topology map to facilitate verification of their completeness and accuracy.

[0062] Computing resources can include virtual machines, containers, serverless functions, and physical servers, used to process computing tasks and execute application code. Storage resources are used to persistently store data and can be categorized according to data access patterns and performance requirements, such as block storage, object storage, and file storage. Network resources are responsible for connecting all cloud resources and controlling how traffic flows; these can include virtual networks, load balancing, content delivery networks, and DNS (Domain Name System) services. Security resources permeate computing, storage, and networking, providing authentication, access control, and threat protection functions; therefore, they can include identity and access management, security groups or network ACLs (Access Control Lists), web application firewalls, and key management services. In practical applications, the various basic resources listed above typically work closely together to form the infrastructure of the first cloud platform.

[0063] Step S32: Based on at least one of the historical fault data determined by heterogeneous spatiotemporal data and preset fault rules, analyze the execution process of the target cloud computing business using interaction and dependency relationships to determine the target fault mode and potential fault points.

[0064] Step S33: Based on the target failure mode and potential failure points, determine the various failure events existing in the first cloud platform and the causal logical relationships between the various failure events.

[0065] In this embodiment, the target fault mode can be determined based on its occurrence frequency and / or impact level. For example, fault modes that occur frequently or have a significant impact (e.g., wide propagation range, high recovery difficulty, and deep business impact) (e.g., fault modes with an occurrence frequency greater than a frequency threshold or an impact level greater than a level threshold; this application does not limit the identification method for these), or critical fault modes on the fault propagation path, can be represented by corresponding types of logic gates. This application does not limit the identification method for the target fault mode. The fault mode can be a manifestation or mechanism of a fault, instantiated into different fault events based on a specific context (e.g., resource configuration, load conditions, etc.). Therefore, the same fault mode may correspond to multiple fault events; for example, "network packet loss" can manifest as latency-sensitive service timeouts or video stream stuttering.

[0066] Based on this, the fault events in this application can be specific instances of faults that have occurred or are occurring, with a definite occurrence time, and can be directly captured by the monitoring system. Potential fault points can be locations or components where faults may occur, without a definite time, but can be discovered through analysis and generate corresponding fault events after being triggered. In order to identify target fault modes and potential fault points in core business processes, this application can be implemented based on the fault prior knowledge (such as historical faults or expert experience) and / or preset fault rules (such as fault judgment rules represented by one or more intelligent operation and maintenance algorithms, such as anomaly detection algorithms, dynamic threshold algorithms, etc.). Among them, the fault prior knowledge can include a set of fault strategies for historical faults determined based on heterogeneous spatiotemporal data captured by the monitoring system, such as triggering a circuit breaker mechanism if the API error rate suddenly increases by 10 times; triggering a capacity expansion check if the CPU load is >80% for 5 minutes, etc. Of course, corresponding fault strategies can also be obtained by encoding based on expert experience, etc. This application does not restrict the source and representation of fault prior knowledge.

[0067] Based on this, by analyzing the failure propagation paths of target failure modes and potential failure points, key propagation nodes (e.g., determined based on topological location and dynamic load) and key failure modes on the failure propagation paths are identified. These include failure modes whose occurrence frequency exceeds a frequency threshold, or whose single failure loss exceeds a loss threshold, and which are located in the minimum cut set. Therefore, they can be judged by identifying corresponding indicators, or by using pre-trained dynamic identification algorithms or quantitative evaluation models. The implementation process is not detailed in this application. Subsequently, by analyzing each failure propagation path, the various failure events participating in the core business process and the causal logical relationships between different failure events can be determined.

[0068] In some embodiments, this application can use techniques such as simulated fault injection to verify the rationality of each target fault mode, such as verifying whether the fault mode will actually occur, whether the impact when it occurs is as expected, and whether our countermeasures are effective, so as to correct unreasonable target fault modes in a timely manner and ensure the robustness of the first cloud platform architecture. Based on this, after step S32, the rationality of each target fault mode can be verified, and unreasonable target fault modes can be adjusted. This application does not limit the implementation method, and the fault resources to be simulated and injected, as well as the content to be verified, can be determined according to the characteristics of each target fault mode. The implementation process is not described in detail in this application.

[0069] Optionally, this application can also utilize heterogeneous spatiotemporal data such as system monitoring and log information to analyze monitoring data and parse log information in real time, combine it with a logical fault tree, identify key factors that may cause faults, trace the root cause of faults, and adjust the settings of key fault modes and potential fault points in the logical fault tree accordingly. Therefore, after executing step S32, source tracing analysis can be performed on each fault event, and the corresponding target fault mode or potential fault point can be adjusted based on the source tracing analysis results. Preferably, this fault tracing process can be implemented around the core business process, but is not limited to it. In the source tracing analysis process, causal reasoning algorithms such as those based on deep learning and Bayesian networks can be used to reason and analyze the correlation and causality of each fault event, locate the root cause of the fault, and verify whether the settings of the corresponding fault modes and potential fault points are correct, so as to adjust the settings that are inconsistent with the source tracing analysis results in a timely manner. This application does not limit the implementation method.

[0070] After adjusting the target fault mode and potential fault points according to, but not limited to, the methods described above, the determined fault events and the causal logical relationships between each fault event can be re-determined or adjusted according to the method described in step S33. This implementation process can be referred to the relevant description of the implementation process in step S33, and will not be elaborated upon here. It is evident that this application verifies the target fault mode and potential fault points through the above-mentioned rationality verification and fault source analysis, ensuring the reliability and accuracy of each event node and relationship used to construct the logical fault tree, thereby improving the reliability and accuracy of the logical fault tree construction.

[0071] In some embodiments, this application can identify each fault event as a different event node, and construct a logical fault tree for the first cloud platform using logic gates of corresponding types for each causal logical relationship. The structure of this logical fault tree can be as follows: Figure 2As shown, but not limited to this. In practical applications, a complete logical fault tree for a first cloud platform can be extremely large and complex, containing hundreds or thousands of basic events and logic gates. Especially for complex first cloud platforms, direct analysis would involve a massive amount of computation, making it difficult to manage and understand. Based on this, in some embodiments, to improve the efficiency of subsequent processing of the logical fault tree, simplify analysis, improve processing quality, and reduce computational complexity, this application proposes to extract subtrees (i.e., logical fault subtrees) from the logical fault tree before subsequent processing and analysis. That is, a large logical fault tree is split into multiple logical fault subtrees, and the subsequent processing of the complete logical fault tree is replaced by processing of each logical fault subtree.

[0072] In one possible implementation, this application can perform common cause failure analysis on the logical fault tree, using various event nodes caused by the same failure cause to construct corresponding logical fault subtrees as the logical fault tree of the first cloud platform, which is then used to generate the target fault tree. Common cause failure (CCF) refers to the phenomenon where multiple components fail simultaneously due to a single root cause; such failures undermine the effectiveness of system redundancy design. Based on this, this application can extract common cause failures into logical fault subtrees or basic events within the logical fault tree, facilitating subsequent modeling or analysis. Optionally, this application can use common cause failure identification methods such as architecture dependency analysis, historical failure event clustering, and chaos engineering testing to determine the common cause of each common cause failure. After identifying the affected components, it can extract the various failure events (event nodes) and their causal logical relationships that led to the common cause (same failure cause) in each component to construct the corresponding logical fault subtree, but is not limited to this. For example, in disaster recovery design scenarios, common cause failures typically include types such as power supply interruptions, software configuration errors, network partitions, and data corruption. For different common cause failures, the corresponding logical fault subtrees can be extracted to express the corresponding common cause failures through a tree structure. The implementation process is similar, and will not be described in detail here.

[0073] Based on this, refer to Figure 4 The flowchart shown in Embodiment 3 of this application illustrates a fault tree generation method for cloud computing services. This embodiment describes an optional implementation method for extracting subtrees (i.e., logical fault subtrees, hereinafter referred to as subtrees) of a logical fault tree through common cause failure analysis. Figure 4 As shown, this optional implementation method may include, but is not limited to:

[0074] Step S41: Identify each fault event in the first cloud platform as a different event node, and construct a logical fault tree for the first cloud platform using the logic gates of the corresponding types of causal logical relationships.

[0075] Step S42: Identify the dependencies between the logical fault tree and the various types of basic resources provided by the first cloud platform;

[0076] Step S43: Based on the dependency relationship, extract the logical fault subtrees related to different types of basic resources from the logical fault tree;

[0077] In the implementation method of extracting subtrees with common cause sources through common cause failure analysis, this application, combined with the above description of various basic resources provided by the First Cloud Platform, can analyze each event node and its relationship in the logical fault tree to determine the basic resources involved in each fault event and the dependency relationship between different basic resources. Based on this, common cause events are identified, that is, events that cause the same failure cause for multiple fault events of the same level (common dependent nodes, the identification method of this application is not limited). The common cause event is taken as the root node of the logical subtree for the current common cause failure. Based on the dependency relationship, each fault event that causes the common cause event is taken as the intermediate node of the subtree. It can be the failure of the combination of affected resources. Further, based on the dependency relationship, the basic events of the specific resource failure manifestation of the intermediate node are taken as the leaf nodes of the subtree to obtain a subtree composed of the root node, intermediate nodes and leaf nodes.

[0078] During subtree extraction, common causes (events with the same failure reason) within the logical fault tree can be identified through analysis of the logical fault tree (i.e., historical prior knowledge). Alternatively, they can be identified based on expert experience (expert prior knowledge), i.e., common cause failure analysis is performed based on expert prior knowledge to extract the corresponding subtrees. Optionally, depending on actual needs, related subtrees can be concatenated as a new subtree for subsequent processing. For example, if analysis reveals that virtual machines run on a host machine, and a host machine failure will cause a batch of virtual machine anomalies, the host machine-related subtrees can be extracted as logical fault tree subtrees. It should be noted that subtree extraction from the fault logic tree can be implemented according to the extraction methods described in steps S42 and S43 above. Subsequently, each logical fault subtree is treated as a logical fault tree, and the target fault tree is generated according to the method described in the context. Preferably, in order to improve the completeness and accuracy of subtree extraction, especially in the extraction of subtrees of the logic fault tree of a large-scale first cloud platform, another extraction method described in steps S44 and S45 can be combined to obtain the subtrees in the logic fault tree, that is, the various subtrees contained in the logic fault tree are extracted by two extraction methods.

[0079] Step S44: Analyze the logic fault tree through simulation environment to determine the failure modes under various basic resource combinations; each basic resource combination includes at least two types of basic resources.

[0080] Step S45: Based on the failure propagation path of the common cause of failure mode in the logical fault tree, construct the corresponding logical fault subtree.

[0081] In this embodiment, the application can also determine the common cause failure subtree or basic event by analyzing failure modes across different types of basic resource combinations. These cross-resource failure modes may include, but are not limited to: failure modes of computing and storage resource combinations, such as the analysis of instances and storage volumes being unavailable together; failure modes of network and security resource combinations, such as network policy errors causing service inaccessibility; failure modes of container and database resource combinations, such as container Pods and database connection pool exhaustion; and failure modes of content delivery network (CDN) and object storage resource combinations, such as data inconsistency between edge nodes and the origin server. It should be noted that cross-resource combinations of failure modes can include, but are not limited to, the two basic resource combinations shown in the examples above, and can also include more types of basic resource combinations, which can be determined based on the fault conditions of different failure scenarios. This application will not provide detailed examples of each type here.

[0082] Optionally, this application can pre-configure common cause triggering sources corresponding to different failure modes (e.g., determined through expert prior knowledge), or it can combine historical prior knowledge analysis to determine them. For example, the common cause triggering source for the failure mode where both the instance and storage volume are unavailable could be a physical host failure / AZ-level power outage; the common cause triggering source for the failure mode where container Pods and database connection pools are exhausted could be a shared virtualization layer failure, etc. Then, failure propagation analysis can be performed based on the common cause triggering sources under different failure modes to determine their corresponding failure propagation paths (causal chains) to construct logical failure subtrees. For example, analyzing the common cause triggering source acting on multiple dependent components that simultaneously fail due to the common cause source can yield failure propagation paths. As can be seen, during the implementation of steps S44 and S45, spatial data of the minimum adaptable scale (such as a scenario of one failure mode) and various types of resource redundancy structures (i.e., combinations of various basic resources) can be constructed through simulation and / or a simulation environment. Subtrees or basic events of each common cause failure can be obtained through Failure Mode and Effects Analysis (FMEA) or qualitative and quantitative analysis of fault trees, as described above. It should be noted that the extraction of subtrees from the fault logic tree can also be directly implemented according to the extraction method described in steps S44 and S45, that is, after executing step S41, steps S44 and S43 can be executed directly.

[0083] In some other embodiments, besides the implementation methods for extracting subtrees from the logical fault tree through the various common-cause failure analysis methods described above, this application can also construct different logical fault subtrees based on the difference data between vulnerability analysis and fault analysis of cloud computing services in heterogeneous spatiotemporal data. Alternatively, after extracting the common-cause failure subtrees according to, but not limited to, the methods described above, the extracted subtrees can be processed (such as integrated / splittered) based on the difference data between vulnerability analysis and fault analysis to construct a logical fault subtree for vulnerability analysis (which can be denoted as the first logical fault tree) and a logical fault subtree for fault analysis (which can be denoted as the second logical fault tree). This application does not limit the implementation method.

[0084] In one possible implementation, based on the above analysis, since the focus of fault analysis (fault diagnosis) and vulnerability analysis of cloud computing services differs, the heterogeneous spatiotemporal data on which the first logical fault tree and the second logical fault tree are constructed, as well as the attribute information used for attribute annotation, will differ. This application can construct the first logical fault tree and the second logical fault tree based on these differences, referring to the logical fault tree construction method described above. The implementation process can be referred to the description in the corresponding part of the above embodiments, and will not be detailed here. It is understood that this application can perform attribute annotation on the first logical fault tree to generate a first target fault tree, and perform attribute annotation on the second logical fault tree to generate a second target fault tree.

[0085] Optionally, this application can determine the necessary attributes to be assigned based on the spatiotemporal data status at the corresponding time, and, in conjunction with the corresponding monitoring time series data, mark the assignment judgment rules and thresholds to the relevant intermediate events and basic events, thereby realizing the attribute labeling of the corresponding subtrees for subsequent analysis activities. Regarding the attribute information content of the event objects labeled in the corresponding logical fault trees during the generation process of these two target fault trees, please refer to the description of the attribute labeling process in the above embodiment; this embodiment will not repeat it here. In practical applications, this application can determine whether to integrate the subtree into the corresponding first or second logical fault tree based on the actual state changes of each intermediate event contained in the corresponding subtree during different reliability analyses of the target fault tree. For example, during fault detection, if certain indicators in the intermediate event exceed the threshold, the corresponding subtree is executed, and the subtree can be integrated with the intermediate event, but it is not limited to this definition of subtree integration.

[0086] Reference Figure 5This is a flowchart illustrating the fault tree generation method for cloud computing services proposed in Embodiment 4 of this application. This embodiment allows for the modification of the logical fault tree / subtree to ensure its reliability and accuracy. Therefore, after constructing the logical fault tree / extracting subtrees from it (which can be used as the logical fault tree to perform the following steps) according to, but not limited to, the method described above, as follows... Figure 5 As shown, it may also include, but is not limited to:

[0087] Step S51: Perform a syntax check on the logical fault tree; the syntax check includes at least one of the following: redundancy structure check, resource dependency check, and common cause event identification check.

[0088] Step S52: In response to the syntax check result being unqualified, correct the logic fault tree;

[0089] In this embodiment, after initially constructing the logical fault tree / subtree according to the methods described in the above embodiments, rigorous syntax checking and optimization can be performed. This includes identifying and correcting potential problems in the logical fault tree, enabling it to more accurately reflect the fault logic relationships of the first cloud platform system and ensuring the accuracy and effectiveness of the corrected logical fault tree. Based on this, this application can primarily perform syntax checks on at least one of the following three aspects: redundant structure, resource dependencies, and common cause event identification (correct identification of common subtrees and common basic events) to ensure the logical completeness, rationality, and feasibility of the fault tree. However, it is not limited to these aspects and can be flexibly configured or adjusted according to actual needs. Redundancy is introduced into the logical fault tree to improve system reliability. It consists of functionally duplicated resources (components or paths) so that if one component fails, a backup component can take over, thereby preventing overall system failure. This application can combine the type of logic gate to identify and process the corresponding redundant structure.

[0090] Optionally, this application can perform a full scan of the logic fault tree. For multiple AND gates (or multiple OR gates), if their input events are exactly the same, these gates can be considered to constitute a redundant structure. In this case, one AND gate (or OR gate) can be retained, and the parent node pointers of these multiple AND gates (or OR gates) can be redirected to the retained logic gate. Other AND gates (or OR gates) can be configured with "fold" / "skip" flags so that when traversing the logic fault tree later, the corresponding logic gate can be skipped based on the flag (i.e., logic gates with the flag are not traversed). For multiple K / N gates, if their (K, N) parameters and input sets are the same, these multiple K / N gates are determined to be a redundant structure and can be merged. For example, one K / N gate can be retained, other redundant K / N gates can be deleted, and the pointers that originally pointed to the deleted K / N gates can be adjusted to point to the retained logic gates. For logic gates that have a parent-child relationship across layers, the BDD (Binary Decision Diagram) method can be used to identify redundant structures across layers. Shared nodes can be merged (i.e., nodes with the same Boolean sub-functions that appear across layers). At this time, the shared nodes can be marked as a subtree, and all parent pointers that originally pointed to non-shared nodes can be mapped to this subtree. The same function can be expressed by this subtree to ensure that the minimum cut set remains unchanged.

[0091] It should be noted that this application does not limit the methods for identifying and processing redundant structures in logical fault trees. For different reliability analysis scenarios—vulnerability analysis and fault analysis—the methods for identifying redundant structures in logical fault trees can differ. This application does not limit the methods for identifying redundant structures in different scenarios. Optionally, for the first logical fault tree used for vulnerability analysis, different subtrees or basic events can be considered the same structure and identified as redundant based on their type or reliability statistics, such as conforming to a Weibull distribution or exponential distribution and having consistent related parameters. For the second logical fault tree used for fault analysis, the structure of different subtrees or basic events can be judged based on other attribute information besides the attribute information corresponding to vulnerability analysis, such as IP, UUID, and judgment thresholds, to determine the redundant structures in the second logical fault tree. Afterwards, methods described above can be used, but are not limited to, to perform a preset processing strategy on the redundant structure after identifying it. The redundancy processing strategy can differ for different types of redundant structures and can be determined based on prior knowledge of redundancy processing; this application does not impose any restrictions. For example, redundant structures (logical redundancy) of the type of repeated subtrees can be extracted into shared subtrees to achieve global reference; redundant structures (logical redundancy) of the type of redundant logic gates can be simplified using Boolean functions or the simplification process described in the above examples; and redundant structures of the type of unmodeled physical redundancy can be supplemented with redundant logic based on the architecture document of the first cloud platform.

[0092] Optionally, for the identification of redundant structures, this application may employ, but is not limited to, graph-based detection methods to identify redundant structures in the logical fault tree. It may also use a minimum cut set verification method to verify the legitimacy of the identified redundant structures, such as verifying whether the number of subtrees or basic events and the causal logical relationships in the redundant structure meet the pre-configured legal conditions for the corresponding type of redundant structure (also known as constraint information on subtrees and logical relationships). These legal conditions can be determined by analyzing spatial data from the first cloud platform, or by analyzing spatial data under different business loads at different time points, to reliably achieve the abstraction and merging of redundant structures in the logical fault tree. Legal conditions may include, but are not limited to: the redundant structure should ensure that the minimum cut set of the critical service is greater than or equal to 2, i.e., at least two independent failure events will lead to system failure; or, the number of subtrees or logical events of AND gates and OR gates is greater than or equal to 2. K / N gates require no less than K subordinate subtrees or logical events, etc. This application does not limit the content of the legal conditions or the method of obtaining them.

[0093] Based on the above analysis, the redundancy structure check of the logic fault tree / subtree can be implemented based on the legal structure corresponding to each type of redundancy (i.e., determined according to legal conditions). For example, a suitable detection algorithm can be selected to detect the redundant structure in the logic fault tree and compare it with the legal structure of the corresponding redundancy type to obtain the corresponding syntax check result. If the syntax check result is unqualified, that is, the detected redundant structure is inconsistent with the legal structure, the corresponding unqualified redundant structure (i.e., illegal redundant structure) in the logic fault tree can be corrected based on the legal structure, such as merging / deleting / supplementing redundant structures, etc., without any restrictions.

[0094] In some embodiments, due to the complex nature of resource calls and dependencies in a cloud computing environment, any unreasonable resource dependencies can lead to deviations in fault diagnosis. Therefore, this application can also perform a rationality check on the dependencies (i.e., resource dependencies) between each fault event represented by the logical fault tree and the basic resources provided by the first cloud platform to correct unreasonable dependencies in the logical fault tree. To this end, this application can compare the resource dependencies defined in the logical fault tree with the resource call records during actual cloud computing operation (which come from the heterogeneous spatiotemporal data of the currently acquired first cloud platform) to verify whether the logical fault tree accurately reflects the actual system operation. For example, in a multi-tenant cloud computing environment, applications from different tenants may share certain public resources. If the logical fault tree incorrectly associates fault events of these public resources with only one tenant's application, this unreasonable resource dependency will directly affect the accuracy of fault diagnosis. Through the above comparative analysis, such problems can be identified and corrected in a timely manner, ensuring that the resource dependencies in the logical fault tree are consistent with the actual system operation. In terms of actual operation, the analysis can be performed by analyzing the currently acquired spatial data or by converting the spatial relationships of the current first cloud platform into an RBD (Reliability Block Diagram), and then performing logical fault tree analysis according to the above method, but it is not limited to this.

[0095] In practical applications, this application can, according to the methods described above, determine whether resource dependencies have a valid structure during the closed-loop detection process of dependencies between various fault events represented by the logical fault tree. If the detected resource dependencies are directed acyclic graphs, they can be determined to be valid structures. Furthermore, resource dependency checks can be performed on the logical fault tree based on pre-configured syntax requirements for different types of resource dependencies (strong dependencies, weak dependencies, and dynamic dependencies, etc.). This includes correcting resource dependencies in the logical fault tree that do not meet syntax requirements (e.g., using OR gates to connect strong dependencies, AND gates to connect weak dependencies, and no time limit labeling). The implementation process includes, but is not limited to, the content of this example.

[0096] In some embodiments, since common subtrees or common basic events may be referenced in multiple locations within a logical fault tree, corresponding identifiers (which can be called common cause event identifiers) are typically assigned to them to ensure reliable and accurate implementation of these references. Furthermore, the usage of these common cause event identifiers can be recorded during the logical fault tree construction process to track the location of each subtree and basic event within the logical fault tree. If a subtree or basic event is found to be referenced in multiple locations without being correctly identified as a common element (common cause event), it should be corrected promptly to ensure the correctness of the common cause event identifiers, thereby improving the accuracy of the logical fault tree and providing a reliable foundation for subsequent fault probability quantification analysis.

[0097] It is evident that the correct identification of common subtrees or common basic events directly affects the accuracy and reliability of their references. Incorrect identification may lead to duplication or omission in fault probability calculations. Therefore, during the syntax checking of the logical fault tree, this application can determine the correctness of corresponding objects in the logical fault tree by checking the common cause event identification. This involves verifying the correct identification of common factor trees in the logical fault tree, specifically verifying the assigned identifiers of common subtrees and common basic events that appear repeatedly in different branches of the logical fault tree, and correcting the assigned identifiers of common subtrees or common basic events that are not correctly identified. Common subtrees requiring identifier allocation can represent multiple top events (or intermediate events) sharing the same fault propagation path, such as multiple microservices depending on the same database cluster; or storage services across availability zones (AZs) sharing the same network middleware. Common basic events can be atomic-level fault events repeatedly referenced at multiple locations in the logical fault tree, indicating that the same underlying fault affects multiple upper-level components. This application does not restrict the content of each common subtree and common basic event in the logical fault tree; it can be determined as appropriate.

[0098] It should be noted that, in the process of performing syntax checking on the logic fault tree, this application may employ one or any combination of two or three of the three implementation methods described above. Furthermore, after correcting the logic fault tree based on the syntax checking results, a target fault tree can be directly generated based on the corrected logic fault tree and heterogeneous spatiotemporal data. This generation process can refer to the descriptions in the corresponding sections of the above embodiments, but is not limited thereto. Depending on actual needs, the verification methods described in steps S53-S54 and / or steps S55-S56 can be executed to improve the completeness and reliability of the corrected logic fault tree.

[0099] Step S53: Based on the spatial data of the first cloud platform currently acquired, verify each fault event represented by the logical fault tree;

[0100] Step S54: In response to the verification result being unqualified, correct the logic fault tree;

[0101] Based on the above analysis of the spatial data of the First Cloud Platform, which describes the physical location and connectivity of various components in the cloud computing environment, such as the geographical location of the data center, the arrangement of server racks, and the connection topology of network devices, this application compares and analyzes the various fault events represented by the logical fault tree with the currently acquired spatial data (which may differ from the spatial tree used to construct the logical fault tree) to verify whether the logical fault tree correctly reflects the spatial relationships of the components. For example, it checks whether the events representing data transmission failures in the logical fault tree accurately correlate with the location relationships of the corresponding network devices and servers within the data center. If any deviations or omissions are found in the spatial relationship description in the logical fault tree, the logical fault tree is immediately corrected to ensure the accuracy of its spatial relationships. This correction may include, but is not limited to: adjusting the corresponding causal logical relationships (connection relationships); adding / removing fault events and configuring / adjusting the corresponding causal logical relationships, etc.

[0102] Based on the above analysis, in one possible implementation, step S63 may include, but is not limited to: determining the location information and connection relationships of different components in the current first cloud platform based on the spatial data of the first cloud platform currently acquired (the acquisition method is similar to the spatial data acquisition method in step S11, and may come from one or more data sources). Optionally, this application can utilize the spatial data to construct a graph-structured spatial data model, where nodes in the graph represent components (devices or resources) at different locations, and edges in the graph represent the connection relationships between corresponding two components, so that the location information and connection relationships of each component can be quickly and completely determined through graph traversal. However, this is not limited to this graph-based representation of spatial data. Afterwards, each fault event represented by the logical fault tree can be matched with the location information of different components to determine whether each fault event in the logical fault tree matches the location information of the corresponding component in the spatial data model. For example, the fault event "server overheating fault" in the logical fault tree should be associated with the location information of the corresponding server in the spatial data model, but this association may cease due to changes in the server's location. This application uses this matching check to promptly detect and correct connection relationships in the logical fault tree that do not conform to the actual spatial relationship. Similarly, this application can also check whether the connection relationship between different fault events in the logical fault tree is consistent with the connection relationship between the corresponding components (corresponding nodes in the figure) in the spatial data model (i.e. the currently obtained actual spatial relationship), so as to promptly discover the connection relationship in the logical fault tree that does not match the actual spatial relationship, such as the connection path between two fault events not existing in space, and promptly and specifically correct the corresponding connection relationship in the logical fault tree.

[0103] Therefore, in one possible implementation, this application can first match the location information of each fault event represented by the logical fault tree with that of different components (components in the spatial data model). Upon successful matching, the consistency verification of the connection relationships between the successfully matched fault events and the connection relationships between the components is performed, correcting any inconsistent connection relationships in the logical fault tree. Upon failure to match, the logical fault tree is corrected based on the failed component and the connection relationships for that component. In another possible implementation, this application can also simultaneously verify the consistency of the connection relationships between the fault events represented by the logical fault tree and the connection relationships indicated by the spatial data while matching the location information of each fault event represented by the logical fault tree with that of the components. Then, based on the matching results of the location information and the consistency verification results of the connection relationships, the logical fault tree is corrected. The implementation processes for each step are similar and will not be detailed in this application.

[0104] In summary, this application can verify the dynamic adaptability of the logical fault tree based on actual changes in spatial data, using methods not limited to those described above. For example, in a cloud computing environment, after constructing the logical fault tree according to the methods described above, changes such as the addition or removal of devices or adjustments to the network topology may cause inconsistencies between the logical fault tree constructed based on the original spatial data and the current actual spatial relationships. The aforementioned checking methods can be used to dynamically and adaptively adjust the logical fault tree to ensure that the adjusted logical fault tree reflects the actual spatial relationships. Optionally, this application can also simulate changes in spatial data based on log information such as the addition or removal of devices and adjustments to the network topology. Afterwards, the checking methods described above can still be used to determine whether the logical fault tree can reflect these changes, so as to update the structure of the logical fault tree in a timely manner. For example, when a new server is added to the first cloud platform and connected to a network switch, the application can verify whether the logical fault tree can correctly add fault events related to that server and adjust the connection relationships with other fault events. In cases where simulations determine that changes in spatial data have led to inaccuracies in the constructed logical fault tree, corresponding inspection prompts / instructions can be output to verify and correct the logical fault tree by executing the logical fault tree inspection steps described above. Furthermore, an inaccurate logical fault tree may also cause subsequent target fault tree modeling to fail; in such cases, logical fault tree error prompts can be output to return to the aforementioned inspection steps for verification and correction. This application does not limit the content of each prompt or its output method; it can be determined as appropriate.

[0105] In some embodiments, during the time-varying process of cloud computing, the spatial data of the first cloud platform may change, which will lead to inaccuracies in the logical fault tree and target fault tree constructed by combining the spatial data. To address this, this application proposes incremental modeling of the logical fault tree and / or target fault tree based on the changes in the spatial data (which may be generated within a preset time interval). The former can generate a new target fault tree based on the new logical fault tree obtained from the incremental modeling. The fault tree previously constructed for the final generated fault tree can be a fuzzy fault tree. Thus, the fuzzy fault tree can be applied to compare and analyze spatial data generated at different times (within different time intervals) to achieve incremental modeling, etc.

[0106] Based on the above analysis, in one possible implementation, this application can construct a current spatial data model with a graph structure based on the spatial data of the first cloud platform currently acquired; perform graph similarity analysis on the current spatial data model and the historical spatial data models of adjacent historical times to obtain the similarity analysis results. For example, a suitable graph similarity calculation method can be selected to perform consistency detection on spatial data models constructed at different times, that is, to perform graph similarity analysis on the graphs of spatial data at different times to determine whether the graphs at different times are consistent and identify the inconsistent parts, that is, the parts with low graph similarity. Optionally, this application can construct multiple subgraphs from the spatial data to perform similarity analysis on the same type of subgraphs at different times to determine whether the subgraphs at different times are consistent, etc. This application does not limit the similarity analysis method between different graphs (subgraphs or complete graphs).

[0107] Subsequently, based on the similarity analysis results, incremental modeling can be performed on the target fault tree to obtain a new target fault tree. In other words, if the similarity analysis determines that the graphs at different times are consistent, it indicates that the spatial data at these two times is essentially unchanged, and incremental modeling is unnecessary. The fuzzy fault tree modeled at the previous time can be used as the fuzzy fault tree at the current time to detect changes in spatial data between the current and next time moments. Conversely, if the spatial data at the current time has changed compared to the previous time, the target fault tree needs to be reconstructed. In this case, the parts with low graph similarity (e.g., less than a preset threshold) identified based on the similarity analysis results—that is, the changed parts of the target fault tree constructed at the previous time—can be locally modeled and updated to achieve incremental modeling of the target fault tree, resulting in a new target fault tree for the current time. If the changed parts are large, to avoid large errors caused by incremental modeling, the logical fault tree constructed at the current time can also be used to regenerate the target fault tree, i.e., full modeling.

[0108] Optionally, in practical applications, this application can perform target fault tree modeling at regular intervals (e.g., m hours, which can be set according to actual needs or experience). During this interval, if the spatial data changes, incremental modeling of the target fault tree can be performed according to the method described above to ensure a balance between modeling quality and efficiency. After a period of time / number of incremental modeling operations, a full modeling operation can be performed to eliminate the errors introduced by incremental modeling and improve the reliability of the target fault tree.

[0109] Step S55: Based on the currently acquired monitoring time-series data and logical fault tree of the first cloud platform, verify each fault event represented by the logical fault tree;

[0110] Step S56: In response to the verification result being unqualified, correct the logic fault tree.

[0111] In the process of checking the logical fault tree by combining monitoring time-series data, the event status of the corresponding fault events in the logical fault tree can be verified by combining the associated monitoring time-series data during / after the above-mentioned spatial data-based check, so as to ensure that the logical fault tree can accurately reflect the system status. The implementation process is not detailed in this application. In some embodiments, the process of checking the rationality / accuracy of the logical fault tree can also be implemented independently based on the currently acquired monitoring time-series data of the first cloud platform. For example, by performing correlation analysis between the monitoring time-series data and each fault event represented by the logical fault tree, it can be verified whether each fault event represented by the logical fault tree is associated with monitoring time-series data. It can also be verified whether the logical fault tree can accurately reflect the actual operating status of the first cloud platform system at different times. The implementation method is not limited in this application.

[0112] In one possible implementation, this application considers the enormous computational burden of processing full-volume monitoring time-series data, which impacts processing efficiency and accuracy. Therefore, this application proposes employing a suitable data sampling strategy to extract a portion of the time-series data from the full-volume monitoring data, integrate it with a logical fault tree, and combine it with corresponding spatial data to evaluate whether the logical fault tree is suitable for fault diagnosis scenarios. It should be noted that the extracted time-series data needs to cover key performance indicators (KPIs) and monitoring metrics, such as server CPU utilization, memory usage, disk I / O, and network traffic, during the operation of the first cloud platform system to ensure verification reliability.

[0113] The aforementioned data sampling strategy can be configured based on factors such as the frequency of fault occurrence, the system's operating cycle, and the importance of the business. For example, a higher sampling frequency is used for critical business systems to ensure that more fault information can be captured; while for non-critical business systems, the sampling frequency is appropriately reduced to decrease the amount of data processing. Therefore, in the process of configuring the data acquisition strategy, the data sampling characteristics of different cloud computing services executed by the first cloud platform can be determined based on heterogeneous spatiotemporal data. These data sampling characteristics may include at least one of the following indicators: fault occurrence frequency, system operating cycle, and business importance. Subsequently, a data sampling strategy for the first cloud platform can be configured based on these data sampling characteristics so that the data sampling strategy can characterize the data sampling frequency of different cloud computing services. The data sampling frequency of the target cloud computing service (critical / core business) is higher than that of the non-target cloud computing service (non-critical / non-core business). This application does not impose any restrictions on the values ​​of each data sampling frequency and can be determined as appropriate.

[0114] Optionally, to further improve the accuracy of subsequent checks on the logical fault tree, this application may also perform preprocessing operations such as cleaning and normalization on a portion of the time-series data extracted from the currently acquired monitoring time-series data to eliminate noise and outliers in the data, ensuring the quality of the preprocessed monitoring time-series data. Then, the preprocessed monitoring time-series data is used to verify the status of each fault event represented by the logical fault tree, so as to correct the logical fault tree and ensure that the corrected logical fault tree can reflect the actual abnormal state of the system in a timely and reliable manner.

[0115] Based on the above, in the implementation of steps S55 and S66, current monitoring time-series data can be sampled from at least one data source based on a pre-configured data sampling strategy. The values ​​of each resource indicator in the monitoring time-series data are then determined. This can be done by preprocessing the sampled monitoring time-series data using methods such as data cleaning and data normalization, before determining the resource indicator values. Afterward, the fault events represented by the logical fault tree can be correlated with each resource indicator value to identify the first fault event in the logical fault tree that is not associated with any resource value. Corresponding marking information is then configured for this first fault event to indicate that it was generated during the construction of the logical fault tree and that no associated resource indicator was actually monitored. Subsequent processing is unnecessary, or relevant personnel can be reminded to conduct manual inspection and processing. Alternatively, in this case, the monitoring data from the corresponding data source can be adjusted to obtain resource indicator values ​​associated with the first fault event, providing data support for the first fault event and facilitating subsequent reliability analysis.

[0116] Optionally, through the above correlation analysis, if it is determined that at least one fault mode is not included in the logical fault tree, and this fault mode was omitted during the construction of the logical fault tree, or if a new fault mode adapted to the logical fault tree is added, the corresponding fault event and causal logic relationship can be added to the logical fault tree to correct it and improve the comprehensiveness and accuracy of the corrected logical fault tree. Of course, if the newly added fault mode is not compatible with the logical fault tree, this correction operation can be avoided to prevent the logical fault tree from becoming less reliable and accurate due to the correction operation, thus affecting subsequent reliability analysis.

[0117] In some embodiments, for logical fault tree checks in reliability analysis scenarios such as fault analysis, in addition to the combined implementation method of one or more checking methods described above, the logical fault tree can also be integrated with the alarm threshold of the monitoring system to achieve necessary analysis of the monitoring data. For example, when the monitoring index exceeds the threshold, it is determined whether the corresponding fault event in the logical fault tree can be accurately triggered (determining whether the logical fault tree covers the fault mode of the corresponding fault event, that is, whether the event state of the fault event is consistent with the fault mode, etc.), thereby ensuring that the logical fault tree corrected accordingly can closely cooperate with the actual monitoring system and reflect the abnormal state of the system in a timely manner. Based on this, in the implementation of the above step S55, the historical alarm information associated with the corresponding fault event represented by the logical fault tree can be verified for consistency based on the actual alarm information determined by the monitoring time-series data of the first cloud platform currently obtained, thereby verifying whether the logical fault tree can accurately reflect the actual fault situation. The historical alarm information may be alarm threshold configuration information extracted from the monitoring system, which may include, but is not limited to, the threshold range, alarm level and associated fault type of each monitoring indicator. This application does not limit its value and can configure or adjust it as appropriate, such as according to the corresponding business needs or alarm needs.

[0118] The extracted alarm threshold configuration information can be associated with corresponding fault events in the logical fault tree. For example, the alarm threshold of "disk remaining capacity less than 10%" can be associated with the "disk space insufficient fault" fault event in the logical fault tree. Simultaneously, during system operation, changes in various monitoring indicators are monitored in real time. When a monitoring indicator exceeds its corresponding alarm threshold, the corresponding fault event in the logical fault tree is triggered, and the occurrence time, duration, and related parameter information of the fault event are recorded. This information is then compared with historical alarm information associated with the corresponding fault events in the logical fault tree to verify whether the logical fault tree accurately reflects the actual fault situation.

[0119] If the triggering conditions of fault events in the logical fault tree are found to be inconsistent with the actual fault phenomena (described by the actual alarm information), i.e., in response to the inconsistency between the historical alarm information and the actual alarm information; or if it is determined that there are missed or false alarms in the fault events in the logical fault tree, i.e., in response to the fault events in the logical fault tree not being associated with historical alarm information, the causal logic relationship of the corresponding fault events represented by the logical fault tree can be corrected in a timely manner. That is, the structure and parameters of the logical fault tree are adjusted, and the triggering conditions and connection relationships of the fault events are optimized so that the corresponding fault events in the logical fault tree are associated with the actual alarm information of the current system (updating historical alarm information to enable the next round of checks on the logical fault tree), ensuring that the corrected logical fault tree can reflect the actual abnormal situation of the current system. After correcting the logical fault tree, the corrected logical fault tree can be attribute-labeled according to, but not limited to, the methods described above, and then combined with heterogeneous spatiotemporal data to generate a corresponding target fault tree for reliability analysis of cloud computing services. The implementation process is not described in detail in this embodiment.

[0120] Optionally, this application can also perform syntax checks on the initially constructed logical fault tree, and correct syntax and logic errors in the logical fault tree based on the syntax check results, and then annotate the attributes of the corrected logical fault tree. Afterwards, based on the spatial data and monitoring time-series data of the currently acquired first cloud platform, the logical fault tree is checked separately. During this check, in addition to verifying and correcting the spatial relationships and fault modes as described above, the attribute information annotated in the logical fault tree can also be verified to correct any inconsistencies in the original attribute information, ensuring that the attribute information annotated in the corrected logical fault tree is accurate and reliable, and consistent with the actual attribute information of the corresponding resources and components in the current first cloud platform. Of course, after correcting the logical fault tree using these two checking methods, for logical fault trees that require integration with operation and maintenance tools, the corrected logical fault tree is annotated with corresponding operation and maintenance information, such as alarm thresholds, object IPs, and related information, according to the attribute annotation method described above. For logical fault trees that integrate operation and maintenance tools, such as anomaly detection algorithms, different monitoring indicators need to be integrated. These monitoring indicators may come from the same system or different systems. This application needs to analyze their combination to improve the completeness and controllability of the logical fault tree used for fault analysis.

[0121] It should be noted that for the initially constructed logical fault tree, after extracting each logical fault subtree according to the method described above, the inspection and correction of each logical fault subtree can also be achieved by combining one or more inspection methods described above. Then, the corrected logical fault subtrees are merged into a logical fault tree to generate the target fault tree. The implementation process will not be elaborated in this application.

[0122] Reference Figure 6 The flowchart shown in Embodiment 5 of this application illustrates the fault tree generation method for cloud computing services. Based on the fault tree generation methods for cloud computing services described in the preceding embodiments, this embodiment describes an optional implementation method for generating a target fault tree for a first cloud platform based on heterogeneous spatiotemporal data and logical fault trees. Figure 6 As shown, this optional implementation method may include, but is not limited to:

[0123] Step S61: Extract unit spatial data from the spatial data of the first cloud platform; the unit spatial data describes the location information and connection relationships of the target resources provided by the first cloud platform.

[0124] In this embodiment of the application, in order to improve efficiency and reduce resource consumption, a method is proposed to construct minimum spatial data (i.e., cell spatial data, the amount of data contained in which this application does not limit) to generate small-sized cell fault trees for reliability analysis. By performing multi-dimensional checks on them (such as the method of directly checking the logic fault tree described in the above embodiment) and tracing back the syntax and logic errors of the logic fault tree, this approach can not only efficiently discover and correct problems in the logic fault tree, but also gradually improve the quality of the logic fault tree, laying a solid foundation for the subsequent generation and analysis of large-scale target fault trees.

[0125] This application uses the most basic spatial data reflecting the core structure and relationships of the system as unit spatial data to form a simplified spatial data model. Optionally, extraction methods such as querying cloud platform databases, configuration management databases (CMDB), or using lightweight network scanning tools can be used to extract only the connection and location information of key servers, core network devices, and main storage systems, thus obtaining the key nodes and connection relationships of the system and forming a minimal spatial data. Therefore, the target resource in step S81 can be key nodes such as key servers, core network devices, and main storage systems; this application does not restrict the type of resource.

[0126] Step S62: The unit space data is fused with the data in the logical fault tree to generate a unit fault tree; the unit fault tree contains the target fault modes and target fault propagation paths for the target cloud computing services supported by the first cloud platform.

[0127] Similar to the implementation process of generating a target fault tree based on logical fault trees and heterogeneous spatiotemporal data described in the above embodiments, this application can integrate the extracted unit space data with the logical fault tree (which can be the initially constructed logical fault tree, or the logical fault tree after the above checks and corrections, etc.). At this time, the parts of the logical fault tree related to the core business (target cloud computing business) process and key system components (target fault modes) can be selected, and those complex branches related to secondary businesses or non-critical components can be ignored to generate a small-sized complete fault tree, denoted as a unit fault tree. This is equivalent to a simplified version of the template fault tree. Compared with the target fault tree generated by the method in the above embodiments, the size is greatly reduced, but it can cover the target fault modes (main fault modes / critical fault modes) and critical fault propagation paths (causal propagation paths, denoted as target fault propagation paths) of the system.

[0128] Step S63: Check the unit fault tree;

[0129] Step S64: Determine whether any abnormal information has been detected. If yes, proceed to step S65; otherwise, proceed to step S66.

[0130] Step S65: Correct the logical fault tree to return to step S61 and continue execution;

[0131] The process of checking the unit fault tree can be carried out by using fault tree analysis tools or algorithms to check the generated small-sized unit fault tree to analyze whether its structure is reasonable, whether the causal logic relationship between each fault event is correct, and whether there are any potential syntax or logic errors. It also checks whether the unit fault tree can accurately reflect the main fault scenarios of the system (the scenario corresponding to the target fault mode). This application does not limit the implementation methods of each check. Optionally, this application can combine one or more of the various checking methods for the initially constructed logical fault tree or its subtrees described in the above embodiments, such as syntax checking, verification combined with spatial data and / or monitoring time-series data, to check the unit fault tree. The implementation process is not elaborated in this application. Based on this, if the unit fault tree check fails according to any of the above checking methods, the obtained abnormal information may include unreasonable structure, incorrect causal logic relationship, syntax or logic error, or failure to cover at least one key fault mode. This application does not limit the content of the abnormal information and can be determined as appropriate.

[0132] Subsequently, the aforementioned anomaly information can be mapped onto the original logic fault tree. By comparison, the specific locations and causes of these anomalies within the logic fault tree can be determined, such as incorrect logic gate connections, inaccurate event definitions, or omissions of critical resource dependencies, thus enabling targeted correction of the logic fault tree. Since the unit fault tree is very small, compared to checking and correcting large-scale logic fault trees, this embodiment's method of constructing unit fault trees for inspection and analysis to correct the logic fault tree significantly reduces the computational load and time cost of inspection and analysis, thereby improving the efficiency of logic fault tree correction.

[0133] Step S66: In response to the detection that the unit fault tree meets expectations, at least based on heterogeneous spatiotemporal data and the modified logical fault tree, generate a target fault tree for the first cloud platform.

[0134] As analyzed above, this application can iteratively correct the logic fault tree according to the method described above. After each round of correction of the logic fault tree, the unit fault tree can be regenerated and checked according to the methods described in steps S61 and S62 above to verify whether the problems checked in the previous round have been resolved. Through multiple iterative corrections, the quality of the logic fault tree is gradually improved until the unit fault tree meets expectations and accurately reflects the causal logic relationship of the fault events in the system. Since the unit fault tree is relatively small in scale, even after multiple rounds of correction, the total computational load and time cost of checking and analyzing the unit fault tree are relatively low, which improves the generation efficiency of the target fault tree. Moreover, compared with directly checking and analyzing the logic fault tree, this application checks and analyzes the small-scale unit fault tree, which improves the efficiency and reliability of problem discovery. Furthermore, through iterative correction, the quality of the logic fault tree is gradually improved, laying a reliable and solid foundation for the subsequent generation and analysis of large-scale target fault trees.

[0135] It should be noted that, considering the differences in generating target fault trees for vulnerability analysis and fault analysis as described above, the process of generating the target fault tree for fault analysis according to the method described in Embodiment Six also requires the addition of differences in operation and maintenance information. The implementation process will not be elaborated upon in this embodiment. In some embodiments, after processing the initial construction of the logical fault tree (i.e., the original logical fault tree), this application can directly modify the original logical fault tree according to the methods described in steps S61-S65 to generate the target fault tree. The implementation process will not be detailed in this application.

[0136] Reference Figure 7This is a flowchart illustrating the fault tree generation method for cloud computing services proposed in Embodiment Six of this application. Since the modified logical fault tree obtained by the method in the above embodiments is relatively large, to reduce the time and space complexity of target fault tree modeling, the large-scale logical fault tree can be pruned before generating the target fault tree. Then, based on the pruned logical fault tree and heterogeneous spatiotemporal data, a target fault tree for the first cloud platform is generated. Based on this, as... Figure 7 As shown, the pruning method for a logic fault tree can include, but is not limited to, at least one of the following steps:

[0137] Step S71: Based on the logical fault tree and heterogeneous spatiotemporal data, predict the distribution and number of target fault modes in the first target fault tree used for vulnerability analysis, or perform quantitative analysis on the fault events in the second target fault tree used for fault analysis, so as to achieve pruning of the logical fault tree.

[0138] In this embodiment, the target fault tree modeling scale is estimated to prune common failures (subtrees, events) that cannot be modeled within a valid timeframe or whose failure probability is less than a specific level. Therefore, this application can integrate logical fault trees and heterogeneous spatiotemporal data from a first cloud platform to combine the logical structure of system faults provided by the logical fault tree with key information such as the physical location, connectivity, and historical operating status of resources covered by the heterogeneous spatiotemporal data. This allows for a comprehensive understanding of the amount and complexity of data required for modeling. For example, in a large cloud computing center, resources such as servers, storage devices, and network devices are widely distributed, and their interconnections are complex. By integrating the logical fault tree with heterogeneous spatiotemporal data, the interaction relationships between various devices and potential fault propagation paths can be clearly seen.

[0139] In practical applications, considering the differences between the different target fault trees (i.e., the first and second target fault trees mentioned above) used for vulnerability analysis and fault analysis respectively, the size of the different target fault trees in these two scenarios can be evaluated separately. Based on the above analysis, since vulnerability analysis focuses more on the overall reliability level and potential risk points of the system, when evaluating the size of the first target fault tree, key failure modes (i.e., target failure modes) that have a significant impact on system reliability can be identified, and the distribution (whether concentrated or dispersed) and quantity (more or less) of these modes in the complete fault tree can be evaluated. This directly affects the pruning strategy and focus of the first target fault tree, prioritizing the pruning of branches that contribute the most to the probability of the top event occurring. This application does not limit the implementation method. It should be noted that pruning in fault tree analysis refers to selectively ignoring certain branches or fault events to simplify the analysis process, and does not necessarily involve deletion.

[0140] Optionally, if the target failure modes in the first target fault tree are concentrated and few in number, such as only a few (e.g., 1-3) critical failure modes located on the same branch or at the same level, an aggressive pruning strategy can be adopted, focusing on the target failure modes, such as removing all other non-critical branches (temporarily ignored in the analysis). If the target failure modes in the first target fault tree are scattered and few in number, such as a small number of critical failure modes distributed on distinct and unrelated branches, a conservative pruning strategy can be adopted, processing in parallel. For example, if no critical branches (representing an independent, complete path that could lead to the top event) can be pruned, the scattered critical branches can be analyzed separately and in depth to take appropriate measures.

[0141] Similarly, if the target failure modes in the first target fault tree are concentrated and numerous, such as a large number of critical failure modes clustered under a specific subsystem or component, a modular pruning strategy can be adopted to achieve overall replacement, such as treating the entire problem area as a "subtree" or "module" for pruning. If the target failure modes in the first target fault tree are dispersed and numerous, such as a large number of critical failure modes evenly distributed across all branches and levels of the fault tree, a hierarchical pruning strategy can be adopted, combined with quantitative analysis using set thresholds. For example, ignoring failure events with a probability less than the preset expected failure rate; calculating at least one importance index for each event, such as criticality (relative importance, which measures the proportion of the event's contribution to system failure), failure impact (which measures how much the failure probability changes when the probability of the event changes slightly), and response priority, prioritizing failure events with the highest index and ignoring those with lower rankings. However, this is not limited to the pruning strategies described in this application.

[0142] In fault analysis (fault diagnosis) scenarios, it is necessary to quickly locate specific faulty nodes. This requires detailed analysis and quantification of each link in the target fault tree that may lead to the fault, i.e., quantitative analysis of fault events in the second target fault tree. For example, when evaluating the scale of fault diagnosis for a distributed storage system, factors such as the distribution of storage nodes, data redundancy strategies, and network connectivity should be comprehensively considered to determine the parts of the target fault tree that require key attention, so as to avoid pruning these parts. Optionally, this application can also combine quantitative indicators such as the probability of occurrence of basic events, repair events, and costs to strategically ignore fault events that contribute little to the top event, simplify the structure, focus on key issues, and the pruning implementation process is not detailed in this application.

[0143] Therefore, in the implementation of step S71, common failures (subtrees, events) that cannot be modeled within the effective time or whose failure probability is less than a certain level can be identified as pruning objects. Since these failures may involve too much detail or contribute little to the overall analysis, leading to a time-consuming and laborious modeling process with limited practicality, pruning these objects improves modeling efficiency and reduces resource consumption. For example, some rare hardware failure modes may have a negligible impact on the overall system reliability, but require significant time and resources to describe in detail during modeling and model / fault tree analysis. This application can perform pruning operations on these parts according to, but not limited to, the methods described above.

[0144] Step S72: Based on the historical inspection data of the first cloud platform in the heterogeneous spatiotemporal data, obtain the predicted failure probability of each event object contained in the logical fault tree, and prune the logical fault tree based on the predicted failure probability.

[0145] The inspection data can include key information such as the system's operating status and fault occurrences over a period of time, which helps to understand the system's actual operating condition and potential fault risks. Therefore, this application can prune the logical fault tree based on the inspection data. Based on this, the cloud platform can be inspected regularly to record historical inspection information such as the device's operating status, performance indicators, and fault alarms. Based on historical inspection data, the probability of a fault occurring after an inspection within a certain period (predicted fault probability) is estimated, i.e., the predicted fault probability of each event object contained in the logical fault tree is obtained. This application does not limit the implementation method of the predicted fault probability evaluation. For example, by analyzing historical inspection data, it can be found that certain devices have a higher probability of failure (predicted fault probability) under specific conditions, thereby allowing for advance countermeasures.

[0146] Next, the logical fault tree can be pruned by combining historical inspection data and predicted failure probabilities. For example, subtrees or events related to common failures with low predicted failure probabilities and minimal impact on fault diagnosis or vulnerability analysis can be removed from the logical fault tree. For instance, if historical inspection data shows that a certain component of a server has never failed in the past year, and the failure of that component has a minimal impact on the overall system operation, then the subtrees related to that component can be removed from the fault tree to simplify its structure, making it more focused on critical failure modes and improving the model's usability and analytical efficiency.

[0147] Step S73: Perform common cause failure analysis on each branch of the logic fault tree, and prune the logic fault tree based on the obtained common cause failure probability of each branch.

[0148] Based on the above description of common-cause failure analysis, this application can identify and process common failure characteristics in the logical fault tree by estimating and simplifying the failure probability analysis of subsystems of different device types, thereby achieving logical fault tree pruning. Based on this, this application can first estimate the failure probability for subsystems of different device types. Optionally, common-cause failure probability estimation can be performed separately for subsystems of different device types in the first cloud platform (such as server subsystems, storage subsystems, network subsystems, etc.). This can be achieved by analyzing the failure modes and historical failure data of each subsystem to calculate its common-cause failure probability. For example, by statistically analyzing the hardware failure records and software error reports of the server subsystem, key indicators such as MTBF and MTTR of the subsystem can be obtained, and then its common-cause failure probability can be estimated.

[0149] Subsequently, based on the assessed common-cause failure probabilities, a simplified analysis and pruning are performed directly on the subsystem. During the simplified analysis, it's not necessary to expand the subsystem into a detailed fault tree structure; instead, the common-cause failure probabilities of the subsystems are used as a whole for simplification. Based on the estimated common-cause failure probabilities, pruning operations are performed on common subtrees or basic events related to subsystems with low common-cause failure probabilities and minimal impact on the overall fault tree analysis. Intermediate events can also be used to simplify the analysis, avoiding excessive detail and improving analysis efficiency and model simplicity. For example, for a storage subsystem with a low common-cause failure probability, it can be included as a whole node in the fault tree without needing to detail its internal failure modes and events. This simplifies the logical fault tree while maintaining its accurate description of the overall system failure behavior, thereby improving the reliability and accuracy of the generated target fault tree.

[0150] Step S74: Perform fuzzy estimation of the failure probability or importance of each event node in the logical fault tree. Based on the estimated fuzzy values ​​corresponding to each event node, prune the logical fault tree. In response to the anti-fuzzing processing result of the generated target fault tree satisfying the pruning condition, prune the target fault tree based on the failure probability obtained from the anti-fuzzing processing until the anti-fuzzing processing result of the pruned target fault tree satisfies the pruning termination condition.

[0151] In real-world business scenarios, many failure probabilities are difficult to estimate. This application proposes using fuzzy and antifuzzy techniques for parameter-based fault and pruning decisions to prune the logical fault tree. Specifically, this application proposes using fuzzy fault tree technology to perform fuzzy estimation of the logical fault tree to handle uncertain information. Optionally, this application can use fuzzy sets and membership functions to represent the failure probability or importance of each event node as a fuzzy number for fuzzy estimation. For example, for a server overheating failure event, its failure probability can be represented as a triangular fuzzy number (e.g., (0.05, 0.1, 0.2)) based on historical data and expert experience, indicating that the failure probability of the event is approximately between 0.05 and 0.2, with the most likely value being 0.1.

[0152] Subsequently, based on the aforementioned fuzzy estimation, preliminary pruning of the logical fault tree is implemented to consider pruning subtrees or events with low fuzzy fault probabilities or those contributing little to the overall fault tree. For example, if the fuzzy fault probabilities of all basic events in a certain subtree are relatively low, the subtree can be removed from the logical fault tree, thereby simplifying the logical fault tree. Then, based on the pruned logical fault tree and heterogeneous spatiotemporal data, a corresponding target fault tree (denoted as the fuzzy fault tree) is generated. After defuzzification processing, the fuzzy fault tree is converted back into a logical fault tree for quantitative analysis. For example, methods such as the centroid method or weighted average method can be used to convert the fuzzy fault probabilities into specific numerical values. This application does not limit the implementation method of the defuzzification processing. For example, for the triangular fuzzy number (0.05, 0.1, 0.2) of the aforementioned server overheating fault event, the centroid method can be used to calculate its defuzzified fault probability value, thereby obtaining a definite numerical value for subsequent analysis.

[0153] After the aforementioned defuzzification process, the resulting logical fault tree can be further pruned. At this stage, based on the defined fault probability values, subtrees or events in the target fault tree with fault probabilities below a certain set threshold are pruned. For example, if the threshold is set to 0.05, subtrees or events with fault probabilities below 0.05 can be considered to have a relatively small impact on the overall system failure and are thus removed from the target fault tree. This combination of fuzzing and defuzzification techniques with pruning operations allows for better handling of uncertain information during the modeling process from the logical fault tree to the target fault tree, while simultaneously optimizing the model's size and complexity, and improving the efficiency and accuracy of fault diagnosis and vulnerability analysis.

[0154] Based on the above analysis, the fact that the defuzzification processing result in step S74 meets the pruning conditions means that the scale of the logical fault tree is still relatively large, and there are objects that need to be pruned. Pruning can continue according to the above method until the data volume (scale) and complexity of the pruned target fault tree reach the expected level, or the cumulative pruning rounds reach the preset number of pruning termination conditions, etc. This application does not elaborate on the implementation method of step S74 using fuzzy and defuzzy techniques. Therefore, this application effectively reduces the algorithm time complexity in traditional parameter estimation methods and fault tree analysis by combining fuzzy and defuzzy techniques with fault tree pruning. Furthermore, through fuzzy fault tree and target fault tree pruning, and secondary pruning, it further reduces the algorithm time and space complexity, improving the efficiency and ease of use of large-scale fault tree construction.

[0155] In some embodiments, during the time-varying process of cloud computing, the heterogeneous spatiotemporal data of the first cloud platform often changes over time, and the accuracy of the heterogeneous spatiotemporal data directly affects the size and complexity of the target fault tree. For example, when the load of a node has been migrated, or the effective business load has been migrated, but the node is still marked as deliverable when parsing cloud platform data, such a node needs to be removed from the spatial data. Correcting the spatiotemporal data can ensure that it accurately reflects the current actual operating status of the system and avoid model bias caused by inaccurate data. In this regard, this application can comprehensively consider the impact of fault repair status on the spatiotemporal data of the cloud platform and update and correct the spatiotemporal data in a timely manner. The target fault tree size can then be compressed based on the corrected heterogeneous spatiotemporal data. In other words, this application can also respond to resource adjustment operations on the first cloud platform (which may occur between the acquisition of the heterogeneous spatiotemporal data and the current time) to correct the heterogeneous spatiotemporal data. Based on the corrected heterogeneous spatiotemporal data, the target fault tree can be pruned, i.e., removing parts that are no longer applicable or inconsistent with actual operating conditions, thereby reducing the size of the target fault tree and lowering the complexity of modeling. For example, if some network connections are found to be no longer in use in the corrected spatiotemporal data, the parts related to these connections can be removed from the target fault tree, simplifying the model structure and improving modeling efficiency.

[0156] In practical applications, multiple cloud platforms are typically deployed in cloud computing environments, and the architectures of different cloud platforms may differ. This application can obtain the first target fault tree and / or the second target fault tree corresponding to each cloud platform using the method described above. To improve construction efficiency and save modeling resources, transfer learning techniques can be used to apply existing fault tree models to new cloud platforms and then adaptively adjust them, which can effectively save modeling time and achieve knowledge reuse.

[0157] Based on the above analysis, referring to Figure 8The flowchart shown in Embodiment 7 of this application illustrates the fault tree generation method for cloud computing services. This embodiment describes the migration and adjustment of logical fault trees across cloud platforms. After obtaining the logical fault tree of the second cloud platform (which can be the directly constructed original logical fault tree or the checked and corrected logical fault tree, etc.) according to the methods described in any of the embodiments above, an optional implementation method for quickly constructing the logical fault tree of the first cloud platform is described, such as... Figure 8 As shown, the method may include:

[0158] Step S81: The constructed logical fault tree for the second cloud platform is identified as the source logical fault tree; the second cloud platform refers to a cloud platform with an architecture similar to that of the first cloud platform.

[0159] Step S82: Perform a similarity analysis between the spatial data of the first cloud platform and the source logical fault tree to determine the fault events to be adjusted, the causal logical relationships to be adjusted, and the attribute information to be adjusted for the source logical fault tree.

[0160] Step S83: Based on the fault events to be adjusted, the causal logical relationships to be adjusted, and the attribute information to be adjusted, the source logical fault tree migrated to the first cloud platform is adjusted to obtain a logical fault tree for the first cloud platform, which is used to generate a target fault tree for the first cloud platform.

[0161] In this embodiment, when migrating a logical fault tree (i.e., the source logical fault tree, which is the logical fault tree of the source cloud platform, i.e., the second cloud platform, whose construction process can refer to the corresponding part of the above embodiment, i.e., the method of directly constructing the logical fault tree of the first cloud platform) to a new cloud platform (i.e., the target cloud platform, such as the first cloud platform), the architectural spatial information of the first cloud platform can be comprehensively analyzed first, such as the physical layout of the first cloud platform (e.g., data center location, server rack arrangement, etc.), network topology (e.g., the connection relationship of switches and routers), resource allocation (e.g., the distribution of computing resources and storage resources), and the dependencies between various components. Then, a similarity analysis, i.e., a comparative analysis, can be performed on the spatial data of the first cloud platform and the source logical fault tree to determine the similarities and differences between the two, so as to determine the parts in the source logical fault tree that need to be adjusted. This process can focus on the differences in resource types, component layouts, connection relationships, etc., to determine which parts can be directly migrated and which parts need to be adjusted or modified.

[0162] In some embodiments, this application may adjust and optimize the source logic fault tree as necessary according to the architectural characteristics and business requirements of the first cloud platform. This includes adding or deleting certain nodes in the source logic fault tree, modifying the type or connection relationship of logic gates, and updating the attributes and parameters of nodes, so that the adjusted logic fault tree can better adapt to the environment of the first cloud platform. Preferably, the adjusted logic fault tree, i.e., the migrated logic fault tree, is verified to ensure that it accurately reflects the fault events and causal logic relationships of the first cloud platform. In this implementation process, historical fault data and monitoring data of the first cloud platform can be used to test and evaluate the adjusted logic fault tree, check for logical errors or missing fault modes, and further improve the logic fault tree based on the verification results. This checking method may refer to, but is not limited to, one or more of the checking methods described in the above embodiments, and will not be detailed here.

[0163] In some embodiments, to improve the efficiency of logical fault tree migration across cloud platforms, the efficiency and reliability of logical fault tree inspection, the efficiency of constructing target fault trees, and the efficiency and reliability of reliability analysis based on target fault trees, this application can extract common subtrees from the logical fault trees. The method for extracting common subtrees can differ in different scenarios, and this application does not limit the method for extracting common subtrees. Based on this, in the scenario of vulnerability analysis, the target of common subtree extraction may include identifying prevalent vulnerability patterns in the system, which may originate from system design or reliability issues of critical resources. In response to a vulnerability analysis request for cloud computing services, historical fault data and system design in heterogeneous spatiotemporal data are analyzed, a first initial fault tree is constructed based on the identified target fault patterns, and common subtrees for vulnerability analysis are extracted from the first initial fault tree.

[0164] In analyzing historical fault data and system design (i.e., system architecture design documents), in-depth analysis of the system design can clarify the key functional modules and core resource components of the system. Simultaneously, historical fault records are compiled, and the frequency, impact range, and associated system components of various faults are statistically analyzed. Combined with historical fault data analysis, fault modes closely related to the system's key functions or core resources are identified. For example, in storage systems across multiple data centers, if storage device offline issues caused by power failures occur frequently, power failure can be considered one of the key fault modes, i.e., the target fault mode. Then, based on the identified key fault modes, a first initial fault tree containing these fault modes is constructed. During this construction process, logic gates (such as AND gates, OR gates, etc.) and basic event nodes can be used to represent the causal logical relationships between fault events. Furthermore, this application can analyze the constructed first initial fault tree to find subtree structures that repeatedly appear in different fault tree models, i.e., common subtrees. These subtrees typically represent common vulnerability patterns in the system. For example, a common subtree related to storage device power failures may contain nodes such as power equipment failures and power line failures, along with corresponding logical relationships. However, it is not limited to this method of extracting common subtrees for vulnerability analysis.

[0165] In fault diagnosis / analysis scenarios, common subtree extraction focuses on identifying common subtrees that have obvious characteristics and diagnostic value when a fault occurs, in order to quickly locate the root cause of the fault. Based on this, in response to fault analysis requests for cloud computing services, various fault modes in heterogeneous spatiotemporal data are classified and statistically analyzed. A second initial fault tree is constructed based on the classification and statistical results of each fault mode, and common subtrees for fault tracing are extracted from the second initial fault tree.

[0166] The analysis and statistics of various failure modes involves collecting various failure cases and diagnostic records that occur during actual system operation. These records can include detailed information such as failure phenomena, diagnostic processes, repair methods, and involved system components. By classifying and organizing these failure cases, representative common failure modes are summarized. For example, in server failure cases, if the server repeatedly fails to operate normally due to operating system crashes, this is identified as a common failure mode. Based on this, a fault tree model, or second initial fault tree, is constructed for fault diagnosis. This model reflects the characteristic manifestations of the failure and the key nodes for diagnosis, enabling rapid location of the root cause. Subsequently, the second initial fault tree is analyzed to identify common subtrees that have obvious characteristics and diagnostic value when the failure occurs. For example, in the server fault tree, common subtrees related to server operating system crashes are extracted. These subtrees contain nodes related to events such as critical operating system process failures and memory overflows. By quickly identifying these common subtrees, the efficiency and accuracy of fault diagnosis can be improved.

[0167] Based on the fault tree generation method for cloud computing services described in the various embodiments above, preferably, referring to... Figure 9 The flowchart shown in Embodiment 8 of this application illustrates the fault tree generation method for cloud computing services. Based on heterogeneous spatiotemporal data of the first cloud platform, a logical fault tree for the first cloud platform is constructed. Then, methods such as... Figure 9 The method shown extracts subtrees in one or more ways to check and correct each subtree using one or more inspection methods, which also facilitates incremental modification of the target fault tree. Of course, this application also directly checks and corrects the constructed logical fault tree using different extracted subtrees. To further improve the reliability and accuracy of logical faults, such as... Figure 9 As shown, this application can also generate a small-sized cell fault tree by constructing the minimum / cell space data, and then check it to correct the logic fault tree. After multiple iterations of correction, the target fault tree can be generated.

[0168] Before generating the target fault tree, to reduce generation efficiency, the logical fault tree can be pruned / compressed using one or more pruning methods described above, but not limited to those used in the above-described pruning approach. This is used to generate the target fault tree for reliability analysis of cloud computing services, avoiding excessively large target models and high generation and analysis time and space complexity, thus improving analysis efficiency and reliability. Since the target fault tree is mapped to the spatial relationship of physical devices, it facilitates maintenance by operations engineers. For example, when the user experience of cloud computing services on the first cloud platform is slow, engineers can obtain the computing nodes of the load user applications, acquire spatial information from OpenStack, and combine this with the JSON file of the generated logical fault tree and the data holding information of the time-series data to generate the target fault tree in graph database A. The generation process can be referred to the description in the corresponding section of the above embodiment. Finally, the time-series data from the monitoring system is accessed, facilitating fault location analysis of other sub-projects by reading the target fault tree in the graph database.

[0169] In practical applications, when constructing fault trees in a cluster environment, the scale of redundant nodes changes continuously with the time-varying environment of cloud computing, causing the fault tree structure to change. To address this, this application uses reliability analysis and software modeling to progressively extract highly repetitive subtrees and model them to obtain a logical fault tree. Based on this, in the process of generating the target model, a logical fault tree with relatively low redundancy can be constructed to quickly complete a cloud computing fault tree suitable for time-varying characteristics. The target fault tree is then dynamically generated according to the cloud computing architecture at different times. The implementation process is not detailed in this application.

[0170] It should be understood that, in conjunction with the above-described solution, several known solutions involving dynamic fault tree construction are relevant, such as the dynamic fault tree construction methods extracted in current research on the security and reliability of cloud-based secure computers, and the security and reliability analysis models constructed accordingly. In known reliability assessment methods for intelligent ship power system monitoring cloud platforms, FTA is used to establish fault tree models for each component of the monitoring cloud platform. In the field of intelligent fault diagnosis for distributed photovoltaic power station equipment, knowledge graphs, fault tree analysis, and Bayesian inference are used to automatically infer complex fault relationships, providing accurate fault causes and solutions. Furthermore, in known solutions for remote diagnosis of new energy vehicles based on cloud computing, the cloud server uses dynamic fault tree analysis for fault location, and finally pushes the fault diagnosis results for remote control. In reliability modeling methods for cloud computing systems considering common-cause faults, the existence probability of simplified state combinations of similar single servers is proposed using the fault tree method, along with current fault labeling schemes in operation and maintenance systems, and known schemes that select reliability models based on the state of virtualized software components to perform reliability calculations. Other solutions, such as the hybrid model of general fault trees and anomaly detection used in containerized microservices, or the hierarchical availability model using fault trees and Markov chains to achieve reliability analysis, involve dynamic fault trees. However, these differ from the control logic of this application, which first constructs a logical fault tree based on heterogeneous spatiotemporal data and then combines it with heterogeneous spatiotemporal data to generate a target fault tree for reliability analysis. Furthermore, they cannot achieve the characteristics of this application, such as cross-platform migration based on an already constructed logical fault tree to improve fault tree construction efficiency. They cannot achieve the technical effects of this application (which can be described in conjunction with the corresponding parts of the above embodiments).

[0171] Reference Figure 10 This is a schematic diagram of a fault tree generation device for cloud computing services provided in an embodiment of this application. The fault tree generation device for cloud computing services may include:

[0172] The heterogeneous spatiotemporal data acquisition module 101 is used to acquire heterogeneous spatiotemporal data from the first cloud platform; the heterogeneous spatiotemporal data includes spatial data and monitoring time series data from different data sources.

[0173] The determination module 102 is used to determine, based on the heterogeneous spatiotemporal data, various fault events existing in the first cloud platform, and the causal logical relationships between the various fault events;

[0174] The logical fault tree modeling module 103 is used to perform fault tree modeling on each of the fault events and the causal logical relationship to obtain at least one logical fault tree.

[0175] The target fault tree generation module 104 is used to annotate the attributes of each event object contained in the at least one logical fault tree based on the heterogeneous spatiotemporal data, and generate a target fault tree for the first cloud platform; the target fault tree is used for reliability analysis of the cloud computing service.

[0176] Those skilled in the art will understand that the functions and technical effects of each module in the above device embodiments, as well as the units that implement the functions, are equivalent to the corresponding steps described in the foregoing method embodiments. For specific implementation details, please refer to the description in the method section. The device embodiments have not been described in detail.

[0177] This application also provides a computer-readable storage medium carrying one or more computer programs. When these programs are executed by a computer device, the computer device can implement any of the fault tree generation methods for cloud computing services provided in this application. The computer-readable storage medium can be any available medium that the computer device can store, or a data storage device such as a training device or data center that integrates one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

[0178] This application also provides a computer program product comprising one or more computer-readable instructions. When the computer-readable instructions are executed on a computer device, the computer device causes the computer device to implement any of the fault tree generation methods for cloud computing services provided in this application. When the computer-readable instructions are loaded and executed on the computer device, all or part of the processes or functions described in this application are generated. The computer device may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer-readable instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer-readable instructions may be transmitted from one website, computer, training device, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means, depending on the actual application scenario.

[0179] Reference Figure 11 This is a schematic diagram of the hardware structure of a computer device suitable for a fault tree generation method in cloud computing services. This computer device can be a cloud server or a terminal device with data processing capabilities. Figure 11As shown, the computer device may include, but is not limited to, at least one communication element 111, at least one memory 112, and at least one processor 113, wherein the at least one communication element 111, at least one memory 112, and at least one processor 113 can communicate with each other via a bus. The bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 11 The bus is represented by a single bidirectional line, but this does not mean that there is only one bus or one type of bus.

[0180] The communication element 111 can be used to acquire heterogeneous spatiotemporal data from the first cloud platform, including spatial data and monitoring time-series data from different data sources. It can also be used to transmit data or instructions between internal components of a computer device. In this embodiment, the communication element 111 may include ports supporting different wireless communication networks, as well as ports for implementing internal communication within the computer device.

[0181] The memory 112 can be used to store multiple computer instructions for implementing the fault tree generation method for cloud computing services proposed in this application embodiment; the processor 113 can load and execute the multiple computer instructions stored in the memory 112 to implement the various steps of the fault tree generation method for cloud computing services proposed in this application embodiment. The implementation process can be referred to the description of the corresponding part of the method embodiment above. In this application embodiment, the memory 112 may include at least one volatile memory, and may also include at least one non-volatile memory or other storage medium. The processor 113 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), or other processors.

[0182] It should be understood that, Figure 11 The structure of the computer device shown does not constitute a limitation on the computer device in the embodiments of this application. In practical applications, the computer device may include more than Figure 11 The more or fewer components shown, or combinations of certain components, are not listed in detail in this application.

[0183] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

Claims

1. A fault tree generation method for cloud computing services, the method comprising: Acquire heterogeneous spatiotemporal data from the first cloud platform; The heterogeneous spatiotemporal data includes spatial data and monitoring time-series data from different data sources; Based on the heterogeneous spatiotemporal data, the various fault events existing in the first cloud platform and the causal logical relationships between the various fault events are determined. Fault tree modeling is performed on each of the fault events and the causal logical relationships to obtain at least one logical fault tree; Based at least on the heterogeneous spatiotemporal data, attribute annotation is performed on each event object contained in the at least one logical fault tree to generate a target fault tree for the first cloud platform. The target fault tree is used for reliability analysis of the cloud computing service.

2. The method according to claim 1, wherein determining, based on the heterogeneous spatiotemporal data, each fault event existing in the first cloud platform, and the causal logical relationship between each fault event, comprises: Based on the heterogeneous spatiotemporal data, the components at each level used by the target cloud computing service supported by the first cloud platform and the interaction relationships between different components are determined, as well as the dependency relationships between the target business links in the target cloud computing service and the basic resources of the first cloud platform. Based on at least one of the historical fault data determined by the heterogeneous spatiotemporal data and the preset fault rules, the execution process of the target cloud computing service is analyzed using the interaction relationship and the dependency relationship to determine the target fault mode and potential fault points. The target failure mode is determined based on the frequency of occurrence and / or the level of impact. Based on the target failure mode and the potential failure points, determine the various failure events existing in the first cloud platform, as well as the causal logical relationships between the various failure events.

3. The method according to claim 2, wherein determining, based on the heterogeneous spatiotemporal data, each fault event existing in the first cloud platform, and the causal logical relationship between each fault event, further includes at least one of the following: Verify the reasonableness of each of the target failure modes and adjust any target failure modes that do not prove reasonable. For each of the aforementioned fault events, a source tracing analysis is performed, and based on the source tracing analysis results, the corresponding target fault mode or potential fault point is adjusted.

4. The method according to claim 1, wherein performing fault tree modeling on each of the fault events and the causal logical relationship to obtain at least one logical fault tree includes: Each of the aforementioned fault events is identified as a different event node, and a logical fault tree for the first cloud platform is constructed using the logical gates of the corresponding types of the aforementioned causal logical relationships.

5. The method according to claim 4, wherein the step of performing fault tree modeling on each of the fault events and the causal logical relationship to obtain at least one logical fault tree further includes at least one of the following: Perform common cause failure analysis on the logical fault tree, and construct corresponding logical fault subtrees using the various event nodes caused by the same failure cause; Based on the differences between vulnerability analysis and fault analysis of cloud computing services in the heterogeneous spatiotemporal data, different logical fault subtrees are constructed. in, Each of the aforementioned logical fault subtrees can serve as a logical fault tree to generate the target fault tree.

6. The method according to claim 5, wherein performing common-cause failure analysis on the logical fault tree and constructing corresponding logical fault subtrees using event nodes caused by the same failure cause includes at least one of the following: Identify the dependencies between the logical fault tree and the various types of basic resources provided by the first cloud platform, and based on the dependencies, extract the logical fault subtrees related to each type of basic resource from the logical fault tree; By analyzing the logical fault tree in a simulated environment, failure modes under various combinations of basic resources are determined. Based on the common cause triggering source of the failure mode, the failure propagation path in the logical fault tree is constructed, and the corresponding logical fault subtree is constructed. Each of the aforementioned basic resource combinations includes at least two types of basic resources.

7. The method according to any one of claims 1-6, further comprising at least one of the following steps: Perform a syntax check on the logical fault tree, and correct the logical fault tree if the syntax check result is unqualified. The syntax check includes at least one of the following: redundancy structure check, resource dependency check, and common cause event identification check; Based on the spatial data of the first cloud platform currently acquired, each fault event represented by the logical fault tree is verified, and in response to the verification result being unqualified, the logical fault tree is corrected. Based on the currently acquired monitoring time-series data of the first cloud platform and the logical fault tree, each fault event represented by the logical fault tree is verified. In response to the verification result being unqualified, the logical fault tree is corrected.

8. The method according to claim 7, wherein performing a syntax check on the logic fault tree and correcting the logic fault tree in response to a failed syntax check result, comprises at least one of the following: Identify redundant structures in the logical fault tree and execute a preset processing strategy on the redundant structures; A reasonableness check is performed on the dependency relationship between each fault event represented by the logical fault tree and the basic resources provided by the first cloud platform, and unreasonable dependencies in the logical fault tree are corrected. The identifiers assigned to the common subtrees and common basic events that appear repeatedly in different branches of the logical fault tree are verified, and the identifiers assigned to the common subtrees or common basic events that are not correctly identified in the logical fault tree are corrected.

9. The method according to claim 7, wherein verifying each fault event represented by the logical fault tree based on the currently acquired spatial data of the first cloud platform, and correcting the logical fault tree in response to a verification result of non-compliance, comprises: Based on the spatial data of the first cloud platform currently acquired, determine the location information and connection relationships of different components in the first cloud platform. The fault events represented by the logical fault tree are matched with the location information of different components; In response to a successful match, a consistency verification is performed on the connection relationships between the successfully matched fault events and the connection relationships between the components, and inconsistent connection relationships in the logical fault tree are corrected. In response to a matching failure, the logical fault tree is corrected based on the component that failed to match and the connection relationships for that component.

10. The method according to claim 7, wherein verifying each fault event represented by the logical fault tree based on the currently acquired monitoring time-series data of the first cloud platform and the logical fault tree, and correcting the logical fault tree in response to a failed verification result, comprises: Based on a pre-configured data sampling strategy, current monitoring time-series data is sampled from at least one data source; The data sampling strategy can characterize the data sampling frequency of different cloud computing services; the data sampling frequency of the target cloud computing service is higher than that of non-target cloud computing services. Determine the values ​​of each resource indicator in the monitoring time series data; The fault events represented by the logical fault tree are correlated with the values ​​of the various resource indicators. Identify a first fault event in the logical fault tree that is not associated with a resource indicator value, configure corresponding tagging information for the first fault event, or adjust the monitoring data of the corresponding data source to obtain a resource indicator value associated with the first fault event.

11. The method according to claim 7, wherein verifying each fault event represented by the logical fault tree based on the currently acquired monitoring time-series data of the first cloud platform and the logical fault tree, and correcting the logical fault tree in response to a verification result of non-compliance, comprises: Based on the actual alarm information determined by the monitoring time-series data of the first cloud platform currently acquired, the consistency verification is performed on the historical alarm information associated with the corresponding fault events represented by the logical fault tree. In response to a discrepancy between the historical alarm information and the actual alarm information, or if the fault event is not associated with historical alarm information, the causal logic relationship of the corresponding fault event represented by the logical fault tree is corrected, and / or the actual alarm information is associated with the fault event.

12. The method according to any one of claims 1-6, wherein the step of attribute labeling of each event object contained in the at least one logical fault tree based at least on the heterogeneous spatiotemporal data includes: From the heterogeneous spatiotemporal data, determine the attribute information corresponding to each category of components in the first cloud platform; The attribute information comes from at least one of the following: the configuration management database of the corresponding component, the manufacturer's manual, and the deployment configuration file. The attribute information from different sources is fused, and at least based on the fused attribute information, the corresponding event objects contained in the logical fault tree are labeled with attributes. The method further includes: Perform compliance verification on all the aforementioned attribute information; Correct non-compliant attribute information, and then, based on the compliant attribute information and the corrected attribute information, annotate the corresponding event objects contained in the logical fault tree with attributes.

13. The method according to any one of claims 1-6, wherein the step of annotating the attributes of each event object contained in the at least one logical fault tree based at least on the heterogeneous spatiotemporal data to generate a target fault tree for the first cloud platform includes: In response to a vulnerability analysis request for cloud computing services, based on the inherent attributes and historical operating data of different types of components in the first cloud platform contained in the heterogeneous spatiotemporal data, attribute annotation is performed on each event object contained in the at least one logical fault tree. The first spatiotemporal data contained in the heterogeneous spatiotemporal data is then fused with the attribute-annotated logical fault tree to generate a first target fault tree for implementing vulnerability analysis of cloud computing services. In response to a fault analysis request for cloud computing services, based on the heterogeneous spatiotemporal data and business operation and maintenance information, attribute annotation is performed on each event object contained in the corresponding logical fault tree. The second spatiotemporal data contained in the heterogeneous spatiotemporal data is then fused with the attribute-annotated logical fault tree to generate a second target fault tree for achieving fault analysis of cloud computing services. The first spatiotemporal data includes multidimensional spatiotemporal data used for vulnerability scoring, and the second spatiotemporal data includes the first spatiotemporal data and dynamic monitoring data.

14. The method according to any one of claims 1-6, wherein at least based on the heterogeneous spatiotemporal data and the logical fault tree, a target fault tree for the first cloud platform is generated, comprising: Extract unit space data from the spatial data; The unit spatial data describes the location information and connection relationships of the target resources provided by the first cloud platform; The unit space data is fused with the logical fault tree to generate a unit fault tree; The unit fault tree includes the target fault modes and target fault propagation paths for the target cloud computing services supported by the first cloud platform. The unit fault tree is checked, and in response to the detection of abnormal information, the logical fault tree is corrected to return to the step of extracting unit space data from the spatial data; In response to the detection that the unit fault tree meets expectations, a target fault tree for the first cloud platform is generated, based at least on the heterogeneous spatiotemporal data and the modified logical fault tree.

15. The method according to any one of claims 1-6, further comprising: The constructed logical fault tree for the second cloud platform is identified as the source logical fault tree; The second cloud platform refers to a cloud platform with an architecture similar to the first cloud platform; A similarity analysis is performed between the spatial data of the first cloud platform and the source logic fault tree to determine the fault events to be adjusted, the causal logic relationships to be adjusted, and the attribute information to be adjusted for the source logic fault tree. Based on the fault events to be adjusted, the causal logic relationships to be adjusted, and the attribute information to be adjusted, the source logic fault tree migrated to the first cloud platform is adjusted to obtain a logic fault tree for the first cloud platform, which is then used to generate a target fault tree for the first cloud platform.

16. The method according to any one of claims 1-6, wherein the method further comprises: In response to a vulnerability analysis request for cloud computing services, historical fault data and system design in the heterogeneous spatiotemporal data are analyzed, a first initial fault tree is constructed based on the identified target fault modes, and a common subtree for vulnerability analysis is extracted from the first initial fault tree. In response to a fault analysis request for cloud computing services, the fault modes in the heterogeneous spatiotemporal data are classified and statistically analyzed. A second initial fault tree is constructed based on the classification and statistical results of each fault mode. A common subtree for fault tracing is extracted from the second initial fault tree. Each extracted common subtree is assigned a corresponding identifier to identify the corresponding common subtree from among multiple subtrees.

17. The method according to any one of claims 1-6, further comprising: In response to the resource adjustment operation of the first cloud platform, the heterogeneous spatiotemporal data is corrected, and the target fault tree is pruned based on the corrected heterogeneous spatiotemporal data. or, Perform at least one of the following pruning operations on the logical fault tree to generate a target fault tree for the first cloud platform based on the pruned logical fault tree and the heterogeneous spatiotemporal data: Based on the logical fault tree and the heterogeneous spatiotemporal data, the distribution and number of target fault modes in the first target fault tree used for vulnerability analysis are predicted, or the fault events in the second target fault tree used for fault analysis are quantitatively analyzed to achieve pruning of the logical fault tree. Based on the historical inspection data of the first cloud platform in the heterogeneous spatiotemporal data, the predicted failure probability of each event object contained in the logical fault tree is obtained, and the logical fault tree is pruned based on the predicted failure probability. Perform common cause failure analysis on each branch of the logic fault tree, and prune the logic fault tree based on the obtained common cause failure probability of each branch. The failure probability or importance of each event node in the logical fault tree is fuzzy estimated. Based on the estimated fuzzy values ​​corresponding to each event node, the logical fault tree is pruned. In response to the anti-fuzzing processing result of the generated target fault tree satisfying the pruning condition, the target fault tree is pruned based on the failure probability obtained from the anti-fuzzing processing until the anti-fuzzing processing result of the pruned target fault tree satisfies the pruning termination condition.

18. The method according to any one of claims 1-6, further comprising: Based on the spatial data currently acquired from the first cloud platform, construct a graph-structured spatial data model for the current space. A graph similarity analysis is performed on the current spatial data model and the historical spatial data models of adjacent historical times that have been constructed, and the similarity analysis results are obtained. Based on the similarity analysis results, incremental modeling is performed on the target fault tree to obtain a new target fault tree.

19. A computer device, the computer device comprising: At least one communication element, at least one memory, and at least one processor, wherein: The communication element is used to acquire heterogeneous spatiotemporal data from the first cloud platform, the heterogeneous spatiotemporal data including spatial data and monitoring time series data from different data sources; The memory is used to store multiple computer instructions; The processor is used to load and execute the computer instructions to perform the following steps: Based on the heterogeneous spatiotemporal data, various fault events existing in the first cloud platform are determined, as well as the causal logical relationships between the various fault events; the fault events are events that may cause abnormal operation of cloud computing services on the first cloud platform. Fault tree modeling is performed on each of the fault events and the causal logical relationships to obtain at least one logical fault tree; Based at least on the heterogeneous spatiotemporal data, attribute annotations are performed on each event object contained in the at least one logical fault tree to generate a target fault tree for the first cloud platform, thereby enabling reliability analysis of the cloud computing service.