Cloud service impact scope determination method and apparatus, cluster, storage medium, and program product
Patent Information
- Application Number
- PCT/CN2025/121811
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-25
- Filing Date
- 2025-09-17
- Publication Date
- 2026-09-03
Smart Images

Figure CN2025121811_03092026_PF_FP_ABST
Abstract
Description
Methods, devices, clusters, storage media, and application products for determining the impact of cloud services
[0001] This application claims priority to Chinese patent application filed on February 25, 2025, with application number 202510220216.7 and titled "Method, Apparatus, Cluster, Storage Medium and Program Product for Determining the Impact Surface of Cloud Services", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of computer technology, and in particular to methods, apparatus, clusters, storage media, and program products for determining the impact of cloud services. Background Technology
[0003] With the development of cloud computing, various cloud services have been widely applied to various industries and fields. However, once the cloud infrastructure equipment such as servers, storage devices, and network devices that provide basic resources for these cloud services fails, it will have a serious impact on the cloud services used by users, such as increased latency or even service interruption. In order to make more accurate decisions on fault recovery, it is necessary to first conduct an accurate and rapid assessment of the impact.
[0004] Currently, when assessing the impact of cloud failures, it is often necessary to rely on real-time monitoring of whether the operating metrics of tenant cloud service instances are normal. However, this method has problems such as high false alarm rate and false negative rate. For example, when network equipment fails or infrastructure failure is severe and real-time data transmission is interrupted, the impact cannot be accurately assessed because operating metrics cannot be obtained. Summary of the Invention
[0005] This application provides a method, apparatus, cluster, storage medium, and program product for determining the impact surface of cloud services. By identifying the dependencies between multiple units in a cloud computing scenario, the affected cloud services are ultimately determined, and the tenants using these cloud services are identified as the impact surface. This method does not rely on real-time data monitoring and can still accurately determine the impact surface even in the event of a network outage, significantly improving the applicability of the impact surface determination method.
[0006] In one aspect, this application provides a method for determining the impact surface of cloud services, wherein at least one data center belongs to an availability zone of a region, each data center in the at least one data center has one or more clusters, the one or more clusters include multiple servers, the multiple servers deploy multiple instances, and the multiple instances provide one or more cloud services; the method includes: receiving an impact surface assessment request, the impact surface assessment request including an identifier of a fault impact unit, the fault impact unit including at least one of the following dimensions: region, availability zone, data center, cluster, physical server, cloud service; determining a target cloud service that depends on the fault impact unit for operation based on dependency information including the identifier of the fault impact unit; the dependency information is used to characterize the dependencies between region, availability zone, data center, cluster, server, instance, and cloud service; determining the impact surface assessment result based on the target cloud service; the impact surface assessment result includes tenants using the target cloud service.
[0007] Understandably, by identifying the dependencies (which can be pre-built) between different units required for cloud service operation, it is possible to accurately identify tenants using cloud services that depend on units affected by a failure, obtain impact assessment results, avoid dependence on the runtime data of cloud service instances, and ensure accurate and efficient determination of the impact even in the event of a severe network failure.
[0008] In one possible implementation, the target cloud service includes one or more of cloud services that directly depend on the fault-affected unit and cloud services that indirectly depend on the fault-affected unit.
[0009] It is understandable that cloud computing scenarios often involve multiple units and form relatively complex resource dependencies. Whether the cloud service directly or indirectly depends on the unit affected by the fault, it will be affected when the unit affected by the fault experiences an anomaly. Therefore, they all belong to the target cloud service, and the accuracy of the impact is determined by this.
[0010] In one possible implementation, the method further includes: providing a first interface including an interactive control for determining the affected unit of the fault; and, in response to receiving a trigger operation on the interactive control, determining the identifier of the affected unit of the fault in the affected area determination request.
[0011] It is understandable that by improving the interactive interface, users can more easily and accurately input the identifier of the fault-affected unit. Since the fault-affected unit in the user's input impact area determination request is the starting unit for subsequent impact area determination, accurately inputting the identifier of the fault-affected unit helps to improve the accuracy of impact area determination.
[0012] In one possible implementation, after determining the impact assessment results based on the target cloud service, the method further includes providing a second interface for displaying the impact assessment results.
[0013] Understandably, displaying the impact assessment results through the interface allows relevant operations and maintenance personnel to intuitively understand the severity of the impact and thus take more accurate recovery measures.
[0014] In one possible implementation, the dependency information includes a directed graph; the directed graph includes vertices and directed edges, where the vertices represent one or more units among regions, availability zones, data centers, clusters, servers, instances, and cloud services, and the directed edges represent unidirectional dependencies between the vertices in the directed graph.
[0015] It is understandable that the points and directed edges in a directed graph can accurately and intuitively represent the unidirectional dependencies between different units in a unit cloud computing scenario, thereby helping computing devices to determine the target cloud service.
[0016] In one possible implementation, determining the target cloud service that depends on the fault-affected unit based on dependency information including the identifier of the fault-affected unit includes: the computing device determining the first point corresponding to the fault-affected unit in the directed graph based on the identifier of the fault-affected unit; traversing the points in the directed graph from the first point according to the target direction until the target point is obtained; the target direction is the direction of the dependent fault-affected unit indicated by the directed edge of the first point; the target point is a point used to characterize the target cloud service.
[0017] Understandably, in a more complex directed graph, computing devices can determine all target points that meet the conditions by traversing the target direction multiple times. In other words, they can find all target cloud services that meet the conditions. This allows them to accurately determine the impact even if the tenant's cloud service instance has not yet shown any abnormalities or has not even started.
[0018] In one possible implementation, the point obtained in the i-th traversal is the point that the directed edge connecting the point obtained in the (i-1)-th traversal points to in the target direction; i is an integer greater than 1.
[0019] Understandably, the computing device strictly follows the target direction in each traversal to ensure the unidirectionality of the dependency relationship and improve the accuracy of determining the impact area.
[0020] In one possible implementation, the points in the directed graph have corresponding label information. The label information is used at least to characterize whether the unit represented by the point in the directed graph has disaster recovery capability. The units represented by the points obtained in the i-th traversal and the (i-1)-th traversal are units that do not have disaster recovery capability.
[0021] It is understandable that a unit with disaster recovery capabilities can continue to ensure the normal operation of its own business even if the unit it depends on fails. In other words, a unit with disaster recovery capabilities can continue to provide normal services to other units that depend on it. Consequently, cloud services that depend on units with disaster recovery capabilities can remain normal. Therefore, when computing devices traverse a directed graph, traversing the points corresponding to units that do not have disaster recovery capabilities based on the label information can improve the accuracy of the final determined impact surface.
[0022] In one possible implementation, each point in the directed graph has corresponding label information. This label information at least characterizes whether the unit represented by that point has disaster recovery capabilities. The target cloud service is a cloud service that depends on units that do not have disaster recovery capabilities. It is understandable that cloud services that depend on units that do not have disaster recovery capabilities will be affected when those units fail. Therefore, using such cloud services as the target cloud service ensures the accuracy of determining the impact surface.
[0023] In one possible implementation, the unit represented by the point corresponding to the tag information has disaster recovery capability when the tag information includes any of the following conditions: the unit's deployment architecture is multiple availability zones (AZ) or multiple regions (GRI); the unit has automatic flow switching capability; the version information corresponding to the unit conforms to the target version information range; and the unit is a stacked switch.
[0024] Understandably, by using the data in the tag information and various target conditions, it is possible to accurately determine whether a unit has disaster recovery capabilities. Units with disaster recovery capabilities can continuously provide support for cloud services that depend on them. Therefore, such tags can improve the accuracy of determining the scope of impact.
[0025] In one possible implementation, the units required for the operation of the first cloud service include one or more of the following: network devices, physical machines, cloud computing clusters, and cloud services.
[0026] It is understandable that network devices, physical machines, cloud computing clusters, and cloud services can cover various types of units in a cloud computing scenario, thereby helping computing devices to improve the coverage of the identified impact area through more comprehensive dependency information.
[0027] In one possible implementation, dependency information is determined based on one or more of the following: the inclusion relationship between regions and availability zones, the inclusion relationship between availability zones and data centers, the deployment relationship of clusters on data centers, the deployment relationship of clusters on servers, the deployment relationship of instances on servers, and the deployment relationship of cloud services on instances.
[0028] Understandably, using a variety of information can help construct a more comprehensive understanding of the dependencies between different units in a cloud computing scenario, thereby improving the accuracy of determining the scope of impact.
[0029] In one possible implementation, the unit is specifically a fault-recoverable unit (FRU), and each FRU has fault recovery capability.
[0030] Understandably, by focusing on the units required for cloud service operation that have fault recovery capabilities, it is possible to improve the efficiency of subsequent fault recovery while ensuring the accuracy of determining the scope of impact.
[0031] Secondly, this application provides a cloud service impact surface determination apparatus, which is used to perform any of the cloud service impact surface determination methods provided in the first aspect above.
[0032] In one possible implementation, this application can divide the cloud service impact area determination device into functional modules according to the method provided in the first aspect above. For example, each function can be divided into a separate functional module, or two or more functions can be integrated into one module. For example, this application can divide the cloud service impact area determination device into a receiving module, a first determining module, and a second determining module, etc., according to function. The descriptions of the possible technical solutions and beneficial effects of the various functional modules described above can be found in the technical solutions provided in the first aspect above or its corresponding possible implementations, and will not be repeated here.
[0033] Thirdly, embodiments of this application provide a computing device that includes a processor and a memory, the processor being coupled to the memory; the memory is used to store computer instructions that are loaded and executed by the processor to enable the computing device to implement the cloud service impact surface determination method provided in the various optional implementations of the first aspect above.
[0034] Fourthly, embodiments of this application provide a computing device cluster, which includes at least one computing device, each computing device including a processor and a memory, the processor being coupled to the memory; the processor of the at least one computing device is used to execute computer instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the cloud service impact surface determination method provided in the various optional implementations of the first aspect above.
[0035] Fifthly, embodiments of this application provide a computer-readable storage medium storing at least one computer program instruction, which is loaded and executed by a processor to implement the cloud service impact surface determination method provided in various optional implementations of the first aspect above.
[0036] Sixthly, embodiments of this application provide a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computing device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computing device to perform the cloud service impact determination method provided in the various optional implementations of the first aspect described above.
[0037] For a detailed description of aspects two through six and their various implementations in this application, please refer to the detailed description in aspect one and its various implementations; and for a detailed analysis of the beneficial effects of aspects two through six and their various implementations in aspect one and its various implementations, please refer to the beneficial effect analysis in aspect one and its various implementations, which will not be repeated here.
[0038] These or other aspects of this application will become more readily apparent in the following description. Attached Figure Description
[0039] Figure 1 is a schematic diagram of the architecture of a cloud computing system provided in an embodiment of this application;
[0040] Figure 2 is a schematic diagram of the architecture of a cloud service impact surface determination system provided in an embodiment of this application;
[0041] Figure 3 is a schematic diagram of the structure of a computing device provided in an embodiment of this application;
[0042] Figure 4 is a flowchart of a cloud service impact surface determination method provided in an embodiment of this application;
[0043] Figure 5 is a schematic diagram of an interactive interface involved in the embodiment shown in Figure 4;
[0044] Figure 6 is a schematic diagram of a dependency information involved in the embodiment shown in Figure 4;
[0045] Figure 7 is a schematic diagram of a label information involved in the embodiment shown in Figure 6;
[0046] Figure 8 is a schematic diagram of the deployment architecture of a cloud service instance involved in the embodiment shown in Figure 6;
[0047] Figure 9 is a flowchart of a traversal pruning method involved in the embodiment shown in Figure 6;
[0048] Figure 10 is a schematic diagram of a rendered dependency information involved in the embodiment shown in Figure 6;
[0049] Figure 11 is a schematic diagram of an impact surface assessment result involved in the embodiment shown in Figure 4;
[0050] Figure 12 is a schematic diagram of a cloud service impact surface determination device provided in an embodiment of this application;
[0051] Figure 13 is a schematic diagram of a computing device provided in an embodiment of this application;
[0052] Figure 14 is a schematic diagram of a computing device cluster provided in an embodiment of this application;
[0053] Figure 15 is a schematic diagram of a connection method between computing device clusters provided in an embodiment of this application. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0055] In this article, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0056] Furthermore, in the description of the embodiments of this application, unless otherwise stated, "multiple" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0057] Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that "first" and "second" are not necessarily different. Meanwhile, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is being used as an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present related concepts in a concrete manner for ease of understanding.
[0058] First, the technical terms involved in the embodiments of this application will be introduced by way of example.
[0059] Cloud computing is an internet-based service model that allows users to access and use shared computing resources and services, such as servers, storage, databases, applications, and services, over a network. These resources are typically virtualized and distributed across the network, providing users with flexible, on-demand computing power.
[0060] Cloud platform: Also known as a cloud computing platform, it's a service based on hardware and software resources, providing computing, networking, and storage capabilities. The platform provider combines the cloud (remote hardware resources) and computing (remote software resources) into a single platform, offering various services to tenants. Cloud platform services can include Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). SaaS involves the service provider deploying application software uniformly on the cloud platform's servers, with tenants subscribing to these services via the internet. PaaS provides a development environment as a service; the service provider offers a development environment, server platform, and hardware resources to tenants, who then develop applications within this environment and share them with other users via the cloud platform and the internet. IaaS involves the service provider offering a cloud infrastructure consisting of multiple servers as a service, integrating memory, input / output devices, storage, and computing resources into a virtual resource pool, providing storage resources and virtualized servers to tenants.
[0061] Cloud service instances generally refer to virtual computing resources provided by cloud service providers. These resources may include a CPU, memory, operating system, network, and disk, allowing tenants to manage and use them as if they were local servers. Furthermore, different instance types offer varying computing and storage capabilities, suitable for different application scenarios. For example, compute-optimized instances are suitable for applications requiring powerful computing capabilities, such as scientific modeling and media transcoding. Memory-optimized instances are suitable for memory-intensive applications, such as big data analytics.
[0062] Cloud service unit: can refer to the unit required for the operation of cloud services. Cloud service units include, but are not limited to, hardware and / or software, such as data centers, server racks, network equipment, physical servers, cloud service clusters, instances (virtual machines), and cloud services, etc.
[0063] Tenant: The basic unit for allocating resources on a cloud service platform. After a user registers an account on the cloud platform, the platform considers that account to correspond to a tenant. Tenants can create and use cloud service instances that meet their own needs through the cloud platform.
[0064] With the development of cloud computing technology and the increasing sophistication of cloud platform capabilities, cloud computing technology is widely used and has penetrated into various industries and fields. When a unit in a cloud computing scenario fails, a rapid and accurate impact assessment is required.
[0065] For example, a comprehensive monitoring system can be established, encompassing monitoring at all levels, including the network, servers, storage, and applications. This monitoring system and alerting mechanisms can promptly detect cloud computing failures. After a failure is detected, fault localization can be performed. For instance, tools such as log analysis and performance monitoring can be used to quickly pinpoint the specific location of the failure. Based on the fault localization results, the cause of the failure can be analyzed: is it a hardware failure, a software error, a configuration error, or another external factor? Furthermore, an impact assessment can be conducted based on the analysis results, and corresponding recovery strategies can be developed to achieve cloud computing failure recovery.
[0066] The impact can include, but is not limited to, cloud services or applications, users, businesses, and data affected by the failure. Impact assessment mainly includes the following types of assessments:
[0067] (1) Service Scope Assessment: Determine which cloud services or applications are affected. Service scope assessment requires a deep understanding of cloud computing architecture in order to accurately identify the service scope that may be affected by the failure point. In this application, cloud services may include, but are not limited to, cloud storage services, cloud computing services, cloud database services, etc.
[0068] (2) User Impact Assessment: Assess the affected user groups, including the number of users, user types (e.g., internal employees, external customers), and the business importance of the users. User impact assessment helps determine the actual extent of the impact of cloud computing failures on user businesses. It should be understood that users can also be referred to as tenants; this application uses the term "tenant" as an example for description, without limitation.
[0069] (3) Business Impact Assessment: Analyze the potential impact of cloud computing outages on business processes, operational efficiency, customer satisfaction, etc. Business impact assessment may require close collaboration with business units to understand the specific impact of the outage on business operations.
[0070] (4) Data Impact Assessment: Check whether the failure resulted in data loss, corruption, or inaccessibility. The data impact assessment requires a comprehensive evaluation of the data's integrity, availability, and security.
[0071] As shown above, taking user-side impact assessment as an example, it is necessary to identify the tenants using the cloud services affected by the failure, and then determine the severity of the failure and the corresponding recovery strategy based on the tenants' attributes. For example, if no tenants are currently using the cloud services affected by the failure, the corresponding recovery strategy can adopt a lower priority and lower cost approach.
[0072] It can be said that quickly and accurately determining the impact surface is crucial in the fault recovery process. A traditional method is to configure a corresponding information collection module on the tenant's cloud service instance. The information collection module collects the target logs and operating indicators on the instance in real time. Based on the collected target logs and operating indicators, fault detection (or anomaly detection) is performed to determine whether the cloud service instance has failed. If the cloud service instance has failed, it means that the cloud service instance is a faulty instance, and the tenant to which the cloud service instance belongs is a tenant in the impact surface. Following this method, the impact surface is finally obtained by collecting information on the tenants to which the faulty instance belongs on a large scale.
[0073] This approach relies on real-time data acquisition from cloud service instances and places high demands on the target logs and operational metrics to be acquired, requiring both high coverage and a low false positive rate. However, in cloud computing scenarios, even when different tenants use the same cloud service, their business operations vary significantly and are constantly evolving. These target logs and operational metrics struggle to adapt in a timely manner, leading to false positives and missed positives. For example, fluctuations in the memory usage of the corresponding cloud service instance (not necessarily a failure) can trigger false positives. Furthermore, this method fails in cases of severe network or infrastructure failures that prevent real-time data transmission, making it impossible to determine the extent of the impact.
[0074] Research has revealed that various cloud services often involve complex resource dependency chains. The operation of any cloud service is inseparable from hardware support, which is one of the characteristics of cloud computing scenarios: involving numerous different types of units. Hardware units include, but are not limited to, physical servers, network devices (such as switches), motherboards, interfaces, and chips. Software units include, but are not limited to, cluster resources built on physical computing devices, virtual machines, instances, and cloud services. In addition, cloud computing scenarios include other types of units, including but not limited to: regions, availability zones (AZs), data centers (DCs), and server rooms. These aforementioned cloud computing units collectively constitute the real-world foundation for the normal operation of cloud services. In other words, a failure in any unit of the resource dependency chain behind a cloud service can affect the normal operation of the cloud service, thereby impacting tenants using the service. For example, when describing the resource dependency chain (or dependency relationship) of a cloud service, it could be that cloud service A runs on a server with device number 123, and this server is located in server room 1 of data center A in availability zone A of region A.
[0075] In view of this, this application proposes a method for determining the impact surface by utilizing the dependency relationships between different units in multiple cloud computing scenarios. This method does not rely on real-time collection of information on tenant cloud service instances, but starts from a unit (which may fail) in the relevant cloud computing scenario, and finally determines the cloud services that depend on the unit based on one or more layers of dependency relationships. The tenants using these cloud services are identified as the impact surface, which significantly improves the accuracy of determining the impact surface.
[0076] In some feasible embodiments, at least one data center belongs to an availability zone of a region, and each data center in the at least one data center has one or more clusters. The one or more clusters include multiple servers, the multiple servers deploy multiple instances, and the multiple instances provide one or more cloud services. The method includes: receiving an impact surface assessment request, the impact surface assessment request including the identifier of a fault impact unit, the fault impact unit including at least one of the following dimensions: region, availability zone, data center, cluster, physical server, cloud service; determining the target cloud service that depends on the fault impact unit based on dependency information including the identifier of the fault impact unit; the dependency information is used to characterize the dependencies between region, availability zone, data center, cluster, server, instance, and cloud service; determining the impact surface assessment result based on the target cloud service; the impact surface assessment result includes tenants using the target cloud service. Since it is not necessary to obtain data metrics on the tenant's cloud service instances in real time, the applicability of the cloud service impact surface determination method can be significantly improved. Even in the event of network interruption, the impact surface can still be accurately and quickly determined. Moreover, determining the impact surface through dependency information is not affected by data jitter, and the probability of false positives and false negatives can be reduced.
[0077] The system architecture of the embodiments of this application will be described exemplarily below.
[0078] Figure 1 is a schematic diagram of the architecture of a cloud computing system provided in an embodiment of this application. As shown in Figure 1, the cloud computing system 100 includes a computing server cluster 110, a storage server cluster 120, a management server cluster 130, a network device cluster 140, and a user terminal 150. The computing server cluster 110, the storage server cluster 120, and the management server cluster 130 communicate with the user terminal 150 through the network device cluster 140.
[0079] The computing server cluster 110 includes one or more computing servers (two computing servers, namely computing server 111 and computing server 112, are shown in Figure 1, but are not limited to two computing servers).
[0080] Computing servers, such as servers and desktop computers, serve as computing resources within the cloud computing system 100. They are used to generate and allocate computing resources based on virtualization technology and user needs. At the hardware level, computing servers are equipped with processors and memory (not shown in Figure 1). The computing functions of the computing server are implemented by the processor running programs in memory. The computing server can also read / write data from various storage servers in the storage server cluster 120 according to user needs.
[0081] The storage server cluster 120 includes one or more storage servers (two storage servers, namely storage server 121 and storage server 122, are shown in Figure 1, but are not limited to two storage servers).
[0082] Storage servers, as storage resources within the cloud computing system 100, such as servers, desktop computers, or storage array controllers and hard disk enclosures, provide logical disk storage and integrated backup services for cloud virtual machines within the cloud computing system 100. In terms of hardware, storage servers are equipped with network interface cards (NICs), processors, and memory. The processor in the storage server processes data from outside the storage server. The NIC controls the access process to the memory, such as controlling address signals, data signals, and various command signals, enabling the storage server to provide the memory as a storage resource to users. Memory is used to store data and may include RAM and / or hard disks. RAM refers to internal memory that directly exchanges data with the processor; RAM can quickly read and write data at any time, serving as temporary data storage for the operating system or other running programs. Unlike RAM, hard disks are slower to read and write data and are typically used for persistent data storage.
[0083] The management server cluster 130 includes one or more management servers (two management servers, namely management server 131 and management server 132, are shown in Figure 1, but are not limited to two management servers).
[0084] The management server is used to manage all computing services, shared storage, and network of the entire cloud computing system 100, and also provides users or administrators with an application program interface (API) for managing the entire node. In this application, the cloud computing system 100 can provide the program product of this application to users by providing an accessible application program interface.
[0085] Network device cluster 140 includes one or more switches and routers, as shown in Figure 1. In this embodiment, network device cluster 140 includes router 141, switch 142, switch 143, switch 144, and switch 145. User terminal 150 is connected to router 141 via the Internet. Router 141 is then connected to switch 143 via switch 142. Switch 143 is connected to each computing server in computing server cluster 110. Switch 144 is connected to each computing server in computing server cluster 110 and each storage server in storage server cluster 120. Switch 145 is connected to each computing server in computing server cluster 110, each storage server in storage server cluster 120, and each management server in management server cluster 130.
[0086] Optionally, the number and type of switches included in the network device cluster 140 can be adjusted according to the needs of the cloud computing system 100. Switches 142, 143, 144, and 145 can be switches with different functions. For example, switch 142 is a core switch, while switches 143, 144, and 145 are switches used to manage specific network segments. For instance, switch 142 can be a core switch, switch 143 can be an internal / external switching network segment switch, switch 144 can be a storage network segment switch, and switch 145 can be a management network segment switch.
[0087] User terminal 150 includes one or more user terminals (two user terminals, namely user terminal 151 and user terminal 152, are shown in Figure 1, but are not limited to two user terminals). The user terminal contains the interfaces and applications required for accessing the cloud computing system 100.
[0088] It is worth noting that Figure 1 is only a schematic diagram and should not be construed as a limitation of this application. The cloud computing system 100 may also include other devices, which are not shown in Figure 1.
[0089] The computing servers and storage servers in the cloud computing system 100 provided in this application can be nodes in a cloud platform. Based on the equipment of the cloud computing system 100 shown in Figure 1, the cloud computing system 100 implements the functions of computing nodes, storage nodes, etc., based on Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS), and provides services (such as computing services, storage services, and network services) to user terminals 150 through these nodes. These nodes can be service nodes (such as computing nodes and storage nodes) in the cloud obtained by virtualizing the resources (such as computing resources and storage resources) of the cloud computing system 100.
[0090] Based on the cloud computing architecture shown in Figure 1, Figure 2 is a schematic diagram of the architecture of a cloud service impact surface determination system provided in an embodiment of this application. The cloud service impact surface determination system 1000 shown in Figure 2 includes at least a computing module 1100, which can execute the cloud service impact surface determination method provided in the embodiment of this application. For example, it can perform calculations based on externally input dependency information and information used to characterize cloud computing units (external data shown in Figure 2), and finally determine the tenants who used potentially affected cloud services as the impact surface.
[0091] Optionally, the impact area determination system 1000 also includes a tag database 1300, which can be used to acquire external data, extract key information from it, and construct and store tag information.
[0092] In the embodiments of this application, the tag information associated with each unit includes at least the identification information of the cloud computing unit. The identification information is unique information that can be used to locate the target unit in the dependency information involving multiple units, such as the device code of the device, the Internet Protocol (IP) address, the identity document (ID) number of the instance, etc.
[0093] Optionally, the tag information also includes deployment information of the cloud computing unit. The deployment information can be used to confirm the form in which the unit is deployed, such as whether it is deployed in multiple regions or multiple availability zones (AZs). Based on the deployment information in the tag information, the computing module 1100 can determine whether the corresponding unit has disaster recovery capabilities and further determine whether other units that depend on the unit are affected, which helps to improve the accuracy and efficiency of determining the impact.
[0094] Furthermore, the tag information can include more attributes according to actual needs, such as: unit type information (e.g., data center, switch, cluster, etc.), main attribute information of the unit (e.g., unit purpose, current status, etc.).
[0095] Optionally, the impact surface determination system 1000 also includes a modeling module 1200. The modeling module 1200 can be used to acquire data from the tag database 1300 and / or external data, including but not limited to configuration information of network devices, topological connection relationships, dependencies between physical computing resources and cloud computing clusters, dependencies between cloud computing clusters and cloud services, and relationships between cloud services and tenants. Based on the acquired data, it constructs dependency information in the cloud computing scenario (e.g., draws a directed graph, where points in the directed graph represent different units, and directed edges in the directed graph represent dependencies between units), thereby providing a data foundation for the computing module 1100 to execute the cloud service impact surface determination method provided in the embodiments of this application.
[0096] The aforementioned influence surface determination system 1000 can run on the computing device shown in FIG3. FIG3 is a schematic diagram of the structure of a computing device provided in an embodiment of this application. The computing device 3000 shown in FIG3 includes at least: a memory 3010, a processor 3020 and a bus 3030.
[0097] The processor 3020 is capable of acquiring input information and dependency information, and determining the impact surface based on the input information and dependency information, etc. The memory 3010 is capable of storing the logic code corresponding to the cloud service impact surface determination method provided in this application embodiment.
[0098] Optionally, the computing device 3000 can be various types of server equipment such as rack servers and full rack servers. The computing device 3000 can also be terminal computing devices such as computers, mobile terminals, tablet computers, laptops, desktop computers, all-in-one computers, personal digital assistants (PDAs), and ultra-mobile personal computers (UMPCs).
[0099] Optionally, the memory 3010 may include random access memory (RAM), read-only memory (ROM), etc., wherein the memory 3010 may run the necessary operating system in its RAM, as well as modules such as a receiving module and a first determining module for executing the cloud service impact surface determination method provided in this application.
[0100] Optionally, the processor 3020 can be a central processing unit (CPU) or other general-purpose processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), digital signal processor (DSP) or other programmable logic device, discrete gate or transistor logic device, discrete hardware unit, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0101] Optionally, bus 3030 can be a peripheral component interconnect (PCI) bus, etc., and this application does not limit the type of bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 3, but this does not mean that there is only one bus or one type of bus. Bus 3030 may include a path for transmitting information between various components of computing device 3000 (e.g., memory 3010 and processor 3020).
[0102] The system architecture and application scenarios described in this application are intended to more clearly illustrate the technical solutions of this application and do not constitute a limitation on the technical solutions provided in this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in this application are also applicable to similar technical problems.
[0103] For ease of understanding, the cloud service impact surface determination method provided in this application embodiment is described below with reference to the accompanying drawings. This cloud service impact surface determination method is applicable to the computing device shown in Figure 3.
[0104] Figure 4 is a flowchart of a cloud service impact surface determination method provided in an embodiment of this application, which specifically includes the following steps:
[0105] S110, The computing device receives an impact surface assessment request.
[0106] The embodiments of this application can be applied to cloud computing scenarios, which include at least one data center belonging to the availability zone of a region, each of the at least one data center having one or more clusters, the one or more clusters including multiple servers, the multiple servers deploying multiple instances, and the multiple instances providing one or more cloud services.
[0107] The impact assessment request received by the computing device includes an identifier for the fault impact unit, which includes at least one of the following dimensions: region, availability zone, data center, cluster, physical server, or cloud service. In other words, the fault impact unit can be any unit required to provide cloud services to tenants in a cloud computing scenario.
[0108] It is worth noting that, in order to improve the efficiency of fault repair, the units involved in the embodiments of this application should be considered to have a certain degree of repairability. For example, they can be fault recoverable units (FRUs). Each FRU has fault recovery capability and can quickly enter the working state within a certain time by executing clearly defined recovery measures and can provide services normally to the outside world.
[0109] The first information obtained by the computing device in step S110 can be the identification information of the unit, such as the number information of the data center, the management IP of the switch, the ID of the elastic compute service (ECS), etc. The specific information is related to the unit it represents and can be used to uniquely locate the corresponding unit among multiple units.
[0110] For example, as shown in Figure 5, which is a schematic diagram of an interactive interface (first interface) related to the embodiment shown in Figure 4, the user can input basic information about the unit affected by the computing device fault, such as the fault object: a switch, and enter the management IP of the switch. Furthermore, the user can input more relevant information to improve the accuracy and efficiency of determining the impact surface, such as the fault range shown in Figure 5, specifically including the region name and availability zone (AZ) name. After filling in the information, the user can trigger the computing device to execute the cloud service impact surface determination method provided in this application embodiment by clicking the "Determine Impact Surface" interactive control in Figure 5.
[0111] S120, the computing device determines the target cloud service that depends on the fault-affected unit for operation based on dependency information including the identifier of the fault-affected unit.
[0112] In the embodiments of this application, dependency information is used to characterize the dependencies between regions, availability zones, data centers, clusters, servers, instances, and cloud services. Specifically, it can be represented by any of a variety of data structures, such as directed graphs, arrays, or linked lists, etc., without limitation.
[0113] When executing step S120, the computing device first determines the unit corresponding to the first piece of information in the dependency relationship information, and then determines other units that depend on that unit based on the dependency relationship information, and so on, until the target cloud service is determined. That is to say, the target cloud service can be either a cloud service that directly depends on the unit affected by the fault or a cloud service that indirectly depends on the unit affected by the fault.
[0114] Specifically, the modeling module in the computing device can build dependency information in advance or in real time based on a variety of information and update it periodically, including but not limited to: the inclusion relationship between regions and availability zones, the inclusion relationship between availability zones and data centers, the deployment relationship of clusters on data centers, the deployment relationship of clusters on servers, the deployment relationship of instances on servers, the deployment relationship of cloud services on instances, as well as the configuration information and topology connection relationship of network devices, the relationship between cloud services and tenants, etc.
[0115] The following explanation uses a directed graph to represent dependency information as an example. A directed graph includes vertices and directed edges. Vertices in a directed graph represent units in a cloud computing architecture. A directed edge in a directed graph is an edge that points from one vertex to another in the directed graph. A directed edge represents a one-way dependency between the two vertices that connect the directed edge. For example, if a directed edge points from vertex A to vertex B, it means that vertex B depends on vertex A.
[0116] Please refer to Figure 6 for details. Figure 6 is a schematic diagram of dependency information involved in the embodiment shown in Figure 4. For example, points 1-4 in Figure 6 represent switches 1-4 respectively; point 5 represents board 1 on switch 4; points 6 and 7 represent chip 1 and chip 2 on board 1 respectively; points 8 and 9 represent port 1 and port 2 of chip 2 respectively; point 10 represents physical machine 1 connected to port 2; point 11 represents physical machine 2 connected to port 1; point 12 represents cluster 1 running on physical machine 1; point 13 represents cluster 2 running on physical machine 2; points 14 and 15 represent ECS 1 and ECS 2 running on cluster 2 respectively; point 16 represents cloud service 1 running on ECS 1; and point 17 represents cloud service 2 running on ECS 2.
[0117] Furthermore, in step S120, the computing device can locate switch 1 in Figure 6 based on the management IP (first information) of the switch input by the user. Next, based on the dependencies shown in Figure 6, the computing device first traverses the points directly connected to the outgoing edges of point 1, determining that the units directly dependent on switch 1 include: switch 2 represented by point 2 and switch 3 represented by point 3. Then, the computing device traverses the points directly connected to the outgoing edges of points 2 and 3, determining that the units directly dependent on switch 2 include: switch 4 represented by point 4. In other words, switch 4 represented by point 4 indirectly depends on switch 1. Afterwards, the computing device traverses the points directly connected to the outgoing edges of point 4... The unit directly dependent on switch 4 is identified as: board 1 on switch 4 represented by point 5; then, the computing device traverses the points directly connected to the outgoing edge of point 5 to identify the unit directly dependent on board 1 as: chip 1 on board 1 represented by point 6 and chip 2 on board 1 represented by point 7, and so on. The computing device traverses layer by layer multiple times until it identifies points 16 and 17 in Figure 6, where cloud service 1 represented by point 16 and cloud service 2 represented by point 17 are the target cloud services.
[0118] To further improve the accuracy and efficiency of computing devices in determining the impact surface, each point in the directed graph has its own label information, as shown in Figure 7. Figure 7 is a schematic diagram of label information involved in the embodiment shown in Figure 6. Specifically, different points can represent different types of units, and the corresponding label information can include different types of information. The specific information can be set according to actual needs. The label information shown in Figure 7 includes: region / availability zone, identifier, deployment architecture, unit type, unit purpose, whether it is under load, and version information.
[0119] For example, referring to Figure 8, which is a schematic diagram of the deployment architecture of a cloud service instance involved in the embodiment shown in Figure 6, it illustrates three different deployment forms of a cloud database service (RDS) instance. Specifically, a shows the single-machine deployment form of RDS, that is, RDS instance 1 is deployed only on Availability Zone 1 (AZ1), and the corresponding tag information records the identifier of RDS instance 1: rds instance id1, deployment architecture: single-machine deployment, unit type: cloud database. Example b illustrates a multi-AZ deployment of RDS, where RDS instance 2 is deployed in both Availability Zone 2 (AZ2) and Availability Zone 3 (AZ3). This means that when a unit in AZ2 fails, RDS instance 2 can quickly switch to AZ3 to continue normal operation, and vice versa. The corresponding tag information records the identifier of RDS instance 2: rds instance id2, deployment architecture: multi-AZ deployment, unit type: cloud database. Example c illustrates a multi-region deployment of RDS, where RDS instance 3 is deployed in both Region 1 (region1) and Region 2 (region2). The corresponding tag information records the identifier of RDS instance 3: rds instance id3, deployment architecture: multi-region deployment, unit type: cloud database.
[0120] Furthermore, cloud services that rely on units with disaster recovery capabilities, and cloud services with disaster recovery capabilities (such as the aforementioned multi-az deployment RDS instance 2), can achieve rapid switching when the units they rely on fail in the cloud computing scenario, ensuring continued normal service provision. Therefore, tenants using such cloud services can be considered unaffected. Thus, when the computing device executes step S120, it can utilize label information to prune the directed graph during traversal, thereby significantly improving the accuracy and efficiency of determining the impact surface.
[0121] See Figure 9 for details. Figure 9 is a flowchart of a traversal pruning method involved in the embodiment shown in Figure 6, including:
[0122] S121, the computing device traverses the points connected by the outgoing edges.
[0123] In this step, if the current computing device has just received the first information and has not yet performed a traversal, it starts from the point corresponding to the fault-affected unit and performs a traversal according to the outgoing edges; if the computing device has already performed i-1 traversals, it starts from the point obtained in the most recent traversal and performs the i-th traversal according to the outgoing edges. Here, i is an integer greater than 1.
[0124] S122, The computing device determines whether the unit has disaster recovery capability based on the tag information.
[0125] In this step, the computing device will determine whether the unit represented by each point has disaster recovery capability based on the label information corresponding to the points obtained in the most recent traversal.
[0126] Specifically, in one possible implementation, if the tag information includes one or more of the following conditions, the unit represented by the point corresponding to the tag information has disaster recovery capability: a) the unit's deployment architecture is multiple availability zones (AZ) or multiple regions (GRI); b) the unit has automatic flow switching capability; c) the version information corresponding to the unit conforms to the target version information range; d) the unit is a stacked switch. The criteria for determining disaster recovery capability can also be adjusted according to the actual situation.
[0127] If the unit has disaster recovery capability, the computing device continues to execute step S123; if the unit does not have disaster recovery capability, the computing device continues to execute step S121 to perform the next traversal until the target cloud service is determined.
[0128] S123, the computing device performs pruning.
[0129] In this step, the computing device no longer traverses the points connected to the outgoing edges of points of units with disaster recovery capabilities. That is, when the computing device repeatedly executes step S121, it only traverses the points connected to the outgoing edges of points of units without disaster recovery capabilities, which significantly improves the efficiency of determining the impact surface.
[0130] In other words, through pruning, the points obtained in the i-th traversal and the (i-1)-th traversal represent units that lack disaster recovery capabilities. Here, i is an integer greater than 1.
[0131] In some feasible embodiments, the computing device can also display dependency information through a display terminal. In order to further improve intuitiveness, the dependency information shown in Figure 10 can be displayed. Figure 10 is a schematic diagram of a rendered dependency information involved in the embodiment shown in Figure 6. In this diagram, each unit is rendered according to its common form, which intuitively distinguishes different units. Since the units shown in Figure 10 and the dependencies between different units can be referred to Figure 6, they will not be described in detail here.
[0132] Furthermore, in response to user interactions (such as clicks, touches, etc.), the unit's tag information can be displayed to help users understand the details of each unit. For example, in Figure 10, in response to a user clicking on a target switch, the switch's tag information is displayed, specifically including: the switch is located in region / availability zone: zone A, identified as IP1 (a management IP), unit type: switch, unit purpose: storage cluster (which can be understood as being used to meet the data forwarding needs in the storage cluster), and whether it is under load: yes. Alternatively, the tag information corresponding to each unit can be displayed by default, and this application does not limit this.
[0133] S130, The computing device determines the impact assessment results based on the target cloud service.
[0134] In this step, the computing device can identify the tenant using the target cloud service based on the identifier of the cloud service, such as the cloud service number or cloud service name. Specifically, it can identify the tenant's ID and other identifier information, and use the obtained tenant ID as the impact surface. Furthermore, after determining the impact surface assessment result, a corresponding interface (second interface) can be provided to intuitively display the impact surface assessment result. See Figure 11 for details. Figure 11 is a schematic diagram of an impact surface assessment result involved in the embodiment shown in Figure 4.
[0135] Specifically, in Figure 11, the user-input fault-affected unit is the identifier of physical machine 2. The computing device then starts from physical machine 2 and ultimately determines the set of tenants using cloud service 1 and the set of tenants using cloud service 2 as the impact surface assessment results. Here, depending on the user's permissions, different levels of detail can be displayed on the interface. For example, ordinary users may not be shown the intermediate process shown in Figure 11; however, when facing operations engineers, to assist them in resolving faults more quickly, the intermediate process can be displayed, or even the entire content in Figure 11 can be displayed, including: physical machines 1-3, cluster 1 running on physical machine 2, instance 1 and instance 2 deployed on cluster 1. Instance 1 has disaster recovery capabilities, therefore subsequent units (such as cloud service 3) are not included in the impact surface determination process, relying instead on cloud services 1 and 2 of instance 2, as well as the set of tenants using cloud service 1 and the set of tenants using cloud service 2. This allows relevant users to intuitively understand the scope of the impact surface, thus accelerating recovery.
[0136] By using the aforementioned steps S110-S130, the reliance on data on the tenant's cloud service instance is avoided when determining the impact area. In other words, even in the event of a network outage, the computing device can still accurately and efficiently determine the impact area, providing a basis for decision-making in subsequent fault recovery.
[0137] It is worth noting that the cloud service impact determination method provided in this application embodiment is applicable not only to cases where a unit (such as the aforementioned switch) experiences a real failure, but also to cases where changes need to be made to the unit (such as driver upgrades or restarts of the aforementioned switch). In other words, even when all cloud computing units are running normally, the impact of unit changes can still be obtained in advance through the cloud service impact determination method provided in this application embodiment. Then, the appropriate window time for unit changes can be determined based on the impact, significantly improving controllability and avoiding unexpected losses. In contrast, traditional methods for collecting instance information can only be applied to cases where certain failures occur and cannot predict the impact.
[0138] The foregoing mainly describes the solutions of the embodiments of this application from a methodological perspective. It is understood that, in order to achieve the above functions, the cloud service impact surface determination device includes at least one of the hardware structures and software modules corresponding to each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.
[0139] This application embodiment can divide the cloud service impact surface determination device into functional units based on the above method example. For example, each function can be divided into its own functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. The unit division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0140] For example, Figure 12 is a schematic diagram of a cloud service impact surface determination device provided in an embodiment of this application. The cloud service impact surface determination device 800 is applied in a computing device, or the cloud service impact surface determination device 800 can be a computing device. The cloud service impact surface determination device 800 includes:
[0141] The receiving module 810 is used to acquire first information, which is used to indicate the fault-affected unit; the fault-affected unit is one of the units required for the operation of the first cloud service in the cloud system architecture; the first cloud service is any cloud service provided by the cloud system architecture.
[0142] The first determining module 820 is used to determine the target cloud service based on the first information and the dependency information; the dependency information is used to indicate the dependency relationships between the units included in the cloud system architecture.
[0143] The second determining module 830 is used to determine the impact area based on the target cloud service; the impact area includes tenants using the target cloud service.
[0144] In one possible implementation, the target cloud service includes one or more of cloud services that directly depend on the fault-affecting unit and cloud services that indirectly depend on the fault-affecting unit.
[0145] In one possible implementation, the device further includes a first interface module for providing a first interface, the first interface including an interactive control for determining the fault-affected unit; in response to receiving a trigger operation on the interactive control, determining the identifier of the fault-affected unit in the impact surface determination request.
[0146] In one possible implementation, the device further includes a second interface module for providing a second interface for displaying the impact surface assessment results.
[0147] In one possible implementation, the dependency information includes a directed graph; the directed graph includes vertices and directed edges, the vertices in the directed graph are used to represent one or more units among the region, the availability zone, the data center, the cluster, the server, the instance, and the cloud service, and the directed edges in the directed graph are used to represent unidirectional dependencies between the vertices in the directed graph;
[0148] The first determining module 820 is further configured to: determine a first point corresponding to the fault-affecting unit in the directed graph based on the identifier of the fault-affecting unit; traverse the points in the directed graph from the first point according to the target direction until a target point is obtained; the target direction is the direction indicated by the directed edge of the first point that depends on the fault-affecting unit; the target point is a point used to characterize the target cloud service.
[0149] In one possible implementation, the point obtained in the i-th traversal is the point to which the directed edge connecting the point obtained in the (i-1)-th traversal points in the target direction; where i is an integer greater than 1.
[0150] In one possible implementation, the points in the directed graph have corresponding label information, which is at least used to characterize whether the unit represented by the point has disaster recovery capability. The units represented by the points obtained in the i-th traversal and the points obtained in the (i-1)-th traversal are units that do not have disaster recovery capability.
[0151] In one possible implementation, the unit represented by the point corresponding to the tag information has disaster recovery capability when the tag information includes any one of the following conditions: the deployment architecture of the unit is multiple availability zones (AZ) or multiple regions (GRI); the unit has automatic flow switching capability; the version information of the unit conforms to the target version information range; and the unit is a stacked switch.
[0152] In one possible implementation, the dependency information is determined based on one or more of the following: the inclusion relationship between the region and the availability zone, the inclusion relationship between the availability zone and the data center, the deployment relationship of the cluster on the data center, the deployment relationship of the cluster on the server, the deployment relationship of the instance on the server, and the deployment relationship of the cloud service on the instance.
[0153] For example, referring to Figure 4, the receiving module 810 can be used to execute S110 as shown in Figure 4, the first determining module 820 can be used to execute S120 as shown in Figure 4, and the second determining module 830 can be used to execute S130 as shown in Figure 4.
[0154] As a feasible example, the cloud service impact determination device 800 provided in this application is implemented through a software module. For example, the software module can be provided to users through a cloud service subscription model, and users can choose different subscription levels according to their needs. Alternatively, the software module can also provide enterprise-level customized services with professional domain customization, interface personalization and extended functions according to the needs of users or enterprises.
[0155] Furthermore, the cloud service impact assessment device 800 provided in this application can also be provided to users as a value-added service, and this application does not limit this. When the cloud service impact assessment device 800 is implemented through a software module, it can also be embedded into other impact assessment software or fault handling systems.
[0156] This application also provides a computing device 100. As shown in FIG13, the computing device 100 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other via the bus 102. The computing device 100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 100.
[0157] Bus 102 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 13, but this does not imply that there is only one bus or one type of bus. Bus 102 can include pathways for transmitting information between various components of computing device 100 (e.g., memory 106, processor 104, communication interface 108).
[0158] The processor 104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0159] Memory 106 may include volatile memory, such as random access memory (RAM). Processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0160] The memory 106 stores executable program code, and the processor 104 executes the executable program code to implement the functions of the aforementioned receiving module, the first determining module, and the second determining module, thereby realizing the cloud service impact surface determination method. That is, the memory 106 stores instructions for executing the cloud service impact surface determination method.
[0161] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0162] As shown in Figure 14, the computing device cluster includes at least one computing device 100. The memory 106 of one or more computing devices 100 in the computing device cluster may store the same instructions for executing the cloud service impact surface determination method.
[0163] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the cloud service impact surface determination method. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for executing the cloud service impact surface determination method.
[0164] The memories 106 in different computing devices 100 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the cloud service impact surface determination device. That is, the instructions stored in the memories 106 of different computing devices 100 can implement the functions of one or more modules among the receiving module, the first determining module, and the second determining module.
[0165] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 15 illustrates one possible implementation. As shown in Figure 15, two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 106 in computing device 100A stores instructions for performing the functions of the receiving module. Simultaneously, the memory 106 in computing device 100B stores instructions for performing the functions of the first determining module and the second determining module.
[0166] It should be understood that the functions of computing device 100A shown in Figure 15 can also be performed by multiple computing devices 100. Similarly, the functions of computing device 100B can also be performed by multiple computing devices 100.
[0167] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform a cloud service impact surface determination method.
[0168] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform a cloud service impact surface determination method, or instruct the computing device to perform a cloud service impact surface determination method.
[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for determining the impact area of cloud services, characterized in that, At least one data center belongs to an availability zone of a region, each of the at least one data center has one or more clusters, the one or more clusters include multiple servers, the multiple servers deploy multiple instances, and the multiple instances provide one or more cloud services; the method includes: Receive an impact assessment request, the impact assessment request including the identifier of the fault impact unit, the fault impact unit including at least one of the following dimensions: region, availability zone, data center, cluster, physical server, cloud service; Based on dependency information including the identifier of the fault-affected unit, the target cloud service that depends on the fault-affected unit is determined; the dependency information is used to characterize the dependencies between the region, the availability zone, the data center, the cluster, the server, the instance, and the cloud service. Based on the target cloud service, the impact assessment results are determined; the impact assessment results include tenants using the target cloud service.
2. The method according to claim 1, characterized in that, The target cloud service includes one or more cloud services that directly depend on the fault-affecting unit and cloud services that indirectly depend on the fault-affecting unit.
3. The method according to claim 1 or 2, characterized in that, The method further includes: A first interface is provided, which includes interactive controls for determining the fault-affected unit; In response to receiving a trigger operation on the interactive control, the identifier of the fault-affected unit in the impact surface determination request is determined.
4. The method according to any one of claims 1-3, characterized in that, After determining the impact assessment results based on the target cloud service, the method further includes: A second interface is provided to display the impact surface assessment results.
5. The method according to any one of claims 1-4, characterized in that, The dependency information includes a directed graph; the directed graph includes vertices and directed edges, the vertices in the directed graph are used to represent one or more units among the region, the availability zone, the data center, the cluster, the server, the instance and the cloud service, and the directed edges in the directed graph are used to represent unidirectional dependencies between the vertices in the directed graph; The step of determining the target cloud service that depends on the fault-affected unit based on dependency information including the identifier of the fault-affected unit includes: Based on the identifier of the fault-affected unit, determine the first point corresponding to the fault-affected unit in the directed graph; Starting from the first point, traverse the points in the directed graph according to the target direction until the target point is obtained; the target direction is the direction indicated by the directed edge of the first point that depends on the fault-affected unit; the target point is a point used to characterize the target cloud service.
6. The method according to claim 5, characterized in that, The point obtained in the i-th traversal is the point that the directed edge connecting the point obtained in the (i-1)-th traversal points to in the target direction; where i is an integer greater than 1.
7. The method according to claim 6, characterized in that, The points in the directed graph have corresponding label information, which is used at least to characterize whether the unit represented by the point has disaster recovery capability. The units represented by the points obtained in the i-th traversal and the points obtained in the (i-1)-th traversal are units that do not have disaster recovery capability.
8. The method according to claim 7, characterized in that, The unit represented by the point corresponding to the tag information has disaster recovery capability when the tag information includes any one of the following conditions: The deployment architecture of the unit is multiple availability zones (az) or multiple regions (region); The unit has automatic flow switching capability; The version information of the unit conforms to the target version information range; The unit is a stacked switch.
9. The method according to any one of claims 1-8, characterized in that, The dependency information is determined based on one or more of the following: the inclusion relationship between the region and the availability zone, the inclusion relationship between the availability zone and the data center, the deployment relationship of the cluster on the data center, the deployment relationship of the cluster on the server, the deployment relationship of the instance on the server, and the deployment relationship of the cloud service on the instance.
10. A device for determining the impact area of a cloud service, characterized in that, The device includes: A receiving module is used to receive an impact assessment request, the impact assessment request including the identifier of the fault impact unit, the fault impact unit including at least one of the following dimensions: region, availability zone, data center, cluster, physical server, cloud service; The first determining module is used to determine the target cloud service that depends on the fault-affected unit based on dependency information including the identifier of the fault-affected unit; the dependency information is used to characterize the dependency relationship between the region, the availability zone, the data center, the cluster, the server, the instance, and the cloud service. The second determining module is used to determine the impact assessment result based on the target cloud service; the impact assessment result includes tenants using the target cloud service.
11. The apparatus according to claim 10, characterized in that, The target cloud service includes one or more cloud services that directly depend on the fault-affecting unit and cloud services that indirectly depend on the fault-affecting unit.
12. The apparatus according to claim 10 or 11, characterized in that, The device further includes a first interface module, used for, A first interface is provided, which includes interactive controls for determining the fault-affected unit; In response to receiving a trigger operation on the interactive control, the identifier of the fault-affected unit in the impact surface determination request is determined.
13. The apparatus according to any one of claims 10-12, characterized in that, The device further includes a second interface module, used for, A second interface is provided to display the impact surface assessment results.
14. The apparatus according to any one of claims 10-13, characterized in that, The dependency information includes a directed graph; the directed graph includes vertices and directed edges, the vertices in the directed graph are used to represent one or more units among the region, the availability zone, the data center, the cluster, the server, the instance and the cloud service, and the directed edges in the directed graph are used to represent unidirectional dependencies between the vertices in the directed graph; The first determining module is also used to, Based on the identifier of the fault-affected unit, determine the first point corresponding to the fault-affected unit in the directed graph; Starting from the first point, traverse the points in the directed graph according to the target direction until the target point is obtained; the target direction is the direction indicated by the directed edge of the first point that depends on the fault-affected unit; the target point is a point used to characterize the target cloud service.
15. The apparatus according to claim 14, characterized in that, The point obtained in the i-th traversal is the point that the directed edge connecting the point obtained in the (i-1)-th traversal points to in the target direction; where i is an integer greater than 1.
16. The apparatus according to claim 15, characterized in that, The points in the directed graph have corresponding label information, which is used at least to characterize whether the unit represented by the point has disaster recovery capability. The units represented by the points obtained in the i-th traversal and the points obtained in the (i-1)-th traversal are units that do not have disaster recovery capability.
17. The apparatus according to claim 16, characterized in that, The unit represented by the point corresponding to the tag information has disaster recovery capability when the tag information includes any one of the following conditions: The deployment architecture of the unit is multiple availability zones (az) or multiple regions (region); The unit has automatic flow switching capability; The version information of the unit conforms to the target version information range; The unit is a stacked switch.
18. The apparatus according to any one of claims 10-17, characterized in that, The dependency information is determined based on one or more of the following: the inclusion relationship between the region and the availability zone, the inclusion relationship between the availability zone and the data center, the deployment relationship of the cluster on the data center, the deployment relationship of the cluster on the server, the deployment relationship of the instance on the server, and the deployment relationship of the cloud service on the instance.
19. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device to cause the computing device cluster to perform the cloud service impact surface determination method as described in any one of claims 1-9.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions; when the computer instructions are executed in a computing device, the computing device performs the cloud service impact surface determination method according to any one of claims 1-9.
21. A computer program product, characterized in that, When the computer program product is run on a computing device, the computing device executes the cloud service impact determination method according to any one of claims 1-9.