Performing recovery process based on red button proxy

By identifying the downtime flag and reconfiguring the communication flow in the cloud platform, the service interruption caused by downtime in the multi-availability area cloud platform is solved, and high availability and rapid recovery of services are achieved.

CN120104401APending Publication Date: 2025-06-06SAP SE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311728414.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-06
Filing Date
2023-12-15
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In multi-availability regional cloud platforms, downtime or problems may lead to service disruption and reduced availability, making it difficult for existing technologies to effectively manage and recover cloud components.

Method used

By identifying the flag of downtime, identifying the entity associated with recovery, and initiating the recovery process, reconfiguring the communication flow in the cloud platform to maintain highly available services.

Benefits of technology

It realizes rapid service recovery in a multi-availability regional cloud platform, ensuring high availability and performance, and reducing the risk of service outages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104401A_ABST
    Figure CN120104401A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including medium-encoded computer program products, for a recovery process on a multi-availability area cloud platform, including executing a request for a red button service from a red button agent to obtain a red flag state associated with a downtime of a component defined for the cloud platform, the red button agent is installed at a first cloud component instance, and the first cloud component instance runs at a first area of a cloud platform comprising a plurality of availability areas; in response to a red sign state of the first cloud component instance received from the red button service, determining that the first cloud component instance is associated with downtime; and executing a recovery process of the first cloud component instance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to computer-implemented methods, software, and systems for data processing in a cloud environment. Background Art

[0002] Software complexity is increasing and leading to changes in the lifecycle management and maintenance of software applications and platform systems. Customer needs are shifting, with increasing requests for flexibility in processes and landscapes and for high availability of access to software resources provided by the underlying platform infrastructure. Failures in network connectivity or the underlying infrastructure can lead to interruptions in the services provided by software applications and degradation of their availability and performance. Summary of the invention

[0003] The present disclosure relates to systems, software, and computer-implemented methods for managing recovery reconfiguration of cloud components when an outage or problem is identified in at least a portion of an availability zone (e.g., one or more segments of a zone or an entire zone) of a multi-availability zone cloud platform. Even if an outage or problem affects at least a portion of the multi-availability zone cloud platform, a recovery process can be applied to maintain highly available services provided by entities running at the multi-availability zone cloud platform.

[0004] In a first aspect, the subject matter described in this specification can be embodied in one or more methods (and one or more non-transitory computer-readable media tangibly encoding a computer program operable to cause a data processing device to perform operations), the method comprising: identifying a selection of a flag from a set of flags defined at a cloud platform comprising multiple availability zones, wherein the flag is selected to identify an outage at a first region of the cloud platform, and wherein each flag in the set of flags is mapped to an entity from a plurality of entities defined for the cloud platform; based on identifying the entity corresponding to the selected flag, determining one or more entities associated with recovering from the outage from the plurality of entities defined for the cloud platform; and in response to determining the one or more entities associated with recovering from the outage, initiating a recovery process to reconfigure communication flows associated with the determined one or more entities at the cloud platform.

[0005] In a second aspect, the subject matter described in the present specification can be embodied in one or more methods (and one or more non-transitory computer-readable media tangibly encoding a computer program operable to cause a data processing device to perform operations), comprising: receiving a selection of a flag from a set of flags defined at a cloud platform comprising multiple availability zones, wherein the selection of the flag is received to trigger recovery execution of an entity running at a first region of the cloud platform and mapped to the flag; determining a type of the entity; in response to determining the type of the entity, activating a load balancer monitor or a central service to generate a corresponding execution plan for recovery; and executing the generated execution plan by the load balancer monitor or the central service to reconfigure, at the cloud platform, communication flows associated with the entity for which recovery execution was triggered.

[0006] In a third aspect, the subject matter described in the present specification can be embodied in one or more methods (and one or more non-transitory computer-readable media that tangibly encode a computer program that is operable to cause a data processing device to perform operations), the method comprising: executing a request from a red button agent to a red button service to obtain a red flag status associated with an outage of a component defined for a cloud platform, wherein the red button agent is installed at a first cloud component instance, and the first cloud component instance runs at a first region of a cloud platform comprising multiple availability regions; in response to receiving the red flag status of the first cloud component instance from the red button service, determining that the first cloud component instance is associated with an outage; and executing a recovery process for the first cloud component instance, wherein executing the recovery process comprises: initiating termination of a cloud component process running on the first cloud component instance; and configuring to send a request directed to the first cloud component instance to a second cloud component instance running at a second region, wherein the second region is a healthy region that is not associated with an outage.

[0007] In a fourth aspect, the subject matter described in the present specification can be embodied in one or more methods (and one or more non-transitory computer-readable media that tangibly encode a computer program that is operable to cause a data processing device to perform an operation), the method comprising: configuring a red button agent at a first entity running at a first region of a cloud platform comprising multiple availability zones; configuring a monitor for evaluating a health state of the first entity, wherein the monitor is configured to perform a health state check by communicating with a health check endpoint provided by the first entity, and wherein the monitor is configured to trigger recovery execution based on detecting a downtime based on an evaluation of communication with the health check endpoint; determining an outage associated with the first entity based on the monitor determining the health state of the first entity based on communication with the health check endpoint or based on selection of a flag at a red button service to notify the red button agent; triggering recovery execution, the recovery execution comprising: stopping a process running by the first entity at the first region; and reconfiguring, at the cloud platform, a communication flow associated with the first entity for which recovery execution was triggered.

[0008] In a fifth aspect, the subject matter described in the present specification can be embodied in one or more methods (and one or more non-transitory computer-readable media that tangibly encode a computer program that is operable to cause a data processing device to perform operations), the method comprising: installing a red button agent at a first cloud component instance of a first cloud component running in a first region of a cloud platform comprising multiple availability zones; executing a request from the red button agent to a red button service to obtain a status of a red flag selected for the cloud platform; in response to receiving the status of the red flag associated with the first cloud component instance, determining that the first cloud component instance is associated with an outage; and executing a recovery process for the first cloud component instance, executing the recovery process comprising: initiating termination of a cloud component process running on the first cloud component instance; and reconfiguring a communication flow directed to the first cloud component to a second cloud component instance running in a second region, wherein the second region is a healthy region that is not associated with the outage, and wherein the first cloud component instance and the second cloud component instance are instances of the same first cloud component running in different regions of the cloud platform.

[0009] In a sixth aspect, the subject matter described in the present specification can be embodied in one or more methods (and one or more non-transitory computer-readable media tangibly encoding a computer program operable to cause a data processing device to perform operations), the method comprising: receiving a selection of a flag from a set of flags defined at a cloud platform comprising multiple availability zones, wherein the flag is selected to identify an outage at a first region of the cloud platform; determining an instance of an entity running at the first region of the cloud platform; determining a state pattern of the running instance of the entity at the cloud platform; in response to determining the state pattern, determining a rule for executing a recovery process for the instance of the entity running at the first region; and in response to determining the rule, executing a recovery process determined based on the state pattern of the running instance of the entity to reconfigure subsequent communications directed to the entity to another one or more instances running at one or more other regions of the cloud platform.

[0010] Similar operations and processes can be performed in a system including at least one processor and a memory communicatively coupled to the at least one processor, wherein the memory stores instructions that, when executed, cause the at least one processor to perform operations. In addition, a non-transitory computer-readable medium storing instructions that, when executed, causes the at least one processor to perform operations can also be envisioned. In other words, although generally described as computer-implemented software embodied on a tangible non-transitory medium that processes and transforms corresponding data, some or all aspects may be computer-implemented methods or further included in corresponding systems or other devices for performing the functions described. Details of these and other aspects and embodiments of the present disclosure are set forth in the drawings and the following description. Other features, objects, and advantages of the present disclosure will be apparent from the specification and drawings and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 An example computer system architecture is shown that can be used to perform implementations of the present disclosure.

[0012] Figure 2 is a block diagram of an example cloud platform environment including multiple availability zones provided with tools and techniques for managing deployment and maintenance of applications and services in different zones according to implementations of the present disclosure.

[0013] Figure 3 is a flow chart of an example method for implementing a "red button" as a solution for triggering a recovery process in accordance with the present implementation.

[0014] Figure 4A is a flow chart of an example method for triggering a recovery process on a cloud platform with multiple availability zones according to an implementation of the present disclosure.

[0015] Figure 4B is a flow chart of an example method for reconfiguring communications of entities running on a multi-availability zone cloud platform according to implementations of the present disclosure.

[0016] Figure 5 is a flow chart of an example method for performing a recovery process based on a type of entity experiencing an outage on a cloud platform with multiple availability zones according to implementations of the present disclosure.

[0017] Figure 6 is a sequence diagram of an example method for performing recovery based on a trigger flag of a centrally managed entity operating at a cloud platform that is a multi-availability zone cloud platform according to an implementation of the present disclosure.

[0018] Fig. 7A is a sequence diagram of an example method for performing recovery based on a trigger flag of an entity running at a multi-availability zone cloud platform managed by a load balancer according to an implementation of the present disclosure.

[0019] Figure 7B is a block diagram of an example system for performing a recovery process in a cloud platform environment including multiple availability zones according to an implementation of the present disclosure.

[0020] Fig. 8A is a sequence diagram of an example method for performing recovery based on a trigger flag of an entity running on a multi-availability cloud platform according to an implementation of the present disclosure, wherein the entity is configured with an agent for monitoring a flag that triggers a recovery process.

[0021] Figure 8B is a flow chart of an example method for performing a recovery process based on a red button agent according to an implementation of the present disclosure.

[0022] Fig. 9 is a schematic diagram of an example computer system that can be used to perform implementations of the present disclosure. DETAILED DESCRIPTION

[0023] This disclosure describes various tools and techniques for managing recovery of cloud components in a multi-availability zone cloud platform.

[0024] In some cases, cloud platforms aim to maintain high availability solutions that can scale and provide services that meet client expectations. Sporadic failures in the underlying infrastructure or network connectivity between components running on the cloud platform may cause outages, which can limit access to the services provided. Configuring a cloud platform with multiple distributed deployments of instances (e.g., applications, platform core services, application services, databases, etc.) may be associated with complex setup and high maintenance costs.

[0025] In some cases, the cloud platform can be constructed to include multiple availability zones (AZs) connected to a high-availability and high-speed network. Typically, each AZ can be an independent data center associated with its own hardware (e.g., associated with a different geographical location), which is connected to other AZs via a high-availability network connection. In some cases, applications can be distributed across one or more AZs to provide high availability of the services provided. Since applications can be executed using different instances running at each of the different AZs and / or hardware nodes (regions or data centers), the risk of not being able to provide services through applications due to downtime can be reduced. In some cases, in order to provide additional availability and reliability, data centers (AZs) can be located in physical locations that are very close to each other.

[0026] In some cases, the cloud platform landscape can be configured to include multiple AZs, wherein an application or service can include multiple instances running in multiple different AZs. A cloud platform can be defined as a public platform including multiple AZs. In some cases, the cloud platform can be accessed from the outside through a single address (e.g., an IP address) as an entry point. The cloud platform can be configured with multiple AZs to ensure that even when a single instance, a segment of a region, or an entire region experiences downtime, the application can be accessed and the application can provide services that can be consumed by a client (e.g., a user or other service or application). According to an implementation of the present disclosure, the cloud platform can be configured with a first region and one or more second regions. In some cases, if a downtime is identified for an instance in the first region, a network request received by an instance running on the first region can be provided to a corresponding instance at the second region. The availability of the services provided can be ensured because the service execution can be routed through a path to access application instances that are not associated with connectivity issues. In some implementations, a path can include several instances organized to exchange requests in a communication flow, where instances can be running at two of the regions, i.e., some instances running at a first region (those instances not affected by the outage) and other instances running at a second region (those instances used for recovery due to a detected outage in their corresponding instances at the first region). Thus, execution of the application, service, and / or database can be independent of problems originating from the underlying infrastructure or problems in one or more AZs where the instances of the application, service, database are running.

[0027] In some cases, the availability zones of the cloud platform can be connected so that if one zone experiences problems such as network downtime, hardware downtime or other problems, the platform can still remain available because there can be at least one or more zones that remain healthy. In some cases, when the cloud platform is configured with a primary zone, and if there is a cloud component that runs its instance in active-passive mode, the primary zone can be the zone where the active instance is running to provide services and is first requested, and the secondary zone (or multiple) can include a passive instance that can be used as a backup in the event of a failure in the primary zone. In some cases, the cloud component can be configured to have multiple instances in an active state (running in active-active mode) so that multiple instances can provide services for requests, and these instances can be distributed in multiple availability zones of the cloud platform. In some cases, the primary zone of the cloud platform can be considered to be a zone that is configured to first (by default) receive requests from the Internet, and then the load balancer in the primary zone can distribute the received requests to the active instances in all zones of the cloud platform.

[0028] In some cases, the cloud platform can be configured to work with multiple availability zones, and multiple availability zones can provide services from one or more regions simultaneously or sequentially. In some cases, the cloud component can be configured to run with different settings or modes for its corresponding instance. In some cases, the cloud component can be instantiated using multiple instances in one or more regions, wherein one or more of these instances can be in active state at the same time. In some examples, the cloud component can run with only one instance as an active instance, or can run with multiple instances on multiple availability zones as active instances, and the active instance can share the load of providing services and / or resources to cloud platform users or customers. For example, a database as a cloud resource that can be provided by a cloud platform can run in active-passive mode, wherein the active instance is deployed in a first region, and the active instance provides services for incoming requests, whether it is a request from outside the platform or a request from an application or service running on the platform. The active instance can perform data replication to the passive instance (or multiple) in the second region (or multiple). In this way, in the event of a problem such as a first instance of a database in a first region being unable to provide service, the passive instance in the second region can be defined as the new current active instance and that instance can continue to serve requests for database resources (e.g., until the first instance, which had experienced the problem, can be repaired and, if in a healthy state, can be restored to the first instance).

[0029] In some cases, a region of a cloud platform can be internally divided into multiple network segments. The segments of a region can correspond to corresponding categories of cloud components. For example, a region can include a segment of a service, a segment of an application, a segment of a database, and other example segment categories.

[0030] In some cases, the cloud platform can provide monitoring of the health of entities defined for the cloud platform. For example, an entity can be an instance of an application, service, or database, but can also include a segment or an entire region of a region. Entities can be defined for the cloud platform based on the criteria selected for monitoring and managing the life cycle of the cloud platform. In some cases, segments can be associated with cloud providers including cloud provider services, while other segments can include services associated with one or more different customers that can deploy those services in these segments. Therefore, different granularities can be defined for entities on the cloud platform, and entities can be managed separately (for example, even if running in the same segment, the services of one customer are handled differently from the services of another customer) or managed in combination with other entities (for example, all core services of the platform can be associated with a single segment, which can be managed as a whole so that the configuration of all core services can be applied in a similar manner). In some cases, based on monitoring, an alarm can be issued if a problem is detected in the first region or a portion thereof. In some cases, the recovery process when a problem is detected involves complex operations and may be time-consuming because complex reconfiguration of the cloud landscape may be required. In addition, the recovery process can also include consideration of future recovery operations so that operations can be transferred back to the instance at the first region once the problem is resolved. According to the techniques described in this application, the recovery process can be applied with fewer requirements and constraints on the configuration of the recovery operation, wherein, in some cases, the recovery operation can be applied only to the first region segment or instance of the application, service, or database.

[0031] The present disclosure provides technology for segmented application of recovery processes in a cloud platform that includes multiple availability zones. Segmented applications can be defined for portions of a cloud platform, wherein, if a problem is detected in a single instance of a service, a recovery process can be triggered based on application logic for performing recovery for the single instance. In some cases, recovery can include operations that are only relevant to a single instance. In other cases, based on considerations of the instance type, recovery can be considered to be applied to a wider entity, such as an entire segment in which a single instance is running. For example, if a core service is detected to be experiencing a problem, the entire core service segment can be considered relevant to performing a recovery process because the core service can be highly important to the service level of the platform and is tightly coupled to other core services running in the core service segment.

[0032] In some cases, the recovery process can be triggered manually or in an automated manner based on monitoring the health of the cloud platform. In some cases, flags can be defined for different entities defined for the cloud platform, where based on the trigger flags, entities of the cloud platform can be notified and the recovery process as defined for the specific entity can be triggered.

[0033] Figure 1 An example architecture 100 is depicted according to an implementation of the present disclosure. In the depicted example, the example architecture 100 includes a client device 102, a client device 104, a network 110, a cloud environment 106, and a cloud environment 108. The cloud environment 106 may include one or more server devices and databases (e.g., processors, memory). In the depicted example, a user 114 interacts with the client device 102, and a user 116 interacts with the client device 104.

[0034] In some examples, client device 102 and / or client device 104 can communicate with cloud environment 106 and / or cloud environment 108 via network 110. Client device 102 can include any suitable type of computing device, such as a desktop computer, a laptop computer, a handheld computer, a tablet computer, a personal digital assistant (PDA), a cellular phone, a network device, a camera, a smart phone, an enhanced general packet radio service (EGPRS) mobile phone, a media player, a navigation device, an email device, a game console, or a suitable combination of any two or more of these devices or other data processing devices. In some implementations, network 110 can include a large computer network connecting any number of communication devices, mobile computing devices, fixed computing devices, and server systems, such as a local area network (LAN), a wide area network (WAN), the Internet, a cellular network, a telephone network (e.g., PSTN), or a suitable combination thereof.

[0035] In some implementations, the cloud environment 106 can include at least one server and at least one data store 120. Figure 1 In the example of , cloud environment 106 is intended to represent various forms of servers, including but not limited to web servers, application servers, proxy servers, network servers and / or server pools. Typically, the server system accepts requests for application services and provides such services to any number of client devices (e.g., client devices 102 via network 110).

[0036] According to implementations of the present disclosure, and as described above, cloud environment 106 can host applications and databases running on a host infrastructure. In some cases, cloud environment 106 can include multiple cluster nodes that can represent physical machines or virtual machines. Hosted applications and / or services can run on VMs hosted on cloud infrastructure. In some cases, an application and / or service can run as multiple application instances on multiple corresponding VMs, where each instance runs on a corresponding VM.

[0037] In some cases, cloud environment 106 and / or cloud environment 108 can be configured in a multi-AZ architecture, wherein the cloud environment can correspond to a data center connected to a highly available network and provide high-speed communication and high network bandwidth for data exchange. In some cases, data centers can be located in close physical proximity to each other. In some cases, in addition to the two cloud environments 106 and 108, a multi-availability zone cloud platform can be defined to provide a public cloud platform that can make it transparent to users and customers to perform operations on multiple AZs. The cloud platform can receive requests to run applications, services, and / or databases that can run on cloud environment 106 and / or cloud environment 108. These applications, services, and databases can be designed, developed, executed, and maintained with respect to different customers and based on accounts configured to execute processes that define applications, services, and databases.

[0038] Figure 2 is a block diagram of an example cloud platform environment 200 including multiple availability zones provided with tools and techniques for managing deployment and maintenance of applications and services in different zones according to implementations of the present disclosure.

[0039] In some cases, cloud platform 200 is a multi-availability zone cloud platform that includes Figure 1 The cloud environment 106 and / or the multiple data centers of the cloud environment 108. Figure 2 Only two AZs are shown in FIG, but the cloud platform 200 can include multiple AZs. In some cases, multiple AZs can be defined as multiple data centers that can execute multiple instances of a single application, service, and / or database in different segments.

[0040] In some cases, the cloud platform 200 includes a first AZ (AZ1) 205 and a second AZ (AZ2) 210. In some cases, the cloud platform 200 provides services through deployed applications (or multiple) and services (or multiple). In some cases, a particular application can be deployed as a single AZ application or a multi-AZ application. In the case where the application is deployed as a multi-AZ application, the application can be deployed using at least one instance in each of the two AZs. In the case where the cloud platform is defined as having a first region and a second region (or multiple), the instance running at the first region can be the instance to which the network call is first routed. In the example cloud platform 200, the first AZ (AZ1) 205 is configured as the first AZ, and therefore, when a request is received from the Internet, the request is routed to AZ1 through the routing layer 201.

[0041] In some cases, the two AZs - AZ1 205 and AZ2 210 - can be implemented as two data centers that are physically relatively close to each other (e.g., having a close physical proximity below a given threshold distance value). In the case where the two AZs are physically close to each other, the two AZs may experience low latency and high-speed interconnection when they communicate (e.g., exchange information and / or requests). In this case, when the two AZs communicate, they can perform data replication and communication between services and application instances located in the two data centers more quickly and reliably.

[0042] In some cases, the multi-instance cloud component can include a first instance running at AZ1 205 and a second instance running at AZ2 210. In some implementations, a load balancer can be defined at each availability zone to dispatch requests received at one zone to corresponding instances of a service, application, or database. In some cases, a load balancer 207 running on a first zone (i.e., AZ1 205) of the cloud platform can receive external requests and dispatch these requests to different services, core services, and applications that can run at any one of the zones 205 and 210. Services 215, core services 212, and applications 220 running on AZ1 205 can provide services to end users and can be coupled to obtain resources from the database segment of the cloud platform. Services 215, core services 212, and applications 220 have corresponding instances at AZ2 210, where, in the event of a problem at an entity, a recovery process can be triggered and the communication flow directed to the instance running on the first zone can be reconfigured to the instance running on the second zone. AZ1 205 includes a database segment where two different databases are running. When databases 217 and 222 are running on the first AZ, they are in the active state. While the instances on the first zone are in the active state, databases 217 and 222 have corresponding instances in the passive state.

[0043] In some implementations, the load balancer 207 of the AZ can process requests received from the Internet and dispatch these requests to any one of the AZs based on dispatch criteria, which can include one or more requirements for processing requests at a specific instance of the application. The criteria can be associated with whether the recovery process is triggered for one or more entities of the cloud platform (such as for AZ1 as the first region), for one of the segments of AZ1 (for example, an application segment for an application including application 220, a core service segment for core service 212, a database segment for any one of databases 217 and 222, or a database segment defined for both of them). In some cases, dispatching requests to instances can be performed in consideration of the load currently experienced by an instance of a service or application. In some cases, an application can have two or more instances running on one AZ (e.g., AZ1 205). Therefore, when a request to access an application is received, the request can be dispatched to the instance of the application with the least number of requests currently being processed. The determination of the instance of the application that processes the request can be based on an evaluation of data associated with multiple instances of the application.

[0044] In some cases, the cloud platform 200 can include a database that can be used by services and applications running on the cloud platform 200. In some cases, high availability of the database at a database segment (e.g., defined as a persistence layer of the cloud platform 200) can be achieved by configuring a redundant setup of database instances in which data is replicated between instances (at data sync 230). In some cases, different DB instances can be located or managed in different AZs. In some cases, and depending on the database, replication can be performed unidirectionally or bidirectionally.

[0045] In some cases, the application and / or service can work primarily with one of the DB instances of a given database, while in other cases, the application and / or service works and interacts with each instance or at least a subset of two or more instances. In some cases, the communication method between the application and / or service and the instance of the database can be based on the capabilities of the DB. By providing flexible configuration of the application or service to work with one or more instances of the database, the processing of requests from the application or service related to the database can be performed without interruption.

[0046] In some cases, only the active load balancer of the load balancers 207 on AZ1 205 can be responsible for handling incoming traffic. In some cases, both AZs can be running and providing resources, where one of the AZs can be associated with an active first-level load balancer that will handle incoming requests to the AZ related to applications and / or services on the cloud platform 200. In the case where one of the AZs is experiencing a region-wide outage, the first-level load balancer instance at the other AZ (i.e., AZ2 210) can be automatically configured to be in active mode (e.g., if it is not an active instance).

[0047] In some cases, the load balancer 207 of AZ1 205 also includes a second-level load balancer managed by the first-level load balancer (the load balancer can be Figure 2 ), and is responsible for routing traffic to a specific application based on, for example, the application location (e.g., URL) of the application.

[0048] In some cases, when an application instance is started at the cloud platform 200, the application instance is registered in a registry (e.g., a load balancer pool) of the second-level load balancer. Such a registry is maintained at both AZs. Based on the registry, the second-level load balancer can route traffic to different application instances. The instance of the second-level load balancer can route received requests to the AZ where the instance of the second-level load balancer resides or to another AZ. This routing is possible because the second-level load balancer registers information about each application (and application instance) running on the cloud platform 200 at each instance of the second-level load balancer.

[0049] In some cases, a flag can be defined to trigger the execution of a recovery process for each entity defined for the cloud platform 200. In some cases, a set of flags can be predefined, wherein each flag can be mapped to an entity defined for the cloud platform. An entity can be a region of a platform, a segment of a region of a platform, a service running on a platform, an application running on a platform, or a database that provides persistence for applications and services running on a platform. In the case of a downtime being detected (e.g., manually detected based on user monitoring of the platform or detected in an automatic manner based on predictive modeling techniques to predict downtime), a recovery process can be triggered by selecting a flag corresponding to the scope of the downtime. For example, if it is determined that the database on AZ1205 is down, a flag for the downtime database can be triggered. This flag triggering can create an event monitored by a cloud platform (e.g., an agent, load balancer, central management service, etc. running on an instance of a service and application), and a corresponding recovery process can be triggered to reconfigure the processing flow only for the portion of the platform (e.g., segment or specific instance) affected by the failure (e.g., database downtime). The technology described in this application can incrementally trigger the recovery process of a component. For example, if a recovery process for segment X is initiated and then a crash (or other problem) is detected at another segment Y, the recovery process for segment Y can also be triggered to run in parallel (or at least partially in parallel) with executing the recovery process for segment X.

[0050] In some cases, recovery techniques according to the present disclosure provide different levels of granularity for applying recovery actions, which can be targeted to resolve issues without requiring recovery of the entire first region to recover a specific segment or a specific cloud component (e.g., a core service). Instead, using recovery techniques as described herein, recovery can flexibly define the scope of entities on the cloud platform that will be associated with the triggered recovery process. This flexibility and segmentation of the portion of the region affected by the recovery process can support rapid transitions from one healthy state "down" to a healthy state, as the number of reconfigurations can be smaller while covering the actions required to return the system to a healthy state.

[0051] Figure 3 is a flow chart of an example method 300 for implementing a "red button" as a solution for triggering a recovery process according to the present implementation. The example method 300 is configured to be used when an application (such as Figure 2 The method 300 defines steps for triggering and executing recovery of an application, which can be performed in the context of triggering and executing recovery, as discussed with respect to the cloud platform 200, with the same as described with respect to Figure 2 The different levels of granularity described.

[0052] In some implementations, it is possible to Figure 2The described technology for managing recovery defines multiple flags for entities defined for a cloud platform. In some cases, selecting a flag from the flags can trigger the execution of a recovery process. The recovery process can be executed based on a specific implementation and configuration at the cloud platform.

[0053] In some cases, the selection of the flag can indicate that the health state of the entity mapped to the flag is modified to a critical state. One or more entities of the plurality of entities can be determined to be associated with reconfiguring the communication flow to recover from the outage. For example, the determination of the one or more entities can be based on an evaluation of the critical state of the entity mapped to the selected flag at an orchestrator component that runs for the cloud platform to manage the communication flow. For example, as described with respect to Figure 6 , Fig. 7A , Fig. 8A and Figure 8B As described, other options for evaluating the changed state based on the selected flag can be applied.

[0054] A data center health service (DCHS) 310 can be defined at the cloud platform, where the service 310 can maintain information about the state of different entities defined for the cloud platform. The DCHS 310 provides an interface that can be called to determine the state of an application. For example, as shown at 315, the state of an application on a first region (DC1) is determined to be healthy (state "ok"), for example, based on the state of the application segment at the first region. In some implementations, recovery can be handled based on logic implemented by an orchestrator 340, which is defined to manage the lifecycle of the cloud platform.

[0055] In some cases, multiple flags can be defined for the cloud platform, as presented in Table 1. The flag can be implemented as a "red button" and provided at a user interface, where it can be selected by a user or based on input from another service. The selection of the button can trigger the corresponding flag. The flag can trigger the execution of the recovery process at different levels of granularity, and thus the segmentation mechanism for application recovery can actually be executed at the cloud platform. Each flag as presented in Table 1 below is an indicator of a problem related to a specific cloud component (e.g., a core service (e.g., an infrastructure element (IEL) segment), a segment (e.g., a service segment), a load balancer (or multiple) (annotated as LB), a database (annotated as DB1 and DB2 at dc1 as an example of different databases), the entire first region (annotated as dc1)). Raising the flag after pressing the corresponding "red button" can trigger a reaction in a service or monitor that monitors the flag state change. Based on the recognition that the flag is raised, the corresponding recovery process can be initiated.

[0056]

[0057] Table 1

[0058] At 320, a flag for an application on the first region (AZ1) is raised based on the selected red button to identify that the application is experiencing a problem. Once the flag is selected at 320, DCHS 310 determines that the state of the application is "dangerous". This change of state at DCHS 310 can trigger orchestrator 340 to initiate execution of a recovery process associated with the application (and with the selected flag). In some cases, the recovery process is initiated by orchestrator 340 by selecting the red button "dc1AppsDown" 320 corresponding to the flag for triggering recovery of the application running on the first AZ1 region, because orchestrator 340 is configured to listen for events created based on the flag (or "red button" selection). For example, recovery process 350 can be triggered by the orchestrator. For example, recovery process 350 can define a set of operations to be performed for failover of the application from AZ1 to AZ2. For example, the set of operations can include:

[0059] Disable all members of AZ1 registered at the load balancer on AZ2;

[0060] Change the status of the application from AZ1 to "unknown" status;

[0061] Disable all cloud controllers in AZ1; and

[0062] Stop working with the load balancer in AZ1.

[0063] In some cases, a cloud controller configured for availability zones stores a configuration that defines details for instantiating new applications, such as the associated hardware and location for starting new application VM(s) (or container(s)) in the cloud platform. The cloud controller can be split between zones, so disabling the controller on AZ1 can prevent new application instances from being started on AZ1. The cloud controller can be configured to work across multiple zones, i.e., during a recovery from AZ1 to AZ2, it may not be necessary to activate the controller in the second zone (e.g., in AZ2). Typically, when a request is made to start N instances of an application (e.g., to serve a larger load), and if the multi-AZ landscape is healthy, the platform's orchestrator can use the cloud controller to start instances in each zone (e.g., dividing them evenly or based on other ratios that can support a better distribution of the load).

[0064] In some cases, if one or more application instances are down in the first region AZ1 of the cloud platform, the corresponding application can continue to operate from its remaining instances in the second region AZ2. In some cases, the application operator may need to manually start additional instances of the application in AZ2 to meet the increased load, or such initialization of new instances can be automated based on the received indication that the instance in the first region AZ1 is down.

[0065] Figure 4A is a flow chart of an example method 400 for triggering a recovery process on a cloud platform with multiple availability zones according to an implementation of the present disclosure.

[0066] In some cases, applications (or other entities, such as services) can be hosted in a cloud environment and can provide services for consumption based on requests (e.g., requests from end users and / or customers). Entities running on a cloud platform can execute logic that includes processing received requests and providing responsive resources or data, dispatching received requests to other entities, querying database entities, accessing external resources to collect data or request services, and other examples of processing logic implemented for a cloud platform.

[0067] In some cases, the example method 400 may be performed in a manner that is capable of Figure 2 The recovery process can be performed at a multi-availability zone cloud platform similar or substantially similar to the multi-availability zone cloud platform 200 of the present invention. The cloud platform can be configured to perform the recovery process based on flag selections that trigger execution related to different parts of the entities defined for the cloud platform. The flags can be presented as shown in Table 1.

[0068] At 410, a selection of a flag from a set of flags defined at a cloud platform including multiple availability zones (e.g., as presented on Table 1) is identified. A flag can be selected to indicate an outage that can limit access to services provided by an entity mapped to the flag.

[0069] Each flag in the set of flags is mapped to an entity in the entities defined for the cloud platform. For example, an entity can be defined to include a region, a segment, an application, a service, a database. The flag is selected to identify an outage at a first region of the cloud platform. For example, the outage at the first region identified can be an outage for an application running at the first region, such as Figure 3 As described.

[0070] In some cases, the selection of the flag can be a manual step performed by a user interacting with a user interface that provides a button for triggering the flag. In some other instances, the selection of the flag can be performed based on a trigger initiated from a service or application that has identified the presence of an outage. In some cases, an outage at the cloud platform can be identified based on monitoring data associated with the health of the cloud platform input at a trained machine learning model. The trained model can use the monitoring data as an input to determine whether an outage exists at an entity running on the cloud platform.

[0071] In some cases, machine learning techniques can be applied to time series data of historical executions of instances running on the cloud platform. Machine learning models can be trained to predict future values ​​of the health status of instances of applications, services, core services, databases, or segments and regions that are defined at the cloud platform. Machine learning models can be trained based on past historical data to determine when a dangerous state of an entity can be expected. The trained model can be a deep neural network. In some cases, predictions of health status can be made based on a combination of several models. Based on such predictions, predicted outages can be determined by obtaining data of the execution of instances running at segments and regions of the cloud platform, and such predictions can be used as input to trigger the selection of flags and thus trigger a recovery process according to the implementation of the present implementation.

[0072] At 420, one or more entities associated with recovering from an outage are determined from the entities defined for the cloud platform. The determination is based on identifying an entity corresponding to the selected flag. In the example of selecting a flag associated with an application running at a first region, the identified entity can be an application segment on the first region. By identifying the application segment as an entity associated with the selected flag, the entity that can be determined to be associated with recovering from an outage can be, for example, an application segment on a second region (e.g., Figure 2AZ2 210 of the application segment. In addition, and depending on the logic used to perform recovery when associated with the application segment, other entities can be defined as being related to recovery. For example, because the problem is associated with an application that consumes a service (such as core service 212 or service 215), the recovery process for the outage of the application segment can be configured to determine that those instances of core service 212 and service 215 are related to the outage, and their corresponding instances running at the second region can be determined to be related to recovery from the outage. For example, a processing flow directed to an application with a "downtime" state can be reconfigured to be directed to an instance at the second region, and the application instance at the second region can consume services from the instance at the second region while the instance of the application at the second region is running. Therefore, processing flows associated with instances of service 215 and core service 212 can also be determined to be related to the outage and set to an "inactive" state to activate their corresponding instances on the second region when the problem is resolved. In some cases, it is possible to activate the corresponding instances on the second region in an orchestrator (such as Figure 3 The orchestrator 340 of the cloud platform can determine the determination of one or more instances associated with recovery from the outage. The orchestrator can be configured to manage communication flows defined at the cloud platform and reconfigure communications associated with entities having instances affected by the outage to instances of entities at another region (or multiple regions). In some cases, when the recovery process is a failover process, calls directed to instances at a first region can be redirected to instances at a second region (when the instances fail over to the second region). In addition, in some cases, the orchestrator can evaluate trigger flags and consider reconfiguring communications in a flow to redirect communications from instances affected by the outage to other instances running at another region or regions instead of running at the first region (e.g., associated with the trigger flag).

[0073] At 430 , in response to determining one or more entities associated with recovery from the outage, a recovery process is initiated to reconfigure communication flows associated with the identified entities at the cloud platform by reconfiguring the communication flows and providing services through the one or more determined entities.

[0074] In some cases, initiation of the recovery process can include determining a new processing flow to replace a previous processing flow that includes an entity mapped to the selected flag. The new processing flow can exclude the entity mapped to the selected flag and replace it with a corresponding entity at another area of ​​the cloud platform, the corresponding entity being defined for recovering the entity mapped to the selected flag. In addition, a service provided by the entity mapped to the selected flag can be disabled, and a request received by the entity can be provided to a corresponding entity at another area.

[0075] In some cases, a notification can be received that an entity that failed over to a second region is successfully running in a first region. In response to receiving the notification, a recovery process can be initiated to reconfigure communication flows and define flows through entities running at the first region of the cloud platform. For example, flows previously defined before the outage was identified can be reconfigured.

[0076] Figure 4B 4 is a flow chart of an example method 450 for reconfiguring communications of entities running on a multi-availability zone cloud platform according to an implementation of the present disclosure. In some cases, entities (e.g., load balancers, applications, databases, or services, among other examples) can be hosted on a cloud platform and can provide services to end users or external applications or platforms. For example, requests for consumption services from services hosted on a cloud platform can be received from end users and / or customers. Entities running on the cloud platform can execute logic that includes processing received requests and providing responsive resources or data, dispatching received requests to other entities, querying database entities, accessing external resources to collect data or request services, and other examples of processing logic implemented for a cloud platform.

[0077] In some cases, multiple entities can be run on a cloud platform, where each entity can be run using one or more instances distributed across one or more availability zones. The entities running on the cloud platform can include one of the following:

[0078] Regional segmentation of cloud platforms;

[0079] · Cloud components running at the regional segments of the cloud platform;

[0080] A region within a cloud platform's multiple availability zones; or

[0081] Load balancers defined for multiple regions of the cloud platform.

[0082] In some cases, various types of entities can be run on the cloud platform at one time or at different times, and a given type of entity can be run using instances at one or more availability zones. In some cases, an entity can be run in a specific state of running instances, such as active-active mode or active-passive mode, in which all instances of an entity are in active mode, and in which one instance is the active instance and other instances can take over the active state in the event that the process of the primary instance is terminated (e.g., due to a crash).

[0083] In some cases, the example method 450 may be performed in a manner that is capable of communicating with Figure 2The cloud platform can be configured to execute a recovery process based on a flag selection that triggers execution related to different parts of the entity defined for the cloud platform. In some implementations, the flags can be presented as shown in Table 1.

[0084] At 455, a selection of a flag from a set of flags defined at a cloud platform including multiple availability zones is received. The selection of a flag can include, for example, Figure 4A 410 of the description for the identification of the selection of the flag. The flag can be selected to identify the downtime at the first area of ​​the cloud platform. In some cases, the selection of the flag can be received (e.g., based on a request for the flag status or as a push notification from the red button service) to trigger a recovery process of an instance running at the first area of ​​the cloud platform and mapped to the flag. In some cases, the selection of the flag can be a selection of the flag as in the example of Table 1. The selection can be performed through a user interface. In some cases, when the flag is selected at the red button service, the red button service can process the selection and determine a portion of the cloud platform affected by the downtime. In some cases, the selection of the flag can be performed by a user of the cloud platform (such as a user who monitors the health of the cloud platform) or by an entity (such as a service for evaluating and / or monitoring the health of the cloud platform). In some cases, the selection of the flag can be performed by notifying the red button service of the received selection user interface.

[0085] In some cases, a flag can be associated with one or more regions of the cloud platform associated with the outage. In some cases, a selected flag can be associated with a regional segment of the cloud platform, and based on such flag selection, it can be determined that entities running at the segment are affected by the outage. Other example associations of flags with portions or the entirety of one or more regions of the cloud platform can be defined and processed in a substantially similar manner.

[0086] At 460, an instance of an entity running at a first region of the cloud platform can be determined. For example, at 455, a first instance of a first application running on a first region can be determined based on the received selection.

[0087] At 465, a state mode of an instance of an entity running at the cloud platform can be determined. As previously discussed, the state mode can be determined as an active-active mode or an active-passive mode.

[0088] At 470, in response to determining the state mode, it is possible to determine rules for performing a recovery process for an instance of an entity running at a first region. In some cases, when the state mode of the running instance of the entity is determined to be an active-passive mode, the determined rules for performing the recovery process include rules for reconfiguring subsequent communications by performing a failover process to redirect subsequent communications directed to the entity to an instance of the entity running in another region of the cloud platform that is not associated with the selected flag. In some cases, when the state mode of the running instance of the entity is determined to be an active-active state, the determined rules for performing the recovery process include rules for reconfiguring subsequent communications by sending requests for services from the entity only to one or more other instances running at one or more regions of the cloud platform, wherein the one or more other regions are not associated with the selected flag.

[0089] At 475 , in response to determining the rule, a recovery process determined based on the state pattern of the running instance of the entity can be executed to reconfigure subsequent communications directed to the entity to another one or more instances running at one or more other regions of the cloud platform.

[0090] Figure 5 is a flow chart of an example method 500 for performing a recovery process based on a type of entity experiencing an outage on a cloud platform with multiple availability zones in accordance with implementations of the present disclosure.

[0091] In some cases, an application (or other entity, such as a service) can be hosted on a server such as Figure 2 The cloud platform 200 of the cloud environment. Entities running on the cloud platform can execute logic that includes processing received requests and providing response resources or data, dispatching received requests to other entities, querying database entities, accessing external resources to collect data or request services, and other examples of processing logic implemented for the cloud platform. The cloud platform can be configured to perform a recovery process based on a flag selection that triggers execution related to the recovery process at the cloud platform. The flag can be presented as in Table 1. Recovery execution can be based on different implementations according to the type of entity indicated as being associated with a downtime (or other failure).

[0092] At 510, a selection of a flag from a set of flags is received. The set of flags is defined at a cloud platform including multiple availability zones. The selection is received to trigger recovery execution of an entity running at a first zone of the cloud platform and mapped to the flag.

[0093] At 520, the type of entity is determined. In some cases, the type of entity can be determined based on the manner in which the life cycle of a particular entity is handled. For example, some applications and databases are managed by a central platform service or subsystem. In those cases, the central platform service can be configured to perform a recovery process when a flag is received to trigger a recovery process for recovering the execution of a failed entity. In another example, some applications or services can be registered for monitoring at a load balancer, and because they are not centrally managed, the central component may not have the tools to perform a recovery process for those applications or services. In another example, some applications or services may include separate logic for handling recovery, wherein such applications or services may be configured with separate agents pre-installed with the provisioning of the application or service, and those agents are able to manage the life cycle of instances (e.g., can stop instances, can redirect communication flows, etc.) without using a central component.

[0094] In some cases, the type of an entity can be determined based on an instance of managing the lifecycle of the entity, where a set of types can be defined for a cloud platform. In some cases, it is possible that a cloud platform includes only a single type of entity and handles their recovery process in a similar manner. Figure 6 (for central components), Fig. 7A and Figure 7B (for load balancers) and Fig. 8A and Figure 8B (For Agents) Describes in more detail the different ways to implement different recovery processes based on the type of entity.

[0095] At 530, in response to determining the type, the load balancer monitor or central service generates a corresponding execution plan for recovery. The execution plan can include a set of steps for recovering from the outage by failing over the marked entity to an instance at a second region and reconfiguring communication flows. For example, the plan for recovery can be such as in Figure 3 A set of steps to be performed at 350 in order to fail over the application.

[0096] In some cases, a first group of entities of the first type can be configured for recovery execution at a central service. The central service can be configured to manage outages associated with entities of the first type at a cloud platform comprising multiple availability zones. In some cases, a second group of entities of the second type can be configured for recovery execution at a load balancer at the cloud platform. Instances of other types can be configured with specific implementations for recovery execution. In those cases, their types can be determined and identified when generating corresponding recovery executions.

[0097] At 540 , the generated corresponding execution plan is executed by the load balancer monitor or the central service to reconfigure the communication flow at the cloud platform associated with the entity for which the recovery execution was triggered.

[0098] Figure 6 is a sequence diagram of an example method 600 for performing recovery based on a trigger flag of a centrally managed entity running at a cloud platform according to an implementation of the present disclosure, wherein the cloud platform is a multi-availability zone cloud platform.

[0099] In some cases, applications (or other cloud components, such as services) can be hosted on a server such as Figure 2 The cloud environment of the cloud platform 200. The entities running on the cloud platform are capable of executing logic, which includes processing received requests and providing response resources or data, dispatching received requests to other entities, querying database entities, accessing external resources to collect data or request services, and other examples of processing logic implemented for the cloud platform.

[0100] The cloud platform can be configured to perform a recovery process based on a flag selection that triggers execution associated with a selected component identified by a raised flag defined at the cloud platform. The flag can be as presented in Table 1. The recovery execution as defined at method 600 involves implementation associated with entities managed by central service 610.

[0101] In some cases, the red button service 605 can be implemented at a cloud platform, where different flags can be defined. The red button service 605 can be substantially similar to Figure 3 DCHS 310, because it can provide similar functionality with respect to identifying a sign and responding with an associated action. Figure 2 and Figure 3 When the selection of the button in question is triggered, an event for the changed state of the entity at the cloud platform is created, and such event can be consumed by the central service 610 .

[0102] For example, at 630, red button service 605 is requested by the central service to provide a response to the status of a flag for the cloud platform managed by central service 610. Central service 610 can obtain a response from which it can determine whether there is a selection of a flag for a cloud component running on the cloud platform. If there is a flag selected for the cloud component, central service 610 can send a request to load balancer 615 of the cloud platform to disable a pool of cloud components 635. In response to disabling the pool of cloud components, load balancer 615 can send a request 640 to stop traffic to cloud component 620 in the first region. Load balancer 615 can send a request to reroute network traffic directed to cloud component 620 in the first region to corresponding cloud component 625 in the second region of the cloud platform.

[0103] Central service 610 can deactivate cloud component 620 running in a first region of the cloud platform and can activate cloud component 625 running in a second region. Cloud component 620 and cloud component 625 are corresponding instances of cloud components (such as applications, services, or databases) whose deployment and lifecycle are managed by central service 610.

[0104] In some cases, a database or persistence service can be considered a type of entity that can be configured to work in active and passive states for its different instances in corresponding regions of a multi-availability cloud platform. While the instance in the first region is servicing the request, the database instance or persistence service instance running in the second region can be active for replication from its corresponding instance in the first region. Therefore, the instance in the second region does not serve any other requests, but is used for replication (e.g., synchronization between data stored in each region). In the event of an outage in the first region, the instance in the second region can be configured to be active and can take over the responsibility of servicing the request when the instance in the first region recovers.

[0105] In some cases, a cloud component as an application can run as multiple instances in active mode in multiple regions. Thus, when traffic to an application instance stops, the instance running in the second region can start receiving rerouted traffic from the load balancer, such as Figure 6 As described in steps 640 and 645.

[0106] By implementing a central service that can handle the lifecycle of a cloud component and can take the necessary actions to perform the recovery process of that component, maintenance of instances of that component can be performed faster because an operator may not need to intervene manually to trigger each task of the recovery plan to be executed. This management of the recovery process automates the process and ensures high availability of the centrally managed components.

[0107] Fig. 7A is a sequence diagram of an example method 700 for performing recovery based on a triggering flag of an entity running at a multi-availability zone cloud platform managed by a load balancer according to an implementation of the present disclosure.

[0108] In some implementations, the red button service 701 can communicate with Figure 6 The red button service 701 can be configured at the cloud platform in the same or substantially similar manner as described above. Figure 2 The cloud platform 200 is the same as or substantially similar to the cloud platform 200.

[0109] The cloud component 703 runs at the first region of the cloud platform. The cloud component 703 can be as follows Figure 2 The described applications, services or core services. Cloud components 703 can be components to be monitored for downtime (or other issues) based on monitors registered at load balancer 705 .

[0110] The load balancer 705 can receive a request from a process 715 running at the cloud component 703 to register the cloud component 703 at the load balancer 705. The load balancer 705 can create a node monitor at 710.

[0111] The red button service 701 can be configured to provide the status of the selected flag. The process 715 requests the status from the red button service at 725 and provides information about the status to the health check endpoint 720 at the cloud component 703 .

[0112] If the flag is selected for component 703 and component 703 is running in the first region, health check endpoint 720 can notify 740 node monitor 710 that cloud component 703 is experiencing downtime. Node monitor 710 can initiate termination of network traffic toward process 715 of cloud component 703.

[0113] In some cases, when a red button flag is defined for a given component and the information is provided to the health check endpoint 720 at 730, the health check endpoint can return an error response at 740 to induce the node monitor 710 to cut off traffic to the cloud component 703 in the first region at 745. The node monitor 710 can be configured to periodically check (requests at 735 and 750) with the health check endpoint 720 to determine the health status of the cloud component 703 because the health check endpoint periodically obtains information about the selected flag from the red button service 701. As long as the node monitor 710 determines that the flag of the component is not pressed, that is, the status of the component is "Ok" rather than "down", "dangerous" or other status indicating an outage or problem, the node monitor 710 forwards (at 755) the network traffic to the process 715. When the red button flag is reset, the health check endpoint 720 can be notified and can provide the status OK to the node monitor 710 to initiate the recovery of traffic to the cloud component 703 in the first region.

[0114] In some cases, the node monitor 710 may be logically separate from the logic of the red button service 701 and how to determine whether to reconfigure communication flows, redirect services, stop services, or perform other actions. The node monitor 710 may be configured to perform actions based on responses obtained from the health check endpoint 720. In this case, the health check endpoint 720 determines the state of the cloud component 703 and notifies the node monitor 710 to react (as in 740) corresponding to the response provided to the node monitor 710 from the health check endpoint 720.

[0115] Using node monitors in the load balancer provides a simple and effective process to handle complex recovery steps. The implementation of health check endpoint 720 is provided by the cloud component. The implemented health check endpoint 720 can be used to perform health checks on various problems or issues that can be identified at the cloud component, which can be specific to the application logic in addition to evaluating the indication of a raised flag of the component or the segment in which the component runs.

[0116] Figure 7B is a block diagram of an example system 765 for performing a recovery process in a cloud platform environment including multiple availability zones according to an implementation of the present disclosure. In some implementations, the recovery process described with respect to the example system 765 can be performed in relation to Fig. 7A The described recovery scenario is performed in which a monitor at a load balancer registers components to handle the execution of the recovery plan.

[0117] In some cases, the data center health service (DCHS) 760 can be configured to run at the cloud platform, such as Fig. 7A The cloud platform discussed in the present disclosure is also related to the cloud platform discussed in the present disclosure. The service instance 785 is running at the first region 775 of the cloud platform. The DCHS 760 can be connected to the Figure 3The same or substantially the same service discussed. In some cases, DCHS 760 is capable of sending a push or pull request to service instance 785 to determine the health status of service instance 785. Service instance 785 is configured with an agent that communicates with the DCHS 760 service. DCHS 760 is capable of determining whether service instance 785 is experiencing downtime. For example, DCHS 760 is capable of determining that there is a downtime based on an evaluation of health status data, wherein the health status data is determined based on monitoring the execution of instances running at different segments or regions of the cloud platform. In another example, the determination of downtime can be based on user input or other application input indicating an identified downtime at the cloud platform. In other examples, a determination can be made based on a combination of manual input and algorithmic evaluation of the input to confirm the downtime. Other examples can include implementing predictive logic to identify expected downtime based on a predictive model, wherein the predictive model is generated based on evaluating historical data from monitoring the cloud platform.

[0118] In some implementations, Fig. 7A The outage at can be determined to be associated with segment A on the first region 775. Service 785 can be registered at a node monitor running on the load balancer 770. A flag can be raised for service segment A (e.g., a red button service, such as red button service 701) of the first region 775. The node monitor 785 of the service can determine that the segment A state has been determined to be dangerous based on monitoring events of the flag state from the red button service. As a result, the node monitor 785 can perform an action to terminate the business direction from the load balancer (LB) 770 to the service 785 because the service is running in the segment A that is experiencing the outage. The node monitor 785 can configure the communication flow so that further communication flows associated with the new network business will be provided to the service instance 790 corresponding to the service instance 785, and the service instance 790 is running at the second region of the cloud platform. The service instance 785 running in the first region is associated with a processing flow in which a service (service B) from the first region sends a request to obtain resources from the service 785. Since service instance 785 has limited network access (e.g., Fig. 7A 745 of the stop network business), so that service B can obtain service from the service instance 790 at the second area 780. The node monitor (such as Fig. 7ANode monitor 710) can continue to monitor the health of instance 785 (by sending checks to a health check endpoint at instance 785 or by communicating with DCHS 760) to determine the status of service instance 785. A DCHS agent can be deployed at service instance 785. The DCHS agent can be configured to handle events associated with a raised red flag, for example, events emitted from a red button service as described in the present disclosure. In some cases, the DCHS agent can be configured to listen for events in a passive mode or actively query DCHS 760 to determine the red flag status and initiate actions related to the service process of service 785. For example, if a flag is raised to indicate an outage at service 785, the DCHS agent can trigger communications to stop the service process on AZ1 775. In this case, the raised flag can stop communications directed to instance 785 at AZ1 775, and communications can be redirected to service instance 790 at a second area AZ2 780. In this case, if DCHS is unable to perform a service process termination at service instance 785, the node monitor at LB 770 can perform a secondary function and act as a backup. For example, DCHS may not be able to perform such an action in the event that the entire runtime of service instance 785 is slowed down due to a hardware problem that would affect the operation of any entity running there. The node monitor at LB 770 can detect the problem (based on the health check endpoints (such as Fig. 7A The service instance 785 may be detected by the service provider in the health check endpoint 720 (a slow response or lack of response), and action may be taken to recover the service instance 785 to the service instance 790 on the second zone 780AZ2.

[0119] When it is determined that the status of service instance 785 is OK (recovered), a recovery process can be triggered, in which case requests to service instance 785 can be restored. For example, service B can be reconfigured to restore requests to service instance 785 instead of service instance 790 used based on the recovery process and reconfiguration of the communication flow.

[0120] By combining the DCHS agent and the node monitor in the load balancer, the risk of experiencing slow execution due to limitations or constraints of the runtime infrastructure in which the cloud component is running can be mitigated. For example, when a component health check endpoint fails to respond, the monitor will time out and the load balancer 770 can automatically disconnect the connection to the cloud instance because there is a problem (e.g., an outage that needs to trigger a recovery process).

[0121] Fig. 8Ais a sequence diagram of an example method 800 for performing recovery based on a flag that triggers an entity running on a multi-availability zone cloud platform according to an implementation of the present disclosure, wherein the entity is configured with an agent for monitoring the flag and triggering the recovery process. In some cases, the multi-availability cloud platform can be substantially similar to Figure 2 Describe the cloud platform.

[0122] In some cases, when outages are identified at the cloud platform, recovery processes can be configured and executed to provide high availability without process and service interruptions. In some cases, it is possible to recover from outages that are not covered by a central authority (e.g., Figure 6 ) and without component registration in the load balancer (e.g. Fig. 7A and 7B In some cases, the red button agent can be installed with the runtime of the cloud component where the cloud component is running (e.g., within a virtual machine, container or container group (pod), and other example infrastructure).

[0123] In some cases, a cloud platform can be hosted on multiple regions and / or zones to support high availability and service level performance, for example, in the event of an outage or other failure in one region and / or zone. Replication of resources, applications, databases, services, or other entities can be performed on all or some of the multiple regions to achieve at least two instances of each entity running in different regions. In such a multi-availability zone cloud platform, for example, through node monitors at a load balancer or through a central service (such as, for example, Figure 6 , Fig. 7A and Figure 7B), the health status of the entity can be centrally monitored, and the configuration for the recovery process (e.g., recovery process) can be implemented. Additionally or alternatively, the configuration step can be performed by a component-specific agent, which can have logic for performing the recovery step to facilitate smooth service from the cloud platform while seamlessly reconfiguring the communication associated with the entity having the instance experiencing the downtime. Based on the reconfiguration of the communication of the entity, the request to the entity can be processed at the corresponding instance of the entity in the healthy area. For example, the configuration step for performing the recovery can include activating a new instance in the healthy area and rerouting the business to the new instance. For example, if the first area of ​​the cloud platform becomes unavailable or unavailable during the downtime, the database running in the healthy area can be promoted from a passive state to an active state. As another example, if the first area is unavailable, the instance of the entity (such as an application running at the first area) may be unreachable. Therefore, the reconfiguration step can include business redirection. For example, Internet business can be re-directed to an application instance in the second area (or multiple) that remains healthy while the first area is experiencing unavailability (e.g., due to downtime). The lack of automated recovery processes for different types of cloud components or complex (manual) processes that are time-consuming to perform may result in negative impacts because service availability may be interrupted, the platform may experience downtime, and user experience may be affected. In some cases, if an instance of an entity is unavailable due to downtime and a recovery process is not implemented in due time, multiple requests to that instance may create inefficiencies in the processing and operation of the platform. This may be associated with time delays in responding to requests received at the platform and may also be associated with data loss.

[0124] In some implementations, it is possible to Figure 6 The red button service 805 can be configured at the cloud platform in the same or substantially similar manner as the red button service 805 described above. Figure 2 The cloud platform 200 is the same as or substantially similar to the cloud platform 200.

[0125] The cloud component 810 is running at the first region of the cloud platform. The cloud component 810 can be as follows Figure 2 The described application, service or core service. The cloud component 810 can be a component that is defined to handle the recovery process alone, and the central component or load balancer is not responsible for monitoring the flag raised for this component.

[0126] In some implementations, the cloud component 810 is configured with a red button agent 815 (which can be substantially similar to Figure 7B) and is capable of executing cloud component process 820. Cloud component process 820 is defined as executing process flows to provide services to other instances on the cloud platform or externally.

[0127] The red button agent 815 can perform a request to the red button service 805 (e.g., periodically, upon receiving an event, according to a defined schedule, etc.) to obtain the status of a red flag selected for the cloud platform. For example, if a downtime is identified for the cloud component 810 in the first region, a flag can be selected for such downtime, and the red button agent 815 can identify the selection of the flag. At 830, it is determined that a flag for downtime of the cloud component has been triggered (e.g., by pressing a red button, as previously described). In response to determining that the flag is triggered, at 835, the red button agent 815 performs a process of stopping the cloud component process 820 (e.g., executing a script) of the cloud component 810. In some cases, rather than performing a process of stopping the cloud component process 820 outside the service, an instruction can be sent to the cloud component process 820 to stop the process. For example, an instruction can be sent to an endpoint of a cloud component instance. In the case where the cloud component process 820 is instructed to stop, the initiation of a stop or deactivation process can be performed by an external process to trigger the stopping of the process, because the service will be deactivated and will not be able to receive an instruction to start itself. In this manner, the instances of cloud component 810 in the first region will stop servicing any requests, and requests will be handled by the instances (or instances) in the healthy region (or regions), e.g., by the instance(s) in the first region. Figure 2 , Figure 3 and the second region processing described in FIG. 4 . Thus, when the flag is changed after the outage of cloud component 810 ends (e.g., as determined by a regular request from red button agent 815 to red button service 805), red button agent 815 can restart cloud component process 820 in the first region to perform recovery operations. In some cases, cloud component 810 provides an interface through which cloud component process 820 can be started and stopped. In some implementations, the interface can be implemented in the form of a start or stop function, wherein when the start or stop function is triggered, the execution of a script can be initiated (e.g., via a start or stop script, as discussed above) to manage the life cycle of cloud component process 820.

[0128] In some cases, the red button agent 815 can be configured to perform an action after tracking the red button flag state. For example, when the red button agent is installed with the component runtime, the red button agent can track other indicators of the health of the cloud component internally, and report these indicators to external monitoring tools and / or actively disable the cloud component instance until the problem is remedied. In some implementations, such other indicators can include CPU load, memory usage, swap file usage, or available disk space, as well as other examples. These indicators can be tracked by the red button agent and / or by other entities running in the cloud component 810. Such indicators are traceable from within the component, and the measurement of such indicators can be performed by an agent running at the infrastructure (e.g., a virtual machine or container) where the component is running. In some cases, external measurements can be performed for the cloud component 810, for example, the response time of the component to the received request. In some cases, the red button agent 815, which is an internal agent running at the cloud component 810, may not be able to directly measure such response time metrics and can obtain them indirectly. Based on the information obtained of one or more indicators of the health of cloud component 810 , red button agent 815 is able to manage the execution of cloud component process 820 , for example, starting and stopping the process in response to determining the health status of cloud component 810 based on notifications from red button service 805 .

[0129] In some additional instances, some cloud components may need to perform specific steps for their recovery. In this case, the red button agent can read a file including instructions related to the specific cloud component and can perform these steps with or without other default steps defined for performing the recovery process.

[0130] Using Red Button Proxy as a component proxy is a centralized solution that implements default behaviors for disabling components in affected regions or region segments during outages, and their adoption requires little effort from component owners. This approach addresses the case where communication with cloud component instances is performed directly through IP addresses and bypasses the load balancer. In some cases, the use of Red Button Proxy can be combined with the registration of monitors in the load balancer (such as Figure 7B The red button agent and the monitor are combined in a load balancer to cover a wider range of outages. By combining the red button agent and the monitor in the load balancer, the risk of experiencing slow execution due to limitations or constraints of the runtime infrastructure in which the cloud components are running can be mitigated. For example, when the component health check endpoint fails to respond, the monitor will time out and the load balancer will automatically direct connections away from the cloud instance that has problems (e.g., an outage that needs to trigger a recovery process).

[0131] Because the red button agent 815 runs in the same runtime infrastructure as the cloud component 810, the red button agent has greater flexibility and also has access to perform further actions, thus having more capabilities for monitoring and managing the cloud component.

[0132] Figure 8B is a flow diagram of an example method 850 for performing a recovery process based on a red button agent according to an implementation of the present disclosure.

[0133] In some cases, example method 850 can be performed in the context of a cloud component running at a first availability zone of a cloud platform that is a multi-region platform. The cloud component can be substantially similar to Fig. 8A The cloud component 810 of the cloud platform. The cloud component can be an instance of a service, application, or database, etc. The cloud component can run as an instance of an entity, where the entity can be associated with other instances running at various availability zones of the cloud platform.

[0134] In some implementations, the red button agent can be configured for a given cloud component instance running at a region of the cloud platform. The red button agent can be substantially similar to the red button agent 815. The red button agent can be installed in the same virtual machine, container or container group, and other examples together with the runtime of the first cloud component instance. In some cases, the red button agent can be configured to query the red button flag to obtain a change in health status, for example, when the component is associated with a downtime (e.g., experiencing downtime itself, or being affected by the downtime of another entity to which the component instance is coupled, and other example associations), the red flag state can be raised. The red button agent can be implemented to have logic for starting and stopping a component instance such as the first cloud component instance based on a change in a determined red button flag (starting the component instance when the downtime ends, and stopping the instance when the downtime is identified). In some cases, the first cloud component instance can be implemented to provide an interface, wherein the red button agent can start and stop the instance through the interface.

[0135] At 855, a request is sent from the red button agent to the red button service to obtain the red flag status associated with the downtime of the component defined for the cloud platform. In some cases, the red button service can be configured to push notifications to the red button agent when the red flag is raised for the corresponding component. In some cases, the red button agent can be configured to pull such information about the status of the red flag, for example, according to a pull schedule (e.g., every 5 seconds or every 2 minutes, among other examples). In some cases, the red button service can be substantially similar to Fig. 8AThe red button service 805. The red button agent can be configured to periodically execute the request, for example, based on a predefined time interval. The red button agent can be installed at a first cloud component instance at a first region of a cloud platform including multiple availability zones.

[0136] In some cases, the red flag state can identify that the downtime is associated with the corresponding component associated with the flag. For example, the flag can be defined at the regional level. In such an example, if the red flag of the first region is raised, the red button service can provide the red flag state to identify and indicate that the entire first region is down. In some other examples, the red flag state can be defined for a component that is a service, application, or database running at a specific region of the cloud platform. In some more examples, the red flag state can be defined for a segment of a region of the cloud platform (e.g., for a network segment dedicated to running a database, application, or service).

[0137] At 860, in response to receiving a red flag status about the first cloud component instance from the red button service, the red button agent determines that the first cloud component instance is associated with an outage. For example, the red flag status can define that the first cloud component instance is experiencing an outage, the network segment in which the first cloud component instance is running is experiencing an outage, or the first region in which the first cloud component instance is running is experiencing an outage. In such an example case, the red button agent can determine that the outage identified (e.g., by a monitoring service, by a status checker, or by a platform administrator) is affecting the first cloud component instance. The red button agent can implement logic for performing a recovery process, which can include reconfiguring (e.g., activating / terminating and / or redirecting) communications and executing instances on the cloud platform so as to maintain the provided services even if one or more entities may be experiencing an outage (e.g., a network outage, hardware problems, downtime, etc.). In some implementations, the reconfiguration may not be performed directly by the red button agent. However, the red button agent can control the start and stop of the process instance in response to the identified outage. When a process instance stops, the cloud platform's load balancer can detect that the instance is no longer responding and direct traffic to other instances of the component running at the healthy zone (eg, we can call these instances healthy).

[0138] In some cases, the cloud platform's load balancer(s) can be configured to provide an interface that can facilitate communication with the red button agent. In some cases, the red button agent can send a request directly to the load balancer(s) to stop service to a specific instance of a component because it is in an unhealthy state (e.g., running at an unhealthy segment or region). In some cases, termination of a specific instance of a component can be performed without direct instructions from the load balancer(s). In some cases, the red button agent and / or the load balancer(s) can inherently detect that an instance is terminated and initiate reconfiguration of communication flows, as previously discussed.

[0139] At 865, a recovery process of the first cloud component instance is performed (eg, performed by or triggered by the red button agent).

[0140] At 870, a recovery process is performed by initiating termination of a cloud component process running on a first cloud component instance, and at 875, a recovery process is performed by configuring a load balancer of the cloud platform to detect that the instance is stopped because it is not responding, and directing the business for the component to another instance (e.g., a healthy instance in a healthy zone (or multiple)) by redirecting the request to a healthy instance. The cloud component process configured at the first cloud component instance is a process flow configured to provide services to other instances running on the cloud platform and / or outside the cloud platform. In some cases, termination of execution of the cloud component process can be performed by a process that runs a script to stop the process. In some cases, initiation of termination can be performed by sending an instruction from a red button agent to a predefined endpoint (e.g., an exposed interface) of the first cloud component instance to stop the process.

[0141] In some cases, the first cloud component can be associated with an active-passive state mode of the component instance (e.g., as described with respect to Figure 4BAs discussed above). When it is determined that an instance of a component is running in a region associated with an identified outage, other instances (or multiple instances) of the component can be determined as available instances at other regions, as informed by the selection of a flag, but they can be in passive mode (not actively processing new requests and, for example, used for backup). These passive instances can be started and used to process further requests for the first cloud component. Thus, requests directed to the first cloud component can be executed at a second cloud component instance running at a second region. The second region is a healthy region not associated with an outage, wherein the first cloud component instance and the second cloud component instance are instances of the same cloud component running at different regions of a cloud platform (e.g., a cloud application). In some cases, when an outage is determined, the first cloud component instance may not be deleted. Instead, the process running at the instance may be terminated without deleting the component. However, since the process will be terminated, another instance is changed to an active state. Based on this reconfiguration, the load balancer (or multiple) of the cloud platform can execute requests associated with the first cloud component by distributing requests for the cloud component to other instances (i.e., the second cloud platform instance).

[0142] In some cases, the first cloud component can be configured to run instances in an active-active mode, and requests directed to the first cloud component can be distributed among one or more other instances of the component according to a distribution algorithm. For example, the distribution algorithm can be based on the Round Robin principle. In those instances, the load balancer of the cloud platform can distribute the requests to those instances that are active and not send requests to instances that are affected by the downtime and have their processes terminated.

[0143] In some cases, the execution of the recovery process as described with respect to steps 865, 870, and 875 can also include the execution of other operations. Typically, the recovery process can be predefined for a component or a group of components. In some cases, when the recovery process is triggered to execute, a file including instructions for executing the recovery process can be read to determine the steps of the process that may be customized and / or generic for a particular component.

[0144] In some cases, the red button service can obtain information for the health status of components on the cloud platform and can restart the first cloud component instance if the status of the first cloud component instance changes and is no longer associated with a red flag (the component is no longer associated with an outage). The start can be initiated by the red button agent upon receiving an indication of the end of the outage via the red flag status provided by the red button service.

[0145] In some cases, the red button agent can be configured for a first cloud component instance in a setup where there is a node monitor running at a load balancer, the load balancer being substantially similar to a server comprising Figure 7BIn this case, the first cloud component instance (e.g., Figure 7B The service instance 785 in the cloud platform can be configured with a node monitor at a load balancer (e.g., load balancer 770) and also has a red button agent installed with the first cloud component instance. The red button agent can register the first cloud component instance at the node monitor, and the node monitor can be configured to monitor the first cloud component instance at the cloud platform.

[0146] In such a configuration, when a node monitor running at a load balancer determines that a first cloud component instance is associated with an outage, the node monitor can be configured to modify a communication flow toward the first cloud component so that a request for the component can be processed at a second cloud component instance at a second region of the cloud platform because the first cloud component instance is affected by the outage and can have terminated process execution. Additionally, execution of the first cloud component instance at the first region can be terminated, and execution of the second cloud component instance at the second region of the cloud platform can be activated.

[0147] In some cases, communications directed to the first cloud component instance can be received directly, and such communications can bypass the load balancer configured for the cloud platform. In those cases, when the instance is associated with an outage, the load balancer will not be able to redirect requests to the first cloud component instance or reconfigure communications associated with the first cloud component. For example, the instance can communicate with the first cloud component instance based on an IP address. In those cases, when the first cloud component instance is determined to be associated with an outage, the red button agent of the first cloud component instance can trigger a recovery process so that even if there is direct communication toward the first cloud component instance during the outage, the communication can be rerouted to a healthy instance for recovery, such as the second cloud component instance discussed above. In some cases, the recovery process can include terminating the affected cloud component instance of the outage, thereby ensuring that communication with the component instance will be stopped and / or avoided.

[0148] In some cases, in certain outage situations, such as a slow runtime environment (e.g., hypervisor) that deploys components and red button agents, the red button agent may not be able to access the received indication of the red button flag status. In those cases, if the red button agent is implemented in a configuration where a node monitor is also registered with a load balancer, the node monitor can determine that the first cloud component instance is not responding or responding slowly, and when a threshold criterion for the response is met, a recovery process can be automatically triggered. The node monitor can be configured to call a health check endpoint provided by the first cloud component instance, and if the response times out, the load balancer can automatically terminate the connection to the first cloud component instance identified as an instance associated with a problem (e.g., an outage), and the services of the first cloud component can be provided by a healthy running instance (such as a second cloud component instance).

[0149] Reference now Fig. 9 , a schematic diagram of an example computing system 900 is provided. The system 900 can be used for the operations described in association with the implementations described herein. For example, the system 900 can be included in any or all of the server components discussed herein. The system 900 includes a processor 910, a memory 920, a storage device 930, and an input / output device 940. The components 910, 920, 930, and 940 are interconnected using a system bus 950. The processor 910 can process instructions for execution within the system 900. In some implementations, the processor 910 is a single-threaded processor. In some implementations, the processor 910 is a multi-threaded processor. The processor 910 can process instructions stored in the memory 920 or on the storage device 930 to display graphical information of a user interface on the input / output device 940.

[0150] The memory 920 stores information within the system 900. In some implementations, the memory 920 is a computer-readable medium. In some implementations, the memory 920 is a volatile memory unit. In some implementations, the memory 920 is a non-volatile memory unit. The storage device 930 can provide large-capacity storage for the system 900. In some implementations, the storage device 930 is a computer-readable medium. In some implementations, the storage device 930 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device. The input / output device 940 provides input / output operations for the system 900. In some implementations, the input / output device 940 includes a keyboard and / or a pointing device. In some implementations, the input / output device 940 includes a display unit for displaying a graphical user interface.

[0151] The described features can be implemented in digital electronic circuits, or in computer hardware, firmware, software, or a combination thereof. The apparatus can be implemented in a computer program product tangibly embodied in an information carrier (e.g., in a machine-readable storage device for execution by a programmable processor), and the method operations can be performed by a programmable processor executing a program of instructions to perform the functions of the described implementation by operating on input data and generating output. The described features can advantageously be implemented in one or more computer programs executable on a programmable system, the programmable system comprising at least one programmable processor coupled to receive data and instructions from a data storage system, at least one input device, and at least one output device, and to send data and instructions to a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used directly or indirectly in a computer to perform some activity or produce some result. A computer program can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0152] As an example, suitable processors for executing instruction programs include both general-purpose and special-purpose microprocessors, as well as the sole processor or one of multiple processors of any type of computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The elements of a computer can include a processor for executing instructions and one or more memories for storing instructions and data. Typically, a computer can also include one or more mass storage devices for storing data files, or be operably coupled to communicate with one or more mass storage devices; such devices include disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory can be supplemented by or incorporated into an ASIC (Application Specific Integrated Circuit).

[0153] To provide interaction with a user, these features can be implemented on a computer having a display device (such as a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to the user and a keyboard and pointing device (such as a mouse or trackball) through which the user can provide input to the computer.

[0154] Features can be implemented in a computer system that includes a back-end component (such as a data server), or includes a middleware component (such as an application server or an Internet server), or includes a front-end component (such as a client computer with a graphical user interface or an Internet browser), or any combination thereof. The components of the system can be connected by any form or medium of digital data communication (such as a communication network). Examples of communication networks include, for example, LANs, WANs, and the computers and networks that form the Internet.

[0155] A computer system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a network (such as the one described). The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0156] In addition, the logic flows depicted in the accompanying drawings do not require the particular order or sequential order shown to achieve the desired results. In addition, other operations can be provided, or operations can be eliminated from the described flows, and other components can be added to or removed from the described systems. Therefore, other implementations are within the scope of the appended claims.

[0157] A number of implementations of the present disclosure have been described. However, it will be appreciated that various modifications may be made without departing from the spirit and scope of the present disclosure. Therefore, other implementations are within the scope of the appended claims.

[0158] In view of the implementation of the above-mentioned subject matter, the present application discloses the following list of examples, wherein one feature of a separate example or more than one feature combination of the examples and optionally combined with one or more features of one or more additional examples are additional examples that also fall within the disclosure of the present application.

[0159] Example

[0160] Although the present application is defined in the accompanying claims, it is to be understood that the invention can also (alternatively) be defined according to the following examples.

[0161] Segmented recovery of cloud components in a multi-availability zone cloud environment

[0162] Example 1. A computer-implemented method comprising:

[0163] identifying a selection of a flag from a set of flags defined at a cloud platform comprising a plurality of availability zones, wherein the flag is selected to identify an outage at a first zone of the cloud platform, and wherein each flag in the set of flags is mapped to an entity from a plurality of entities defined for the cloud platform;

[0164] Based on identifying the entity corresponding to the selected indicia, determining one or more entities associated with recovering from the outage from a plurality of entities defined for the cloud platform; and

[0165] In response to determining one or more entities associated with recovering the outage, a recovery process is initiated to reconfigure, at the cloud platform, communication flows associated with the determined one or more entities.

[0166] Example 2. The method of example 1, wherein each entity of the plurality of entities is defined as one of:

[0167] Regional segmentation of cloud platforms;

[0168] cloud components running at the segments of the regions of the cloud platform;

[0169] A region in multiple availability zones of a cloud platform; or

[0170] Load balancers defined for multiple regions of the cloud platform.

[0171] Example 3. A method according to any of the preceding examples, wherein selection of a flag indicates modification of a health state of an entity mapped to the flag to a dangerous state, and wherein determining one or more entities associated with recovery from a plurality of entities defined for a cloud platform comprises: determining one or more entities based on an evaluation of the dangerous state of the selected flag mapped to the entity at an orchestrator component running for the cloud platform to manage communication flows.

[0172] Example 4. A method according to any of the preceding examples, wherein the multiple entities include different cloud component types, the different cloud component types include applications, services, and databases, and wherein each type of cloud component is associated with an instance of the corresponding component type running in each corresponding region in multiple availability zones.

[0173] Example 5. A method according to any of the preceding examples, wherein selection of the flag is received in response to determining an outage in one or more specific segments from a first region in the cloud platform, and wherein one or more of the determined entities are segments from another region of the cloud platform, wherein the other region is a second region in the cloud platform.

[0174] Example 6. The method of any of the preceding examples, wherein starting the recovery process comprises:

[0175] determining a new processing flow to replace a previous processing flow including the entity mapped to the selected indicia, wherein the new processing flow excludes the entity mapped to the selected indicia and replaces the entity with a corresponding entity at another region of the cloud platform, wherein the other region is defined for recovery that reconfigures communications of the entity mapped to the selected indicia; and

[0176] Services provided by an entity mapped to the selected flag are disabled, and requests to be received by the entity are redirected to a corresponding entity at another region.

[0177] Example 7. The method of any of the preceding examples, wherein starting the recovery process comprises:

[0178] A new processing flow is determined to replace a previous processing flow that includes an entity mapped to the selected marker, wherein the new processing flow excludes the entity mapped to the selected marker and includes a processing flow that includes only instances of the mapped entity executed at another area of ​​the cloud platform, the other area being defined for recovery of the entity mapped to the selected marker.

[0179] Example 8. A method according to any of the preceding examples, comprising:

[0180] Flags are configured at the cloud platform, wherein each flag, when triggered, generates an event associated with an outage at one or more entities at the cloud platform.

[0181] Example 9. A method according to any of the preceding examples, wherein multiple entities defined at the cloud platform are configured to monitor events associated with a trigger flag from a set of flags and initiate a recovery process corresponding to the entity associated with the trigger flag.

[0182] Example 10. The method according to any of the preceding examples further includes: before receiving a selection of a flag mapped to an entity at the cloud platform, identifying an outage at the cloud platform based on monitoring data associated with a health status of the cloud platform input at a trained model, wherein the outage limits access to services provided by the entity mapped to the flag selected at the cloud platform.

[0183] Example 11. The method according to any of the preceding examples, further comprising:

[0184] receiving notification that an entity mapped to the selected indicia is operating successfully at the first region after the outage has been identified and resolved; and

[0185] In response to the received notification, a recovery process is initiated for reconfiguring the communication flow and defining the flow through the entity operating at the first region of the cloud platform.

[0186] Example 12. A system comprising:

[0187] one or more processors; and

[0188] One or more computer readable memories, coupled to one or more processors and having instructions stored thereon, the instructions executable by the one or more processors to perform the method according to any one of Examples 1 to 11.

[0189] Example 13. A non-transitory computer-readable medium coupled to one or more processors and having instructions stored thereon, which, when executed by the one or more processors, cause the one or more processors to perform the method of any one of Examples 1 to 11.

[0190] Mechanism for implementing reconfiguration of cloud components across multiple availability zones

[0191] Example 1. A computer-implemented method comprising:

[0192] receiving a selection of a flag from a set of flags defined at a cloud platform including a plurality of availability zones, wherein the selection of the flag is received to trigger execution of a recovery of an entity running at a first zone of the cloud platform and mapped to the flag;

[0193] Determine the type of entity;

[0194] In response to determining the type of the entity, activating a load balancer monitor or a central service to generate a corresponding execution plan for recovery; and

[0195] The generated execution plan is executed by the load balancer monitor or the central service to reconfigure, at the cloud platform, communication flows associated with the entity for which the recovery execution was triggered.

[0196] Example 2. The method of example 1, wherein a flag is selected to identify an outage at the cloud platform, and wherein each flag in a set of flags is mapped to an entity from entities defined for the cloud platform.

[0197] Example 3. The method according to any one of Examples 1 or 2, further comprising:

[0198] configuring a first set of entities of the first type for recovery execution at a central service to manage outages associated with the entities of the first type at a cloud platform comprising a plurality of availability zones; and

[0199] A second set of entities of the second type are registered for recovery execution at the load balancer at the cloud platform.

[0200] Example 4. The method according to Example 3, further comprising: when determining that the type of the entity is the first type:

[0201] sending, by the central service, an instruction to a load balancer defined for the cloud platform to restrict network access to an entity at the first region;

[0202] reconfiguring a previously defined communication flow towards an entity to a corresponding entity at a second region of the cloud platform; and

[0203] Execution of an entity at a first region of the cloud platform is terminated, and execution of a corresponding entity at a second region of the cloud platform is activated.

[0204] Example 5. The method according to Example 3, comprising:

[0205] configuring an entity of the second type to obtain health status information about a trigger flag at the cloud platform; and

[0206] When determining that the type of the entity is the second type:

[0207] creating, by the load balancer of the entity, a health status load balancer monitor for the entity; registering, at the load balancer monitor, a second entity corresponding to the entity and running at a second region on the cloud platform;

[0208] Gets information about the entity's health status from the entity.

[0209] Example 6. A method according to any of Examples 1-5, wherein each entity is defined as one of: a regional segment of the cloud platform, a cloud component running at a segment of a region of the cloud platform, a region in multiple availability zones of the cloud platform, or a load balancer defined for multiple regions of the cloud platform.

[0210] Example 7. A method according to any one of Examples 1-6, wherein the entity includes different cloud component types, the different cloud component types include applications, services, and databases, wherein each type of cloud component is associated with an instance of the entity type, and the instances of the entity type run in corresponding regions in multiple availability zones.

[0211] Example 8. A system comprising:

[0212] one or more processors; and

[0213] One or more computer readable memories, coupled to one or more processors and having instructions stored thereon, the instructions executable by the one or more processors to perform the method of any one of Examples 1 to 7.

[0214] Example 9. A non-transitory computer-readable medium coupled to one or more processors and having instructions stored thereon, which, when executed by the one or more processors, cause the one or more processors to perform the method of any one of Examples 1 to 7.

[0215] Red Button Agent based recovery process

[0216] Example 1. A computer-implemented method comprising:

[0217] performing a request from a red button agent to a red button service to obtain a red flag status associated with an outage of a component defined for a cloud platform, wherein the red button agent is installed at a first cloud component instance, the first cloud component instance running at a first region of a cloud platform comprising a plurality of availability zones;

[0218] In response to receiving a red flag status of the first cloud component instance from the red button service, determining that the first cloud component instance is associated with an outage; and

[0219] Executing a recovery process of the first cloud component instance, wherein executing the recovery process includes:

[0220] initiating termination of a cloud component process running on the first cloud component instance; and

[0221] Configuring to send requests directed to the first cloud component instance to a second cloud component instance running at a second region, wherein the second region is a healthy region not associated with the outage.

[0222] Example 2. The method of Example 1, wherein the first cloud component instance and the second cloud component instance are instances of the same cloud component running at different regions of the cloud platform.

[0223] Example 3. The method according to Example 1 or Example 2, comprising:

[0224] In response to determining that the outage is over, initiating a first cloud component instance to start at the first region to perform a recovery operation.

[0225] Example 4. A method according to any of the preceding examples, wherein the first cloud component instance is configured to execute the cloud component process as a process flow that provides services to other instances running on the cloud platform and / or outside the cloud platform.

[0226] Example 5. A method according to any of the preceding examples, wherein initiating termination includes executing a process for stopping a cloud component process running on the first cloud component instance based on executing a script.

[0227] Example 6. The method of any of the preceding examples, wherein initiating termination comprises:

[0228] An instruction to stop is sent to the cloud component process, wherein the instruction is sent to a predefined endpoint of the first cloud component instance.

[0229] Example 7. A method according to any of the preceding examples, comprising:

[0230] Configure the red button agent to track the health indicator of the first cloud component instance to an external monitoring tool.

[0231] Example 8. The method of any of the preceding examples, wherein performing a recovery process of the first cloud component instance comprises:

[0232] As part of a recovery process of the first cloud component instance, a file including instructions for execution is read.

[0233] Example 9. A method according to any of the preceding examples, comprising:

[0234] Execution of the first cloud component instance at the first region is terminated by the red button agent.

[0235] Example 10. A method according to any of the preceding examples, comprising:

[0236] In response to determining, by a node monitor running at the load balancer, that the first cloud component instance is associated with an outage, the node monitor is configured for the first cloud component instance of the cloud platform to:

[0237] A previously defined communication flow toward the first cloud component instance is reconfigured so that a second cloud component instance at a second region of the cloud platform processes requests directed to the first cloud component.

[0238] Example 11. The method of example 10, wherein, when the first cloud component instance is operating in an active-passive state of the running instance of the first cloud component, reconfiguring the previously defined communication flow includes redirecting a request directed toward the first cloud component instance to the second cloud component instance, wherein the method includes:

[0239] Based on determining that the first cloud component instance is terminated, activating execution of a second cloud component instance at a second region of the cloud platform.

[0240] Example 12. The method of Example 10, wherein the red button agent registers the first cloud component instance at the node monitor, and wherein the node monitor is configured to monitor the first cloud component instance at the cloud platform.

[0241] Example 13. A method according to any of the preceding examples, wherein the red button agent is running in the same runtime infrastructure as the first cloud component instance.

[0242] Example 14. A computer-implemented method comprising:

[0243] configuring a red button agent at a first entity operating at a first region of a cloud platform comprising a plurality of availability zones;

[0244] configuring a monitor for evaluating a health status of the first entity, wherein the monitor is configured to perform a check of the health status by communicating with a health check endpoint provided by the first entity, and wherein the monitor is configured to trigger recovery execution based on detecting a downtime based on the evaluation of the communication with the health check endpoint;

[0245] Based on the monitor determining a health status of the first entity based on communication with a health check endpoint or based on selection of a flag at a red button service to notify a red button agent, determining a downtime associated with the first entity; and

[0246] Triggering the resumption of execution, which includes:

[0247] stopping a process running at the first region by the first entity; and

[0248] A communication flow associated with the first entity for which the resumption execution was triggered is reconfigured at the cloud platform.

[0249] Example 15. A system comprising:

[0250] one or more processors; and

[0251] One or more computer readable memories, coupled to one or more processors and having instructions stored thereon, the instructions executable by the one or more processors to perform the method of any of Examples 1 to 14.

[0252] Example 16. A non-transitory computer-readable medium coupled to one or more processors and having instructions stored thereon, which, when executed by the one or more processors, cause the one or more processors to perform the method of any one of Examples 1 to 14.

[0253] Recovering cloud components in a multi-availability zone cloud environment

[0254] Example 1. A computer-implemented method comprising:

[0255] installing a red button agent at a first cloud component instance of a first cloud component running at a first region of a cloud platform comprising a plurality of availability zones;

[0256] performing a request from the red button proxy to the red button service to obtain the status of the red flag selected for the cloud platform;

[0257] In response to receiving a status of a red flag associated with the first cloud component instance, determining that the first cloud component instance is associated with an outage; and

[0258] Executing a recovery process of the first cloud component instance, wherein the recovery process includes:

[0259] initiating termination of a cloud component process running on the first cloud component instance; and

[0260] Reconfigure communication flows for the first cloud component to a second cloud component instance running at a second region, wherein the second region is a healthy region not associated with the outage, and wherein the first cloud component instance and the second cloud component instance are instances of the same first cloud component running at different regions of the cloud platform.

[0261] Example 2. The method according to Example 1, comprising:

[0262] In response to determining that the outage is over, initiating a first cloud component instance to start at the first region to perform a recovery operation.

[0263] Example 3. A method according to any of the preceding examples, wherein the first cloud component instance is configured to execute the cloud component process as a process flow that provides services to other instances running on the cloud platform and / or outside the cloud platform.

[0264] Example 4. A method according to any of the preceding examples, wherein initiating termination includes executing a process for stopping a cloud component process running on the first cloud component instance based on executing a script.

[0265] Example 5. The method of any of the preceding examples, wherein initiating termination comprises:

[0266] An instruction to stop is sent to the cloud component process, wherein the instruction is sent to a predefined endpoint of the first cloud component instance.

[0267] Example 6. A method according to any of the preceding examples, comprising:

[0268] The red button agent is configured to track the health indicator of the first cloud component instance to an external monitoring tool.

[0269] Example 7. The method of any of the preceding examples, wherein performing a recovery process of the first cloud component instance comprises:

[0270] As part of a recovery process for the first cloud component instance, a file including instructions for execution is read.

[0271] Example 8. A method according to any of the preceding examples, comprising:

[0272] In response to determining, by a node monitor running at the load balancer monitor, that the first cloud component instance is associated with an outage, the node monitor is configured for the first cloud component instance configured for the cloud platform to:

[0273] A previously defined communication flow toward the first cloud component instance is reconfigured to a second cloud component instance at a second region of the cloud platform.

[0274] Example 9. The method of Example 8, wherein the red button agent registers the first cloud component instance at the node monitor, and wherein the node monitor is configured to monitor the first cloud component instance at the cloud platform.

[0275] Example 10. A system comprising:

[0276] one or more processors; and

[0277] One or more computer readable memories, coupled to one or more processors and having instructions stored thereon, the instructions executable by the one or more processors to perform the method according to any one of Examples 1 to 9.

[0278] Example 11. A non-transitory computer readable medium coupled to one or more processors and having instructions stored thereon, the instructions, when executed by the one or more processors, causing the one or more processors to perform a method according to any one of Examples 1 to 9.

[0279] Reconfigure communications for entities running on multi-availability zone cloud platforms

[0280] Example 1. A computer-implemented method comprising:

[0281] receiving a selection of a flag from a set of flags defined at a cloud platform comprising a plurality of availability zones, wherein the flag is selected to identify an outage at a first zone of the cloud platform;

[0282] determining an instance of an entity running at a first region of the cloud platform;

[0283] determining a state mode of a running instance of an entity at a cloud platform;

[0284] In response to determining the state pattern, determining rules for performing an instance recovery process of an entity running at the first region; and

[0285] In response to determining the rule, a recovery process determined based on the state pattern of the running instance of the entity is executed to reconfigure subsequent communications directed to the entity to another one or more instances running at one or more other regions of the cloud platform.

[0286] Example 2. The method of Example 1, wherein selection of a flag is received to trigger a recovery process of an instance running at a first region of the cloud platform and mapped to the flag.

[0287] Example 3. The method of example 1 or example 2, wherein a plurality of entities are running on a cloud platform, wherein each entity is running with one or more instances distributed across one or more availability zones, and wherein each entity from the plurality of entities is defined as one of the following:

[0288] Regional segmentation of cloud platforms;

[0289] cloud components running at the segments of the regions of the cloud platform;

[0290] A region in multiple availability zones of a cloud platform; or

[0291] Load balancers defined for multiple regions of the cloud platform.

[0292] Example 4. A method according to Example 3, wherein the multiple entities include different cloud component types, the different cloud component types include applications, services, and databases, and wherein each type of cloud component is associated with an instance of the corresponding component type running in a corresponding region in multiple availability zones.

[0293] Example 5. A method according to Example 3, wherein each cloud component type of an entity is associated with a corresponding state mode of a running instance of the entity of that type, and wherein the state mode of the cloud component type is an active-active mode or an active-passive mode.

[0294] Example 6. A method according to any of the preceding examples, wherein, when the state mode of a running instance of an entity is determined to be an active-passive mode, the rules determined for performing a recovery process include rules for reconfiguring subsequent communications by performing a failover process to redirect subsequent communications directed to the entity to an instance of the entity running in another area of ​​the cloud platform that is not associated with the selected flag.

[0295] Example 7. A method according to any of the preceding examples, wherein, when the state mode of the running instances of the entity is determined to be an active-active state, the rules determined for performing the recovery process include rules for reconfiguring subsequent communications by sending requests for services from the entity only to another one or more instances running in one or more regions of the cloud platform, wherein the another one or more regions are not associated with the selected flag.

[0296] Example 8. A system comprising:

[0297] one or more processors; and

[0298] One or more computer readable memories, coupled to one or more processors and having instructions stored thereon, the instructions executable by the one or more processors to perform the method of any one of Examples 1 to 7.

[0299] Example 9. A non-transitory computer-readable medium coupled to one or more processors and having instructions stored thereon, the instructions, when executed by the one or more processors, causing the one or more processors to perform a method according to any one of Examples 1 to 7.

Claims

1. A computer-implemented method, include: performing a request from a red button agent to a red button service to obtain a red flag status associated with an outage of a component defined for a cloud platform, wherein the red button agent is installed at a first cloud component instance, the first cloud component instance running at a first region of a cloud platform comprising a plurality of availability zones; In response to receiving a red flag status of the first cloud component instance from the red button service, determining that the first cloud component instance is associated with an outage; and Executing a recovery process of the first cloud component instance, wherein executing the recovery process includes: initiating termination of a cloud component process running on the first cloud component instance; and The method is configured to send a request directed to the first cloud component instance to a second cloud component instance running at a second region, wherein the second region is a healthy region not associated with the outage.

2. The method according to claim 1, in, The first cloud component instance and the second cloud component instance are instances of the same cloud component running at different regions of the cloud platform.

3. The method according to claim 1 or 2, wherein include: In response to determining that the outage is over, initiating a first cloud component instance to start at the first region to perform a recovery operation.

4. The method according to any one of the preceding claims, in, The first cloud component instance is configured to execute the cloud component process as a process flow that provides a service to other instances running on the cloud platform and / or external to the cloud platform.

5. The method according to any one of the preceding claims, in, Initiating termination includes performing a process for stopping a cloud component process running on the first cloud component instance based on executing the script.

6. The method according to any one of the preceding claims, in, Initiating termination includes: An instruction is sent to the cloud component process to stop, wherein the instruction is sent to a predefined endpoint of the first cloud component instance.

7. A method according to any one of the preceding claims, wherein include: The red button agent is configured to track the health indicator of the first cloud component instance to an external monitoring tool.

8. The method according to any one of the preceding claims, in, The recovery process of performing the first cloud component instance includes: As part of a recovery process of the first cloud component instance, a file including instructions for execution is read.

9. A method according to any one of the preceding claims, wherein include: Execution of the first cloud component instance at the first region is terminated by the red button agent.

10. A method according to any one of the preceding claims, wherein include: In response to determining, by a node monitor running at the load balancer, that the first cloud component instance is associated with an outage, the node monitor is configured for the first cloud component instance of the cloud platform to: A previously defined communication flow toward the first cloud component instance is reconfigured so that a second cloud component instance at a second region of the cloud platform processes requests directed to the first cloud component instance.

11. The method according to claim 10, in, When the first cloud component instance is operating in an active-passive state of the running instance of the first cloud component, reconfiguring the previously defined communication flow includes redirecting a request directed toward the first cloud component instance to the second cloud component instance, wherein the method includes: Based on determining that the first cloud component instance is terminated, activating execution of a second cloud component instance at a second region of the cloud platform.

12. The method according to claim 10, in, The red button agent registers the first cloud component instance at the node monitor, and wherein the node monitor is configured to monitor the first cloud component instance at the cloud platform.

13. The method according to any one of the preceding claims, in, The red button agent is running in the same runtime infrastructure as the first cloud component instance.

14. A system, include: one or more processors; as well as One or more computer readable memories, coupled to the one or more processors and having instructions stored thereon, the instructions being executable by the one or more processors to perform the method according to any one of claims 1 to 13.

15. A non-transitory computer readable medium coupled to one or more processors and having instructions stored thereon, the instructions, when executed by the one or more processors, causing the one or more processors to perform the method according to any one of claims 1 to 13.