Systems and methods for managing failover regions
By providing a system for managing network-based failover services, the problems of failover management complexity and lack of customer control in the prior art are solved, and simplified failover management and high availability applications are achieved.
Patent Information
- Application Number
- CN202180020717.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-27
- Filing Date
- 2021-03-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-03-25
AI Technical Summary
The existing mechanisms for managing failover are too complex, increasing customer design efforts and lacking the characteristics that provide customer visibility and control over the mechanism.
Provides a system for managing network-based failover services that coordinates the design and execution of failover workflows, maintains data integrity, and enables customers to manually or automatically repair unavailable failover areas.
Simplifies the failover management process, provides better customer control and visibility, ensuring high availability and data integrity of applications.
Smart Images

Figure CN115698954B_ABST
Abstract
Description
Background Art
[0001] Generally, network-based computing is a method of providing access to information technology resources through services such as web services, where the hardware and / or software used to support these services can be dynamically scaled to meet the service demands at any given time. In network-based computing, elasticity refers to the network-delivered computing resources that can be scaled up and down by a network service provider to adapt to the changing requirements of users. For example, the elasticity of these resources may be in terms of processing power, storage, bandwidth, etc. Elastic computing resources can be automatically delivered on demand, thus dynamically adapting to changes in resource requirements on or within a given user system. For example, a client can use a web service to host a large online streaming service and use elastic resources to set it up such that the number of web servers streaming content to users expands during peak viewing hours to meet the bandwidth requirements and then shrinks when the system usage is low.
[0002] Clients typically lease, rent, or otherwise pay for access to the elastic resources accessed through web services and thus do not have to purchase and maintain the hardware and / or software that provides access to these resources. This provides several benefits, including allowing users to quickly reconfigure their available computing resources in response to the changing needs of the enterprise and enabling the network service provider to automatically scale the provided computing service resources based on usage, traffic, or other operational requirements. Compared with the relatively static infrastructure of an on-premises computing environment, this dynamic nature of network service computing services requires a system architecture that can reliably reallocate its hardware according to the changing needs of its customer base and the demand for network-based computing services.
[0003] In network-based computing, the locations where applications can be hosted and / or partitioned can be described as regions and / or availability zones. Each region includes a geographical area separate from other regions and includes multiple isolated availability zones. Each region can be isolated from all other regions in a network-based computing system. An availability zone is an isolated location within a region. Each region consists of several availability zones, and each availability zone belongs to a single region. Moreover, each availability zone is isolated, but the availability zones within a particular region are connected by low-latency links. When an application is distributed across multiple availability zones, instances may be launched in different availability zones so that your application can continue to run if one of the instances fails (e.g., by allowing another instance in another availability zone to handle requests for the application). Brief Description of the Drawings
[0004] Various features will now be described with reference to the following drawings. In all the drawings, reference numerals may be reused to indicate corresponding relationships between the elements being referred to. The drawings are provided to illustrate examples described herein and are not intended to limit the scope of the disclosure.
[0005] Figure 1 A schematic diagram of a network service provider in which various embodiments according to the present disclosure may be implemented is depicted.
[0006] Figure 2 An exemplary workflow depicting interactions for managing the availability of a failover service is depicted.
[0007] Figure 3 An exemplary client interface allowing a client to select how to manage a failover service is depicted.
[0008] Figure 4 An exemplary schematic diagram depicting the implementation of a regional information processing service according to illustrative aspects of the present disclosure is depicted.
[0009] Figure 5 An exemplary schematic diagram depicting the implementation of a failover management service according to illustrative aspects of the present disclosure is depicted.
[0010] Figure 6 A flowchart showing a failover management routine implemented by a failover service according to illustrative aspects of the present disclosure. Detailed Description
[0011] In the following description, various examples will be described. For purposes of explanation, specific configurations and details will be set forth in order to provide a thorough understanding of the examples. However, it will be apparent to those skilled in the art that the examples may be practiced without specific details. Additionally, well-known features may be omitted or simplified in order not to obscure the described examples.
[0012] Generally, aspects of the present disclosure relate to the management of network-based failover services in network-based computing systems. In network-based computing systems, a customer may design an application that is partitioned across various isolated computing systems (referred to as "availability zones" or regions). When so partitioned, each of the various zones or regions hosts a partition of the application that is equivalent to other partitions of the application.
[0013] In the event of a failure in one of the described zones or regions, partitions of the application hosted by other zones or regions provide redundancy or failover, allowing the application to continue operating based on resources in the other zones or regions. More specifically, aspects of the present disclosure relate to providing a network-based failover service that enables predictable, controlled, and reliable failover. The network-based failover service helps manage one or more failover zones to be available or ready when the current or designated zone fails. The network-based failover service can identify target failover zones and utilize processing rules to determine which target failover zones can be characterized as "available for" failover based on information such as capacity, readiness, etc. Further, for target failover zones that have been characterized as "unavailable" or not otherwise characterized as "available", the network-based failover service can further implement a repair process to modify or supplement. The repair process can be illustratively implemented manually or automatically and can be customized to allow one or more failover zones to achieve an available characterization. An application can be considered highly available when such a failure of an application partition does not impede the operation of the application in other partitions or have a negative impact on the data integrity of data associated with the application (i.e., when the failover workflow ensures that network requests, etc. are properly redirected or directed to backup partitions), since the partitions keep the application generally available.
[0014] Existing mechanisms for managing failover are overly complex, significantly increasing the design effort required by customers, and lack features that provide customers with visibility and control over the mechanisms. The present disclosure addresses these types of issues by providing a system for managing a network-based failover service (sometimes referred to as a "failover service") that better coordinates failover workflow design and execution while maintaining the data integrity of data associated with application partitions to achieve a highly available application. The system for managing the failover service described herein supports a wide range of failover use cases. For example, the failover service supports use cases where a primary application partition runs at a customer (or other) site and disaster recovery (DR) is set up in the cloud, a primary application partition runs in the cloud and DR is set up locally, and a primary application partition and DR are set up in the cloud or locally.
[0015] The network-based failover service of the present disclosure improves on the errors of existing mechanisms in a number of ways. The system for managing the failover service of the present disclosure enables a customer to manually repair an unavailable failover region such that it meets the requirements to be considered available in the event of a failover. As described above, in some embodiments, the network-based failover service automatically repairs a failover based on certain readiness requirements set by a client. The system for managing the failover service notifies the client of available failover regions, which may be specifically identified or characterized according to custom rules provided by the user. In some embodiments, the list of rules may be at least partially based on status information derived from a primary or default region. As an illustrative example, one rule in the list of rules may correspond to matching or exceeding the number of partitions hosted in the primary region. Accordingly, the network-based system will identify target failover regions that meet the established partition threshold, identify target failover regions that do not meet the established partition threshold, and repair one or more target threshold regions by increasing the number of partitions. Additional details regarding each of these benefits are provided below.
[0016] These and other aspects of the present disclosure will now be described with respect to certain examples and embodiments, which are intended to be illustrative but not limiting of the present disclosure. While, for purposes of illustration, the examples and embodiments described herein will focus on specific computations and algorithms, those skilled in the art will appreciate that the examples are merely illustrative and are not intended to be limiting.
[0017] Figure 1 An exemplary computing environment 100 is depicted, in which a network service provider 110 provides network-based services to a client device 102 via a network 104. As used herein, the network service provider 110 implements a network-based service 110 (sometimes referred to simply as the “network-based service 110” or “service 110”) and refers to a large shared pool of network-accessible computing resources (such as computing, storage, or networking resources, applications, or services) that may be virtualized or bare metal. The network service provider 110 may provide convenient on-demand network access to a shared pool of configurable computing resources that can be programmatically provisioned and released in response to customer commands. These resources can be dynamically provisioned and reconfigured to adjust to variable loads. The concepts of “cloud computing” or “network-based computing” can thus be considered to be both the applications delivered as services over the network 104 and the hardware and software in the network service provider 110 that provide those services.
[0018] As Figure 1As shown, network service provider 110 is illustratively divided into a number of zones 112A through 112D. Each zone 112 may be geographically isolated from other zones 112. For example, the geographical location of zone 112A may be on the east coast of the United States, the geographical location of zone 112B may be on the west coast of the United States, the geographical location of zone 112C may be in Europe, the geographical location of zone 112D may be in Asia, and so on. Although four zones 112 are shown in Figure 1 , network service provider 110 may include any number of zones. Each zone 112 illustratively communicates via a network, which may be a private network of system 110 (e.g., a private circuit, leased line, etc.) or a public network (e.g., the Internet).
[0019] In Figure 1 , each zone 112 is also shown as being divided into a number of subzones 120 (subzones 120A through 120L across all zones 112), which may also be referred to as availability zones or availability regions. Each subzone 120 illustratively represents a computing system that is isolated from the systems of other subzones 120 such that the likelihood of a large-scale event, such as a natural or man-made disaster, affecting the operation of all (or any two) subzones 120 in a zone is reduced. For example, the computing resources of each subzone 120 may be physically isolated by being distributed at selected distances throughout a zone 112 to reduce the likelihood of a large-scale event affecting the performance of all (or any two) subzones 120. Additionally, the computing resources of each subzone 120 may be associated with independent power and are thus electrically isolated from the resources of other subzones 120 (although the resources may still communicate with each other via a network, which may involve transmitting electrical signals for communication rather than power), independent cooling systems, independent in-zone networking resources, etc. In some instances, subzones 120 may be further isolated by restricting the operation of computing resources between subzones 120. For example, virtual machine instances in a subzone 120 may be limited to using the storage resources, processing resources, and communication links in that subzone 120. Constraining cloud or network-based computing operations between subzones may limit the "blast radius" of any failure within a single subzone 120, thereby reducing the likelihood that such a failure will inhibit the operation of other subzones 120. Illustratively, the services provided by network service provider 110 may generally be replicated within a subzone 120 such that a client device 102 may (if it so chooses) fully (or almost fully) utilize network service provider 110 by interacting with a single subzone 120.
[0020] As Figure 1As shown, each zone 120 communicates with other zones 120 via a communication link. Preferably, the communication links between zones 120 represent a high-speed dedicated network. For example, zones 120 may be interconnected via dedicated fiber optic lines (or other communication links). In one embodiment, the communication links between zones 120 are fully or partially dedicated to communication between the zones and are separate from other communication links of the zones. For example, each zone 120 may have one or more fiber optic connections to each other zone, as well as one or more separate connections to other regions 112 and / or network 104.
[0021] Each zone 120 within each region 112 is illustratively connected to network 104. Network 104 may include any suitable network, including an intranet, the Internet, a cellular network, a local area network, or any other such network or a combination thereof. In the illustrated embodiment, network 104 is the Internet. Protocols and components for communicating via the Internet or any of the other aforementioned types of communication networks are known to those skilled in the art of computer communication and thus need not be described in more detail herein. Although system 110 is shown in Figure 1 as having a single connection to network 104, multiple connections may exist in various embodiments. For example, each zone 120 may have one or more connections to network 104 that are different from those of other zones 120 (e.g., one or more links to an Internet exchange point that interconnects different autonomous systems on the Internet).
[0022] Each of regions 112A through 112D includes endpoints 125A through 125D, respectively. Endpoints 125A through 125D may include computing devices or systems through which a customer's application may access network-based services 110. Information provided to one of endpoints 125 may be propagated to all other endpoints 125. Each region 112 may include more than one endpoint 125, or each region 112 may not even include an endpoint 125.
[0023] Continuing to refer to Figure 1, the network service provider 110 also includes a regional information processing service 130 and a failover region management service 140. As will be described in more detail below, the regional information processing service 130 can be configured to determine a set of target regions that can be designated as the primary region and one or more target failover regions for an individual customer or a set of customers. For example, the regional information processing service 130 can process customer-specific criteria to determine which region will be designated as the primary region. The regional information processing service 130 can further select target failover regions based on the selection criteria described herein. The failover region management service 140 can be configured to receive a set of target failover regions and characterize the availability of at least some portions of one or more of the target failover regions based on the application of one or more processing rules. Illustratively, each processing rule can correspond to the identification of a parameter and one or more thresholds associated with the identified parameter. These parameters correspond to resource configurations or performance metrics that define the ability of a target region to be considered an available failover region. The processing rules can be configured by the customer, the network service provider, or a third party. Additionally, the processing rules can be derived in part based on the attributes or parameters of the designated primary region (e.g., matching the current attributes of the designated primary region). The failover region management service 140 can also implement a processing engine that can implement processing in response to a determined list of available or unavailable failover regions. The processing engine can illustratively implement one or more repair processes that can attempt to modify or supplement target regions that were not determined to be available based on the previous application of the processing rules. The processing engine can also implement a readiness process that can be used to determine whether previously determined available failover regions are operationally ready or operable to function in a failover capacity. The results of the failover process (e.g., repair or readiness processing) can be used to modify or update the list of available failover regions.
[0024] The client computing device 102 can include any network-equipped computing device, such as a desktop computer, laptop computer, smart phone, tablet computer, e-reader, gaming console, etc. A user can access the network service provider 110 via the network 104 to view or manage their data and computing resources, as well as use websites and / or applications hosted by the network service provider 110. For example, a user can access an application having partitions hosted by zone 120A (e.g., the primary partition) in region 112A and zone 120L (e.g., the secondary partition) in region 112D.
[0025] According to an embodiment of the present disclosure, an application having partitions hosted in different zones may be able to withstand a failure in one of zone 120 or region 112, in which one of the partitions is operating. For example, if the primary partition hosted in zone 120A experiences a failure, any requests that would typically be handled by the primary partition in zone 120A may instead be routed to and handled by a secondary partition running in zone 120L. Such a failure may result in a failover scenario where the operation of the primary partition is transferred to the secondary partition for handling. The failover scenario may involve manual actions by the customer associated with the application to request communication routing from the primary partition to the secondary partition and the like. However, embodiments of the present disclosure may also provide a managed failover service with high availability for an application having partitions hosted in different zones, which enables the customer's application to withstand zone or region failures with reduced or minimal interaction from the customer during a failover scenario while maintaining data integrity during such failures and failovers.
[0026] Figure 2 An exemplary workflow 200 depicting the interaction of a regional information processing service 130, a failover region management service 140, and a client device 102 according to an illustrative embodiment to determine and manage failover region availability is shown. As Figure 2 shown, at (1), the regional information processing service 130 determines a primary region and a set of target failover regions. The regional information processing service 130 may include components for determining a primary region, a list of target failover regions, and a list of processing rules. In one embodiment, the regional information processing service 130 may generate or obtain a list of regions based on geographic or network proximity (e.g., regions within a defined radius). For example, the regional information processing service 130 may be configured to provide a list of regions located within 500 miles of a specified location or set of locations. In some implementations, the regional information processing service 130 may be configured to provide a list of regions located within the same country as the user. In some implementations, the regional information processing service 130 periodically updates the list of rules, the list of failover regions, and the designation of the primary region. For example, the regional information processing service 130 may update once per hour. In some implementations, the regional information processing service 130 may update when indicated by the client. In some implementations, the regional information processing service 130 may update periodically and when indicated by the client.
[0027] In another embodiment, the regional information processing service 130 may also determine or identify a set of primary regions or target regions based on the application of selection criteria related to the attributes or characteristics of the regions. For example, the regional information processing service 130 may identify or select the region hosting the largest number of partitions as the primary region. The regional information processing service 130 may also identify one or more additional regions as having the smallest number of partitions to be used as potential failover regions. Illustratively, the minimum number of partitions for selecting a partition as a failover region need not correspond to the desired number of partitions, since the failover region management service 140 may repair the target region to increase the number of partitions. In other examples, the regional information processing service 130 may also consider network capacity when selecting a set of target failover regions based on measured network traffic or executed instructions / processes, measured load or utilization availability, error rate, attributed financial cost, infrastructure, workload location, etc. Illustratively, the client 102 may select any parameter related to determining a set of target regions. The network service provider 110 may also specify one or more parameters, such as a list of minimum requirements. For example, the network service provider 110 may specify minimum requirements in terms of capacity and measured load for selecting a primary region or a target failover region.
[0028] At (2), the regional information processing service 130 transmits a list of regions to the failover region management service 140. At (3), the regional information processing service 130 transmits a set of availability processing rules that allow the failover region management service 140 to determine or characterize the availability of a set of target failover regions. As described above, individual processing rules may include the identification of one or more parameters (or combinations of parameters) and corresponding one or more thresholds characterizing the availability of individual target regions. Illustratively, the same parameters and thresholds may determine whether a region is available or unavailable (e.g., a region that matches or exceeds the threshold). In other embodiments, the processing rules may include a first parameter threshold for determining availability and a second parameter threshold for determining unavailability. In this embodiment, different parameters may be used in combination with the region selection criteria previously processed by the regional information processing service 130 or the repair process implemented by the failover region management service 140. For example, if the regional information processing service 130 has not filtered out any regions, the second threshold parameter may be set to filter out any regions that cannot be repaired by the failover region management service 140.
[0029] At (4), the failover region management service 140 determines the number of available failover regions and at (5) transmits these regions to the client 102. As described above, the failover region management service 140 may apply the processing rules to the set of target failover regions to identify a set of available failover regions, a set of unavailable regions, or a combination or subset thereof.
[0030] At (6), the failover zone management service 140 can implement one or more additional processes in response to the availability or unavailability of the identified set of zones. Such response processes can include self-healing, where the failover zone management service 140 automatically attempts to configure one or more zones that have been characterized as unavailable in a manner that allows the zones to subsequently be characterized as available. In some embodiments, self-healing can include fixing capacity issues of the failover zones. For example, self-healing can include increasing the capacity of a zone, where increasing the capacity causes the zones to be ready such that they are available for failover if an event occurs. In some embodiments, self-healing can include fixing the configuration of the failover zones. For example, self-healing can include changing the configuration of one or more zones such that they are available for failover. The automatic or self-healing can be restricted or configured by the failover zone management service 140 based on client programs / limitations (such as defining cost limitations or the degree of allowable change). In other embodiments, as described herein, the failover zone management service 140 can also perform a readiness check to verify that the target failover zones are currently running and are capable of being used as failover zones.
[0031] At (7), the failover zone management service 140 can wait for a client response from the client 102. A list of available failover zones and a list of unavailable failover zones can be provided to the client 102. An interface can be provided to the client 102 to select one or more of the unavailable failover zones to be repaired such that the one or more unavailable failover zones become one or more available failover zones. For example, as Figure 3 further shown herein, a client interface can be provided to the client 102 that details the available failover zones and the unavailable zones. In other embodiments, the client 102 can also specify priority information that helps determine which potential unavailable zone to repair.
[0032] Illustratively, at (8), the client 102 transmits the client response to the failover zone management service 140. The failover zone management service 140 can be configured to perform the specified repair corresponding to the client response. The client response can include any set of instructions related to the state of one or more zones. In some embodiments, the client response can provide one or more zones to be repaired such that the one or more zones meet each rule in a list of rules. In some embodiments, the client response can include a modification to the list of rules, where the client 102 provides one or more rules to be included in the list of rules.
[0033] At (9), the failover region management service 140 may transmit an updated list of available failover regions or other configuration information to the region information processing service 130. The updated failover region list may include updates based on successful repairs or pass / fail of readiness tests. The failover region management system 106 may be configured to update the list of available failover regions and provide this information to the region information system 102. The failover region management system 106 may also be configured to update the rules list based on the client response in (7). The failover region management 106 may then be configured to provide the updated rules list to the region information system 102. The region information system 102 may then store the updated list of available regions and the updated rules list.
[0034] Figure 3 An exemplary client interface 300 for managing failover services is depicted. The client interface 300 may enable customers whose applications are hosted by a network service provider 110 to create dependency trees and failover workflows for their applications. The dependency tree may map and track the upstream and downstream dependencies of the customer's application to determine the steps to take in a failover to ensure data integrity between application partitions and the continued availability of the application. Additionally, the failover service may map the upstream and / or downstream dependencies of sub-applications of the customer application. Based on the mapped partitions and dependencies, the failover service may coordinate partition or node failovers for any one of the individual applications in a sequential manner. In some embodiments, the dependencies may include other applications or services that provide data and requests.
[0035] In some embodiments, the interface 300 is also used to identify failover workflows to be triggered based on the failover status and / or other conditions. The dependency tree and workflow may be created when the customer designs and creates the application or after the application has been created and partitioned. Such dependency trees and failover workflows may enable the failover service to provide visibility into the specific dependencies of the application. For example, by enabling the customer to see the upstream and downstream dependencies of their application, the customer may better understand what steps or sequence of actions are required during a failover of an application partition or node to ensure the availability of the application and the data integrity of the associated data, and may accordingly generate a failover workflow. Thus, the customer may be able to more easily generate a workflow, including the steps or sequence of actions required in the event of a failover, compared to when the dependency tree is not available.
[0036] In some embodiments, such failover workflows can be triggered manually by a customer or automatically by a failover service based on the failover status of an application partition or node. By tracking application dependencies and corresponding workflows, the failover service enables a customer to coordinate failover procedures for an application in a secure, reliable, and predictable manner such that data integrity and application availability are maintained.
[0037] In some embodiments, a customer uses a failover service to model their applications and / or cells of their applications. As used herein, a cell can represent a partition, node, or any unit of an application that may be a point of failure or may experience a failure. A customer can use the model of the failover service to define the sequence of steps required during failover across one or more applications based on, for example, a dependency tree. For example, if a customer detects a failure in the primary partition of an application, the customer can trigger an auto-scaling step to scale the application in a secondary partition, and then the customer can trigger a traffic management service to redirect client traffic to the secondary partition. Such control enables a customer to manage a distributed multi-tier application in a controlled, reliable, and predictable manner. In some embodiments, the traffic management service can route traffic to the best application endpoint based on various parameters related to the performance of the application. In some embodiments, a customer can generate a workflow to include the actions identified above in the event of a failure such that the actions are automatically executed by the failover service.
[0038] Similarly, the failover service can provide such control to a customer to configure workflows (e.g., including traffic routing actions using a traffic management service and / or a Domain Name System (DNS) service) implemented based on state changes of an application partition or node. In some embodiments, a customer can also configure metadata with state changes for an application partition or node. For example, an application partition or node state change can trigger a failover or change in endpoints or traffic weights for each zone or region of a traffic management service and / or a DNS service (also referred to herein as a routing service), which may support automation of failover workflows and / or step sequences.
[0039] As described herein, the failover service for a customer application enables the customer to generate a failover workflow for the application that identifies one or more actions or steps to take if the primary partition of the application experiences a failure. Thus, as described above, the failover workflow can include steps to take to ensure the continued operation of the application and maintain data integrity through individual partition failures. For example, the workflow can include the identification of a secondary partition that serves as a backup (e.g., becomes the new primary partition) of the previous primary partition when the previous primary partition experiences a failure. The failover workflow can also define the state to which the primary partition transitions when it experiences a failure. Although this document refers to primary and secondary partitions, the failover service and failover workflow can equally apply to primary and secondary nodes.
[0040] The client interface 300 can include a first client interface 302 that is used to represent the current region being used by the client application. The first client interface 302 can include the name of the region that the client application is currently using. The first client interface 302 can also include the number of partitions being implemented in a certain region. The first client interface 302 can contain other information related to one or more regions that are being actively used by the client at a given moment.
[0041] The client interface 300 can include a second client interface 304 that is used to represent the failover regions available to the user. The second client interface 304 can provide the client with information related to the failover regions. For example, the second client interface 304 can provide the name of the region or zone, the location of the region or zone, and the endpoints. In addition, the second client interface 304 can be configured to provide information related to the failure rate, downtime, or any other factor related to the regions that can be used to select a region for failover.
[0042] The client interface 300 may include a third client interface 306 for representing client input, where the client may select one or more options to be executed by the client interface 300. The third client interface 306 may first include a designation of a main area. The third client interface 306 may select the area to be designated as the main area at least partially based on the area hosting the largest number of partitions related to the application. In some embodiments, the main area may be selected by the client according to other factors including the designation. For example, the application may be provided to the client to select the area to be selected as the main area. In certain configurations, the client may periodically update the main area. The third client interface 306 may include one or more areas as the designation of available failover areas. The available failover areas may correspond to one or more areas that meet each rule in the rule list. The available areas may also correspond to a list of areas that have been previously designated as available. The third client interface 306 may be configured to periodically update the list of available failover areas and the main area. For example, the third client interface 306 may be configured to update the available failover areas and the main area hourly. Additionally, the third client interface 306 may be configured to update the available failover areas and the main area based on input provided by the client. For example, the client may instruct the third client interface 306 to update the available failover areas based on the client pressing a refresh button located in the third client interface 306.
[0043] The third client interface 306 may include one or more areas as the designation of unavailable failover areas. The unavailable failover areas or zones may correspond to one or more areas or zones that do not meet at least one availability rule in the rule list. The unavailable failover areas may also correspond to a list of areas that have been previously designated as unavailable. The third client interface 306 may include information detailing why one or more areas are unavailable failover areas. The third client interface 306 may include a description of one or more unavailable failover areas. The third client interface 306 may include a description of the repair steps that may be taken to repair one or more unavailable failover areas. The third client interface 306 may be configured to periodically update the list of unavailable failover areas. For example, the third client interface 306 may be configured to update the unavailable failover areas hourly. Additionally, the third client interface 306 may be configured to update the unavailable failover areas based on input provided by the client. For example, the client may instruct the third client interface 306 to update the unavailable failover areas based on the client pressing a refresh button located in the third client interface 306.
[0044] The third client interface 306 may include an action list corresponding to each region of the client. Each region corresponding to the client may include one or more regions within the radius of the client. Each region corresponding to the client may include one or more regions that the client has pre-selected for possible failover. The action list may include a list of actions that the third client interface 306 may cause to be performed on the corresponding region. One possible action may be to make the region a primary region. The third client interface 306 may be instructed to mark the region as a primary region based on client input. In addition, a possible action may be to make a previously unavailable failover region an available failover region. For example, the third client interface 306 may detect that the primary region is hosting 15 partitions, and the first region can only host 10 partitions. The third client interface 306 may then determine that the first region is an unavailable failover region because it cannot meet the capacity requirements of the primary region. The third client interface 206 may make the first region an available failover region by increasing the capacity of the first region to 15 or more partitions according to the client's input. The third client interface 306 may include other options for client communication, including but not limited to a "cancel" button and an "accept all changes" button.
[0045] In some embodiments, the client interface 300 may include one or more other client interfaces for representing more information about the region hosting the client application. The client interface 300 may include a fourth client interface representing the currently hosted application. The fourth client interface may include information about the number of regions hosting each application. The fourth client interface may include information about the status of each application. The client interface 300 may include a fifth client interface representing one or more clusters associated with the client.
[0046] Figure 4 Depicted is a general architecture of a computing device configured to execute a zone information processing service 130 according to some embodiments. Figure 4 The general architecture of the regional information processing service 130 depicted in FIG. 1 includes an arrangement of computer hardware and software that can be used to implement various aspects of the present disclosure. The hardware can be implemented on a physical electronic device, as will be discussed in more detail below. The regional information processing service 130 may include Figure 4 Those more (or fewer) elements shown in . However, in order to provide an enabling disclosure, it is not necessary to show all of these generally conventional elements. In addition, Figure 4 The general architecture shown in can be used to implement Figure 1 One or more of the other components shown in .
[0047] As shown in the figure, the area information processing service 130 includes a processing unit 402, a network interface 404, a computer-readable medium drive 406, and an input / output device interface 408, all of which can communicate with each other via a communication bus. The network interface 404 can provide connectivity to one or more networks or computing systems. The processing unit 402 can thus receive information and instructions from other computing systems or services via the network. The processing unit 402 can also communicate with a memory 410 and also provide output information for an optional display via the input / output device interface 408. The input / output device interface 408 can also accept input from an optional input device (not shown).
[0048] The memory 410 can contain computer program instructions (which are grouped into units in some embodiments), and the processing unit 402 executes the computer program instructions to implement one or more aspects of the present disclosure. The memory 410 corresponds to one or more layers of memory devices, including (but not limited to) RAM, 4D XPOINT memory, flash memory, magnetic storage devices, etc.
[0049] The memory 410 can store an operating system 414, which provides computer program instructions used by the processing unit 402 in the general management and operation of the failover service. The memory 410 can also include computer program instructions and other information for implementing aspects of the present disclosure. For example, in one embodiment, the memory 410 includes a user interface unit 412, which generates a user interface (and / or instructions for the user interface) for display on a computing device via a navigation and / or browsing interface such as a browser or an application installed on the computing device. In addition to or in combination with the user interface unit 412, the memory 410 can also include a target area determination component 416, which is configured to detect and generate a list of areas and a list of rules. The memory 410 can also include a rule configuration component 418 to manage the implementation of availability rules.
[0050] Figure 5 Depicts a general architecture of a computing device configured to perform a failover area management service 140 according to some embodiments. Figure 5 The general architecture of the failover area management service 140 depicted in includes an arrangement of computer hardware and software that can be used to implement aspects of the present disclosure. The hardware can be implemented on a physical electronic device, as will be discussed in more detail below. The failover area management service 140 can include more (or fewer) elements than those shown in Figure 5 However, in order to provide a disclosure that can be implemented, it is not necessary to show all these generally conventional elements. In addition, Figure 5 The general architecture shown in can be used to implement Figure 1One or more of the other components shown in
[0051] As shown, the failover region management service 140 includes a processing unit 502, a network interface 504, a computer-readable medium drive 506, and an input / output device interface 508, all of which can communicate with each other via a communication bus. The network interface 504 can provide connectivity to one or more networks or computing systems. The processing unit 502 can thus receive information and instructions from other computing systems or services via the network. The processing unit 502 can also communicate with a memory 510 and also provide output information for an optional display via the input / output device interface 508. The input / output device interface 508 can also accept input from an optional input device (not shown).
[0052] The memory 510 can contain computer program instructions (which are grouped into units in some embodiments), and the processing unit 502 executes the computer program instructions to implement one or more aspects of the present disclosure. The memory 510 corresponds to one or more layers of memory devices, including (but not limited to) RAM, 3D XPOINT memory, flash memory, magnetic storage devices, etc.
[0053] The memory 510 can store an operating system 514, which provides computer program instructions used by the processing unit 502 in the general management and operation of the failover service. The memory 510 can also include computer program instructions and other information for implementing aspects of the present disclosure. For example, in one embodiment, the memory 510 includes a user interface unit 512, which generates a user interface (and / or instructions for the user interface) for display on a computing device via a navigation and / or browsing interface such as a browser or an application installed on the computing device. In addition to or in combination with the user interface unit 512, the memory 510 can also include a target region availability determination component 516, which is configured to detect and generate a list of regions and a list of rules. The memory 510 can also include a failover region processing engine component 518 to manage the implementation of processes such as a repair or readiness process.
[0054] Figure 6 is a flowchart depicting an exemplary routine 600 for managing a failover service. For example, the routine 600 can be executed by the failover region management service 140.
[0055] Routine 600 begins at block 602, where the failover region management service 140 obtains a list of failover regions. The list of failover regions can include one or more failover regions. The list of failover regions can correspond to regions previously designated as available failover regions. The list of failover regions can correspond to all regions within a certain area or a portion of all regions. The list of failover regions can be provided by a client for input into the failover service. The list of failover regions can be detected by examining each region in which the client partition is running. In some embodiments, the list of failover regions can include regions in which no client partition is currently running.
[0056] Routine 600 then continues at block 604, where the failover region management service 140 obtains a list of rules. The list of rules 306 can include one or more rules, where one or more rules can be associated with one or more region parameters. The list of rules 306 can correspond to rules that must be satisfied for a region to be considered an available failover region. The list of rules 306 can be provided in whole or in part by a client. The list of rules 306 can be provided in whole or in part based on a determination by the region information system 306. For example, the region information system 306 can determine that the client is running seven regions and the region with the highest running partition is region X, which is running 20 partitions. The region information system 306 can then determine that one rule is that a region must have the ability to run 20 partitions to be considered an available failover region.
[0057] Routine 600 continues at block 606, where the region information processing service 130 must obtain a list of available failover regions. The list of available failover regions can include one or more available failover regions. The failover region management service 140 can obtain the list of available failover regions by receiving the list from the client. The failover region management service 140 can obtain the list of available failover regions by listing which regions in the list of regions satisfy each rule in the list of rules. In some embodiments, the failover region management service 140 can obtain a list of available failover regions corresponding to a previous list of available failover regions.
[0058] Routine 600 continues at block 608, where the failover region management service 140 must first determine a list of unavailable failover regions. The failover region management service 140 can obtain the list of unavailable failover regions by listing which regions in the region list do not satisfy one or more of the rules in the rule list. In some embodiments, the failover region management service 140 can obtain the list of unavailable failover regions by listing which regions in the region list are not in the list of available failover regions. The failover region management service 140 must then determine one or more rule engines that are configured to operate on one or more available failover regions. The one or more rule engines can include one or more of a repair engine and a readiness engine. The repair engine can be configured to repair one or more unavailable failover regions such that the one or more unavailable failover regions satisfy each rule in the rule list. The readiness engine can be configured to prepare the failover service such that one or more of the unavailable failover regions are placed in an available location.
[0059] At decision block 610, a test is made to determine whether to update the list of available regions. If so, at block 612, the list of available regions can be updated by the target regions that were previously indicated as unavailable but have been successfully repaired. In other embodiments, the updated list of available regions can be updated to remove previously available target regions that were not successful in the readiness process. At decision block 614, a test is made to determine whether to repeat routine 600. As described above, the trigger event can correspond to timing information, a manual selection, or other established events such as a client input event, a reduction in the capacity of the primary region, or any other event. Routine 600 can repeat to block 602.
[0060] Depending on the embodiment, certain actions, events, or functions of any of the processes or algorithms described herein can be executed in a different sequence, can be added, combined, or left out altogether (e.g., not all described operations or events are necessary to practice the algorithm). Additionally, in some embodiments, the operations or events can be executed simultaneously rather than sequentially by, for example, multithreading, interrupt processing, or one or more computer processors or processor cores or on other parallel architectures.
[0061] The various illustrative logical blocks, modules, routines, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware or as a combination of electronic hardware and executable software. To clearly illustrate this interchangeability, the various illustrative components, blocks, modules, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or as software running on the hardware depends on the particular application and design constraints imposed on the overall system. For each particular application, the described functionality may be implemented in a different manner, but such implementation decisions should not be construed as causing a departure from the scope of the present disclosure.
[0062] In addition, the various illustrative logical blocks and modules described in connection with the embodiments disclosed herein can be implemented or performed by a machine designed to perform the functions described herein, such as a similarity detection system, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof. The similarity detection system can be or include a microprocessor, but in an alternative, the similarity detection system can be or include a controller, a microcontroller, or a state machine configured to estimate and communicate prediction information, combinations thereof, and the like. The similarity detection system can include circuitry configured to process computer-executable instructions. Although described herein primarily in terms of digital technology, the similarity detection system can also primarily include analog components. For example, some or all of the prediction algorithms described herein can be implemented in an analog circuit or a mixed analog and digital circuit. The computing environment can include any type of computer system, including (but not limited to) a computer system based on a microprocessor, a mainframe computer, a digital signal processor, a portable computing device, a device controller, or a computing engine within an appliance, to name a few.
[0063] The elements of the methods, processes, routines, or algorithms described in connection with the embodiments disclosed herein can be embodied directly in hardware, in software modules executed by a similarity detection system, or in a combination of both. The software modules can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of non-transitory computer-readable storage medium. The exemplary storage medium can be coupled to the similarity detection system such that the similarity detection system can read information from, and write information to, the storage medium. In an alternative, the storage medium can be integral with the similarity detection system. The similarity detection system and the storage medium can reside in an ASIC. The ASIC can reside in a user terminal. In an alternative, the similarity detection system and the storage medium can reside as discrete components in a user terminal.
[0064] Unless otherwise specifically stated or otherwise understood in the context in which it is used, conditional language such as "able to", "can", "may", "could", "for example", etc., as used herein, generally intends to convey that certain embodiments include certain features, elements or steps while other embodiments do not. Thus, such conditional language generally does not intend to imply that one or more embodiments require, in any way, the features, elements, and / or steps, or that one or more embodiments must include logic for determining whether to include or perform these features, elements, and / or steps in any particular embodiment, with or without other input or cues. The terms "comprising", "including", "having", etc. are synonymous and are used in an open-ended, inclusive sense and do not exclude additional elements, features, acts, operations, etc. Further, the term "or" is used in its inclusive sense (and not in its exclusive sense) such that when used, for example, to connect a list of elements, the term "or" means one, some, or all of the elements in the list.
[0065] Unless otherwise specifically stated, disjunctive language such as the phrase "at least one of X, Y, or Z" should be understood in the context as is commonly used to present that items, elements, etc. can be X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language generally does not and should not imply that certain embodiments require the presence of at least one of X, at least one of Y, or at least one of Z.
[0066] Unless otherwise explicitly specified, articles such as "a" or "an" generally should be understood to include one or more of the items being described. Thus, phrases such as "a device configured to..." are intended to include one or more of the recited devices. Such one or more of the recited devices may also be configured jointly to perform the stated recitation. For example, "a processor configured to perform recitations A, B, and C" may include a first processor configured to perform recitation A working in conjunction with a second processor configured to perform recitations B and C.
[0067] While the foregoing detailed description has shown, described, and pointed out novel features as applicable to various embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms shown may be made without departing from the spirit of the present disclosure. As will be recognized, some embodiments described herein may be embodied in forms that do not provide all of the features and benefits set forth herein, as some features may be used or practiced separately from other features. The scope of certain embodiments disclosed herein is indicated by the appended claims rather than by the foregoing description. All changes within the meaning and range of equivalents of the recited claims will be included within the scope of the recited claims.
[0068] The various exemplary embodiments of the present disclosure may be described by the following clauses:
[0069] Clause 1. A system for managing failover regions, the system comprising:
[0070] One or more computing devices associated with a regional failover management system, wherein the regional failover management system is configured to:
[0071] Obtain a list of target failover regions corresponding to failover regions in a region that is the same as or adjacent to the primary region;
[0072] Obtain a first processing rule, the first processing rule being associated with capacity information attributed to the primary region;
[0073] Process the obtained list of target failover regions to determine a list of available failover regions, the list of available failover regions including one or more target failover regions associated with capacity information that matches or exceeds at least one of the capacity information associated with the primary region;
[0074] Process the obtained list of target failover regions to determine a list of unavailable failover regions, the list of unavailable failover regions including one or more target failover regions associated with capacity information that does not exceed the capacity information associated with the primary region; and
[0075] Perform a repair operation, the repair operation being configured to increase the capacity of the determined one or more unavailable failover regions associated with capacity information that does not exceed the capacity information associated with the primary region.
[0076] Clause 2. The system according to Clause 1, wherein the regional failover management system is further configured to:
[0077] Perform tests on the one or more available failover regions to generate a determination of the readiness of the regional failover management system;
[0078] Indicate the readiness of the regional failover management system; and
[0079] Update the list of available failover regions at least in part based on the readiness of the regional failover management system.
[0080] Clause 3. The system according to Clause 1 or Clause 2, wherein the regional failover management system automatically performs a repair operation on the determined one or more unavailable failover regions associated with capacity information that does not exceed the capacity information associated with the primary region.
[0081] Clause 4. The system according to any one of Clauses 1 to 3, wherein the regional failover management system is configured to:
[0082] Obtain one or more additional processing rules corresponding to the primary region, the one or more additional processing rules defining separate parameters and associated thresholds; and
[0083] Process the obtained list of target failover regions by applying the one or more additional processing rules to determine one or more available failover regions.
[0084] Clause 5. A system for managing failover regions, the system comprising:
[0085] One or more computing devices associated with a regional failover management system, wherein the regional failover management system is configured to:
[0086] For an identified application, identify a primary region and a list of target failover regions;
[0087] Obtain a list of processing rules, the processing rules defining at least one associated parameter and threshold associated with characterizing the list of target failover regions;
[0088] Determine a list of available failover regions based on applying the obtained list of processing rules to the identified list of target failover regions; and
[0089] Perform at least one rule engine operation in response to the determined list of available failover regions.
[0090] Clause 6. The system according to Clause 5, wherein the regional failover management system is further configured to characterize one or more failover regions as unavailable based on applying the obtained list of processing rules to the identified list of target failover regions.
[0091] Clause 7. The system according to Clause 6, wherein at least one rule engine operation includes repairing one or more failover regions characterized as unavailable.
[0092] Clause 8. The system according to Clause 7, wherein the at least one rule engine operation further includes repairing one or more available failover regions from the list of available failover regions based on the proximity of the defined failover events.
[0093] Clause 9. The system according to Clause 7 or Clause 8, wherein the at least one rule engine operation further includes repairing one or more available failover regions from the list of available failover regions at least in part based on at least one of the failure rate, cost, availability, workload location, infrastructure, or latency of the one or more available failover regions.
[0094] Clause 10. The system according to any one of Clauses 5 to 9, wherein the at least one rule engine operation includes:
[0095] Performing a readiness check on the one or more available failover regions;
[0096] Providing an indication of the readiness of the one or more available failover regions; and
[0097] Updating the list of available failover regions at least in part based on the indication of the readiness of the one or more available failover regions.
[0098] Clause 11. The system according to any one of Clauses 5 to 10, wherein the one or more processing rules correspond to capacity.
[0099] Clause 12. The system according to any one of Clauses 5 to 11, wherein the one or more processing rules correspond to error rate.
[0100] Clause 13. The system according to any one of Clauses 5 to 12, wherein the list of processing rules is generated by a third-party user.
[0101] Clause 14. The system according to Clause 13, further comprising determining the regional capacity of the one or more target failover regions, wherein the at least one rule engine operation includes repairing one or more available failover regions from the list of available failover regions at least in part based on the list of processing rules and the regional capacity.
[0102] Clause 15. The system according to any one of Clauses 5 to 14, wherein the regional failover management system is configured to run at a predetermined interval.
[0103] Clause 16. The system according to any one of Clauses 5 to 15, wherein the regional failover management system is configured to run at least in part based on the occurrence of an event, wherein the event can be a client input event, a reduction in the capacity of the primary region, or any other event.
[0104] Clause 17. A computer-implemented method for managing a group of regions, wherein the group of regions includes a primary region and at least one failover region, the method comprising:
[0105] Obtain a list of target failover regions;
[0106] Define a list of available failover regions including one or more target failover regions, wherein the list of available failover regions is at least partially based on applying processing rules that define an availability metric to the list of one or more target failover regions; and
[0107] In response to the defined list of available failover regions, determine one or more operations on the one or more target failover regions.
[0108] Clause 18. The method according to clause 17, further comprising:
[0109] Perform a readiness check on the defined list of one or more available failover regions;
[0110] Provide an indication of the readiness of the one or more available failover regions; and
[0111] Update the list of available failover regions at least partially based on the indication of the readiness of the one or more available failover regions.
[0112] Clause 19. The method according to clause 17 or clause 18, further comprising:
[0113] Define a list of unavailable failover regions including one or more target failover regions, wherein the list of unavailable failover regions is at least partially based on applying processing rules that define an availability metric to the list of one or more target failover regions;
[0114] Generate a repair recommendation based on defining the one or more unavailable failover regions; and
[0115] Repair the one or more unavailable failover regions by performing repair operations at least partially based on the repair recommendation.
[0116] Clause 20. The method according to clause 19, further comprising updating the list of available failover regions in response to successful repair of one or more unavailable failover regions.
[0117] Clause 21. The method according to any one of clauses 17 to 20, wherein the one or more operations on the one or more target failover regions in response to the defined list of available failover regions includes determining the list of target failover regions corresponding to a target criterion.
[0118] Clause 22. The method according to any one of clauses 17 to 21, further comprising:
[0119] Define a primary region, where the primary region is the region with the largest number of partitions among a group of regions; and
[0120] Define the list of failover regions, where the list of failover regions is the regions among the group of regions associated with a smaller number of partitions.
[0121] Clause 23. The method according to any one of Clauses 17 to 22 further includes: restricting a capacity check on the list of failover regions.
[0122] Clause 24. A system for managing failover regions, the system includes:
[0123] One or more computing devices associated with a regional failover management system, where the regional failover management system is configured to:
[0124] Obtain a specification of a set of primary regions corresponding to an identified application;
[0125] For a separate primary region in the specified set of primary regions, obtain a list of target failover regions corresponding to failover regions in regions that are the same as or adjacent to the separate primary region;
[0126] Obtain a plurality of separate primary region processing rules, where the separate primary region processing rules are associated with capacity information belonging to at least one separate primary region in the specified set of primary regions;
[0127] Process the obtained list of target failover regions to determine a list of available failover regions for a separate primary region in the specified set of primary regions, where the list of available failover regions includes one or more target failover regions associated with capacity information that matches or exceeds at least one of the capacity information associated with at least one separate primary region in the specified set of primary regions;
[0128] Process the obtained list of target failover regions to determine a list of unavailable failover regions for a separate primary region in the specified set of primary regions, where the list of unavailable failover regions includes one or more target failover regions associated with capacity information that does not exceed the capacity information associated with at least one separate primary region in the specified set of primary regions;
[0129] Execute a repair operation, where the repair operation is configured to increase the capacity of one or more unavailable failover regions in the determined list of unavailable failover regions.
[0130] Clause 25. The system according to Clause 24, where the regional failover management system is further configured to:
[0131] Perform tests on one or more available failover regions to generate a determination of the readiness of the regional failover management system;
[0132] Indicate the readiness of the regional failover management system; and
[0133] Update the list of available failover regions for the one or more corresponding primary regions, at least in part based on the readiness of the regional failover management system.
[0134] Clause 26. The system according to clause 24 or clause 25, wherein the regional failover management system automatically performs the repair operation on one or more unavailable failover regions for the one or more primary regions.
[0135] Clause 27. The system according to any one of clauses 24 to 26, wherein the regional failover management system is further configured to:
[0136] Obtain one or more additional processing rules corresponding to one or more of the primary regions, the one or more additional processing rules defining separate parameters and associated thresholds; and
[0137] Process the obtained list of target failover regions by applying one or more additional processing rules to determine one or more available failover regions for each primary region.
[0138] Clause 28. The system according to any one of clauses 24 to 27, wherein the list of available failover regions includes one or more target failover regions associated with capacity information that matches or exceeds at least one of the capacity information associated with two or more primary regions in a specified set of primary regions.
[0139] Clause 29. A system for managing failover regions, the system comprising:
[0140] One or more computing devices associated with a regional failover management system, wherein the regional failover management system is configured to:
[0141] For an identified application, identify a plurality of primary regions and a corresponding list of target failover regions, wherein a separate primary region in the plurality of primary regions corresponds to one or more target failover regions in the corresponding list of target failover regions;
[0142] Obtain a plurality of individual master area processing rules, where individual master areas in the plurality of master areas respectively correspond to one or more of the plurality of processing rules, and the individual master area processing rules define at least one associated parameter and threshold associated with characterizing a corresponding target failover area list;
[0143] Based on applying one or more individual master area processing rules to one or more target failover areas corresponding to the respective master area, determine an available failover area list for each master area in the master area list; and
[0144] Execute at least one rule engine operation in response to at least one available failover area list.
[0145] Clause 30. The system according to Clause 29, wherein the area failover management system is further configured to characterize one or more target failover areas as unavailable relative to one or more of the plurality of master areas at least in part based on applying one or more individual master area processing rules in the obtained list of processing rules to the identified corresponding target failover area list.
[0146] Clause 31. The system according to Clause 29 or Clause 30, wherein the area failover management system is further configured to:
[0147] Characterize a plurality of target failover areas in the identified corresponding target failover area list as unavailable relative to a first subset of the master areas in the plurality of master areas at least in part based on applying one or more individual master area processing rules in the obtained list of processing rules to the identified corresponding target failover area list; and
[0148] Characterize a plurality of target failover areas in the identified corresponding target failover area list as available relative to a second subset of the master areas in the plurality of master areas at least in part based on applying one or more individual master area processing rules in the obtained list of processing rules to the identified corresponding target failover area list.
[0149] Clause 32. The system according to any one of Clauses 29 to 31, wherein the area failover management system is further configured to:
[0150] Characterize a first plurality of target failover areas in the identified corresponding target failover area list as available relative to a first subset of the plurality of master areas at least in part based on applying one or more individual master area processing rules in the obtained list of processing rules to the identified corresponding target failover area list; and
[0151] Characterize the first plurality of target failover regions in the identified corresponding target failover region list as available relative to a second subset of the plurality of primary regions, at least in part based on applying one or more processing rules from the obtained list of processing rules to the identified corresponding target failover region list.
[0152] Clause 33. The system according to any one of Clauses 29 to 32, wherein the at least one rule engine operation includes repairing one or more target failover regions characterized as unavailable relative to one or more of the plurality of primary regions.
[0153] Clause 34. The system according to any one of Clauses 29 to 33, wherein the at least one rule engine operation further includes repairing one or more available failover regions from one or more lists of available failover regions relative to one or more of the plurality of primary regions, at least in part based on the proximity of the defined failover events.
[0154] Clause 35. The system according to any one of Clauses 29 to 34, wherein the at least one rule engine operation includes repairing one or more available failover regions from one or more lists of available failover regions relative to one or more of the primary region lists, at least in part based on at least one of the failure rate, cost, availability, workload location, infrastructure, or latency of the one or more available failover regions.
[0155] Clause 36. The system according to any one of Clauses 29 to 35, wherein the at least one rule engine operation includes:
[0156] Performing a readiness check on one or more available failover regions from one or more lists of available failover regions relative to one or more of the primary region lists;
[0157] Providing an indication of the readiness of the one or more available failover regions relative to the one or more primary regions; and
[0158] Updating the one or more lists of available failover regions at least in part based on the indication of the readiness of the one or more available failover regions relative to the one or more primary regions.
[0159] Clause 37. The system according to any one of Clauses 29 to 36, wherein one or more of the plurality of processing rules correspond to capacity.
[0160] Clause 38. The system according to any one of Clauses 29 to 37, wherein one or more of the plurality of processing rules correspond to error rate.
[0161] Clause 39. The system according to any one of Clauses 29 to 38, wherein one or more of the plurality of processing rules are generated by a client.
[0162] Clause 40. A computer-implemented method for managing a set of regions, wherein the set of regions includes a plurality of primary regions and at least one failover region, the method comprising:
[0163] Obtaining a list of a plurality of primary regions;
[0164] Obtaining a list of target failover regions;
[0165] Defining a list of available failover regions based on the list of target failover regions for individual primary regions among the plurality of primary regions, wherein the list of available failover regions is at least partially based on applying individual primary region processing rules that define an availability metric to the one or more lists of target failover regions, wherein the individual primary region processing rules correspond to the respective primary regions; and
[0166] Determining one or more operations on one or more target failover regions in response to the defined list of available failover regions.
[0167] Clause 41. The method according to Clause 40, further comprising:
[0168] Performing a readiness check on the defined list of available failover regions for each of the one or more primary regions;
[0169] Providing an indication of the readiness of one or more available failover regions in the defined list of available failover regions for each of the one or more primary regions; and
[0170] Updating the list of available failover regions for each of the one or more primary regions at least partially based on the indication of the readiness of the one or more available failover regions.
[0171] Clause 42. The method according to Clause 40 or Clause 41, further comprising:
[0172] Defining a list of unavailable failover regions based on the list of target failover regions for individual primary regions among the plurality of primary regions, wherein each list of unavailable failover regions is at least partially based on the application of the individual primary region processing rules;
[0173] Generating a repair recommendation based on defining the list of unavailable failover regions for each of the one or more primary regions;
[0174] Repair one or more unavailable failover regions by performing repair operations, at least in part based on the repair suggestions.
[0175] Clause 43. The method according to any one of Clauses 40 to 42 further includes updating the list of available failover regions for one or more corresponding primary regions in response to successful repair of one or more unavailable failover regions.
[0176] Clause 44. The method according to any one of Clauses 40 to 43, wherein the one or more operations on one or more target failover regions in response to the defined list of available failover regions includes determining a list of target failover regions corresponding to a target criterion for a corresponding primary region.
[0177] Clause 45. The method according to any one of Clauses 40 to 44 further includes:
[0178] Defining the plurality of primary regions, wherein the plurality of primary regions are the regions with the largest number of partitions in a set of regions; and
[0179] Defining the list of target failover regions, wherein the list of target failover regions are the regions in the set of regions associated with a smaller number of partitions.
[0180] Clause 46. The method according to any one of Clauses 40 to 45 further includes restricting a capacity check on the list of target failover regions.
[0181] Clause 47. The method according to any one of Clauses 40 to 46, wherein the list of available failover regions corresponds to failover regions common to two or more of the plurality of primary regions.
[0182] Clause 48. The method according to any one of Clauses 40 to 47, wherein the list of available failover regions corresponds to failover regions unique to a primary region among the plurality of primary regions.
[0183] Clause 49. A system for managing primary regions associated with an application, the system includes:
[0184] One or more computing devices associated with a region management system, wherein the region management system is configured to:
[0185] Obtain a list of a plurality of primary regions corresponding to an application;
[0186] For a separate primary region among the plurality of primary regions, obtain a separate primary region processing rule corresponding to the capacity information of the separate primary region;
[0187] Process the obtained individual master region processing rules to identify an increase in capacity for additional individual master regions among the multiple master regions, the identified increase in capacity corresponding to a set of individual master region processing rules; and
[0188] Perform a repair operation that is configured to increase the capacity of at least one of the identified one or more additional master regions in response to the identified increase in capacity of the multiple master regions.
[0189] Clause 50. The system according to clause 49, wherein the region management system automatically performs the repair operation on the one or more unavailable master regions in the determined list of unavailable master regions.
[0190] Clause 51. The system according to clause 49 or clause 50, wherein the region management system is further configured to:
[0191] Obtain one or more additional processing rules corresponding to the master regions in the obtained list of master regions, the one or more additional processing rules defining individual parameters and associated thresholds; and
[0192] Process the obtained list of master regions by applying the one or more additional processing rules to determine one or more available master regions.
[0193] Clause 52. The system according to any one of clauses 49 to 51, wherein the region management system performs the repair operation based on a variable increase in the capacity for the multiple master regions.
[0194] Clause 53. A system for managing regions, the system comprising:
[0195] One or more computing devices associated with a region management system, wherein the region management system is configured to:
[0196] For an identified application, identify a list of master regions;
[0197] Obtain a plurality of processing rules, wherein each master region in the identified list of master regions corresponds to one or more of the plurality of processing rules;
[0198] Determine a capacity requirement for a set of master regions based on applying the obtained list of processing rules to the identified list of master regions; and
[0199] In response to the determined capacity requirement of the set of master regions, perform at least one rule engine operation.
[0200] Clause 54. The system according to Clause 53, wherein the area management system is further configured to characterize one or more of the identified list of primary areas as available based on applying one or more of the obtained processing rules to the one or more primary areas.
[0201] Clause 55. The system according to Clause 54, wherein the at least one rule engine operation includes repairing the one or more primary areas based on the determined capacity.
[0202] Clause 56. The system according to any one of Clauses 53 to 55, wherein the at least one rule engine operation includes repairing one or more of the primary areas from the determined list of primary areas based on the proximity of the defined event.
[0203] Clause 57. The system according to any one of Clauses 53 to 56, wherein the at least one rule engine operation includes repairing one or more of the available primary areas from the determined list of available primary areas based at least in part on at least one of the failure rate, cost, availability, workload location, infrastructure, or latency of the one or more available primary areas.
[0204] Clause 58. The system according to any one of Clauses 53 to 57, wherein at least one of the plurality of processing rules corresponds to capacity.
[0205] Clause 59. The system according to any one of Clauses 53 to 58, wherein at least one of the plurality of processing rules corresponds to error rate.
[0206] Clause 60. The system according to any one of Clauses 53 to 59, wherein at least one of the plurality of processing rules is generated by a client.
[0207] Clause 61. The system according to any one of Clauses 53 to 60, wherein the area management system is further configured to determine the area capacity of one or more of the available primary areas in the list of available primary areas, and wherein the at least one rule engine operation includes repairing one or more of the available primary areas from the list of available primary areas based at least in part on the plurality of processing rules and the area capacity.
[0208] Clause 62. The system according to any one of Clauses 53 to 61, wherein the area management system is configured to run at a predetermined interval.
[0209] Clause 63. The system according to any one of Clauses 53 to 62, wherein the area management system is configured to operate at least in part based on the occurrence of an event, where the event can be a client input event, a reduction in capacity of one of the main areas in the identified list of main areas, or any other event.
[0210] Clause 64. A computer-implemented method for managing a set of areas, where the set of areas includes a plurality of main areas, the method comprising:
[0211] defining a list of a plurality of main areas that support applications,
[0212] identifying a failover capacity modification for the plurality of main areas, where the failover capacity modification is at least in part based on applying a processing rule that defines an availability metric to individual main areas in the list of the plurality of main areas; and
[0213] determining one or more operations on one or more of the main areas in the obtained list of main areas corresponding to the defined list of available main areas.
[0214] Clause 65. The method according to Clause 64, the method further comprising:
[0215] defining a list of unavailable main areas, where the list of unavailable main areas is at least in part based on applying a processing rule that defines an availability metric to each main area in the obtained list of main areas;
[0216] generating a repair recommendation based on defining the list of unavailable main areas; and
[0217] repairing one or more of the unavailable main areas in the list of unavailable main areas by performing repair operations at least in part based on the repair recommendation.
[0218] Clause 66. The method according to Clause 65, further comprising updating the defined list of available main areas in response to successful repair of the one or more unavailable main areas.
[0219] Clause 67. The method according to any one of Clauses 64 to 66, where the determined one or more operations on one or more of the main areas in the obtained list of main areas include determining the list of main areas corresponding to a target criterion.
[0220] Clause 68. The method according to any one of Clauses 64 to 67, the method further comprising restricting a capacity check on the obtained list of main areas.
[0221] It should be emphasized that many changes and modifications can be made to the above-described embodiments, and the elements of these changes and modifications should be understood to be among other acceptable examples. In this document, all such modifications and variations are intended to be included within the scope of the present disclosure and are protected by the following claims.
Claims
1. A regional failover management system, which comprises: a computer-readable memory that stores executable instructions; and one or more computer processors that communicate with the computer-readable memory to manage the failover of applications partitioned across multiple isolated regions, wherein the one or more computer processors are configured to execute the executable instructions to at least: obtain a list of target failover regions corresponding to failover regions in a region that is the same as or adjacent to the primary region, wherein each of the failover regions and the primary region hosts a partition of the application and is geographically isolated from each other; obtain a first processing rule, the first processing rule being associated with capacity information associated with the primary region, wherein the capacity information associated with the primary region identifies a first quantity of partitions of the application hosted by the primary region; process the list of target failover regions to determine a list of available failover regions, the list of available failover regions including one or more target failover regions associated with capacity information that matches or exceeds at least one of the capacity information associated with the primary region; process the list of target failover regions to determine a list of unavailable failover regions, the list of unavailable failover regions including one or more target failover regions associated with capacity information that does not exceed the capacity information associated with the primary region, wherein a second quantity of partitions of the application that each of the one or more target failover regions associated with capacity information that does not exceed the capacity information associated with the primary region can host does not exceed the first quantity of partitions; and perform a repair operation, the repair operation being configured to increase the capacity of one or more unavailable failover regions in the list of unavailable failover regions associated with capacity information that does not exceed the capacity information associated with the primary region, wherein, in response to performing the repair operation, the one or more unavailable failover regions become one or more available failover regions to be associated with capacity information that matches or exceeds at least one of the capacity information associated with the primary region.
2. The regional failover management system according to claim 1, wherein the one or more computer processors are configured to execute further executable instructions to at least: perform tests on one or more available failover regions in the list of available failover regions to generate a determination of the readiness of the regional failover management system; indicate the readiness of the regional failover management system; and update the list of available failover regions at least partially based on the readiness of the regional failover management system.
3. The regional failover management system according to claim 1, wherein the one or more computer processors automatically perform a repair operation on the one or more unavailable failover regions associated with capacity information that does not exceed the capacity information associated with the primary region.
4. The regional failover management system according to claim 1, wherein the one or more computer processors are configured to execute further executable instructions to at least: Obtain one or more additional processing rules corresponding to the primary region, the one or more additional processing rules defining separate parameters and associated thresholds; and Process the target failover region list by applying the one or more additional processing rules to the target failover region list to update the available failover region list.
5. A regional failover management system, which Comprises: A computer-readable memory storing executable instructions; And One or more computer processors, the one or more computer processors communicating with the computer-readable memory to manage the failover of applications partitioned across multiple isolated regions, wherein the one or more computer processors are configured to execute the executable instructions to at least: For the application, identify a primary region and a target failover region list, wherein each region of the target failover region list and each of the primary regions host partitions of the application and are geographically isolated from each other; Obtain a list of processing rules, each processing rule of the list of processing rules defining at least one associated parameter and threshold associated with characterizing the target failover region list, wherein a first processing rule of the list of processing rules is associated with capacity information associated with the primary region, and the capacity information associated with the primary region identifies a first quantity of partitions of the application hosted by the primary region; Process the target failover region list based on applying the list of processing rules to the target failover region list to determine an available failover region list, the available failover region list including one or more target failover regions associated with capacity information that matches or exceeds the capacity information associated with the primary region, wherein one or more target failover regions in the target failover region list are one or more unavailable failover regions, associated with capacity information that does not exceed the capacity information associated with the primary region, and wherein a second quantity of partitions of the application that each of the one or more target failover regions in the target failover region list can host does not exceed the first quantity of partitions; and Execute at least one rule engine operation in response to the determined available failover region list to increase the capacity of the one or more target failover regions in the target failover region list, wherein, in response to executing the at least one rule engine operation, the one or more target failover regions in the target failover region list become one or more available failover regions to be associated with capacity information that matches or exceeds the capacity information associated with the primary region.
6. The regional failover management system according to claim 5, wherein one or more computer processors are configured to execute further executable instructions to at least: characterize one or more of the failover regions in the target failover region list as unavailable target failover regions based on applying the processing rule list to the target failover region list.
7. The regional failover management system according to claim 6, wherein the at least one rule engine operation includes repairing one or more of the target failover regions in the target failover region list characterized as unavailable target failover regions.
8. The regional failover management system according to claim 7, wherein the at least one rule engine operation further includes repairing one or more available failover regions from the available failover region list based on the proximity of the failover events.
9. The regional failover management system according to claim 7, wherein the at least one rule engine operation further includes repairing one or more available failover regions from the available failover region list at least in part based on at least one of the failure rate, cost, availability, workload location, infrastructure, or latency of the one or more available failover regions.
10. The regional failover management system according to claim 5, wherein the one or more computer processors are configured to execute further executable instructions to at least: perform a readiness check on one or more available failover regions from the available failover region list; provide an indication of the readiness of the one or more available failover regions; and update the available failover region list at least in part based on the indication of the readiness of the one or more available failover regions.
11. The regional failover management system according to claim 5, wherein one or more of the processing rules of the processing rule list correspond to capacity.
12. The regional failover management system according to claim 5, wherein one or more of the processing rules of the processing rule list correspond to error rate.
13. The regional failover management system according to claim 5, wherein the processing rule list is generated by a third-party user.
14. The regional failover management system according to claim 13, wherein the one or more computer processors are configured to execute further executable instructions to at least determine the regional capacity of one or more of the target failover regions in the target failover region list, wherein the at least one rule engine operation includes repairing one or more available failover regions from the available failover region list at least in part based on the processing rule list and the regional capacity.
15. The regional failover management system according to claim 5, wherein the regional failover management system is configured to execute further executable instructions to run at a predetermined interval.
16. The regional failover management system according to claim 5, wherein the one or more computer processors are configured to execute further executable instructions to operate at least in part based on the occurrence of an event, where the event can be a client input event, a reduction in the capacity of the primary region, or any other event.
17. A computer-implemented method for managing failover of an application partitioned across multiple isolated regions, where the multiple isolated regions include a primary region and at least one failover region, and where each of the primary region and the at least one failover region hosts a partition of the application and is geographically isolated from each other, the method comprises: obtaining a list of target failover regions; processing the list of target failover regions to define a list of available failover regions, where the list of available failover regions is at least in part based on applying one or more processing rules that define an availability metric to the list of target failover regions, where a first processing rule among the one or more processing rules is associated with capacity information related to the primary region, the capacity information related to the primary region identifying a first quantity of partitions of the application hosted by the primary region, where the list of available failover regions includes one or more target failover regions associated with capacity information that matches or exceeds the capacity information related to the primary region, one or more target failover regions in the list of target failover regions being one or more unavailable failover regions, associated with capacity information that does not exceed the capacity information related to the primary region, where a second quantity of partitions of the application that each of the one or more target failover regions in the list of target failover regions can host does not exceed the first quantity of partitions; and in response to the defined list of available failover regions, performing one or more operations to increase the capacity of the one or more target failover regions in the list of target failover regions, where, in response to performing the one or more operations, the one or more target failover regions in the list of target failover regions become one or more available failover regions associated with capacity information that matches or exceeds the capacity information related to the primary region.
18. The method according to claim 17, further comprising: performing a readiness check on one or more available failover regions in the list of available failover regions; providing an indication of the readiness of the one or more available failover regions; and updating the list of available failover regions at least in part based on the indication of the readiness of the one or more available failover regions.
19. The method according to claim 17, further comprising: Define a list of unavailable failover regions including one or more target failover regions from the list of target failover regions, where the list of unavailable failover regions is at least partially based on applying the one or more processing rules to the list of target failover regions; Generate a repair recommendation based on defining the list of unavailable failover regions; And At least partially based on the repair recommendation, repair one or more unavailable failover regions in the list of unavailable failover regions by performing repair operations.
20. The method according to claim 19, further comprising updating the list of available failover regions in response to successful repair of the one or more unavailable failover regions.
21. The method according to claim 17, wherein the one or more operations for modifying the capacity of the one or more target failover regions in the list of target failover regions in response to the list of available failover regions include determining the list of target failover regions corresponding to a target criterion.
22. The method according to claim 17, further comprising: Define a primary region, where the primary region is the region among the multiple isolated regions that hosts the largest number of partitions of the application; And Define the list of target failover regions, where the list of target failover regions is the region among the multiple isolated regions that hosts a number of partitions of the application smaller than the largest number of partitions.
23. The method according to claim 17, further comprising: Limit the capacity check of the list of target failover regions.
24. A system for managing failover regions, the system comprising: A data storage medium that stores region specifications; And One or more computer hardware processors for managing the failover of an application partitioned across multiple isolated regions and communicating with the data storage medium, where the one or more computer hardware processors are configured to execute computer-executable instructions to at least: Obtain a specification for a set of primary regions for the application, where each of the primary regions hosts a partition of the application and is geographically isolated from each other; For a separate primary region in the set of primary regions, obtain a list of target failover regions for the application, the list of target failover regions corresponding to failover regions in a region that is the same as or adjacent to the separate primary region; Obtain a plurality of separate primary region processing rules, where a separate primary region in the set of primary regions corresponds to one or more processing rules of the plurality of separate primary region processing rules, and the separate primary region processing rules are associated with capacity information belonging to at least one separate primary region in the set of primary regions; Processing the list of target failover regions based on the application of one or more individual primary region processing rules of the plurality of individual primary region processing rules to determine a list of available failover regions for individual primary regions in the set of primary regions, the list of available failover regions including one or more target failover regions associated with capacity information that matches or exceeds at least one of the capacity information associated with at least one individual primary region in the set of primary regions; Processing the list of target failover regions based on the application of one or more individual primary region processing rules of the plurality of individual primary region processing rules to determine a list of unavailable failover regions for individual primary regions in the set of primary regions, the list of unavailable failover regions including one or more target failover regions associated with capacity information that does not exceed the capacity information associated with at least one individual primary region in the set of primary regions; And Performing a repair operation configured to increase the capacity of one or more unavailable failover regions in the list of unavailable failover regions.
25. The system according to claim 24, wherein the one or more computer hardware processors are configured to execute further computer-executable instructions to at least: Perform tests on one or more available failover regions in the list of available failover regions for the individual primary regions to generate a determination of the readiness of the regional failover management system; Indicate the readiness of the regional failover management system; And Update the list of available failover regions for the individual primary regions at least in part based on the readiness of the regional failover management system.
26. The system according to claim 24, wherein the one or more computer hardware processors are configured to execute further computer-executable instructions to at least automatically perform the repair operation to increase the capacity of the one or more unavailable failover regions.
27. The system according to claim 24, wherein the one or more computer hardware processors are configured to execute further computer-executable instructions to at least: Obtain one or more additional processing rules corresponding to one or more of the primary regions, the one or more additional processing rules defining individual parameters and associated thresholds, wherein processing the list of target failover regions to determine the list of available failover regions for individual primary regions in the set of primary regions is further based on the application of the one or more additional processing rules.
28. The system according to claim 24, wherein the list of available failover regions further includes one or more target failover regions associated with capacity information that matches or exceeds at least one of the capacity information associated with two or more primary regions in the set of primary regions.
29. A system for managing failover regions, the system Comprises: A data storage medium that stores region specifications; And One or more computer hardware processors for managing failover of applications partitioned across multiple isolated regions and communicating with the data storage medium, wherein the one or more computer hardware processors are configured to execute computer-executable instructions to at least: For an identified application, identify a plurality of primary regions and a corresponding list of target failover regions, wherein a separate primary region among the plurality of primary regions corresponds to one or more target failover regions in the corresponding list of target failover regions, and wherein each of the primary regions and the target failover regions hosts a partition of the application and is geographically isolated from each other; Obtain a plurality of separate primary region processing rules, wherein a separate primary region among the plurality of primary regions corresponds to one or more of the plurality of processing rules, and the separate one or more processing rules define at least one associated parameter and threshold associated with characterizing the corresponding list of target failover regions; Process the corresponding list of target failover regions to determine a list of available failover regions for each primary region in the list of primary regions based on applying one or more separate primary region processing rules to one or more target failover regions corresponding to the respective primary region; And Execute at least one rule engine operation in response to at least one list of available failover regions, the at least one rule engine operation including increasing the capacity of one or more target failover regions with respect to one or more primary regions in the list of primary regions.
30. The system of claim 29, wherein the one or more computer hardware processors are configured to execute further computer-executable instructions to at least partially characterize one or more target failover regions as unavailable with respect to one or more primary regions among the plurality of primary regions based on applying one or more separate primary region processing rules from a list of processing rules to the corresponding list of target failover regions.
31. The system of claim 29, wherein the one or more computer hardware processors are configured to execute further computer-executable instructions to at least: Characterize a plurality of target failover regions in the corresponding list of target failover regions as unavailable with respect to a first subset of primary regions among the plurality of primary regions at least partially based on applying one or more separate primary region processing rules from a list of processing rules to the corresponding list of target failover regions; and Characterize the plurality of target failover regions in the corresponding list of target failover regions as available with respect to a second subset of primary regions among the plurality of primary regions at least partially based on applying one or more separate primary region processing rules from the list of processing rules to the corresponding list of target failover regions.
32. The system of claim 29, wherein one or more computer hardware processors are configured to execute further computer-executable instructions to at least: Characterize a first plurality of target failover regions in the corresponding list of target failover regions as available relative to a first subset of the plurality of primary regions, at least in part based on applying one or more individual primary region processing rules from a list of processing rules to the corresponding list of target failover regions; and Characterize the first plurality of target failover regions in the corresponding list of target failover regions as available relative to a second subset of the plurality of primary regions, at least in part based on applying one or more processing rules from the list of processing rules to the corresponding list of target failover regions.
33. The system of claim 29, wherein, to perform the at least one rule engine operation, the one or more computer hardware processors are configured to execute further computer-executable instructions to repair one or more target failover regions characterized as unavailable relative to one or more of the plurality of primary regions.
34. The system of claim 29, wherein, to perform the at least one rule engine operation, the one or more computer hardware processors are configured to execute further computer-executable instructions to repair at least one available failover region from one or more lists of available failover regions relative to one or more of the plurality of primary regions, at least in part based on the proximity of the defined failover event.
35. The system of claim 29, wherein, to perform the at least one rule engine operation, the one or more computer hardware processors are configured to execute further computer-executable instructions to repair at least one available failover region from one or more lists of available failover regions relative to one or more of the plurality of primary regions, at least in part based on at least one of the failure rate, cost, availability, workload location, infrastructure, or latency of one or more available failover regions.
36. The system of claim 29, wherein, to perform the at least one rule engine operation, the one or more computer hardware processors are configured to execute further computer-executable instructions to at least: Perform a readiness check on one or more available failover regions from one or more lists of available failover regions relative to one or more of the plurality of primary regions; Provide an indication of the readiness of the one or more available failover regions relative to the one or more primary regions; And Update the one or more lists of available failover regions at least in part based on the indication of the readiness of the one or more available failover regions relative to the one or more primary regions.
37. The system of claim 29, wherein one or more of the plurality of processing rules correspond to capacity.
38. The system of claim 29, wherein one or more of the plurality of processing rules correspond to error rate.
39. The system according to claim 29, wherein one or more of the plurality of processing rules are generated by a client.
40. A computer-implemented method for managing failover of an application partitioned across a plurality of isolated regions, wherein the plurality of isolated regions include a plurality of primary regions and at least one failover region, the method comprises: obtaining a list of a plurality of primary regions for the application; obtaining a list of target failover regions for the application, wherein each of the plurality of primary regions and target failover regions hosts a partition of the application and is geographically isolated from each other; defining, based on the list of target failover regions, a list of available failover regions for a respective primary region of the plurality of primary regions, wherein the list of available failover regions is at least partially based on applying a respective primary region processing rule that defines an availability metric to the one or more lists of target failover regions, wherein the respective primary region of the plurality of primary regions corresponds to one or more of a plurality of respective primary region processing rules; determining one or more operations configured to increase the capacity of one or more target failover regions in response to the list of available failover regions, relative to one or more primary regions of the list of primary regions; and performing the one or more operations.
41. The method according to claim 40, the method further comprises: performing a readiness check on the list of available failover regions for a respective primary region of the plurality of primary regions; providing an indication of readiness of one or more available failover regions in the list of available failover regions for a respective primary region; and updating the list of available failover regions for a respective primary region at least partially based on the indication of readiness of the one or more available failover regions.
42. The method according to claim 40, which further comprises: defining, based on the list of target failover regions, a list of unavailable failover regions for a respective primary region of the plurality of primary regions, wherein the list of unavailable failover regions is at least partially based on applying the respective primary region processing rule to the one or more lists of target failover regions; generating a repair recommendation based on defining the list of unavailable failover regions for a respective primary region of the plurality of primary regions; and repairing one or more unavailable failover regions by performing a repair operation at least partially based on the repair recommendation.
43. The method according to claim 42, which further comprises: updating the list of available failover regions for a respective primary region in response to successful repair of the one or more unavailable failover regions.
44. The method according to claim 40, wherein the one or more operations configured to increase the capacity of the one or more target failover regions in response to the list of available failover regions, relative to one or more primary regions of the list of primary regions comprises: Determine that the list of target failover regions corresponds to a target criterion for a corresponding primary region.
45. The method according to claim 40, further comprising: defining the plurality of primary regions, where the plurality of primary regions are the regions with the largest number of hosted partitions among the plurality of isolated regions; and defining the list of target failover regions, where the list of target failover regions is the region that hosts a number of partitions smaller than the largest number of partitions among the plurality of isolated regions.
46. The method according to claim 40, further comprising restricting a capacity check on the list of target failover regions.
47. The method according to claim 40, where the list of available failover regions corresponds to failover regions common to two or more of the plurality of primary regions.
48. The method according to claim 40, where the list of available failover regions corresponds to failover regions unique to a primary region among the plurality of primary regions.
49. A system for managing primary regions associated with an application, the system comprising: a data storage medium that stores region designations; and one or more computer hardware processors for managing failover of an application partitioned across a plurality of isolated regions and communicating with the data storage medium, where the one or more computer hardware processors are configured to execute computer-executable instructions to at least: obtain a list of a plurality of primary regions corresponding to the application, where each of the plurality of primary regions hosts a partition of the application and is geographically isolated from each other; for each individual primary region among the plurality of primary regions, obtain an individual primary region processing rule corresponding to the number of partitions of the application hosted by the specific individual primary region; based on obtaining the individual primary region processing rules for each individual primary region among the plurality of primary regions, process the individual primary region processing rules of each individual primary region to determine the capacity requirements of the plurality of primary regions; and perform a repair operation configured to increase the capacity of at least one additional individual primary region among the plurality of primary regions in response to the capacity requirements of the plurality of primary regions.
50. The system according to claim 49, where the one or more computer hardware processors are configured to execute further computer-executable instructions to at least automatically perform the repair operation.
51. The system according to claim 49, where the one or more computer hardware processors are configured to execute further computer-executable instructions to at least: obtain one or more additional processing rules corresponding to the list of the plurality of primary regions, the one or more additional processing rules defining individual parameters and associated thresholds; and process the list of the plurality of primary regions by applying the one or more additional processing rules to each of the plurality of primary regions to determine one or more available primary regions.
52. The system according to claim 49, wherein the one or more computer hardware processors are configured to execute further computer-executable instructions to perform the repair operation at least based on a variable increase in the capacity for the plurality of primary regions.
53. A system for managing regions, the system comprising: a data storage medium that stores region designations; and one or more computer hardware processors for managing failover of applications partitioned across a plurality of isolated regions and communicating with the data storage medium, wherein the one or more computer hardware processors are configured to execute computer-executable instructions to at least: for an identified application, identify a list of primary regions, where each primary region hosts a partition of the application and is geographically isolated from each other; obtain a plurality of processing rules, where each primary region in the list of primary regions corresponds to one or more of the plurality of processing rules; determine a capacity requirement for the list of primary regions based on applying the plurality of processing rules to the list of primary regions; and in response to the capacity requirement of the list of primary regions, perform at least one rule engine operation configured to increase the capacity of a specific region.
54. The system according to claim 53, wherein the one or more computer hardware processors are configured to execute further computer-executable instructions to characterize one or more of the primary regions in the list of primary regions as available at least based on applying one or more of the plurality of processing rules to the one or more primary regions.
55. The system according to claim 54, wherein to perform the at least one rule engine operation, the one or more computer hardware processors are configured to execute further computer-executable instructions to repair the one or more primary regions at least based on the capacity requirement.
56. The system according to claim 53, wherein to perform the at least one rule engine operation, the one or more computer hardware processors are configured to execute further computer-executable instructions to repair one or more of the primary regions from the list of primary regions at least based on the proximity of the defined events.
57. The system according to claim 53, wherein to perform the at least one rule engine operation, the one or more computer hardware processors are configured to execute further computer-executable instructions to repair at least one or more of the available primary regions from the list of primary regions at least in part based on at least one of the failure rate, cost, availability, workload location, infrastructure, or latency of one or more available primary regions.
58. The system according to claim 53, wherein at least one of the plurality of processing rules corresponds to capacity.
59. The system according to claim 53, wherein at least one of the plurality of processing rules corresponds to error rate.
60. The system according to claim 53, wherein at least one of the plurality of processing rules is generated by a client.
61. The system according to claim 53, wherein the one or more computer hardware processors are configured to execute further computer-executable instructions to at least determine a regional capacity of one or more of the primary regions in the primary region list, and wherein, to perform the at least one rule engine operation, the one or more computer hardware processors are configured to execute further computer-executable instructions to at least partially repair one or more of the primary regions from the primary region list based on the plurality of processing rules and the regional capacity.
62. The system according to claim 53, wherein the one or more computer hardware processors are configured to execute further computer-executable instructions to run at least at a predetermined interval.
63. The system according to claim 53, wherein the one or more computer hardware processors are configured to execute further computer-executable instructions to run at least partially based on the occurrence of an event, where the event can be a client input event, a reduction in the capacity of one of the primary regions in the primary region list, or any other event.
64. A computer-implemented method for managing failover of an application partitioned across multiple isolated regions, where a set of regions includes a plurality of primary regions, the method comprises: for an identified application, identifying a plurality of primary region lists, where each of the plurality of primary regions hosts a partition of the application and is geographically isolated from each other; obtaining a plurality of processing rules, where each primary region in the identified plurality of primary region lists corresponds to one or more of the plurality of processing rules; based on applying the plurality of processing rules to the identified plurality of primary region lists, determining a capacity requirement for the plurality of primary regions, the plurality of processing rules corresponding to matching or exceeding the number of partitions of the application hosted by individual primary regions in the plurality of primary region lists; performing a repair operation configured to increase the capacity of one or more unavailable failover regions in response to the determined capacity requirement of the plurality of primary regions.
65. The method according to claim 64, the method further comprises: defining a list of unavailable primary regions, where the list of unavailable primary regions is at least partially based on the application of the plurality of processing rules; generating a repair recommendation based on defining the list of unavailable primary regions; and at least partially based on the repair recommendation, repairing one or more of the unavailable primary regions in the list of unavailable primary regions by performing a repair operation.
66. The method according to claim 65, which further comprises: updating a list of a plurality of available primary regions in response to successful repair of the one or more unavailable primary regions.
67. The method according to claim 65, which further comprises determining a list of the plurality of primary regions corresponding to a target criterion.
68. The method according to claim 65, the method further comprises restricting a capacity check of the list of the plurality of primary regions.
Citation Information
Patent Citations
Application resource acquisition based on failover
CN103034570A
Semi-automatic failover
CN106716972A