Automated failover with preemptive blanking

A preemptive failover system using machine learning and heuristic models addresses downtime and data consistency issues by automatically transferring workloads to a safer secondary system, ensuring continuous operation and reducing recovery complexities.

JP2026503104APending Publication Date: 2026-01-27HITACHI VANTARA LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025540882
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-02-10
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Traditional failover techniques in storage systems result in downtime and data consistency issues due to reactive failovers triggered after primary system failures, complicating recovery and potentially causing dual processing workloads.

Method used

Implementing a preemptive failover system that uses machine learning and heuristic models to detect potential threats, allowing automated or administrative initiation of failovers to a secondary system with lower threat levels, ensuring seamless workload transfer and minimizing downtime.

Benefits of technology

The preemptive failover system reduces downtime and data loss by proactively transferring workloads to a safer location, maintaining system continuity and avoiding dual system complications during recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026503104000001_ABST
    Figure 2026503104000001_ABST
Patent Text Reader

Abstract

In some examples, a computing device receives first information related to a state of a first computing system and second information related to a state of a second computing system. The computing device may determine, based at least on the first information, that a threat level related to a potential failure in the first computing system meets a threat level condition for performing a preemptive failover. Based at least on the second information, the computing system may determine that the threat level in the second computing system is lower than the threat level in the first computing system. Based at least on determining that the threat level in the second computing system is lower than the threat level in the first computing system, the computing device transmits an instruction to the first computing system to initiate a preemptive failover from the first computing system to the second computing system.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of data storage. [Background technology]

[0002] To maintain continuity and high availability of applications and data, storage systems and other computing systems may include failover technologies that enable workloads to be transferred from a first site to a second site in the event of a failure at the first site. A failover operation may typically be triggered when a first computing system fails, or may be manually triggered by an administrative user, such as after a failure occurs. Thus, traditional failover techniques can result in periods of downtime, as even automatically triggered failovers based on a failed system may require transfer and recovery periods, and can also introduce data consistency issues. Additionally, because a failover typically occurs after a primary system goes down, the recovery process can become even more complicated when the primary site is restored to operational status, such as when both sites may attempt to process the workload. Summary of the Invention [Means for solving the problem]

[0003] In some implementations, a computing device receives first information related to a state of a first computing system and second information related to a state of a second computing system. The first computing system can communicate with the second computing system over a network. The computing device may determine, based at least on the first information, that a threat level related to a potential failure in the first computing system meets a threat level condition for performing a preemptive failover from the first computing system. Based at least on the second information, the computing system may determine that the threat level in the second computing system is lower than the threat level in the first computing system. Based at least on determining that the threat level in the second computing system is lower than the threat level in the first computing system, the computing device transmits an instruction to the first computing system to initiate a preemptive failover from the first computing system to the second computing system. [Brief explanation of the drawings]

[0004] The detailed description is set forth with reference to the accompanying drawings, in which the left-most digit(s) of a reference number identifies the drawing in which the reference number first appears. Use of the same reference number in different drawings indicates similar or identical items or features.

[0005] [Figure 1] 1 illustrates an example architecture of a system capable of detecting events and performing preemptive failover according to some implementations.

[0006] [Figure 2] 1 illustrates an exemplary logical arrangement of a failover program configured to monitor events and preemptively perform a failover when appropriate, according to some implementations of the present disclosure.

[0007] [Figure 3] FIG. 1 is a flow diagram illustrating an example process for preemptively performing a failover according to some implementations.

[0008] [Figure 4] 1 illustrates selected components of an exemplary computing system that may be used to implement some of the functionality of the systems described herein. DETAILED DESCRIPTION OF THE INVENTION

[0009] Some implementations herein are directed to techniques and arrangements for generating and employing machine learning and / or heuristic models to detect events that may threaten the operation of a computing system and preemptively triggering a failover process based at least on detecting the events. Examples of events that may trigger a preemptive failover process may include events detected based on one or more sensors associated with the computing system, such as a temperature rise, a power loss, a network outage, a security breach, intruder detection, a fire alarm, detection of excessive humidity, and the occurrence of an earthquake. Further examples of events that may trigger a preemptive failover process may include events detected based on external state information received from one or more computing devices, such as a light restriction or power outage warning, a tornado warning, a tsunami warning, a hurricane warning, lightning, or a heat wave. Other examples of possible events may include notification of insufficient staffing availability, notification of a security risk (e.g., local unrest, war, etc.), or an unexpected change in operational costs (e.g., energy costs, cooling costs).

[0010] The failover process may be a fully automated failover performed between a primary system and a secondary system. For example, based on a detected threat level at a first computing system, the processing workload at the first computing system may be preemptively transferred to a second computing system at a different site location. For example, the system may determine that the threat level at the first computing system meets a threshold or otherwise warrants a failover to the second computing system. Furthermore, in implementations herein, before determining to perform a preemptive failover to the second computing system, the system may check the current threat level at the second computing system and any other available computing systems to ensure that the threat level at the destination computing system is lower than the threat level at the first computing system. For example, examples herein may take into account an indication of where the current workload is being processed and an assessment of the relative threat levels at the locations of different computing systems within the overall system.

[0011] In some examples, a human administrative user may first be prompted to approve the preemptive failover process before it is executed. In other examples, administrative user approval is not required before initiating the preemptive failover process, and the failover process may be initiated automatically, such as based on a decision made by a failover decision engine. For example, the failover decision engine may include a machine learning model and / or a heuristic model configured to decide to initiate a preemptive failover from a first computing system to a second computing system based at least on a comparison of an indication of the current threat level at each computing system. Furthermore, in some cases, a preemptive failover herein may instead be referred to as a switchover that is performed based on an automated comparison and weighting of the current threat level at each computing system within an overall system that includes multiple computing systems located at different geographic locations.

[0012] For illustrative purposes, some example implementations are described in an environment of multiple computing systems in communication with each other and monitored for threats to enable preemptive failover from one computing system to another based on a detected threat level. However, implementations herein are not limited to the specific examples provided and may be extended to other types of computing system architectures, other types of threats, other types of storage environments, other types of client configurations, other types of data, etc., as will become apparent to those skilled in the art in light of the disclosure herein.

[0013] 1 illustrates an exemplary architecture of a system 100 capable of detecting events and performing preemptive failover according to some implementations. The system 100 includes multiple computing systems 102, each including one or more computing devices 104. For example, a first computing system 102(1) includes one or more computing devices 104(1), a second computing system 102(2) includes one or more computing devices 104(2), and a third computing system 102(3) includes one or more computing devices 104(3). In some examples, the multiple computing devices 104 in each computing system 102 may form a cluster of computing devices, or the like, although implementations are not limited to any particular configuration of computing systems 104.

[0014] The computing systems 102 may be located at individual sites 105 and may be able to communicate with each other via one or more networks 106. For example, a first computing system 102(1) may be physically located at a first site 105(1) in a first geographic location, a second computing system 102(2) may be physically located at a second site 105(2) in a second geographic location remote from the first geographic location, and a third computing system 102(3) may be physically located at a third site 105(3) remote from the first and second geographic locations. For example, the second geographic location and the third geographic location may be sufficiently distant from the first geographic location and from each other, such as in another city, another state, another country, etc., so that a disaster or other disturbance affecting the first computing system 102(1) at the first site 105(1) cannot affect the second computing system 102(2) at the second site 105(2) or the third computing system 102(3) at the third site 105(3), or vice versa. This topology therefore provides redundancy for the data stored on computing systems 102(1)-102(3) to avoid catastrophic loss of data.

[0015] In some examples, computing system 102 may also communicate with one or more client devices 108, one or more management devices 110, and one or more information computing devices 111 over one or more networks 106. For example, computing system 102 may include access nodes, server nodes, management nodes, and / or other types of service nodes that provide storage services to client devices 108 to enable client devices 108 to store data with one or more computing systems 102 and perform other management and control functions, as discussed further below. Client devices 108, management devices 110, and information computing devices 111 may be any of various types of computing devices, as discussed further below.

[0016] The one or more networks 106 may include any suitable network, including a wide area network such as the Internet, a local area network (LAN) such as an intranet, a wireless network such as a cellular network, a local wireless network such as Wi-Fi, and / or a short-range wireless network such as BLUETOOTH®, a wired network including Fibre Channel, optical fiber, Ethernet, or any other such network, a direct wired connection, or any combination thereof. Thus, the one or more networks 106 may include both wired and / or wireless communication technologies. The components used for such communication may depend, at least in part, on the type of network, the selected environment, or both. Protocols for communicating over such networks are well known and will not be discussed in detail herein. By way of example, the network 106 may include a private network such as a LAN, a storage area network (SAN), or a Fibre Channel network. Additionally, the network 106 may include a public network, which may include the Internet, or a combination of public and private networks. Implementations herein are not limited to a particular type of network as the network 106.

[0017] In some examples, computing system 102 may be configured to provide storage services and data management services, respectively, to client users 112 via client devices 108. As some non-limiting examples, client users 112 may include users who perform functions for a business, enterprise, organization, government agency, academic institution, etc., which in some examples may involve storing very large amounts of data. Additionally, in some examples, client users 112 may include administrative users who use client devices 108 to manage one or more of computing systems 102. Nevertheless, implementations herein are not limited to any particular use or application of system 100 and other example systems and configurations described herein. For example, in some examples, client devices 108 may not be included or may be an entirely different type of client device.

[0018] Each client device 108 may be any suitable type of computing device, such as a desktop, laptop, tablet computing device, mobile device, smartphone, wearable device, terminal, and / or any other type of computing device capable of transmitting data over a network. Client users 112 may be associated with client devices 108 by individual user accounts, user login credentials, etc. Further, client devices 108 may be configured to communicate with computing system 102 through one or more networks 106, a separate network, or any other suitable type of communications connection. Numerous other variations will be apparent to those skilled in the art having the benefit of this disclosure.

[0019] In some implementations, each client device 108 may include a respective instance of a client application 114 that may execute on the client device 108, such as to communicate with web applications 116 executable on one or more of the computing devices 104 of the computing system 102. The client application may be configured to submit user data for storage on the computing system 102 and / or receive stored data from the computing system 102 through data instructions such as write operations, read operations, and delete operations. In some cases, the application 114 may include or operate through a browser, while in other cases, the application 114 may include any other type of application having communication capabilities that enable it to communicate with the web application 116 or other applications on the computing system 102 over one or more networks 106.

[0020] Additionally, administrator device 110 may be any suitable type of computing device, such as a desktop, laptop, tablet computing device, mobile device, smartphone, wearable device, terminal, server, and / or any other type of computing device capable of transmitting data over a network. For example, administrative user 113 may be associated with a particular administrator device 110, such as via a particular administrator account, administrator login credentials, or the like. Furthermore, administrator device 110 may be able to communicate with service computing device 102 via one or more networks 106, a separate network, or any other suitable type of communications connection.

[0021] Additionally, each management device 110 may include a respective instance of a management application 115 executable on the administrator device 110, such as to communicate with a management module of a web application 116 executable on the computing device 104, such as to send management instructions for managing the system 100, and / or to send management data for storage by the computing system 102, and / or to receive stored management data from the computing system 102 via management instructions, etc. In some cases, the management application 115 may include or operate via a browser, while in other cases, the management application 115 may include any other type of application having communication capabilities that enable it to communicate with the web application 116 over one or more networks 106. Furthermore, for administrative users 113, the web application 116 may provide remote management functions. Alternatively, in other examples, any of a number of other types of software arrangements may be employed to perform these functions, as would be apparent to one of ordinary skill in the art having the benefit of this disclosure.

[0022] The information computing devices 111 may be servers, such as web servers, or any other suitable type of computing device capable of providing information over a network. For example, the information computing devices 111 may provide local state data 117, such as weather forecasts, government information, news feeds, etc., associated with one or more of the computing systems 102, such as to provide information indicating events likely to trigger a failover. Examples of local state information may include warnings about lighting restrictions or power outages, tornadoes, tsunamis, hurricanes or similar severe storms, lightning, heat waves, or notifications of security risks (e.g., localized unrest, war, localized emergencies, evacuations, disasters, etc.). In some examples, multiple information computing devices 111 may be configured to provide one or more of the types of information discussed above as data feeds, push notifications, responses to periodic requests, etc.

[0023] Computing systems 102(1), 102(2), and 102(3) may be capable of executing instances of storage programs 120(1), 120(2), and 102(3), respectively, which may include, access, or be capable of executing instances of replication programs 122(1), 122(2), and 122(3), respectively. Additionally, in the illustrated example, at least third computing system 102(3) may execute failover program 124 configured to perform monitoring. In some examples, third computing system 124(3) may be a cloud-based computing system operating on computing devices provided by a commercial computing service provider, such as AMAZON WEB SERVICES and / or SIMPLE STORAGE SERVICE, MICROSOFT AZURE, GOOGLE CLOUD, ORACLE CLOUD, IBM CLOUD, etc. Additionally, in some examples, first computing system 102(1) and second computing system 102(2) may be privately owned by a company or other entity, may be operated on a commercial computing platform, or may comprise a combination thereof. Alternatively, in some examples, third computing system 102(3) may also be privately owned or maintained, such as by the same entity as computing systems 102(1) and 102(2). Additionally, in some examples, failover program 124 may not execute on third computing system 102(3), but may execute on a fourth computing system or other suitable computing device (not shown) that is geographically separate from computing systems 102(1)-102(3) to ensure that operation of failover program 124 can withstand the failure of any one of computing systems 102(1)-102(3). As yet another alternative, the failover program may run on each of computing systems 102(1)-102(3) or on any one of computing systems 102(1)-102(3). Additional details of the failover program 124 are discussed below.

[0024] The storage program 120 may provide access to stored data 130 on each computing system 102. In some examples, the stored data 130 may be stored data objects that include object data for each data object and corresponding metadata associated with the object data. However, in other examples, other types of data may be included as stored data 130. As a result, implementations herein are not limited to any particular type of data as stored data 130.

[0025] Storage program 120(1) may access, store, and manage stored data 130(1), storage program 120(2) may access, store, and manage stored data 130(2), and storage program 120(3) may access, store, and manage stored data 130(3). For example, storage program 120 may receive data from client device 108, store the data as stored data 130 on one or more storage devices associated with the respective computing system 102, and / or retrieve and send requested data to client device 108, such as in response to a client read request.

[0026] Additionally, storage program 120 may include, execute, access, or coexist with data replication program 122, which may be configured to perform data replication between computing system 102(1) and computing system 102(2), at least from first computing system 102(1) to second computing system 102(2), as indicated by data replication link 128. Additionally, in some examples, computing system 102(1) may perform replication to third computing system 102(3), or another computing system (not shown). In some cases, second computing system 102(2) and third computing system 102(3) may also be configured to perform replication to other computing systems 102 and / or other computing systems not shown in this example. Furthermore, data replication program 122 may configure computing systems 102(1)-102(3) to perform asynchronous or synchronous data replication between at least computing systems 102(1) and 102(2).

[0027] The computing systems 102 may each include a sensor 131. For example, a first computing system 102(1) may include a sensor 131(1) that provides sensor data 132(1), a second computing system 102(2) may include a sensor 131(2) that provides sensor data 132(2), and a third computing system 102(3) may include a sensor 131(4) that provides sensor data 132(5). The system may include sensors 131(3) that provide a temperature sensor, a smoke detector, a seismometer, a power sensor, an alarm system, a humidity sensor, a network monitor, etc. Implementations herein are not limited to a particular type of sensor 131.

[0028] The sensor data 132 may be transmitted to the failover program 124. For example, in some cases, the computing system 102 may transmit the sensor data 132 to the computing device running the failover program 124. In other examples, some or all of the sensors 131 may be configured to communicate directly with the computing device running the failover program 124 over one or more networks 106. In either case, in the illustrated example, the computing device 104(3) running the failover program 124 at the third computing system 102(3) receives sensor data 131(1) indicative of conditions at the first computing system 102(1) and receives sensor data 131(2) indicative of conditions at the second computing system 102(2). The computing device 104(3) running the failover program 124 at the third computing system 102(3) may receive sensor data 132(3) from a sensor 131(3) located at the third computing site 105(3).

[0029] Additionally, as described above, computing device 104(3) executing failover program 124 at third computing system 102(3) may receive local state data 117 from information computing device 111. Local state data 117 may include local state data 117 for first computing system 102(1) at first site 105(1), local state data 117 for second computing system 102(2) at second site 105(2), and local state data 117 for third computing system 102(3) at third site 105(3).

[0030] In the illustrated example of FIG. 1 , it is assumed that one of the computing devices 104(3) executes failover program 124 to receive first sensor data 132(1), second sensor data 132(2), and third sensor data 132(3), and to further receive local state data 117 for first computing system 102(1), second computing system 102(2), and third computing system 102(3). In addition, computing device 104(3) may receive other information that may be relevant to failover considerations, such as information technology (IT) system information. Examples of IT system information may include information related to an insufficient staffing level, notifications of security hardware or software breaches, such as viruses, malware, denial-of-service attacks, ransomware attacks, etc. In some cases, at least a portion of the IT system information may be provided by administrative user 113, such as via management device 110.

[0031] As discussed further below with respect to FIG. 2, failover program 124 may identify from the received information any events that are likely to pose a threat to any of computing systems 102(1)-102(3). For each identified event, the system may associate a threat level with the event. Based on the individual threat levels associated with each computing system 102(1)-102(3), computing device 104(3) may execute failover program 124 to determine whether any of the threat levels are high enough to justify performing a preemptive failover from one of computing systems 102 to another of computing systems 102. In some cases, the failover program may include or have access to a machine learning model trained to make a decision based on the individual threat levels determined for each of computing systems 102 to determine whether to initiate a preemptive failover.

[0032] In this example, assume that local state information 117 of first computing system 102(1) indicates that a planned power outage is expected to be implemented for the geographic location in which first computing system 102(1) is located. Further, assume that sensor information of second computing system 102(2) indicates that a temperature associated with the second computing system is increasing, but not at a level high enough to meet a threshold corresponding to requesting a failover. Consequently, based on determining that the threat level at the first system meets the conditions for performing a failover from the first system, and further based on determining that the threat level at the second system is lower than the threat level at the first system and does not meet the conditions for performing a failover, computing device 104(3) may send failover command 140 to at least first computing system 102(1) to instruct the first computing system to perform a failover workload processing to second computing system 102(2).

[0033] As a result, the second computing system 102(2) may receive a failover from the first computing system 102(1) and may begin processing workloads previously processed by the first computing system 102(1). For example, the second computing system 102(2) may begin receiving and responding to requests from client devices 108 previously serviced by the first computing system 102(1). In addition, the second computing system 102(2) may begin replicating to the third computing system 102(3), as shown at 142, and replicating back to the first computing system 102(1) so long as the first computing system 102(1) remains operational. This can reduce the execution time of a failback from the second computing system 102(2) to the first computing system 102(1), such as when the threat of a planned power outage has ended. Furthermore, one example is described above, and numerous modifications will be apparent to those skilled in the art having the benefit of this disclosure.

[0034] 2 illustrates an exemplary logical configuration of a failover program 124 configured to monitor events and preemptively perform a failover when appropriate, according to some implementations herein. The failover program 124 may be executed by a computing device, such as one of the computing devices 104 in one of the computing systems 102 discussed above with reference to FIG. 1. For example, in some examples, the failover program 124 may be executed by a computing device 104(3) in a third computing system 102(3). Alternatively, or additionally, the failover program 124 may be executed by the first computing system 102(1), the second computing system 102(2), the management computing device 110, and / or any other suitable computing device capable of communicating with the computing systems 102 over one or more networks 106.

[0035] Failover program 124 may receive multiple feeds of information from multiple sources, including sensor information from each computing system 102 and local state information for each computing system 102. Failover program 124 may perform a ranking of computing systems 102 based at least in part on the likelihood of failure of each computing system 102(1)-102(3), as discussed above with respect to Figure 1. Each different type of data received in these individual data feeds 202 may be treated differently according to its individual type.

[0036] 2, the failover program 124 may receive multiple data feeds 202, such as an environmental data feed 204, an Internet of Things (IoT) data feed 206, and an IT system feed 208. Additionally, as another source of information, the failover program 124 may execute an analytics engine 210. For example, the analytics engine 210 may be a trained machine learning model and / or heuristic model that may receive many pieces of data for use in making decisions regarding the state of each computing system 102(1)-102(3). The data received by the analytics engine for consideration may include data received via the multiple data feeds 204, 206, and 208, as well as other information relevant to determining whether a preemptive failover should be performed. In the case of machine learning models, the analytics engine may be any of a number of types of machine learning models, such as artificial neural networks, e.g., self-organizing neural networks, recurrent neural networks, convolutional neural networks, modular neural networks, deep learning neural networks, generative adversarial networks, etc., as well as predictive models, decision trees, classifiers, regression models such as linear regression models, support vector machines, probabilistic models such as Markov models and hidden Markov models, etc.

[0037] As an example, the machine learning model may be a neural network or other suitable machine learning model trained using information obtained from numerous past failover and non-failover situations as training data. For example, the information may include sensor data, data feeds, environmental conditions, and numerous other data points related to past failovers and non-failovers. The model is trained and validated to recognize situations that may lead to a failover. For example, while several environmental factors or other data points individually may not normally be recognized as a concern, the trained machine learning model may determine that, collectively, these environmental factors and data points indicate an increased likelihood that an incident may occur to the extent that a threat threshold is exceeded.

[0038] The received raw data of each data feed 202 may be converted by a respective external event adapter 212, 214, and 216, each configured to receive an external data feed in a particular format and normalize the output for further processing. For example, assume that the received environmental data feed 204 includes streaming weather information for each of the computing systems 102(1)-102(3) discussed above with respect to FIG. 1. The external event adapter 212 may convert the received weather data feed of each computing system 102 into a standard output, such as predicted outside temperature, predicted wind speed, predicted flooding, predicted lightning strikes, etc. As an example, each external event adapter may include an application programming interface (API), which may include an interface configured to receive certain types of data streams and convert each type of received data stream into a structured or standardized data format.

[0039] The external event adapters 212, 214, and 216 provide their respective outputs to noise filters 222, 224, and 226, respectively. Thus, there may be a noise filter 222 for environmental data, a noise filter 224 for Internet of Things data, and a noise filter 226 for IT system data. For example, each noise filter 222-226 may receive standardized data as input from its associated external event adapter 212-216, respectively, and apply defined filter rules 230 to generate information about an event. As an example, assume that the noise filter 222 receives temperature as standardized data from the external event adapter 212 for environmental data. Based on the data received from the corresponding external event adapter 212, and based on the received temperature value, and in some cases, based on a change (or lack of change) in the temperature compared to other recently received temperatures for the same computing system 100, the noise filter 222 may apply one or more of the defined filter rules 230 to determine whether an event has occurred that may be related to the initiation of a preemptive failover. For example, the noise filter 222 may determine whether the temperature exceeds a defined temperature threshold specified by the defined filter rules 230, and if so, whether the temperature exceeded the defined temperature threshold for a specified time threshold.

[0040] The defined filter rules 230 may identify conditions under which a preemptive failover should be considered and may further attribute a severity level to each event. For example, the higher the severity level, the greater the urgency to perform a preemptive failover. The outputs of the noise filters 222, 224, and 226 may be provided to threat profilers 232, 234, and 236, respectively, and the output of the analysis engine 210 may be provided to a threat profiler 238 for engine output analysis. For example, each threat profiler 232, 234, and 236 may receive input from its corresponding noise filter 222, 224, and 226, respectively, which may include information about detected events and the severity levels of the events. Similarly, the threat profiler 238 may receive output from the analysis engine 210, which may indicate the likelihood of one or more failover events. The threat profiler may apply one or more of the defined filter rules 230 to assign a threat profile to the identified events. For example, each threat profile assigned to a corresponding event may indicate a threat level at which the associated computing system 102 is likely to fail based on the corresponding event. Accordingly, each of threat profilers 232, 234, 236, and 238 may receive and rank individual events for threat profiling and assignment of individual threat levels.

[0041] The threat profilers 232-238 may provide the generated threat profiles and event information to the failover decision engine 240. The failover decision engine 240 may generate a failover instruction, indicated at 242, to initiate a failover process based on an indicated threat level for one of the computing systems 102, indicating that a failover should be performed for that computing system 102. For example, the failover decision engine 240 may generate a failover instruction to create a failover event based on weighing the threat level for each computing system 102 relative to the threat levels for the other computing systems 102. In this manner, the failover decision engine 240 may ensure that a workload is hosted on the most appropriate computing system 102 based on comparing the individual current threat levels to failure for each individual computing system 102. Thus, the failover decision engine 240 is aware of the current threat level at each of the computing systems 102 and can therefore avoid performing a failover from a first site having a lower threat level to a second site having a higher threat level. On the other hand, if one of the computing systems 102 has a high threat level and another computing system 102 has a low threat level, the failover decision engine 240 may preemptively initiate a preemptive failover to the system having the lower threat level.

[0042] In some examples, failover decision engine 240 may include a trained machine learning model, such as any of the example machine learning models discussed above with respect to analytics engine 210. In this case, however, rather than being trained to identify events, failover decision engine 240 is trained to determine whether to generate a failover instruction based on receiving information about a number of different event types and corresponding events for the computing system and the corresponding threat level for each different event. For example, the machine learning model may be trained and validated using training data related to past good and bad failover decisions, which in some cases may be based on both automated and human-commanded failovers.

[0043] Alternatively, in another example, the failover decision engine 240 may include a heuristic-based decision-making program that applies multiple heuristic rules to determine whether to generate a failover command. For example, the heuristic-based decision-making program may include multiple rules that indicate whether a failover should be performed, such as: If the fire alarm is activated for more than 30 seconds, it is considered a fire. If a power loss is detected, the priority is "high" If a power loss is detected and the UPS battery is below 25 percent, the priority is "Critical" If seasonal statistics indicate periods of high usage and computer cluster capacity below 80 percent, then the priority is "medium" (e.g., if some machines in the cluster are down and Black Friday sales start in 2 hours, then fail over to another site). Many other heuristic rules will be apparent to one of ordinary skill in the art having the benefit of this disclosure.

[0044] The failover program 124 can make an informed decision regarding whether to perform a preemptive failover based on a comparison of the current threat levels at each of the computing systems monitored by the failover program 124. For example, this can allow the primary computing system to place workloads on the computing system with the most consistent state to ensure no data loss. Additionally, the failover program 124 may reduce or eliminate outages, such as users being unable to access data, that might otherwise occur due to a failure in a computing system that results in a reactive failover and corresponding switchover time. Thus, the preemptive failover techniques herein may be performed based at least in part on ranking the relative threat levels at each of the monitored computing systems, thereby avoiding performing a failover to a computing system with an equal or greater threat level.

[0045] Additionally, after recovery, the primary system herein recognizes that a failover process has been performed to another computing system, which can eliminate the complexity of potentially having two systems that believe they are currently the primary system for running the workload. For example, this complication can be particularly problematic if connectivity between the primary computing system and the computing system that was the target of the failover has not yet been re-established. Furthermore, in some examples, the techniques herein may be extended to consider staffing shortages, cost of service, workload balancing, and / or managing geographic peak workload demands.

[0046] FIG. 3 is a flow diagram illustrating an example process 300 for preemptively performing a failover according to some implementations. The process is illustrated as a collection of blocks in a logical flow diagram, which represents a series of operations, some or all of which may be implemented by hardware, software, or a combination thereof. In a software context, the blocks may represent computer-executable instructions stored on one or more computer-readable media that, when executed by one or more processors, programs the processors to perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the blocks are described should not be construed as limiting. Any number of the described blocks can be combined in any order and / or in parallel to implement a process or alternative processes, and not all of the blocks need be performed. For purposes of illustration, the process is described with reference to the environments, frameworks, and systems described in the examples herein; however, the process may be implemented within a wide variety of other environments, frameworks, and systems. In some cases, process 300 may be performed at least in part by a computing device, such as in one or more of computing systems 102 or by any other suitable computing device capable of communicating with computing system 102 .

[0047] At 302, the computing device may receive first information related to one or more states of a first computing system, second information related to one or more states of a second computing system, and third information related to one or more states of a third computing system.

[0048] At 304, the computing device may convert at least a portion of the received first information, second information, and / or third information data into a structured data format.

[0049] At 306, the computing device may identify an event based on the data in the structured format exceeding a data threshold.

[0050] At 308, based at least on identifying the event, the computing device may determine a threat level corresponding to the event for the corresponding computing system.

[0051] At 310, based on at least the first information, the computing device may determine that a threat level associated with a potential failure at the first computing system meets a threat level condition for performing a preemptive failover from the first computing system.

[0052] At 312, based on at least the second information, the computing device may determine that the threat level at the second computing system is lower than the threat level at the first computing system.

[0053] At 314, based at least on the threat level at the second computing system being lower than the threat level at the first computing system, the computing device may send an instruction to the first computing system to initiate a preemptive failover from the first computing system to the second computing system.

[0054] The example processes described herein are merely example processes provided for illustrative purposes. Numerous other variations will be apparent to those skilled in the art in light of the disclosure herein. Furthermore, while the disclosure herein describes some example frameworks, architectures, and environments suitable for implementing the processes, the implementations herein are not limited to the specific examples shown and described. Furthermore, the disclosure provides various example implementations as described and illustrated. However, as known or will become known to those skilled in the art, the disclosure is not limited to the implementations described and illustrated herein, and may extend to other implementations.

[0055] FIG. 4 illustrates selected example components of a computing system 102 that may be used to implement some of the system functionality described herein. The computing system 102 includes one or more computing devices 104, which may include servers or other types of computing devices, which may be embodied in various ways. Additionally, in some examples, the computing devices 104 may include or be in communication with one or more storage systems, storage controllers, network-attached storage, storage arrays, storage area networks, etc., to store stored data 130. For example, in the case of a server, the programs, other functional components, and data may be implemented on a single server, a cluster of servers, a server farm or data center, a cloud-hosted computing service, etc., although other computer architectures may also or alternatively be used. Multiple computing devices 104 may be collocated or separately located and may be organized, for example, as virtual servers, server banks, and / or server farms. The described functionality may be provided by servers of a single entity or company, or by servers and / or services of multiple different entities or companies.

[0056] In the illustrated example, computing device 104 may include or be associated with one or more processors 402, one or more computer-readable media 404, and one or more communication interfaces 406. Each processor 402 may be a single processing unit or multiple processing units and may include single or multiple arithmetic units or multiple processing cores. Processor 402 may be implemented as one or more central processing units, microprocessors, microcomputers, microcontrollers, digital signal processors, graphics processing units, system-on-chip processors, state machines, logic circuits, and / or any device that manipulates signals based on operational instructions. By way of example, processor 402 may include one or more hardware processors and / or logic circuits of any suitable type that are specially programmed or configured to execute the algorithms and processes described herein. Processor 402 may be configured to fetch and execute computer-readable instructions stored on computer-readable media 404, which can program processor 402 to perform the functions described herein.

[0057] The computer-readable medium 404 may include volatile and non-volatile memory and / or removable and non-removable media implemented in any type of technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. For example, the computer-readable medium 404 may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, optical storage, solid-state storage, magnetic tape, and magnetic disk storage, or any other medium that can be used to store desired information and that can be accessed by a computing device.

[0058] Depending on the configuration of the computing system 102, the computer readable medium 1004 may be a tangible, non-transitory medium, provided that the non-transitory computer readable medium 404 is referred to as excluding media such as energy, carrier signals, electromagnetic waves, and / or signals themselves. In some cases, the computer readable medium 404 may be co-located with the computing system 102, while in other cases, the computer readable medium 404 may be partially remote from the computing system 102.

[0059] The computer-readable medium 404 may be used to store any number of functional components executable by the processor 402. In many implementations, these functional components include instructions or programs executable by the processor 402 that, when executed, specifically program the processor 402 to perform the operations attributed to the computing system 102 herein. The functional components stored on the computer-readable medium 1004 may include a replication program 122, a failover program 124, each of which may include one or more computer programs, applications, modules, executable code, or portions thereof. Furthermore, although these programs are shown together in this example, in some examples, these programs may be separate programs, and during use, some or all of these programs may execute on separate computing devices 104 of the respective computing systems 102.

[0060] As mentioned above, in some examples, the failover program 124 may include or have access to one or more machine learning models 408 that may be trained to perform the functions described above for the analysis engine 210 and / or the failover decision engine 240. Additionally, the failover program 124 may include or have access to external event adapters 212-216, noise filters 222-226, threat profilers 232-238, and defined filter rules 230.

[0061] Additionally, computer-readable medium 404 may store data, data structures, and other information used to perform the functions and services described herein. For example, computer-readable medium 404 may at least temporarily store one or more data structures including sensor data 132 and local state data 117. Computer-readable medium 404 may also store data 130, such as stored in one or more storage systems or other storage devices, as discussed above.

[0062] Computing system 102 may include or maintain other functional components and data, which may include programs, drivers, etc., as well as other data used or generated by the functional components. Additionally, computing system 102 may include many other logical, programmatic, and physical components, of which the ones described herein are merely examples relevant to the discussion herein.

[0063] The one or more communication interfaces 406 may include one or more software and hardware components for enabling communication with various other devices, such as over one or more networks 106 and 107. For example, the communication interface 406 may enable communication over one or more of a LAN, the Internet, a cable network, a cellular network, a wireless network (e.g., Wi-Fi) and a wired network (e.g., Fibre Channel, Fiber Optic, Ethernet), a direct connection, short-range communication such as BLUETOOTH®, and the like, as additionally listed elsewhere herein.

[0064] Various instructions, methods, and techniques described herein may be considered in the general context of computer-executable instructions, such as computer programs and applications, stored on a computer-readable medium herein and executed by a processor. In general, the terms programs and applications may be used interchangeably and may include instructions, routines, scripts, modules, objects, components, data structures, executable code, etc. for performing particular tasks or implementing particular data types. These programs, applications, and the like may be executed as native code, or may be downloaded and executed in a virtual machine or other just-in-time compilation execution environment, etc. Typically, the functionality of programs and applications may be combined or distributed as desired in various implementations. Implementations of these programs, applications, and techniques may be stored on computer storage media or transmitted over some form of communication medium.

[0065] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.

Claims

1. one or more processors configured with executable instructions to receive information related to a first computing system and a second computing system, the first computing system being in communication with the second computing system over a network, the one or more processors comprising: receiving, by the one or more processors, first information related to one or more states of the first computing system and second information related to one or more states of the second computing system; determining, by the one or more processors, based at least on the first information, that a threat level associated with a potential failure at the first computing system meets a threat level condition for performing a preemptive failover from the first computing system; determining, by the one or more processors, based at least on the second information, that a threat level at the second computing system is lower than the threat level at the first computing system; and transmitting, by the one or more processors, instructions to the first computing system to initiate a preemptive failover from the first computing system to the second computing system based at least on determining that the threat level at the second computing system is lower than the threat level at the first computing system. configured to perform operations including: system.

2. The system of claim 1 , wherein the first information relating to the one or more conditions of the first computing system comprises sensor data received from a plurality of sensors associated with the first computing system.

3. 2. The system of claim 1, wherein the first information related to the one or more states of the first computing system includes local state information received from a computing device over a network, the local state information including at least one of a local weather forecast for a location of the first computing system or information related to events occurring in proximity to the location of the first computing system.

4. 2. The system of claim 1, wherein the first information relating to the one or more conditions of the first computing system includes at least one of staffing shortage information relating to the first computing system, software infringement relating to the first computing system, or hardware infringement relating to the first computing system.

5. The operation is converting at least a portion of the received first information data relating to one or more states of the first computing system into a structured format; and identifying an event based on a portion of the data in the structured format exceeding a threshold value of the data; The system of claim 1 further comprising:

6. The operation is determining a threat level corresponding to the event for the first computing system based at least on identifying the event based on a portion of the data in the structured format exceeding a threshold value of the data; The system of claim 5 further comprising:

7. 2. The system of claim 1, further comprising: a machine learning model trained to determine, based at least on determining that the threat level at the second computing system is lower than the threat level at the first computing system, sending the instruction to the first computing system to initiate the preemptive failover from the first computing system to the second computing system.

8. 2. The system of claim 1, wherein the operations further comprise determining the threat level associated with the potential failure in the first computing system based on a plurality of defined rules including a plurality of thresholds corresponding to a plurality of the conditions of the first computing system.

9. 2. The system of claim 1, wherein sending the instruction to the first computing system to initiate a preemptive failover from the first computing system to the second computing system causes the first computing system to at least partially transfer a processing workload to the second computing system.

10. 10. The system of claim 9, wherein subsequent to transferring the processing workload to the second computing system, the second computing system replicates data related to the workload to a third computing system and the first computing system.

11. 2. The system of claim 1, wherein the operation of determining that the threat level at the second computing system is lower than the threat level at the first computing system further comprises determining that the threat level at the second computing system does not meet a threat level condition for performing a preemptive failover from the second computing system.

12. receiving, by one or more processors, first information relating to one or more states of a first computing system and second information relating to one or more states of a second computing system, the first computing system being capable of communicating with the second computing system over a network; determining, by the one or more processors, based at least on the first information, that a threat level associated with a potential failure at the first computing system meets a threat level condition for performing a preemptive failover from the first computing system; determining, by the one or more processors, based at least on the second information, that a threat level at the second computing system is lower than the threat level at the first computing system; and transmitting, by the one or more processors, instructions to the first computing system to initiate a preemptive failover from the first computing system to the second computing system based at least on determining that the threat level at the second computing system is lower than the threat level at the first computing system. A method comprising:

13. the first information relating to the one or more states of the first computing system; sensor data received from a plurality of sensors associated with the first computing system; local state information received from computing devices over a network, the local state information including at least one of a local weather forecast for a location of the first computing system or information related to events occurring in proximity to the location of the first computing system; staffing shortage information relating to the first computing system; information relating to the infringement of software on said first computing system; or information relating to a hardware compromise in said first computing system; The method of claim 12 , comprising at least one of:

14. A non-transitory computer-readable medium having stored thereon instructions executable by one or more processors, comprising: receiving, by the one or more processors, first information related to one or more states of a first computing system and second information related to one or more states of a second computing system, the first computing system being capable of communicating with the second computing system over a network; determining, by the one or more processors, based at least on the first information, that a threat level associated with a potential failure at the first computing system meets a threat level condition for performing a preemptive failover from the first computing system; determining, by the one or more processors, based at least on the second information, that a threat level at the second computing system is lower than the threat level at the first computing system; and transmitting, by the one or more processors, instructions to the first computing system to initiate a preemptive failover from the first computing system to the second computing system based at least on determining that the threat level at the second computing system is lower than the threat level at the first computing system. a non-transitory computer-readable medium for causing the one or more processors to perform operations including:

15. the first information relating to the one or more states of the first computing system; sensor data received from a plurality of sensors associated with the first computing system; local state information received from computing devices over a network, the local state information including at least one of a local weather forecast for a location of the first computing system or information related to events occurring in proximity to the location of the first computing system; staffing shortage information relating to the first computing system; information relating to the infringement of software on said first computing system; or information relating to a hardware compromise in said first computing system; 15. The non-transitory computer-readable medium of claim 14, comprising at least one of:

Citation Information

Patent Citations

  • Failover system of server by urgent earthquake report, its method, and server

    JP2009211623A

  • Mesh Network Routing Based on Asset Availability

    JP2018528693A

  • Aircraft network cybersecurity apparatus and methods

    JP2021010161A

  • Brink of failure and breach of security detection and recovery system

    US20050050377A1

  • Cluster architecture for network security processing

    US20160164764A1