Determining data center availability
By assessing data center availability through multiple communication connections and diverse protocols, the method ensures reliable detection of data center unavailability, facilitating timely corrective actions to prevent unintended application behavior.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- TRADING TECHNOLOGIES INTERNATIONAL INC
- Filing Date
- 2024-06-24
- Publication Date
- 2026-07-29
AI Technical Summary
Data centers may become unavailable to external computing entities due to network connectivity loss or hardware/software failures, leading to uncontrollable application behavior or complete cessation of function, necessitating rapid and reliable availability determination to enable corrective actions.
A method involving a computing system that determines the state of multiple communication connections between a data center and multiple other data centers, using heartbeat messages and diverse connection types (e.g., UDP and TCP) to assess availability, ensuring reliable detection of connectivity issues and enabling autonomous corrective actions.
Enables rapid and accurate determination of data center availability, reducing the risk of false negatives and allowing prompt corrective actions such as failover or shutdown, thereby maintaining system stability and functionality.
Smart Images

Figure 2026525220000001_ABST
Abstract
Description
Technical Field
[0001] A data center (DC) typically includes a group of computing devices that provide computing power for running applications, and the functions of such applications. A DC may be used when it is not practical or unrealistic to run one or more resource-intensive applications on an individual computing entity, such as a personal computer. In such cases, the resource-intensive applications, or the functions of those applications, may instead be run in the data center, thereby leveraging the computing power of the data center. Alternatively, additionally, a DC can be used when it is beneficial to minimize the communication latency between a computing device that runs one or more applications and another computing device. In such cases, it is possible to run an application using a plurality of computing devices installed in a common DC with the other computing device. Thereby, the physical separation between computing devices is minimized, and thus the communication latency between computing devices is essentially minimized. In any case, the operation of an application in a DC may be controlled by one or more computing entities outside the DC.
[0002] Certain embodiments are disclosed with reference to the following drawings.
Brief Description of the Drawings
[0003] [Figure 1] FIG. 1 is a block diagram showing a computing device according to a particular embodiment. [Figure 2] FIG. 2 is a block diagram showing an example of a system to which a particular embodiment is applicable. [Figure 3] FIG. 3 is a block diagram showing an exemplary system in which an exemplary method of application control in a data center can be implemented. [Figure 4] Figure 4 is a flowchart illustrating one example of a method for determining the availability of a data center. [Figure 5] Figure 5 is a block diagram showing an example of an electronic trading system in which a specific embodiment may be employed. [Figure 6] Figure 6 is a block diagram showing a more detailed example of an electronic trading system in which a particular embodiment may be employed.
[0004] Specific embodiments will be better understood when read in conjunction with the accompanying illustrative drawings. However, it should be understood that embodiments are not limited to the configurations and means shown in the accompanying drawings. [Modes for carrying out the invention]
[0005] The disclosed embodiments relate, in general, to a data center (DC), and more particularly to a method and system for determining the availability of the DC to one or more computing entities (physical entities) located outside the DC. A computing system within the DC may be configured to receive and / or transmit communications to one or more computing entities. For example, certain operations of a computing system may be remotely controlled by one or more computing entities located outside the DC. For example, one or more external computing entities may be connected to the DC via a network and be able to send commands to applications running on the DC to control the operation of those applications. This control may include, for example, stopping or modifying functions performed by one or more applications. This control may be useful, for example, to prevent one or more applications from operating in an undesirable manner.
[0006] Under certain circumstances, a data center (DC) may become unavailable to external computing entities. This can occur if the DC loses its network connectivity, or if some disaster or failure, such as a hardware or software failure, disrupts communication between the DC and external computing entities. In such situations, the DC may lose its ability to send and receive communications with one or more computing entities outside of it. Therefore, for example, applications running on the DC may become uncontrollable by the external computing entity because commands from the external computing entity may not reach the applications running on the DC. In certain cases of such situations, an application may continue to run on the DC without the control of the external computing entity, or a new application may start running. This situation is also known as "islanding" of the DC and can cause applications to behave unintentionally. In other similar cases, applications on the DC may cease to function completely. For example, this could be caused by a power outage or hardware / software failure on the DC.
[0007] For the reasons stated above, it may be beneficial to determine the availability of the data center (DC) to external computing entities. For example, if the DC becomes unavailable under the circumstances described above, it may be desirable to take corrective action. For instance, it may be desirable to restore the remote control capabilities of applications running on the DC and / or prevent applications running on the DC from operating unattended. It may be beneficial to quickly determine that the DC is unavailable and take prompt corrective action. However, a false determination that the DC is unavailable could result in measures being taken that unnecessarily disrupt the DC's functionality.
[0008] This specification discloses embodiments including software running on hardware, but it should be noted that these embodiments are merely illustrative and should not be construed as limiting. For example, all or part of these hardware and software components may be embodied by hardware alone, software alone, firmware alone, or any combination of hardware, software, and / or firmware. Thus, a particular embodiment may be implemented in other ways.
[0009] I. Brief Description of Specific Embodiments A particular embodiment provides a method for determining an indication of data center availability for one or more computing entities located outside the data center. This method includes a computing system determining a first state of a first group of communication connections, which includes one or more communication connections between a first data center and a second data center. The method further includes a computing system determining a second state of a second group of communication connections, which includes one or more communication connections between the first data center and a third data center different from the second data center. The method further includes a computing system determining, based on the first and second states, an indication of the availability of the first data center for one or more computing entities located outside the first data center.
[0010] These features enable reliable determination of the availability of a first data center that one or more computing entities are attempting to access. By considering the state of each communication connection between the first data center and at least two other data centers, the availability of the first data center can be determined reliably and robustly. This method reduces the risk of incorrectly determining that the first data center is unavailable. For example, if communication between the first data center and one other data center is interrupted, the cause may not be a failure in the first data center, but rather a problem specific to the communication between those data centers. In an example of this method, the first, second, and third data centers may form a so-called quorum, and the determination of the availability of the first data center may be based on the state of each communication connection between the first data center and each of the other data centers in the quorum. For example, the first data center may be determined to be unavailable (e.g., failed) only if each of the other data centers in the quorum agrees that their connection to the first data center has been lost.
[0011] In a particular embodiment, the method determines a third state of a third group of communication connections, which includes one or more groups of communication connections between a first data center and a fourth data center distinct from the second and third data centers, and determines an indication of the availability of the first data center based on the third state. Determining the availability of the first data center based on the third state can increase the reliability of the determination. For example, the possibility of incorrectly detecting that the first data center is unavailable can be reduced by adding the fourth data center to the data center quorum and taking into account additional information regarding the communication status between the first and fourth data centers. For example, if the first, second, and third states indicate that the first data center has lost communication with the second and third data centers but maintains good connectivity with the fourth data center, the first data center may be determined to remain available.
[0012] In certain embodiments, determining an indication of the availability of the first data center includes determining that the first data center is unavailable to one or more computing entities outside the first data center in response to the status of each of a group of communication connections, and in response to each of the group of communication connections indicating that it is inoperable. These functions enable rapid and reliable notification that the first data center is unavailable when it is determined that the communication connections between the first data center and the second and third data centers have been lost. Thus, the likelihood of incorrectly determining that the first data center is unavailable is reduced, because the first data center is determined to be unavailable only when communication connections to both the second and third data centers are lost, and not, for example, when communication between the first data center and one of the other data centers is lost.
[0013] In a particular embodiment, the state of a predetermined group of communication connections among a plurality of communication connection groups is determined based on the state of one or more communication connections within that predetermined group. Therefore, for example, the state of the communication connection group between the first data center and the second data center can be determined based on the state of the individual communication connections between the first data center and the second data center.
[0014] In a particular embodiment, each of a group of communication connections includes two or more communication connections. That is, each group of communication connections between two specific data centers may include two or more communication connections. This makes it possible to determine the state of a particular group of communication connections based on two or more communication connections. This provides additional information to base the determination on, enabling a more reliable determination of the communication state between specific data centers. For example, the state of communication between a first data center and a second data center may be determined based on the state of a first communication connection between the first and second data centers and the state of a second communication connection between the first and second data centers.
[0015] In certain embodiments, determining the state of a predetermined group of communication connections from among a plurality of communication connection groups includes determining the state of each of two or more communication connections within the predetermined group, and determining that the predetermined group is inoperable in response to the determination that each of two or more communication connections within the predetermined group is inoperable. This makes it possible to reduce instances in which communication between specific data centers is incorrectly determined to be lost. For example, considering a group of communication connections between a first data center and a second data center, it may be determined that the first connection is not functioning while the second connection is functioning. In this situation, it may be determined that communication between the first data center and the second data center remains operational. For example, a problem specific to the first communication connection or a particular communication mode used for that connection may occur, but if at least one communication connection between the first data center and the second data center remains operational, this is not considered to indicate a loss of communication between the two data centers. In some embodiments, at least two of the two or more communication connections are of different types. For example, one connection may be a UDP connection while the other is a TCP connection. This brings diversity to the communication connection group, making it possible to more reliably determine that even if all communication connections within that group fail, the problem is not due to a specific communication method, but rather a result of a complete disruption of communication between data centers.
[0016] In a particular embodiment, each of a plurality of communication connection groups includes a first communication connection for receiving heartbeat messages from a predetermined data center, and determining the state of a predetermined communication connection group among the plurality of communication connection groups includes determining the state of the first communication connection of the predetermined communication connection group by monitoring the reception of heartbeat messages. For example, heartbeats may be sent from the first data center to each of the other data centers. Heartbeats may be sent periodically, for example, every 30 seconds. The other data centers monitor the reception of heartbeats from the first data center and determine the state of the first communication connection that each data center has to the first data center based on whether or not a heartbeat has been received. For example, if the second data center does not receive a predetermined number of consecutive heartbeats (e.g., three consecutive), it may determine that the first communication connection to the first data center is down. This can be an effective means of determining the state of the first communication connection. For example, a loss of communication that continues for a considerable period of time can be reliably detected, but one or two heartbeat losses may be ignored so as not to unnecessarily determine that communication has been lost. Other data centers may also send their own heartbeats to each other and to the first data center. The first data center can monitor the reception of these heartbeats, for example, in the same manner as described above, and thereby determine the status of each first communication channel from the other data centers to the first data center.
[0017] In certain embodiments, each first communication connection includes a subscription to one or more multicast channels on which heartbeat messages are transmitted. For example, a designated data center may use UDP to transmit heartbeat messages. For example, heartbeat messages may be transmitted via LBM (Latin Sea Busters Messaging). By transmitting heartbeats over multicast channels, heartbeat messages can be efficiently transmitted to each other data center.
[0018] In a particular embodiment, each of a plurality of communication connection groups includes a second communication connection to enable location inspection at a given data center. Furthermore, determining the state of each of the plurality of communication connection groups includes determining whether the location at the given data center can be successfully inspected using the second communication connection. This provides an effective method for monitoring additional or alternative communication connections between data centers.
[0019] In certain embodiments, each second communication connection includes a TCP connection. As described above, in certain embodiments, each of a given group of communication connections may include different types of communication connections. If each group of communication connections includes a first communication connection and a second communication connection, the first communication connection is a UDP connection and the second communication connection is a TCP connection. This provides redundancy to the group of communication connections and can reduce the possibility that a failure specific to a particular type of communication connection may lead to a false determination that communication between a particular data center has failed.
[0020] In certain embodiments, the determination of the status of each of the multiple communication connection groups and the determination of the availability of the first data center are performed at the first data center. This may provide a reliable means for determining whether the first data center is available to one or more external computing entities. For example, the first data center may determine that it is unavailable if it determines that the status of each of the multiple communication connection groups between itself and other data centers is inoperable. The first data center operates autonomously and can determine its own operational status, and can determine that it is unavailable if it is unable to receive commands from external computing entities.
[0021] In certain embodiments, the method includes initiating one or more of the following in response to a determination that the first data center is unavailable to one or more computing entities outside the first data center: a shutdown process to shut down one or more applications running in the first data center; a disaster recovery process to operate one or more applications running in the first data center in disaster recovery mode. Various countermeasures may be taken in response to a determination that the first data center is unavailable. Since the availability of the first data center can be determined at the first data center, the need to rely on operators to initiate countermeasures in the event of a power outage or islanding at the first data center can be reduced. For example, the first data center can initiate a shutdown or disaster recovery process in response to a computing system within the first data center determining that the first data center has become unavailable to one or more external computing entities. Thus, problems that would otherwise require operators to manually take countermeasures (such as shutting down power to the first data center) in response to failures related to the first data center can be addressed. Since these operations can be performed autonomously at the first data center, they can be performed more quickly than if an operator had to manually initiate the operation. This reduces the risk of applications within the primary data center behaving unintended without control from one or more external computing entities.
[0022] In certain embodiments, the determination of the state of each of the plurality of communication connection groups and the determination of the availability of the first data center are performed in a data center different from the first data center. Thus, the availability of the first data center can be monitored by a computing entity within a data center external to the first data center. This can be advantageous in situations where the first data center becomes unavailable. For example, since the data center in which the computing entity is located remains in communication with an external computing entity, actions that may not be able to be initiated at the first data center (such as actions to notify other entities of the loss of availability of the first data center) can be initiated.
[0023] In certain embodiments, the determination of the state of each of the plurality of communication connection groups and the determination of the availability of the first data center are performed in the third data center. In such embodiments, the availability of the first data center can be monitored in the third data center that forms part of a system including the first data center and the second data center. This enables the third data center to coordinate the response to be taken when the first data center becomes unavailable and effectively communicate the loss of availability of the first data center to the second data center.
[0024] In a particular embodiment, determining the state of at least one predetermined group of communication connections from a plurality of groups of communication connections includes obtaining information about one or more communication connections of the predetermined group from data centers connected to the first data center by one or more communication connections of the predetermined group, and determining the state of the predetermined group of communication connections based on the obtained information. Thus, a data center performing the method can collect information about the state of communication connections between the first data center and other data centers and determine the availability of the first data center based on this information. For example, each data center can determine the state of the communication connections between itself and the first data center and store this state in a location accessible to a data center performing the method (e.g., a second data center). This information becomes accessible to the data center performing the method and is used to reliably determine the availability of the first data center. Thus, for example, a computing entity in a third data center may determine that the first data center is unavailable if the third data center determines that it has lost communication from the first data center, and further determines that other data centers in the quorum have also lost communication from the first data center.
[0025] In a particular embodiment, the method includes initiating a failover process to failover the functions of the first data center to a failover data center in response to the determination that the first data center is unavailable to one or more computing entities located outside the first data center. The data center that determined the first data center was unavailable can therefore reliably and quickly initiate a failover of the functions of the first data center to the failover data center.
[0026] In certain embodiments, the start of the failover process includes sending a message from the data center that has determined that the first data center is unavailable to the failover data center, instructing the failover data center to fail over the functions of the first data center. The data center that executes this approach can continue to maintain communication with other data centers within the quorum even if the first data center is unavailable. For example, the first data center may become unavailable due to a software or hardware failure in the first data center, and there may be no impact on other data centers within the quorum. Therefore, the data center may, for example, initiate the failover of the functions of the first data center to a failover data center that is another data center within the quorum. In other examples, the failover data center may be a data center that is not part of the quorum.
[0027] In certain embodiments, the failover data center is the second data center. In such an example, the third data center causes the functions of the first data center to fail over to the second data center. Since all of the first, second, and third data centers are within the quorum, the third data center can effectively monitor both the first data center and the second data center and determine the timing when the first data center becomes unavailable. It can also confirm whether the second data center is in a state where it can replace the functions of the first data center. Therefore, the third data center can effectively initiate the failover if necessary.
[0028] In certain embodiments, a data center different from the first data center, which determines the status of each of a group of communication connections and the availability of the first data center, performs these determinations in response to a determination that another data center is unable to perform these determinations. For example, the other data center can determine the status of each of a group of communication connections and, in response to the other data center determining that another data center, such as a third data center, is unable to perform these determinations, it can perform the availability determination of the first data center. This provides redundancy for the data center functions, allowing for the determination of the availability of the first data center and, for example, coordinating failover of the functions of the first data center. In some such examples, the additional data center may be a fourth data center. In other examples, the further data center may be a data center that is neither the first, second, third, nor fourth data center.
[0029] In a particular embodiment, the failover process includes, for each of the one or more applications running in the first data center, running the application in the failover data center, and modifying the behavior of the existing application in the failover data center to include one or more functions of that application. This method allows the failover data center to take over the functions of the first data center if it is determined to be unavailable.
[0030] A particular embodiment provides a computing system comprising memory and one or more processors, the one or more processors configured to determine a first state of a first group of communication connections including one or more communication connections between a first data center and a second data center, a second state of a second group of communication connections including one or more communication connections between the first data center and a third data center different from the second data center, and, based on the first and second states, determine an indication of the availability of the first data center to one or more computing entities outside the first data center. As described above, this makes it possible to reliably determine the availability of the first data center in an effective manner.
[0031] A particular embodiment provides a data center. The data center includes a computing system. The computing system comprises memory and one or more processors, the one or more processors configured to determine a first state of a first group of communication connections including one or more communication connections between a first data center and a second data center, a second state of a second group of communication connections including one or more communication connections between the first data center and a third data center different from the second data center, and, based on the first and second states, determine an indication of the availability of the first data center to one or more computing entities outside the first data center. As described above, this enables the computing system to reliably and effectively determine the availability of the first data center. The data center including the computing system may be, for example, the first data center itself. This allows for a rapid / or automatic response if the first data center becomes isolated. Alternatively, the data center including the computing system may be another data center other than the first data center, such as a third data center or a second data center. This allows the availability of the first data center to be reliably and effectively determined by the other data center. This other data center can, for example, take actions such as failing over the functions of the first data center to the other data center as a countermeasure.
[0032] A particular embodiment provides a tangible, computer-readable storage medium containing instructions to be executed by one or more processors of a computing system. The instructions include determining a first state of a first group of communication connections, which includes one or more communication connections between a first data center and a second data center; determining a second state of a second group of communication connections, which includes one or more communication connections between the first data center and a third data center different from the second data center; and determining an indication, based on the first and second states, that shows the availability of the first data center to one or more computing entities outside the first data center. As described above, this makes it possible to reliably determine the availability of the first data center in an effective manner.
[0033] II. Examples of Computing Devices Figure 1 shows a block diagram of an exemplary computing device 100. Computing device 100 may be used to carry out a particular embodiment described herein. Other computing devices may be used in other examples. Computing device 100 includes a communication bus 110, a processor 112, memory 114, a network interface 116, an input device 118, and an output device 120. The processor 112, memory 114, network interface 116, input device 118, and output device 120 are connected to the communication bus 110. Computing device 100 is connected to an external network 140, such as a local area network (LAN) or a wide area network (WAN) like the Internet. Computing device 100 is connected to the external network 140 via the network interface 116. Computing device 100 may include additional, different, or fewer components. For example, multiple communication buses (or other types of component interconnections), multiple processors, multiple memory devices, multiple interfaces, multiple input devices, multiple output devices, or any combination thereof may be provided. As another example, computing device 100 may not include an input device 118 or an output device 120. As yet another example, one or more components of computing device 100 may be integrated into a single physical element such as a field-programmable gate array (FPGA) or a system-on-a-chip (SoC).
[0034] The communication bus 110 may include channels, electrical or optical networks, circuits, switches, fabrics, or other mechanisms for communicating data between components within the computing device 100. The communication bus 110 is communicatively connected to any component of the computing device 100 and can transfer data between them.
[0035] The processor 112 can be any suitable processor, processing unit, or microprocessor. The processor 112 may include, for example, one or more general-purpose processors, digital signal processors, application-specific integrated circuits, FPGAs, analog circuits, digital circuits, programmable processors, and / or combinations thereof. The processor 112 may be a multi-core processor. A multi-core processor may contain multiple processing cores of the same or different types. The processor 112 may be a single device or a combination of devices, such as one or more devices associated with a network or distributed processing system. The processor 112 can support various processing strategies, such as multiprocessing, multitasking, parallel processing, and / or remote processing. Processing may be performed locally or remotely, and may be moved from one processor to another. In certain embodiments, the computing device 100 is a multiprocessor system and therefore may include one or more additional processors communicably connected to the communication bus 110.
[0036] The processor 112 is capable of executing logic and other computer-readable instructions encoded in one or more tangible media, such as memory 114. In this specification, logic recorded in one or more tangible media includes instructions that can be executed by the processor 112 or another processor. Such logic may be stored, for example, as part of software, hardware, integrated circuits, firmware, and / or microcode. Such logic may be received from an external communication device via a communication network 140. The processor 112 can execute such logic to perform the functions, operations, or tasks described herein.
[0037] The memory 114 may be one or more tangible media, such as computer-readable storage media. Computer-readable storage media include a variety of volatile and non-volatile storage media, such as random access memory, read-only memory, programmable read-only memory, electrically programmable read-only memory, electrically erasable read-only memory, flash memory, any combination thereof, or other tangible data storage devices. In this specification, the term “non-temporary or tangible computer-readable media” is explicitly defined as including all types of computer-readable media, excluding propagated signals. The memory 114 may include any type of mass storage device, such as a hard disk drive, optical media, magnetic tape, or magnetic disk.
[0038] Memory 114 may include one or more memory devices. For example, memory 114 may include cache memory, local memory, mass storage, volatile memory, non-volatile memory, or a combination thereof. Memory 114 may be adjacent to, part of, programmed with, networked with, and / or remotely located to the processor 112. Thus, for example, data stored in memory 114 may be retrieved and processed by the processor 112. Memory 114 may store instructions that can be executed by the processor 112. These instructions are executed to perform one or more operations or functions described herein.
[0039] Memory 114 can store an application 130 that implements the disclosed technology. In certain embodiments, the application 130 may be accessed from or stored in different locations. The processor 112 can access the application 130 stored in memory 114 and execute computer-readable instructions contained in the application 130.
[0040] The network interface 116 may include one or more network adapters. The network adapters may be wired or wireless network adapters. The network interface 116 may allow the computing device 100 to communicate with an external network 140. The computing device 100 can communicate with other devices via the network interface 116 using one or more network protocols such as Ethernet, Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), or wireless network protocols such as Wi-Fi, Long-Term Evolution (LTE) protocol, or other appropriate protocols.
[0041] The input device 118 may include positional input devices such as a mouse, touchpad, or touchscreen; input devices such as a keyboard, buttons, or switches; and / or other human-machine interface devices. The output device 120 may include a display device, which may be a liquid crystal display (LCD), a cathode ray tube (CRT), a light-emitting diode (LED) display (such as an organic EL display), or other suitable display device.
[0042] In certain embodiments, during the installation process, the application may be transferred from the input device 118 and / or the network 140 to memory 114. If the computing device 100 is running or preparing to run the application 130, the processor 112 can retrieve instructions from memory 114 via the communication bus 110.
[0043] III. System Examples Figure 2 is a block diagram illustrating the schematic of an exemplary system 200. System 200 includes a first data center (DC) 210, a second DC 220, a third DC 230, and a fourth DC 240. DCs 210, 220, 230, and 240 are connected to each other via their respective networks. Specifically, the first DC 210 and the second DC 220 are connected via the first network 261. The first DC 210 and the third DC 230 are connected via the second network 262. The first DC 220 and the fourth DC 240 are connected via the third network 263. The second DC 220 and the third DC 230 are connected via the fourth network 264, the second DC 220 and the fourth DC 240 are connected via the fifth network 265, and the third DC 230 and the fourth DC 240 are connected via the sixth network 266.
[0044] Each data center (DC) 210, 220, 230, and 240 is equipped with a corresponding computing system, which includes components for maintaining and monitoring communication between DCs (210, 220, 230, and 240). Specifically, the first DC 210 includes a first computing system 211, which includes a first monitoring device (monitor) 212, a first database 214, and a first detector 216. The second DC 220 includes a second computer system 221, which includes a second monitoring device 222, a second database 224, and a second detector 226. The third DC 230 includes a third computing system 231, which includes a third monitoring device 232, a third database 234, and a third detector 236. The fourth DC240 includes a fourth computing system 241, which includes a fourth monitoring device 242, a fourth database 244, and a fourth detector 246. The first monitoring device 212 is communicably connected to the first database 214 and the first detector 216, for example, via a local area network. The internal components of the second, third, and fourth computing systems 221, 231, and 241 are connected in a similar manner.
[0045] The first computing system 211 of the first DC210 may be implemented by or include one or more of the computing devices 100 described above with reference to Figure 1. Either or both of the first monitoring device 212 and the first detector 216 may be implemented as software running on one or more computing devices of the first computing system 211 (e.g., the computing devices 100 described above with reference to Figure 1). For example, the first monitoring device 212 and the detector 216 may run on different computing devices within the computing system 211 (e.g., the computing devices 100 described above with reference to Figure 1). However, this is not necessarily required, and other configurations are possible. The first database 214 is stored in the storage device of the first computing system 211. For example, the first database 214 may be stored in the memory of a computing device of the first computing system 211 (e.g., the memory 114 of the computing device 100 described above with reference to Figure 1). The second, third, and fourth computing systems 221, 231, and 241 may have any of the features described above for the first computing system 211.
[0046] The system further includes a first computing entity 250 and a second computing entity 260. The first computing entity 250 and the second computing entity 260 are communicateable with the first DC 210 via a seventh network 267 and an eighth network 268, respectively. Exemplaryly, one or more of the first to eighth communication networks 261-268 may include one or more of the following: a local area network, a wide area network, a multicast network, a wireless network, a virtual private network, an internal network, a cellular network, a peer-to-peer network, a connection point, a dedicated line, the Internet, a shared memory system, and / or a dedicated network. One or more of the first to eighth communication networks 261-268 may be different from each other in whole or in part, or they may be the same network. DCs 210, 220, 230, and 240 may be located in different locations, for example, in different cities or countries. In some examples, system 200 may be one of several such systems. For example, system 200 is located in a first geographical area, and similar systems may be provided in other geographical areas. When functioning correctly, the first DC 210 is configured to communicate with the first computing entity 250 and the second computing entity 260. For example, the first DC 210 can run one or more applications (not shown in Figure 2) that are controllable by the first computing entity 250. These applications may cause information to be sent and received between the first computing entity 250 and / or the second computing entity 260. In some examples (as shown in Figure 2), the first computing entity 250 and the second computing entity 260 may be located outside of the first DC 210.However, in other examples (not shown), one or both of the first and second computing entities 250, 260 may be located in the same location as part of the first DC 210, i.e., the first computing system 211.
[0047] Each monitoring device 212, 222, 232, and 242 is configured to maintain and monitor a set of communication connections between itself and each other monitoring device 212, 222, 232, and 242 (which may be referred to as peers in this specification). Each of the multiple communication connection sets enables the transmission of information from any of the monitoring devices in DC 210, 220, 230, and 240 to the remaining monitoring devices in DC 210, 220, 230, and 240 via the respective networks connecting both DCs. For example, one communication connection set enables the transmission of information from the first monitoring device 212 to the second monitoring device 222 via the first network 261, and a different communication connection set enables the transmission of information from the second monitoring device 222 to the first monitoring device 212 via the first network 261. Similarly, another group of communication connections enables the transmission of information from the first monitoring device 212 to the third monitoring device 232 via the second network 262, and yet another group of communication connections enables the transmission of information from the third monitoring device 232 to the first monitoring device 212 via the second network 262. Again, similarly, yet another group of communication connections enables the transmission of information from the first monitoring device 212 to the fourth monitoring device 242 via the third network 263, and yet another group of communication connections enables the transmission of information from the fourth monitoring device 242 to the first monitoring device 212 via the third network 263.
[0048] Each of multiple groups of connections may contain multiple connections of the same or different types. For example, a particular group of connections may contain two connections using different protocols, such as a UDP connection and a TCP connection. Furthermore, or alternatively, each connection in a group of connections may be transmitted via separate paths, for example, through separate VLANs.
[0049] As an example, a group of communication connections that enable each monitoring device 212, 222, 232, 242 to transmit information to a specific peer each includes a first communication connection for transmitting heartbeats. For example, each monitoring device 212, 222, 232, 242 can transmit heartbeats via its respective multicast channel. Each monitoring device 212, 222, 232, 242 can also subscribe to the corresponding multicast channel of each peer and receive heartbeats from the peer. A subscription made by monitoring devices 212, 222, 232, 242 to one of its peer's multicast channels can be considered a first communication channel for information transmission from monitoring devices 212, 222, 232, 242 that transmit heartbeats to monitoring devices 212, 222, 232, 242 that subscribe to receive heartbeats. For example, the subscriptions of the second, third, and fourth monitoring devices 222, 232, and 242 to the multicast channel from which the first monitoring device 212 transmits heartbeats can each be considered as first communication connections from the first monitoring device 212. Each multicast channel is constructed over UDP using, for example, Latency Busters Messaging (LBM). For example, each monitoring device 212, 222, 232, and 242 can transmit its own heartbeats using its own LBM topic and can subscribe to peer topics to receive those heartbeats. Each monitoring device 212, 222, 232, and 242 can transmit heartbeats periodically. For example, each monitoring device 212, 222, 232, and 242 can multicast heartbeats at regular intervals, such as 1 minute, 30 seconds, 20 seconds, 10 seconds, or 5 seconds.
[0050] Each monitoring device 212, 222, 232, and 242 is configured to record the fact of receipt and a timestamp of the time of receipt in a specific location within its respective database when it receives a heartbeat from one of its peers. For example, considering the operation of the first monitoring device 212 in this regard, the first monitoring device 212 is configured to receive the second heartbeat from the second monitoring device 222, the third heartbeat from the third monitoring device 232, and the fourth heartbeat from the fourth monitoring device 242 via subscriptions to each multicast channel. When the first monitoring device 212 detects that it has received the second heartbeat from the second monitoring device 222, it records the fact of receipt and a timestamp of the received second heartbeat in the first database 214 at a specific location associated with the second monitoring device 222. The first monitoring device 212 may also store the timestamp of the last heartbeat received from the second monitoring device 222 at that location. Similarly, the first monitoring device 212 can record reception records and reception time stamps for the third heartbeat received from the third monitoring device 232 and the fourth heartbeat received from the fourth monitoring device 242, respectively, in predetermined locations in the first database 214. The second monitoring device 222, the third monitoring device 232, and the fourth monitoring device 242 operate in a similar manner, recording the reception of heartbeats from the peer in their respective databases 222, 234, and 244. As will be described in more detail below, each detector 216, 226, 236, and 246 is configured to determine the state of the first communication connection established by the associated monitoring device with the peer, based on the heartbeat reception records from the peer.
[0051] In some examples, each of a group of communication connections that enables the transmission of information from any of the monitoring devices 212, 222, 232, and 242 to a specific peer also includes a corresponding second communication connection. Each second communication connection allows a specific monitoring device (212, 222, 232, and 242) to inspect a location on a specific peer node (e.g., a location in the database associated with that peer node). For example, a second communication connection from the first monitoring device 212 to the second monitoring device 222 allows the first monitoring device 212 to inspect a location in the second database 224. The second communication connections may be read-only connections. The second communication connections may include, for example, TCP connections between each monitoring device.
[0052] Each monitoring device 212, 222, 232, and 242 is configured to write an indication to its respective databases 214, 224, 234, and 244 indicating whether or not it can inspect a given location in the database of its corresponding peer. For example, the first monitoring device 212 is configured to monitor the status of a second communication connection from the second monitoring device 222 to the first monitoring device 212, which includes attempts to inspect a specific location on the second database 224 (e.g., periodically). This inspection may include read-only operations by the first monitoring device 212. If the first monitoring device 212 can inspect a location in the second database 224, it writes an indication to the location in the first database 214 associated with the second monitoring device 222 that the location can be inspected. However, if the first monitoring device 212 is unable to read a location in the second database 224, the first monitoring device 212 writes an indication to the first database 214 that it could not examine that location. Similarly, the first monitoring device 212 determines whether it can examine locations in the third database 234 and the fourth database 244, and writes information indicating the results of these examination attempts to the respective locations in the first database 214 associated with the third monitoring device 232 and the fourth monitoring device 242. The second, third, and fourth monitoring devices 222, 232, and 242 operate in a similar manner, monitoring and recording whether they can examine the relevant locations in the peer's database. Thus, each monitoring device 212, 222, 232, and 242 records the readability of the peer's locations in the database from its own perspective in the associated databases 214, 224, 234, and 244, respectively.
[0053] For example, databases 214, 224, 234, and 244 each store data at locations identified by their corresponding paths. Similar to a file system, a path to a specific location may consist of a hierarchical sequence of path elements. Therefore, data is written to or read from a specific path within one of the databases 214, 224, 234, or 244. Each monitoring device 212, 222, 232, and 242 can maintain its own set of locations stored in its associated databases 214, 224, 234, and 244. Each monitoring device 212, 222, 232, and 242 can set specific locations in the databases of its corresponding peers as targets for monitoring (watching) and can write to the corresponding locations in its own database that a particular connection is active while monitoring (watching) is not triggered. If a location in the database of one of the monitored peers becomes unreadable, the monitoring device may receive notification through monitoring (watching). The monitoring device then records in a specific set of locations in its database that the connection to that particular peer is not working. Each location where a monitoring device records the reception of a heartbeat from a peer may be part of the same set of locations. For example, the first database 214 may have a first location to record the time of the most recent heartbeat received from the second monitoring device 222, and a second location to record the readability of the second database 224 by the first monitoring device 212, in order to record the connection status from the second monitoring device 222 to the first monitoring device 212. The first database 214 may also have first and second locations to record the status of the third monitoring device 232 and the fourth monitoring device 242, respectively, in relation to the first monitoring device 212. The second, third, and fourth databases 224, 234, and 244 may each have a similar set of location information to record the status of the communication connection between the second, third, and fourth monitoring devices 222, 232, and 242 and their corresponding monitoring devices.
[0054] Information stored in the databases of specific DCs 210, 220, 230, and 240 may be used to determine the status of communication connections between DCs. As will be explained in more detail below, this information may be used to determine the availability of a particular DC. For example, the availability of the first DC 210 to one or more computing entities outside of the first DC 210 (which may include the first computing entity 250 and / or the second computing entity 260) can be determined.
[0055] In one example, the first detector 216 is configured to determine the status of a group of communication connections from the second monitoring device 222 to the first monitoring device 212 based on information stored in the first database 214, wherein the information pertains to the following communication connections: a first communication connection in which the second monitoring device 222 sends a heartbeat to the first monitoring device 212; and a second communication connection in which the first monitoring device 212 can inspect the relevant section in the second database 224.
[0056] With respect to the first communication connection, the first detector 216 is configured to determine that the first communication connection from the second monitoring device 222 is not functioning if the first monitoring device 212 does not receive a heartbeat from the second monitoring device 222 for a predetermined period (sometimes called a timeout threshold). The timeout threshold may be set, for example, as an integer multiple of the expected interval between heartbeat transmissions by the second monitoring device 222. Therefore, if, for example, one, two, three, four, or more consecutive heartbeats are not received from the second monitoring device 222, it may be determined that the timeout threshold has been exceeded. This determination may be based on the time of the most recently received heartbeat and the expected interval between heartbeats. In one example, the timeout threshold that the first detector 216 applies to heartbeats from the second monitoring device 222 is 3 heartbeat periods, and the expected interval between heartbeat transmissions is 30 seconds, so the timeout threshold is 90 seconds. If the reception of heartbeats from the second monitoring device 222 is interrupted for 90 seconds, the first detector 216 determines that the first communication connection is not functioning. Otherwise, if at least one heartbeat is received from the second monitoring device 222 before the timeout period expires, the first detector 216 determines that the first communication connection from the second monitoring device 222 is still operational. In some examples, the result of determining the status of the first communication connection is written to the first database 214.
[0057] The first detector 216 is configured to determine the status of the second communication channel from the second monitoring device 222 to the first monitoring device 212, based on the information stored in the first database 214, to determine whether the first monitoring device 212 can inspect the corresponding location in the second database 224. In some examples, the result of determining the status of the second communication connection is written to the first database 214.
[0058] The detector 216 is configured to determine the overall state of the group of communication channels from the second monitoring device 222 to the first monitoring device 212, using the determined states of the first and second communication channels from the second monitoring device 222 to the first monitoring device 212. That is, if both the first and second communication channels are inoperable, the state of the entire group of communication channels is determined to be inoperable. From the perspective of the first DC 210, this indicates that it can no longer receive communication from the second DC 220. However, if only one of the first or second communication channels is determined to be inoperable, the first detector 216 determines that the state of the entire group of communication channels is operational.
[0059] In the same manner as described above for the communication connection group from the second monitoring device 222 to the first monitoring device 212, the first detector 216 is configured to determine the status of the communication connection group from the third monitoring device 232 to the first monitoring device 212, and the status of the communication connection group from the fourth monitoring device 242 to the first monitoring device 212.
[0060] The first detector 216 is configured to determine the availability of the first DC 210 based on the status determined for each of the communication connection groups from the second, third, and fourth monitoring devices 222, 232, and 242 to the first monitoring device 212. The first detector 216 may determine that the first DC 210 is unavailable if all communication connection groups configured to receive information from equivalent devices are inoperable. Specifically, the first detector 216 determines that the first DC 210 is unavailable if it determines that the following requirements are met: the time elapsed since the first monitoring device 212 received a heartbeat from the second monitoring device 222 exceeds the timeout threshold and the first monitoring device 212 is unable to check the location in the second database 224; the time elapsed since the first monitoring device 212 received a heartbeat from the third monitoring device 232 exceeds the timeout threshold and the first monitoring device 212 is unable to check the location in the third database 234; and the time elapsed since the first monitoring device 212 received a heartbeat from the fourth monitoring device 242 exceeds the timeout threshold and the first monitoring device 212 is unable to check the location in the fourth database 244.
[0061] However, if at least one of the multiple communication connection groups to the second, third, and fourth monitoring devices 222, 232, and 242 remains operational, the first detector 216 determines that the first DC 210 remains available. Therefore, if there is an indication that the first DC 210 maintains communication with at least one of the other DCs 220, 230, and 240, the first DC 210 is determined to be available. The determination by the first detector 216 that the first DC 210 is unavailable may result from a complete interruption of communication from the second, third, and fourth DCs 220, 230, and 240 to the first DC 210. This may occur, for example, due to a software or hardware malfunction in the first DC 210, but nevertheless, the functionality of the first DC 210 (including the first detector 216) remains operational. This state is called "islanding" of the first DC 210. In such a situation, unless some action is taken, the application on the first DC210 may continue to operate without external control. This can be undesirable when the first DC210 is unavailable, for example, as commands from external computing entities, such as the first computing entity 250, may not reach the application running on the first DC210. For example, a scenario may occur where the first DC210 loses communication with the first computing entity 250, but maintains communication with the second computing entity 260. This may be because the second computing entity 260 is located in the same place as the first DC210, allowing it to maintain communication with the first DC210 even when communication between the first DC210 and other external entities is lost. In such a scenario, the application on the first DC210 can continue to operate while maintaining communication with the second computing entity 260, without being controlled by the first computing entity 250. Determining the availability of the first DC210 based on the display of the communication status from the first DC210 to all its peers allows for a robust and reliable availability determination.For example, by determining that the first DC210 is unavailable only when it is determined that communication to all peers DC220, 230, and 240 has been lost, the risk of a false determination that the first DC210 is unavailable can be reduced. For example, the risk of a false determination that the first DC210 is unavailable based on the loss of a specific communication channel can be reduced.
[0062] In response to the first detector 216 determining that the first DC210 is unavailable, various actions may be performed. For example, it may initiate the shutdown of one or more applications running on the first DC210. This may be done to reduce the risk that applications on the first DC210 may continue to run and operate in an undesirable manner after the first DC210 becomes unavailable. According to the specific examples described herein, the detection of the first DC210 being unavailable by the first detector 216 may be automatic and therefore obtained in a quick and efficient manner. This automatic determination that the first DC210 is unavailable makes it possible to notify components on the first DC210 of this fact even in situations where communication between the first DC210 and external entities is interrupted. Therefore, measures can be taken quickly and efficiently on the first DC210 without, for example, needing to obtain instructions from entities outside the first DC210 or requiring manual intervention on the first DC210. This reduces the risk described above, namely the risk that applications may continue to operate uncontrollably on the first DC210. In some cases, as will be discussed later with reference to Figure 3, a notification may be sent to one or more applications (not shown in Figure 2) running on DC210 if DC210 is determined to be unavailable to one or more computing entities. For example, this instruction may cause the application to operate in disaster recovery mode, which may include stopping or preventing the execution of one or more functions of the application.
[0063] Under certain circumstances, the first DC210 may become unavailable. For example, this could occur if the first computing system 211 cannot maintain an operational state due to a loss of power supply to the first DC210 or a hardware failure in the first DC210. In such cases, the first DC210 may not be able to determine its own unavailable state. However, in such cases, since the functions of the first DC210 may cease to work, it may not be necessary for the first DC210 to take measures such as shutting down or correcting its functions.
[0064] Alternatively, in addition to the availability of the first DC210 being determined by the first computing system 211 in the first DC210, the availability of the first DC210 may, exemplarily, be determined in one or more of the other DCs 220, 230, and 240. As such an example, the third computing system 231 in the third DC230 is configured to monitor the availability of the first DC210. In this example, the third detector 236 is configured to monitor the status of each of several communication connections for information transmission from the first monitoring device 212 to the other monitoring devices 222, 232, and 242. If the third detector 236 determines that all communication connections from the first monitoring device 212 to its peer devices are inoperable, the third monitoring device 232 determines that the first DC210 is unavailable.
[0065] More specifically, in this example, the third monitoring device 232 is configured to receive heartbeats from the first monitoring device 212 via a first communication channel through a second network 262. These heartbeats are recorded in the third database 234 as described above, and, similar to the method described above for the first detector 216, the third detector 236 is configured to detect when the first communication channel transmitting heartbeats from the first monitoring device 212 to the third monitoring device 232 has become inoperable by determining when the timeout threshold has been exceeded. The third monitoring device 232 is configured to record the readability of the locations in the first database 214 in the second database 224. This recording is done via a second read-only communication connection that allows the third monitoring device 232 to inspect the locations in the first database 214. The third detector 236 is configured to use this information, which has been stored in the third database 234 by the third monitoring device 232, to determine the overall state of the group of communication channels configured for the third monitoring device 232 to receive information from the first monitoring device 212.
[0066] As described above, the second monitoring device 222 and the fourth monitoring device 242 are also configured to receive heartbeats from the first monitoring device 212 via their respective first communication channels and to record the reception of these heartbeats in their respective databases 224 and 244. Similarly, the second monitoring device 222 and the fourth monitoring device 242 are each configured to record an indication in their respective databases 224 and 244 indicating whether or not their location in the first database 214 can be checked. In order to determine the status of the communication connection groups from the first monitoring device 212 to the second monitoring device 222 and the communication connection groups from the first monitoring device 212 to the fourth monitoring device 242, the third monitoring device 232 is configured to access these records in the second and fourth databases 224 and 244.
[0067] For example, the third monitoring device 232 obtains from the second database 224 the heartbeat records received by the second monitoring device 222 from the first monitoring device 212. The third monitoring device 232 also obtains from the second database 224 whether the location in the first database 214 is readable by the second monitoring device 232. The third monitoring device 232 can obtain from the second monitoring device 222 information regarding its connection group to the first monitoring device 212 and store this in the third database 234. Based on this information, the third detector 236 determines the status of the communication connection group from the first monitoring device 212 to the second monitoring device 222. Similarly, in this example, the third monitoring device 232 obtains from the fourth database 232 the heartbeat records received by the fourth monitoring device 242 from the first monitoring device 212, and a record of whether the location in the first database 214 determined by the fourth monitoring device 232 is readable. The third monitoring device 232 stores this information in the third database 234. Based on this information, the third detector 236 can determine the status of the communication connection group from the first monitoring device 212 to the fourth monitoring device 232.
[0068] Based on the determination of the communication connection status from the first monitoring device 212 to the second, third, and fourth monitoring devices 222, 232, and 242, the third detector 236 determines the availability of the first DC 210. As described above, similar to how the first detector 216 determines the availability of the first DC 210, in this exemplary embodiment, the third detector 236 determines that the first DC 210 is unavailable only if it determines that all communication connections between the first monitoring device 212 and its peers are inoperable.
[0069] In response to the third detector 236 determining that the first DC 210 is unavailable, one or more actions may be performed in the third DC 230. Such actions may be performed, for example, by the third monitoring device 232. For example, such actions may include initiating a failover of the functions of the first DC 210 to a failover DC (e.g., the second DC 220). In one example, the third monitoring device 232 initiates a failover of the functions of the first DC 210 to the second DC 220 by sending a message to the second monitoring device 222. For example, this message notifies the second monitoring device 222 that the first DC 210 has become unavailable and instructs the second monitoring device 222 to take over specific functions that the first DC 210 was performing. This allows the second DC 220 to quickly receive notification that the first DC 210 is unavailable and begin to take over the associated functions from the first DC 210. The functionality initiated on the second DC 220 may include, for example, running one or more applications configured to be controlled by the first computing entity 250 and to communicate with the second computing entity 260, on behalf of the first DC 210. The second monitoring device 222 may, for example, start a version of one or more applications previously running on the first DC 210 in response to receiving a message from the third monitoring device 232. Alternatively, the second monitoring device 222 may modify one or more applications running on the second DC 220 to include one or more functions of applications that were running on the first DC 210. For example, the second monitoring device 222 may start or modify an application running on the second DC 220 in order to communicate with the first and / or second computing entities 250, 260. Thus, an example of a failover DC that takes over the functionality of the first DC determined to be unavailable is described below with reference to Figure 3.
[0070] Instead of automatically failing over the functions of DC1210 to another DC, it is also possible to manually handle the failover from DC1210. For example, the third monitoring device 232 issues an alarm in response to determining that DC1210 is unavailable. Upon receiving such an alert, the user can manually initiate a failover procedure and, for example, fail over the functions of DC1210 to DC220.
[0071] For example, if at least one of the multiple communication connection groups from the first monitoring device 212 to its peers remains operational, the third detector 236 does not determine that the first DC 210 is unavailable. However, if the third detector 236 determines that one or more of the communication connection groups from the first monitoring device 212 to any of its peers are inoperable, measures may be taken. For example, the third detector 236 may issue a warning that communication between the first DC 210 and one or more peers may have been lost.
[0072] In some cases, the third monitoring device 232 may determine that it has lost connectivity with one or more peers based on the status of each of its group of communication connections with peers. In such cases, the third monitoring device 232 may issue an alarm so that the user can investigate the situation and take corrective action.
[0073] In the example above, these techniques are applied to determine the availability of the first DC 210 in system 200. However, it will be understood that similar techniques may be applied additionally or alternatively to determine the availability of individual DCs within system 200. For example, techniques similar to those used by DC 210 to determine its availability to one or more external computing entities may be applied to each of DCs 210, 220, 230, and 240. For example, the second detector 226 of the second DC 220 can determine the availability of the second DC 220 based on the status of each of the multiple communication connection groups from the first, third, and fourth monitoring devices 212, 232, and 242 to the second monitoring device 222. Furthermore, the third detector 236 can determine the availability of the second DC 220 and / or the fourth DC 240 in a similar manner to how the availability of the first DC 210 is determined. For example, if the second DC220 functions as a failover DC for a function performed by the first DC210, the third detector 236 can monitor the availability of the second DC220 in addition to the first DC210. In this way, the third detector 230 can determine that the second DC220 remains available to replace the function of the first DC210 in the event that the first DC210 becomes unavailable. The action taken by the third monitoring device 232 in response to the determination that the first DC210 is unavailable may depend on whether the third detector 236 determines that the second DC220 is available. For example, the third monitoring device 232 can initiate a failover to the second DC220 only if it determines that the second DC220 is available.
[0074] For example, the functional redundancy of the third DC230 described above may be provided by one or more additional DCs, such as the fourth DC240. For example, as described above, the third detector 236 determines the availability of the first DC210 and, in response to determining that the first DC210 is unavailable, performs one or more actions, such as initiating a failover of the functionality of the first DC210 to a failover DC, such as the second DC220. However, in this example, a further DC, such as the fourth DC240, may determine the availability of the first DC210 in a similar manner to the third DC230 described above. For example, the fourth detector 246 and the fourth monitoring device 242 may function to determine the availability of the first DC210 in a similar manner to the third monitoring device 232 and the third detector 236 described above. In such a case, in response to determining that the first DC210 is unavailable, the additional DC may perform one or more actions, such as initiating a failover of the functionality of the first DC210 to a failover DC, such as the second DC220. This redundancy ensures that even if the 3rd DC230 becomes unavailable, the system will detect the unavailability of the 1st DC210 and take appropriate countermeasures, such as initiating a failover from the 1st DC210 to the 2nd DC220.
[0075] For example, a further DC, such as DC4240, can provide the functions of DC3230 in response to its determination that DC3230 has become unavailable. In these cases, the availability determination and failover initiation functions of DC3230 can be failed over to yet another DC, such as DC4240. This ensures that the functions are provided even if DC3230 becomes unavailable. The availability of DC3230 can be determined by any suitable method, including the techniques described above with reference to Figure 2. For example, DC4240 can determine the availability of DC3230 in the same or similar way as DC3230 determines the availability of DC1210 as described above. As another example, DC3230 may also be part of a group of DCs (not shown) including yet another DC (not shown). This other DC (not shown) determines the availability of DC3230 and initiates a failover of the functions of DC3230 described above to DC4240. This other DC determines the availability of DC3230, and DC3230 determines the availability of DC1210 and initiates a failover of DC1210's functions to DC220, in the same manner as described above, and initiates a failover of DC3230's functions to DC4240.
[0076] For example, DC230 (3rd DC) may constitute part of the cloud computing resources provided in the first geographical region, while further DCs, such as DC240 (4th DC), may constitute part of the cloud computing resources provided in a second, different geographical region. For example, the first and second geographical regions may be different parts of the same country or continent, or even different countries or continents. This allows the functionality of DC230 (3rd DC) described above to fail over to a further DC (e.g., DC240) of the cloud computing resources in the second geographical region if the cloud computing resources in the first geographical region become unavailable. Cloud computing resources in the second geographical region may be relatively isolated from factors that could cause the cloud computing resources in the first geographical region to become unavailable, such as regional power outages or communication failures. For example, cloud computing resources provided in different geographical regions may be provided by independent infrastructure. This can improve the reliability of the functionality intended to be provided by DC230 (3rd DC).
[0077] IV. Application Control in Data Centers As described above, in some cases, after it is determined that the first DC210 is unavailable to one or more external computing entities, one or more actions may be performed in the first DC210. An example of such behavior performed in the first DC210 will be described in more detail with reference to Figure 3.
[0078] Figure 3 shows an exemplary system 300. System 300 includes a first data center (DC) 310, a second DC 320, a first computing entity 330, a second computing entity 340, a third computing entity 350, and a database entity 360. The first DC 310 includes a first computing system 311. The first computing system 311 includes a first monitoring device 312, an internal communication network 313, a first application 314, a second application 315, a database 316, and a shared library 317. The first DC 310 may be the first DC 210 described above with reference to Figure 2. The first computing system 311 may be the first computing system 211 of the first DC 210 described above with reference to Figure 2. The first monitoring device 312 may be the first monitoring device 212 of the first DC 210 described above with reference to Figure 2. The database 316 may be the first database 214 described above with reference to Figure 2. The second DC320 includes a second computing system 321, which includes a second monitoring device 322, an internal communication network 323, a third application 324, and a fourth application 325. The second DC320 may be the second DC220 described above with reference to Figure 2. The second computing system 321 may be the second computing system 221 of the second DC220 described above with reference to Figure 2. The second monitoring device 322 may be the second monitoring device 222 of the second DC220 described above with reference to Figure 2.
[0079] The first computing entity 330 communicates with the first DC 310 via the first communication network 361. The first DC 310 communicates with the second computing entity 340 via the second communication network 362. The first DC 310 communicates with the second DC 320 via the third communication network 363. The second DC 320 communicates with the second computing entity 340 via the fourth communication network 364. The second DC 320 communicates with the third computing entity 350 via the fifth communication network 365. The first computing entity 330 communicates with the second DC 320 via the sixth communication network 366. The second DC 320 communicates with the database entity 360 via the seventh communication network 367. For example, one or more of the first to seventh communication networks 361 to 367 may include one or more of the following: local area networks, wide area networks, multicast networks, wireless networks, virtual private networks, internal networks, cellular networks, peer-to-peer networks, connection points, dedicated lines, the Internet, shared memory systems, and / or dedicated networks. One or more of the first to seventh communication networks 361 to 367 may be different from one another, or they may be the same network.
[0080] In a manner similar to that described in Figure 2, the first computing system 311 of the first DC 310 may be implemented by or include one or more of the computing devices 100 described in Figure 1. Either or both of the first application 314 and the second application 315 may be implemented as software that runs on one or more computing devices of the first computing system 311 (for example, the computing devices 100 described above with reference to Figure 1). For example, the first application 314 and the second application 315 may run on different computing devices within the computing system 311 (for example, the computing devices 100 described above with reference to Figure 1). However, this is not necessarily required, and other configurations are possible.
[0081] The first application 314 and the second application 315 are configured to perform one or more functions. For example, these functions may include processing data, routing data to one or more other applications, establishing one or more communication connections with computing entities, and / or sending data to and / or receiving data from computing entities. In this example, the first application 314 is configured to process data and route the processed data to the second application 315 via the internal communication network 313. The second application 315 is configured to establish one or more connections, such as TCP connections, with the second computing entity 340 via the second communication network 362. The second application 315 receives data from the first application 314, processes that data further, and sends the further processed data to the second computing entity 340 via the established connections. The second computing entity 340 can take action based on the data sent from the second application 315. Thus, the functions performed by the first application 314 and / or the second application 315 may trigger actions by the second computing entity 340. In some examples (as shown in Figure 3), the second computing entity 340 may be located outside the first DC 310. However, in other examples (not shown), the second computing entity 340 may be located as part of the first DC 310, i.e., in the same location as the first computing system 311.
[0082] Either or both of the first application 314 and the second application 315 can be remotely controlled by the first computing entity 330. The first computing entity 330 is located outside the first DC 310. Under normal or assumed operating conditions, applications 314 and 315 of the first DC can be remotely controlled by the first computing entity 330. For example, the first computing entity 330 can send commands to one or more applications 314 and 315 via the first communication network 361 to control the operation of applications 314 and 315. This control may include, for example, starting the execution of applications 314 and 315 on the first DC 310, starting, changing, or stopping functions performed by applications 314 and 315, and shutting down applications 314 and 315. For example, it may be impossible, impractical, or otherwise undesirable to store or run applications 314 and 315 on the first computing entity 330. In this case, apps 314 and 315 are instead stored and executed by the first computing system 311 of the first DC 310, but operate under the control of the first computing entity 330. As another example, it may be beneficial or desirable for apps 314 and 315 to run on a computing device located physically close to the second computing entity 340. This minimizes the communication latency between apps 314 and 315 and the computing entity 340. The first computing entity 330 may not be located physically close to the second computing entity 340. However, the first computing system 311 may be located physically close to the second computing entity 340, or in the same location as the second computing entity 340. In this case, apps 314 and 315 are stored and executed by the first computing system 311 of the first DC 310 to minimize the communication latency between apps 314 and 315 and the second computing entity 340.These operate under the control of the first computing entity 330.
[0083] However, as explained in Figure 2, a situation may arise in which the first computing entity 330 becomes unable to use the first DC 310. In such a situation, applications 314 and 315 may become uncontrollable by the first computing entity 330. For example, this phenomenon may occur if the first DC 310 loses its connection to the first communication network 361, or if some disaster or failure occurs, such as a hardware or software failure, and commands from the first computing entity 330 cannot reach applications 314 and 315. If applications 314 and 315 become uncontrollable by the first computing entity 330, they may continue or start execution on the first DC 310 without being controlled by the first computing entity 330. For example, the second application 315 may continue or start sending data to the second computing entity 340 via a communication connection established on the second communication network 362, and the first computing entity 330 may be unable to change or stop this. This can be an undesirable situation. This is because it causes the second computing entity 340 to take actions that it would have avoided if the first computing entity 330 had been controlling applications 314 and 315.
[0084] The first monitoring device 312 is configured to obtain an indication that applications 314 and 315 in the first DC 310 have become uncontrollable by a first computing entity 330 located outside the first DC 310. The first monitoring device 312 can obtain an indication in response to a determination that the first DC 310 has become unavailable to one or more computing entities located outside the first DC 310, for example, following one of the examples described above with reference to Figure 2. For example, the first monitoring device 312 can obtain an indication of a determination from the first detector 216 (not shown in Figure 3) in Figure 2. The first detector 216 can determine, for example, that the first DC 210 is unavailable and store an indication of this determination in a database accessible to the first monitoring device 312 (for example, database 316).
[0085] The first monitoring device 312 is configured to determine that applications 314 and 315 should operate in disaster recovery mode based on an indication that applications 314 and 315 have become uncontrollable by the first computing entity 330. For example, as will be described in more detail below, depending on the applications 314 and 315, operation in disaster recovery mode may include stopping or preventing the execution of one or more functions of applications 314 and 315, thereby mitigating the impact of applications 314 and 315 continuing or starting the execution of one or more functions without external control by the first computing entity 330. Exemplaryly, the determination by the first monitoring device 312 that applications 314 and 315 should operate in disaster recovery mode may be made in response to a determination that applications 314 and 315 have become uncontrollable by the first computing entity 330. This can minimize the time that applications 314 and 315 are operating in normal operating mode but are uncontrollable by the first computing entity 330.
[0086] The first monitoring device 312 is configured to transmit an indication to applications 314 and 315 that they are operating in disaster recovery mode when it is determined that applications 314 and 315 on the first DC 310 are operating in disaster recovery mode. As will be explained in more detail below, the first monitoring device 312 transmits this indication to applications 314 and 315 that are already running on the first DC 310 at the time the determination is made, and to applications 314 and 315 that start running on the first DC 310 after the determination is made. This ensures that applications 314 and 315 on the first DC 310 will operate in disaster recovery mode even if they start running after the determination is made.
[0087] In this example, the first monitoring device 312 is configured to write data to the database 316. The database 316 is stored in the storage device of the first computing system 311. For example, the database 316 may be stored in the memory of a computing device of the first computing system 311 (for example, the memory 114 of the computing device 100 mentioned above, referring to Figure 1). For example, the computing device on which the database 316 is stored may be different from the computing device implementing the first application 314, the second application 315, and / or the first monitoring device 312. However, this is not necessarily required, and other configurations are possible.
[0088] In this example, database 316 stores data in a similar manner to the first database 214 of the first DC 210 described above, at locations identified by their respective paths. Database 316 includes a specific disaster recovery (DR) path that records data indicating whether applications 314 and 316 are operating in disaster recovery mode. For example, the DR path might record "false" when applications 314 and 315 are operating normally (i.e., not in disaster recovery mode) and "true" when applications 314 and 315 are operating in disaster recovery mode. Database 316 is configured to provide a callback to shared library 317 whenever the data recorded in the DR path is updated. Specifically, shared library 317 subscribes to database 316 and receives update information whenever the data on the DR path is updated. Whenever the data recorded in the DR path is updated, database 316 sends a callback to shared library 317 indicating that an update has occurred and the updated data recorded in the DR path.
[0089] In response to the determination in the first DC310 that applications 314 and 315 should operate in disaster recovery mode, the first monitoring device 312 writes data to the database 316 indicating that applications 314 and 315 should operate in disaster recovery mode. Specifically, the first monitoring device 312 accesses the DR path in the database 316 and writes data indicating that applications 314 and 315 should operate in disaster recovery mode. Specifically, the first monitoring device 312 can overwrite the "false" recorded in the database's DR path to "true". The database 316 sends a callback to the shared library 317 to notify it that the DR path has been updated, specifically that the data in the DR path has been updated to "true". Thus, the instruction that applications 314 and 315 should operate in disaster recovery mode is transmitted to the shared library 317.
[0090] The shared library 317 is stored in the storage of the first computing system 311. For example, the shared library 317 may be stored in the memory of a computing device of the first computing system 311 (for example, the memory 114 of the computing device 100 described above, referring to Figure 1). Exemplaryly, the computing device in which the shared library 317 is stored may be different from the computing device implementing the first application 314, the second application 315, the first monitoring device 312, and / or the database 316. However, this is not necessarily required, and other configurations are possible. The shared library 317 is used by applications 314 and 315 in the first DC 310. For example, the library 317 may be a library shared by applications 314 and 315. The shared library 317 is accessible from applications 314 and 315. The shared library 317 is a location in storage configured to be accessed by applications 314 and 315. Applications 314 and 315 may be configured to access shared libraries 314 and 315 when they start running on the first DC 310 and while they are running on the first DC 310.
[0091] Shared library 317 contains disaster recovery (DR) functionality that applications 314 and 315 are configured to call. In one example, applications 314 and 315 are each configured to call the DR functionality in shared library 317 when they start running on the first DC 310. Applications 314 and 315 can also be configured to call the DR functionality in shared library 317 while they are running on the first DC 310 (for example, periodically during execution). Specifically, each application 314 and 315 includes code to call the DR function in shared library 317. For example, the DR function might be "IsDREnabled()". Each application 314 and 315 is configured to operate in normal operating mode if the application 314 and 315 call the DR function and the argument is set to "false", i.e., the DR function is "IsDREnabled(false)". However, if apps 314 and 315 call the DR function and the argument is set to "true," that is, if the DR function is "IsDREnabled(true)," apps 314 and 315 will operate in disaster recovery mode. Shared library 317 may be configured to modify the DR function in response to receiving a callback from database 316. Specifically, when shared library 317 receives a callback from database 316, it may be configured to write the argument of the DR function to "true" as a response indicating that the data in the DR path on database 316 has been updated to "true."
[0092] Therefore, in this example, if the first monitoring device 312 determines that applications 314 and 315 in the first DC 310 should operate in disaster recovery mode, the first monitoring device 312 writes the data recorded in the DR path of the database 316 to indicate "true". As a result, the database 316 sends a callback to the shared library 317, notifying it that the data in the DR path has been updated to "true". As a result, the shared library 317 writes the argument of its DR function to "true". Applications 314 and 315 that are already running when the determination is made, and applications 314 and 315 that will start running after the determination, call the DR function from the shared library 317. The indication that applications 314 and 315 should operate in disaster recovery mode is communicated to applications 314 and 315 that are already running when the determination is made, and to applications 314 and 315 that will start running in the first DC 310 after the determination. The data representing the indication can be written to a location in a storage device (i.e., a shared library) that is configured to be accessed when apps 314 and 315 start running on the first DC, and similarly when apps 314 and 315 are running on the first DC 310. In some configurations, functions (i.e., DR functions) that apps 314 and 315 call from a library (i.e., shared library 317) when they start running on DC 310 and while they are running on DC 310 are modified to represent the indication.
[0093] In this example, the first monitoring device 312 communicates with the first application 314 and the second application 315 via the internal communication network 313. The internal communication network 313 may be, for example, a local area network. The internal communication network 313 may be configured for multicast messaging. For example, the first application 314 and the second application 315 can subscribe to multicast messages that the first monitoring device 312 sends via the internal communication network 313. In this example, the first monitoring device 312 is configured to send a message to one or more applications 314, 315 already running in the data center (DC) that contains data indicating that applications 314, 315 should be running in disaster recovery mode, in response to a determination that applications 314, 315 should be running in disaster recovery mode. For example, this message may be a multicast message. For example, applications 314, 315 may be configured to subscribe to receive multicast messages from the first monitoring device 312 when they start running. Therefore, applications 314 and 315, which are already running at the time the determination is made, receive a multicast message from the first monitoring device instructing them to operate in disaster recovery mode. Consequently, at the time the determination is made, the indication is transmitted to applications 314 and 315, which are already running on DC310. Upon receiving such a message from the first monitoring device 312, applications 314 and 315 are configured to operate in disaster recovery mode.
[0094] For example, the first monitoring device 310 can transmit an indication to applications 314 and 315 already running on the first DC 310 via a message transmitted through the shared library 317 and the internal communication network 313. In this case, applications 314 and 315 may operate in disaster recovery mode depending on which signal they receive first. For example, if applications 314 and 315 call the DR function from the shared library 317 after the DR function has been modified to represent the indication, but before receiving a message representing the indication, applications 314 and 315 may operate in disaster recovery mode in response to the call to the DR function. On the other hand, if applications 314 and 315 receive a message representing the indication between calls to the DR function from the shared library 317, and before the DR function has been modified to represent the indication, applications 314 and 315 may operate in disaster recovery mode in response to the message reception.
[0095] Each application 314, 315 is configured to obtain an indication transmitted by the first computing system 311 of the first DC 310, namely an indication that applications 314, 315 should operate in disaster recovery mode. For example, for applications 314, 315 that start execution after the first monitoring device 312 determines that applications 314, 315 should operate in disaster recovery mode, applications 314, 315 can obtain the indication by accessing the location in the storage device where the data representing the indication is stored and reading the data at the start of execution. For example, when applications 314, 315 start execution, they may access the shared library 317 and call a DR function (with the argument written as "true") from the shared library 317. When the first monitoring device 312 determines that applications 314 and 315 should operate in disaster recovery mode, applications 314 and 315 that are already running can obtain the indication by accessing the corresponding location in the storage device to read the data representing the indication while they are running. For example, applications 314 and 315 may access the shared library 317 while they are running and call the DR function (with the argument written as "true") from the shared library 317. Alternatively, applications 314 and 315 can obtain the indication by receiving a message from the first monitoring device 312 via the internal communication network 313 while they are running.
[0096] Each app 314, 315 is configured to run in disaster recovery mode after receiving an indication that apps 314, 315 should run in disaster recovery mode. For example, depending on the app 314, 315, running in disaster recovery mode may include stopping or preventing the execution of one or more functions of the app.
[0097] For example, as mentioned above, during normal operation, the first application 314 processes data and routes the processed data to the second application 315 via the internal communication network 313. The second application 315 receives the processed data from the first application 314 and can process that data further. The second application 315 establishes one or more communication connections with the second computing entity 340 via the second communication network 362 and sends data to the second computing entity 340 via the communication connections.
[0098] In this example, if the second application 315 is already running on the first DC 310, running the second application 315 in disaster recovery mode may include stopping further data processing by the second application 315 and / or terminating one or more established communication connections. Alternatively, if the execution of the second application 315 is started after the determination, running the second application 315 in disaster recovery mode may include preventing further data processing by the second application 315 and / or preventing the second application 315 from establishing a communication connection to the second computing entity 340. In either case, this stops the second application 315 from sending data to the second computing entity 340 and therefore prevents the second computing entity 340 from performing actions based on data sent by the second application 315 while applications 314 and 315 are out of control by the first computing entity 330. If the second computing entity 340 allows only one set of communication connections at a time, terminating or preventing the communication connection may allow application 325 on the second DC 320 to establish a communication connection with the second computing entity 340. This may facilitate an effective failover of the second application 315 on the first DC 310, as will be discussed later. Furthermore, if additional data processing should only be performed once, stopping or preventing the additional processing by the second application 315 allows application 325 on the second DC 320 to perform that additional processing instead. This may facilitate an effective failover of the second application 315 on the first DC 310, as will be discussed later.
[0099] In this example, if the first application 314 is already running on the first DC 310, running the first application 314 in disaster recovery mode may include stopping the first application 314 from processing data and / or stopping the first application 313 from routing data to the second application 315. Alternatively, if the first application 314 is to start running after the determination, running the first application 314 in disaster recovery mode may include preventing the first application 314 from processing data and / or preventing the first application 314 from routing data to the second application 315. In either case, this prevents the first application 314 from sending data to the second application 315, and as a result, prevents the second application 315 from further processing such data and sending the processed data to the second computing entity 340. Therefore, while apps 314 and 315 are not under the control of the first computing entity 330, actions by the second computing entity 340 based on processing and / or routing performed by the first app 314 are prevented. If data processing and / or routing should only be performed once, stopping or preventing the processing or routing in the first DC 314 allows app 324 in the second DC 320 to perform that processing and / or routing instead. Thus, the second DC 320 can enable effective failover of the first app 314 in the first DC 310, as will be described later.
[0100] As another example, running the first application 314 and / or the second application 315 in disaster recovery mode may include shutting down applications 314 and 315. For example, if applications 314 and 315 are already running on the first DC 310, running the applications in disaster recovery mode may include shutting down applications 314 and 315 to stop their functions from running. If applications 314 and 315 start running on the first DC 310, operating the applications in disaster recovery mode may include shutting down applications 314 and 315 before their functions can run. This prevents any actions from being taken by the second computing entity 340 based on the functions that applications 314 and 315 perform while applications 314 and 315 are out of control by the first computing entity 330. If a function should only be executed once (i.e., by a single application), shutting down apps 314 and 315 allows apps 324 and 325 on the second DC320 to execute the function instead. This can facilitate an effective failover of apps 314 and 315 on the first DC310, as will be discussed later.
[0101] Next, we will discuss the second DC320 described above. The second DC320 may be the second DC220 in Figure 2. The second DC320 may be located in a different city or country than the first DC310. The second DC310 may be physically close to the third computing entity 350. For example, the second computing system 321 may be located in the same place as the third computing entity 350 in the second DC320. The second computing system 321 may be implemented by or include one or more of the computing devices 100 described above with reference to Figure 1. One or more of the third application 324, the fourth application 325, and the second monitoring device 322 may be implemented as software running on one or more computing devices of the second computing system 311 (for example, the computing devices 100 described above with reference to Figure 1). For example, the third application 324, the fourth application 325, and the second monitoring device 322 may run on different computing devices of the second computing system 321 (for example, the computing device 100 mentioned above, referencing Figure 1). However, it will be understood that this is not necessarily required, and other configurations are possible.
[0102] The third application 324 and the fourth application 325 perform functions. In the example, the functions configured to be performed by the third application 324 and the fourth application 325 may be identical or similar to the functions configured to be performed by the first application 314 and the second application 315, respectively. For example, the third application 324 and the fourth application 325 may be configured to perform the same or similar functions for the third computing entity 350, similar to how the first application 314 and the second application 315 are configured to perform functions for the second computing entity 340, respectively. The third application 324 and the fourth application 325 are controllable by one or more computing entities outside the second DC 320. For example, the third application 324 and the fourth application 325 are remotely controllable by the first computing entity 330.
[0103] As will be explained in more detail below, if the first application 314 and the second application 315 become uncontrollable by the first computing entity 330, the third application 324 and / or the fourth application 325 of the second DC 320 may provide failover for the first application 314 and / or the second application 315 of the first DC 310, respectively. In addition to the functions that the third application 324 and / or the fourth application 325 would normally perform, for example, the third application 324 and / or the fourth application 325 may perform the functions of the first application 314 and / or the second application 315, respectively. This may provide failover for the first application 314 and / or the second application 315, respectively. In another example, the third application 324 and / or the fourth application 325 may start running on the second computing system 321 to provide failover for the first application 314 and / or the second application 315, respectively.
[0104] The second monitoring device 322 is configured to obtain an indication that applications 314 and 315 in the first DC 310 have become uncontrollable by the first computing entity 330 located outside the first DC 310. As illustrated in Figure 2, the process of obtaining an indication that applications in the first DC 310 have become uncontrollable by the first computing entity 330 includes the second monitoring device 322 receiving an indication that the first DC 310 is unavailable, for example, in the form of a message from a third DC not shown in Figure 3 (such as the third DC 230 in Figure 2). Alternatively, the method of obtaining information indicating that applications in the first DC 310 have become uncontrollable by the first computing entity 330 may include the second monitoring device 322 determining that the first DC 310 is unavailable, which is also similar to the method described above in the specific example illustrated with reference to Figure 2.
[0105] The second monitoring device 322 is configured to determine, based on the acquired indication, that the second DC 320 should provide failover for the first DC 310. For example, as will be described in more detail below, the third application 324 and / or the fourth application 325 may provide failover for the first application 314 and / or the second application 315, respectively. In the example, the second monitoring device 322's determination that the second DC 320 should provide failover for the first DC 310 may be in response to a determination that applications 314 and 315 in the first DC 310 have become uncontrollable by the first computing entity 330. This makes it possible to minimize the time over which failover is provided. In the example, if the second DC 320 is determined to provide failover for the first DC 310, the second monitoring device 322 may determine that the second DC 320 is the DC designated to provide failover for the first DC 310 (from among several further DCs not shown in Figure 3).
[0106] The second monitoring device 322 is configured to operate the third application 324 and / or the fourth application 325 on the second DC 320 in failover mode in response to the determination that the second DC 320 will provide failover for the first DC 310, thereby providing failover for the first application 314 and / or the second application 315, respectively. This provides failover for the first application 314 and / or the second application 315 on the first DC 310. For example, operating applications 324 and 325 in failover mode may involve communicating an indication to applications 324 and 325 that they are operating in failover mode in order to provide failover for applications 314 and 315 on the first DC 310. For example, this indication is communicated by issuing a command to applications 324 and 325 to operate in failover mode. As another example, this indication may be communicated to apps 324 and 325 by sending a message containing data representing the indication via the internal communication network 323. For example, at the time the determination is made, either or both of apps 324 and 325 may not be running on the second DC 320. In these examples, getting apps 324 and 325 to operate in failover mode may involve issuing a command to start apps 324 and 325 from running. In the example, the command and / or message may include identification of the first DC 310 so that apps 324 and 325 can provide failover.
[0107] The third application 324 and / or the fourth application 325 are configured to operate in failover mode to provide failover for the first application 314 and / or the second application 315, respectively, in response to receiving an indication from the second monitor 322. For example, when the third application 324 operates in failover mode, it may include performing one or more functions of the first application 314. For example, the third application 324 performs data processing and / or routing of processed data, which is originally performed by the first application 314. As another example, when the fourth application 325 operates in failover mode, it may include performing one or more functions of the second application 315. For example, the fourth application 325 can establish a communication connection with the second computing entity 340, which is originally established by the second application 315.
[0108] In some examples, a third application 324 and / or a fourth application 325 may query the database entity 360 to identify information related to the operation of the first application 314 and / or the second application 315, respectively. The third application 324 and / or the fourth application 325 use this information to perform the functions of the first application 314 and / or the second application 315, respectively, and thus provide their failover. For example, the third application 324 and / or the fourth application 325 may query the database entity 360 to identify information about the most recent operation of the first application 314 and / or the second application 315, respectively. For example, as mentioned above, the second application 315 of the first DC may have established multiple communication connections with the second computing entity 340 via the second communication network 362. However, as part of the disaster recovery mode in the first DC 310, these communication connections may have been severed by the second application 315. The fourth application 325 queries the database entity 360 to identify information indicating the communication connections that the second application 315 established (or should have established) with the second computing entity 340, in order to provide failover for the second application 315, and then re-establishes these communication connections with the second computing entity 340 via the fourth communication network 364. As another example, as mentioned above, the first application 314 may process and output data, but as part of the disaster recovery mode in the first DC 310, it may cease to process or output data. The third application 324 can query the database entity 360 to identify information indicating the most recent output of the first application 314, in order to provide failover for the first application 315. This allows the third application 324 to infer the state of the first application 315 at the time of its shutdown and start up in this state. The third application 324 processes data based on this state and therefore continues to perform the functions that the first application 314 would normally provide.
[0109] In some examples, the third application 324 and / or the fourth application 325 may perform, in whole or in part, one or more functions that are also performed by the first application 314 and / or the second application 315. In these examples, running the third application 324 and / or the fourth application 325 in failover mode may include enabling such functions. For example, under normal operation, the first application 314 processes the first data and routes the processed first data to the second application 315, but this may be stopped when operating in disaster recovery mode. Under normal operation, the third application 324 processes the first data in the same way as the first application 314, but does not route the processed first data to the fourth application 325. In these examples, running the third application 324 in failover mode may include activating the routing of processed second data by the third application 324 to the fourth application 325. This allows for particularly rapid failover of the functions of the first application 314 to the third application 324.
[0110] As described above, the operation of the second monitoring device 322 and applications 324 and 325 enables the second DC 320 to provide failover for the first DC 310. Applications 324 and 325 of the second DC 320 are remotely controllable by the first computing entity 330. For example, the second data center 320 can communicate to the first computing entity 330 via the sixth communication network 366 that the second DC 320 (specifically its applications 324 and 325) is providing failover for the first DC 310 (specifically its applications 314 and 315). The second DC 320 is available to the first computing entity 330, and therefore its applications 324 and 325 are remotely controllable by the first computing entity 330. For example, if the first communication network 361 is not functioning, or if the components of the first DC 310 are affected by a failure or disaster, applications 314 and 315 may become uncontrollable by the first computing entity 330, while the second communication network 366 and the second DC 320 may function normally. This allows failover applications 324 and 325 on the second DC 320 to be remotely controlled by the first computing entity 330. Therefore, even if applications 314 and 315 on the first data center 310 become uncontrollable by the first computing entity 330, the functionality of applications 314 and 315 on the first DC 310 can continue, particularly by applications 324 and 325 on the second DC, in a manner controllable by the first computing entity 330.
[0111] V. Variations Referring to Figure 2, in the exemplary system 200 described above, system 200 includes four data centers (DCs) 210, 220, 230, and 240. However, this is not necessarily required, and other exemplary systems operating similarly may include only three DCs or four or more. For example, the fourth DC 240 may be omitted from system 200. In such an example, the availability of the first DC 210 can be determined in the same way as described above, except that it is based on the state of each of the multiple communication connection groups between the first DC 210, the second DC 220, and the third DC 230. In such an example, the first detector 216 may determine that the first DC 210 is unavailable to one or more computing entities outside of the first DC 210 if both the communication connection group from the second monitoring device 222 to the first monitoring device 212 and the communication connection group from the third monitoring device 232 to the first monitoring device 212 are inoperable. Similarly, the third detector 236 may determine that the first DC 210 is unavailable based on the status of the communication connection group from the first monitoring device 212 to the second monitoring device 222 and the communication connection group from the first monitoring device 212 to the third monitoring device 232. The third detector 236 may determine the status of each of these communication connection groups from the first monitoring device 212 in the same manner as described above. In such an example, the third detector 226 may be configured to determine that the first data center 210 is unavailable to one or more computing entities outside the data center 210 if both the communication connection group from the first monitoring device 212 to the second monitoring device 222 and the communication connection group from the first monitoring device 212 to the third monitoring device 232 are inoperable.
[0112] In the exemplary system 200, each of the multiple communication connection groups that enables any of the monitoring devices 212, 222, 232, and 242 to transmit information to other monitoring devices 212, 222, 232, and 242 includes a first communication connection and a second communication connection of different types. However, this is not necessarily the case; for example, as briefly mentioned above, each of the multiple communication connection groups may include only a single communication connection or two or more communication connections. Furthermore, multiple communication connection groups may include the same or different number of communication connections to each other. In addition, a particular communication connection group may include multiple communication connections of the same type. For example, a communication connection group for transmitting information from any of the monitoring devices 212, 222, 232, and 242 to another monitoring device 212, 222, 232, and 242 may include a first communication connection for transmitting a first heartbeat and a second communication connection for transmitting a second heartbeat. In such an example, the first and second communication connections may be transmitted via separate paths, such as separate VLANs.
[0113] In the example system 200, the third detector 236 is configured to determine the availability of the first DC 210, and the third monitoring device 234 is configured to obtain information about the first and second communication connections from the first monitoring device 212 to the second monitoring device 222 from the second database 224, and information about the first and second communication connections from the first monitoring device 212 to the fourth monitoring device 242 from the fourth database 244. The third detector 226 can then determine the overall state of the communication connection group from the first monitoring device 212 to the second monitoring device 222 and the communication connection group from the first monitoring device 212 to the fourth monitoring device 242. However, this is not necessarily required, and instead, one or more states of these communication connection groups may be determined by the second DC 220 or the fourth DC 240 and transmitted to the third monitoring device 232. For example, the second detector 226 may be configured to determine the overall state of the communication connection group from the first monitoring device 212 to the second monitoring device 222. This state can be stored in the second database 224 by, for example, the third monitoring device 232 and accessed by the third computing system 231. Similarly, the fourth detector 246 may be configured to determine the overall state of the communication connection group from the first monitoring device 212 to the fourth monitoring device 242. This state can be stored in the fourth database 244 and accessed by the third computing system (for example, the third monitoring device 232).
[0114] In a specific example of the system 200 described above, the third computing system 231 may function as a data center coordinator by determining the availability of the first data center 210 and / or other data centers 220, 240. However, this is not necessarily required, and for example, the second computing system 221 or the fourth computing system 241 could instead function as such a coordinator by monitoring the availability of the first data center 210 and taking measures such as initiating a failover to another data center if the first data center 210 becomes unavailable. For example, as described above, the fourth computing system 241 may function as such a coordinator by failing over the coordinator function of the third computing system 231 belonging to the third data center 230 if it determines that the third data center 230 has become unavailable. In other examples, none of the DC220, 230, and 240 may take on such a coordinating role. Instead, for example, the DC220, 230, and 240 may communicate with each other to determine and agree on the availability of the first DC210 and decide on countermeasures to take if the first DC210 becomes unavailable.
[0115] In the example system 200, the specific function of determining the availability of a data center is performed by specific components of computing systems 211, 221, 231, and 241. For example, monitoring devices 212, 222, 232, and 242 maintain their respective sets of communication connections with one another, and detectors 216, 226, 236, and 246 are configured to perform the determination of the availability of a particular data center. However, this is not necessarily the case, and more generally, the function of determining the availability of a data center, and the function of taking some countermeasure if, for example, the data center is determined to be unavailable, may be performed by computing systems 211, 221, 231, and 241, or any of their components.
[0116] Referring to Figure 3, in the exemplary system 300 described above, the first DC 310 has two applications 314 and 315. However, this is not necessarily the case, and it goes without saying that the first DC 310 can have any number of applications, i.e., one or more applications. In the exemplary system 300, applications 314 and 315 are remotely controllable by the first computing entity 330. However, it will be understood that one or more applications are remotely controllable by one or more computing entities outside the first DC 310 (the first computing entity 330 being one example). In the exemplary system 300, the first monitor 312 receives an indication that remote control of one or more applications 314 and 315 by one or more external computing entities 330 has become impossible, determines that applications 314 and 315 should operate in disaster recovery mode, and transmits an indication to applications 314 and 315 that they should operate in disaster recovery mode. However, this is not necessarily the case, and more generally, it goes without saying that computing system 311 or any of its components may perform these functions.
[0117] In the example system 300, indications are transmitted to applications 314 and 316 via messaging through the internal communication networks 314 and 315 via the shared library 317. However, this is not necessarily required, and it will be understood that in other examples, either one or the other could be used. Furthermore, in some examples, other methods may be used to transmit indications to applications 314 and 315 instead. For example, applications 314 and 315, rather than the shared library 317, could be configured to receive callbacks from DB 316 (specifically the DR path of DB 316). In this case, when the first monitoring device 312 writes data representing an instruction to the DR path of DB 316, DB 316 may send a callback to each of applications 314 and 315. Applications 314 and 315 can each be configured to operate in disaster recovery mode in response to receiving a callback indicating that applications 314 and 315 should operate in disaster recovery mode. As another example, each application 314, 315 may be configured to reference the DR path of DB316, for example, when it starts running on the first DC310 or while the application is running on the first DC310 (for example, periodically). In this case, the DR path of DB316 is an example of a location in storage that can be configured for applications to access. In this case, each application 314, 315 may be configured to access this location in memory and read the data written there. If the data indicates that applications 314, 315 are running in DR mode, then applications 314, 315 will run in DR mode. Other configurations are also possible.
[0118] In the example system 300, the second DC 320 has two apps 324 and 325. However, this is not necessarily the case, and it goes without saying that the second DC 320 can have any number of apps, i.e., one or more apps 324 and 325. Furthermore, in the example system 300, it is explained that the third app 324 provides failover for the first app 315 and the fourth app 325 provides failover for the second app 325, but this is not necessarily the case, and more generally, one or more apps 324 and 325 can provide failover for one or more apps 314 and 315 of the first DC 310. In the example system 300, the second monitoring device 322 receives an indication that one or more applications 314, 315 have become inoperable remotely by one or more external computing entities 330, determines that the second DC 320 should provide failover for the first DC 310, and puts one or more applications 324, 325 into failover mode. However, this is not necessarily required, and more generally, the computing system 321 or any of its components may perform these functions.
[0119] VI. Exemplary Methods Figure 4 is a flowchart of a method according to one embodiment of the present disclosure. This method may be performed by a computing system, such as the first computing system 211 shown in Figure 2, in accordance with any of the examples described above with reference to Figures 1 to 3. Exemplary, the computing system includes a processor (e.g., processor 112 in Figure 1) and memory (e.g., memory 114 in Figure 1). This method may be performed by the processor. The processor may be configured to perform this method. When executed, the memory can store instructions that cause at least one processor of the computing system (e.g., processor 112 in Figure 1) to perform this method or the functions defined by this method. For example, the memory may store an application containing instructions (e.g., application 130 in Figure 1). This method is for determining whether a data center is available to one or more computing entities located outside the data center. For example, this method can be applied to determine whether the first DC210 shown in Figure 2 is available to one or more computing entities outside of the first DC210, such as the first computing entity 250, the second computing entity 260, and one or more of the second, third, and fourth DC220, 230, and 240. This method can be run in the first data center or in another location, such as the third DC230 shown in Figure 2.
[0120] This method includes, in block 402, determining a first state of a first group of communication connections, which includes one or more communication connections between a first data center and a second data center, using a computing system. For example, when this method is performed in the first data center, it may include determining the state of a group of communication connections for information transmission from the second data center to the first data center, as described above. When this method is performed in a data center other than the first data center, it may include determining the state of a group of communication connections for information transmission from the first data center to other data centers, as described above. For example, when this method is performed in the second data center, it may include determining the state of a group of communication connections for information transmission from the first data center to the second data center.
[0121] The determination of the first state of the first group of communications in block 402 can be performed, for example, based on the state of each of one or more communications in the first group of communications, as described above. For example, as described above, the first group of communications includes two or more communications, and the first state may be determined to be inoperable depending on whether each of the two or more communications in the first group of communications is determined to be inoperable. For example, the first communications may include a first communications connection for receiving heartbeat messages from another data center, as described above. For example, if this method is performed by the first data center, the first group of communications may include a first communications connection for receiving heartbeat messages from the second data center to the first data center, as described above. For example, if this method is performed by the second data center, the first group of communications may include a first communications connection for receiving heartbeat messages from the first data center to the second data center. The heartbeat is transmitted over a multicast channel, and the first communications connection appears to include a subscription to the multicast channel. The first group of communications may include a second communications connection for inspecting a location in another data center, as described above. For example, if the method is performed by the first data center, the second communication connection is to enable the first data center to inspect locations within the second data center. Conversely, if the method is performed by the second data center, the second communication connection is to enable the second data center to inspect locations within the first data center. As described above, the second communication connection can include a TCP connection.
[0122] This method, in block 404, includes the computing system determining a second state of a second group of communication connections, which includes one or more communication connections between a first data center and a third data center different from the second data center. For example, when this method is performed in the first data center, it may include determining the state of a group of communication connections for information transmission from the third data center to the first data center, as described above. For example, when this method is performed in a data center other than the first data center (e.g., the second data center), it may include determining the state of a group of communication connections for information transmission from the first data center to the third data center, as described above. The determination of the second state may include, for example, any of the features described above with respect to the determination of the first state in block 402.
[0123] As described above, when this method is performed in a data center other than the first data center, determining the state of at least one of several communication connection groups may include, for example, obtaining information about one or more communication connections of that communication connection group from a data center connected to the first data center by one or more communication connections of that communication connection group. For example, when this method is performed in the second data center, block 404 may include obtaining information about one or more communication connections between the third data center and the first data center from the third data center at the second data center.
[0124] This method, in block 406, includes the computing system determining, based on a first state and a second state, an indication that the first data center is available to one or more computing entities outside the first data center. For example, this may include determining that the first data center is unavailable to one or more computing entities outside the first data center, based on a first state and a second state indicating that the first and second sets of communication connections are inoperable, as described above. As described above, in some examples, this method may include determining further states of one or more sets of further communication connections among a plurality of sets of communication connections between the first data center and one or more further data centers. The determination of the availability of the first data center may also be based on these further states. For example, the determination of the availability of the first data center may be further based on a fourth state of a fourth set of communication connections between the first data center and a fourth data center.
[0125] As explained above, this method may include taking action in response to a determination of the availability of the first data center. For example, as described above, if blocks 402 to 406 are running in the first data center and the first data center is determined to be unavailable, this method may include initiating one or more of the following: a shutdown process to shut down one or more applications running in the first data center, and a disaster recovery process to run one or more applications running in the data center in disaster recovery mode. As another example, as mentioned above, if blocks 402 to 406 are running in a data center other than the first data center, this method may include initiating a failover process to switch the functions of the first data center to the failover data center. As described above, the failover data center may be, for example, the third data center.
[0126] VII. Examples of Electronic Trading Systems Figure 5 is a block diagram showing an example of an electronic trading system 500 in which a particular embodiment may be employed. The system 500 includes a trading device 510, a gateway 520, and an exchange 530. The trading device 510 is communicable with the gateway 520. The gateway 520 is communicable with the exchange 530. In this specification, the expression “in a state of communication” includes direct communication and / or indirect communication through one or more intermediary components. The trading device 510, the gateway 520, and / or the exchange 530 may include one or more computing devices 100 as shown in Figure 1. The exemplary electronic trading system 500 shown in Figure 5 is communicable with additional components, subsystems, and elements to provide additional functionality and capabilities without departing from the scope of the teachings and disclosures described herein.
[0127] During operation, the trading device 510 may receive market data from the exchange 530 via the gateway 520. The trading device 510 may send messages to the exchange 530 via the gateway 520. Users can use the trading device 510 to monitor market data and / or decide to send order messages to the exchange 530 to buy or sell one or more tradable objects. The trading device 510 can perform trading actions using the market data, such as sending order messages to the exchange 530. For example, the trading device may or may not require user input to perform trading actions.
[0128] Market data may include market data about tradable assets. For example, market data may include inside market, market depth, last trading price (LTP), last trading quantity (LTQ), or a combination of these. Inside market refers to the market for a tradable asset at a specific point in time, including the best bid and best ask or offer (as the inside market can fluctuate over time). Market depth refers to the tradable quantity at price levels that include the inside market and price levels that are outside the inside market. Market depth can result in "gaps" at prices where there are no orders for quantity.
[0129] Inside markets and price levels associated with market depth can include not only prices but also derived and / or calculated representations of value. For example, a value level may be presented as the net change from the opening price. Another example is that a value level may be provided as a value calculated from prices in two other markets. Yet another example is that a value level may include an integrated price level.
[0130] Tradable assets refer to anything that can be traded. For example, a certain quantity of tradable assets may be bought and sold at a certain price. Tradable assets include, for example, financial instruments, stocks, options, bonds, futures contracts, currencies, warrants, funds, derivatives, securities, commodities, swaps, interest rate products, index-linked products, tradable events, commodities, or combinations thereof. Tradable assets may include commodities listed and / or managed by an exchange, user-defined commodities, combinations of physical and synthetic commodities, or combinations thereof. Synthetic tradable assets may exist that correspond to or are similar to real-world tradable assets.
[0131] An order message is a message that includes a trading order. A trading order refers to, for example, a command to place a buy or sell order for a tradable asset, a command to initiate order management based on a defined trading strategy, a command to change, modify or cancel an order, instructions to an electronic exchange related to an order, or a combination thereof.
[0132] The trading device 510 may include one or more electronic computing platforms. For example, the trading device 510 may include a desktop computer, a mobile terminal, a laptop computer, a server, a portable computing device, a trading terminal, an embedded trading system, a workstation, an algorithmic trading system such as a "black box" or "gray box" system, a computer cluster, or a combination thereof. As another example, the trading device 510 may include a single-core or multi-core processor capable of communicating with memory or other storage media configured to store one or more computer programs, applications, libraries, computer-readable instructions, etc., in a form accessible for execution by the processor.
[0133] For example, trading device 510 includes computing devices such as personal computers and mobile devices that communicate with one or more servers, and these computing devices and one or more servers collectively constitute trading device 510. For example, trading device 510 may consist of computing devices and one or more servers that jointly run the TT® platform, an electronic trading platform provided by Trading Technologies International, Inc. ("Trading Technologies") located in Chicago, Illinois. For example, one or more servers may run a part of the TT platform, such as a web server, and the computing devices may run another part of the TT platform, such as a part that provides user interface functions on a web browser. The computing devices and servers can communicate with each other, for example, using browser session requests and responses or WebSockets, to implement the TT platform. As another example, trading device 510 includes computing devices such as personal computers and mobile devices, which may be running applications such as TT Desktop or TT Mobile. These are all electronic trading applications provided by Trading Technologies. As another example, trading device 510 may be one or more servers running trading tools such as ADL®, AUTOSPREADER®, AUTOTRADER®, and / or MD-TRADER®, provided by Trading Technologies Inc.
[0134] The trading device 510 may be controlled or otherwise used by a user. In this specification, the term “user” includes, but is not limited to, a person (e.g., a trader), a trading group (e.g., a group of traders), or an electronic trading device (e.g., an algorithmic trading system). One or more users may be involved in controlling or otherwise using the trading device.
[0135] The trading device 510 may include one or more trading applications. In this specification, a trading application means an application that facilitates or improves electronic trading. A trading application provides one or more electronic trading tools. For example, a trading application stored in the trading device may run to organize and display market data in one or more trading windows. Another example is a trading application that may include an automated spread trading application that provides spread trading tools. Another example is a trading application that may include an algorithmic trading application that automatically processes algorithms and performs specific actions such as issuing orders, modifying existing orders, and deleting orders. Another example is a trading application that may provide one or more trading screens. A trading screen may provide one or more trading tools that enable interaction with one or more markets. For example, trading tools may allow users to retrieve and view market data, set order entry parameters, send order messages to exchanges, deploy trading algorithms, and / or monitor positions while executing various trading strategies. The electronic trading tools provided by a trading application may be always available, or they may only be available in specific configurations or operating modes of the trading application.
[0136] Transaction applications can be implemented using computer-readable instructions stored on computer-readable media and executable by a processor. Computer-readable media may include various types of volatile and non-volatile storage media. These include, for example, random-access memory, read-only memory, programmable read-only memory, electrically programmable read-only memory, electrically erasable read-only memory, flash memory, any combination thereof, or other tangible data storage devices. In this specification, the term “non-transient or tangible computer-readable media” is expressly defined as including all types of computer-readable storage media, excluding propagated signals.
[0137] One or more components or modules of a trading application may be loaded onto the computer-readable media of the trading device 510 from another computer-readable medium. For example, a trading application (or an update to a trading application) may be stored by the manufacturer, developer, or publisher on one or more CDs, DVDs, or USB drives and then loaded onto the trading device 510, or loaded onto a server from which the trading device 510 retrieves the trading application. As another example, the trading device 510 may receive a trading application (or an update to a trading application) from a server, for example, via the internet or an internal network. The trading device 510 may receive a trading application or update upon request from the trading device 510 (e.g., "pull delivery") and / or without request from the trading device 510 (e.g., "push delivery").
[0138] The trading device 510 may be configured to send order messages. For example, an order message may be sent to the exchange 530 via the gateway 520. As another example, the trading device 510 may be configured to send order messages to a simulated exchange in a simulated environment that does not perform real-world trading.
[0139] Order messages may be sent at the user's request. For example, a trader can send an order message using the trading device 510, or manually enter one or more parameters for a trade order (e.g., order price and / or quantity). As another example, an automated trading tool provided by a trading application may calculate one or more parameters for a trade order and automatically send an order message. In some cases, the automated trading tool may prepare the order message to send but not actually send it without user approval.
[0140] Order messages may be transmitted in one or more data packets or via a shared memory system. For example, order messages may be transmitted from the trading unit 510 to the exchange 530 via the gateway 520. The trading unit 510 can communicate with the gateway 520 using a local area network, wide area network, multicast network, wireless network, virtual private network, internal network, cellular network, peer-to-peer network, point of presence, dedicated line, internet, shared memory system, and / or dedicated network.
[0141] Gateway 520 may include one or more electronic computing platforms. For example, Gateway 520 may be implemented as one or more desktop computers, handheld devices, laptops, servers, portable computing devices, trading terminals, embedded trading systems, workstations with single-core or multi-core processors, algorithmic trading systems such as "black box" or "gray box" systems, computer clusters, or any combination thereof.
[0142] The gateway 520 facilitates communication. For example, the gateway 520 can perform protocol conversion on data communicated between the trading device 510 and the exchange 530. The gateway 520 can, for example, convert order messages received from the trading device 510 into a data format understandable to the exchange 530. Similarly, the gateway 520 can, for example, convert market data in an exchange-specific format received from the exchange 530 into a format understandable to the trading device 510. As will be discussed later with reference to Figure 6, in some embodiments, the gateway 520 can communicate with a cloud service, and the cloud service may support the functions of the gateway 520 and / or the trading device 510.
[0143] Gateway 520 may include trading applications similar to those described above that facilitate or improve electronic trading. For example, Gateway 520 may include a trading application that tracks orders from trading instruments 510 and updates the order status based on order confirmations received from the exchange 530. Alternatively, Gateway 520 may include a trading application that integrates market data from the exchange 530 and provides it to trading instruments 510. Alternatively, Gateway 520 may include a trading application that provides risk handling, calculates implied inequality, handles order processing, handles market data processing, or a combination thereof.
[0144] In certain embodiments, the gateway 520 communicates with the exchange 530 using a local area network, a wide area network, a multicast network, a wireless network, a virtual private network, an internal network, a cellular network, a peer-to-peer network, a point of presence, a dedicated line, the internet, a shared memory system, and / or a dedicated network.
[0145] Exchange 530 may be owned, operated, managed, or used by an exchange entity. Examples of exchange entities include the CME Group, the Chicago Board Options Exchange, the Intercontinental Exchange, and the Singapore Exchange. Exchange 530 is an electronic exchange, including an electronic trading system (e.g., computers, servers, and other computing devices), configured to enable the buying and selling of tradable assets that are offered for trading by the exchange. Exchange 530 may include separate entities, for example, an entity that lists and / or manages tradable assets and an entity that receives and matches orders. Exchange 530 may include, for example, an electronic communications network ("ECN").
[0146] The exchange 530 is configured to receive order messages and match opposing trade orders for the buy and sell of tradable assets. Unmatched trade orders may be posted for trading by the exchange 530. Once a buy or sell order is accepted and confirmed by the exchange, it remains a valid order until it is executed or canceled. If only a portion of the order quantity is executed, the partially executed order remains a valid order. Trade orders may include trade orders received from, for example, trading devices 510 or other devices that communicate with the exchange 530. Typically, for example, the exchange 530 communicates with various other trading devices (similar to trading device 510) that provide and match trade orders.
[0147] The exchange 530 is configured to provide market data. This market data is provided in one or more messages or data packets, or through a shared memory system. For example, the exchange 530 may publish a data feed to subscribed devices such as the trading device 510 or the gateway 520. This data feed may contain market data.
[0148] System 500 may include additional, different, or fewer components. For example, System 500 may include multiple trading devices, gateways, and / or exchanges. As another example, System 500 may include middleware, firewalls, hubs, switches, routers, servers, exchange-specific communication equipment, modems, security management devices, and / or other communication devices such as encryption / decryption devices.
[0149] Exemplary, gateway 520 may be provided by or include first DC210 in any of the examples described above with reference to Figures 1 to 4. For example, gateway 520 may be configured to provide the same or similar functionality as first DC210, more specifically first computing system 211 of first DC210, according to any of the examples described above with reference to Figures 1 to 4. In some examples, gateway 520 may be provided by or include any of the other data centers (DCs) in the examples described above with reference to Figures 1 to 4, such as second data center 220 in Figure 2 or second data center 320 in Figure 3. For example, if first DC210 is operating normally, gateway 520 may be provided by or include first DC210. However, if the first data center 210 becomes unavailable to the first computing entity 250, the first data center 210 may fail over to the second data center 220, at which point the gateway 520 will be provided by or include the second data center 220 instead. The exchange 530 may be provided by the second computing entity 260, for example, according to one of the examples described above with reference to Figure 2, and / or by the second computing entity 340 described above with reference to Figure 3.
[0150] In the example, Gateway 520 implements an application that can be controlled by one or more computing entities outside of Gateway 520. For example, as will be discussed later with reference to Figure 6, Gateway 520 may include a strategy engine application, a risk server application, and an order connector application. In the example where Gateway 520 is provided by the first DC 210 as described above, if Gateway 520 becomes unavailable to one or more computing entities outside of Gateway 520, and therefore the application is determined to be uncontrollable by one or more external computing entities, Gateway 520 can be configured to run the application in disaster recovery mode as described above. In the example where the second data center 220 places Gateway 520 to provide failover for the first data center 210, as described above, if the application in the first DC 210 is determined to be uncontrollable by one or more external computing entities, Gateway 520 runs the application in failover mode to provide failover for the application in the first DC 210.
[0151] VIII. Specific Electronic Trading Systems Figure 6 is a block diagram showing an example of an electronic trading system 600 in which a particular embodiment may be employed. The electronic trading system 600 includes a trading device 610, a hybrid cloud system 620, and an exchange 630. The trading device 610 may be the same as or similar to the trading device 510 described above with reference to Figure 5. The exchange 630 may be the same as or similar to the exchange 530 described above with reference to Figure 5. The hybrid cloud system 620, or one or more of its components, may provide one or more functions of the gateway 520 described above with reference to Figure 5. That is, the functions of the gateway 520 described above with reference to Figure 5, or some or more of its functions, may be included in the hybrid cloud system 620.
[0152] The hybrid cloud system 620 includes a cloud service 640 and a data center 660. In the example shown in Figure 6, the cloud service 640 and its components are separate from the data center 660. However, in other examples (not shown), some or all of the components and / or functions of the cloud service 640 may be implemented in the data center 660 instead. In such examples and other cases, the electronic trading system 600 may not include the cloud service 640. In such examples, one or more of the functions of the gateway 520 described above with reference to Figure 5 may be provided by the data center 660 or one or more of its components alone.
[0153] To provide low latency for time-dependent processes, the data center 660 may be located within the same facility as the exchange 630 or in its vicinity. Therefore, functions of the hybrid cloud system 620 that are time-constrained or where lower latency with respect to the exchange 630 is beneficial may be performed by the data center 660. Generally, functions of the hybrid cloud system 620 that are not time-constrained and do not benefit from low latency with respect to the exchange 630 may be performed by the cloud service 640. The hybrid cloud system 620 enables the electronic trading system 600 to provide relatively low latency with respect to the exchange 630 while maintaining scalability for time-unconstrained functions.
[0154] In the example in Figure 6, the trading device 610 communicates with the cloud service 640 via the first network 671. For example, the first network 671 may be a wide-area network such as the internet using a Hypertext Transfer Protocol (HTTP) connection. The trading device 610 communicates with the data center 660 via the second network 672. For example, the trading device 610 can communicate with the data center 660 via a virtual private network or using a secure WebSocket or TCP connection. The first network 671 and the second network 672 may be the same network. The data center 660 communicates with the cloud service 640 via the third network 673. For example, the data center 660 can communicate with the cloud service 640 via a private network or a virtual private network (VPN) tunnel. The third network 673 may be the same as the first network 671 and / or the second network 672. The data center 660 communicates with the exchange 630 via the fourth network 674. For example, data center 660 can communicate with exchange 630 using a local area network, wide area network, multicast network, wireless network, virtual private network, internal network, cellular network, peer-to-peer network, point of presence, dedicated line, internet, shared memory system, and / or dedicated network. The fourth network 674 may be the same as the first network 671, the second network 672, and / or the third network 673.
[0155] The cloud service 640 may be implemented as a virtual private cloud, which may be provided by a logically separated section of the overall web service cloud. In this example, the cloud service 640 includes a web database 641 and associated web server 642, a product database 643 and associated product data server (PDS) 644, a user-configured database 645 and associated user-configured server 646, and a transaction database 647 and associated transaction server 648.
[0156] The trading device 610 can communicate with the web server 642. For example, the trading device 610 may run a web browser, referred to herein as a browser, and establish a browsing session with the web server 642. This may occur after appropriate domain name resolution to the IP address of the cloud service 640, and / or after appropriate authentication between the trading device 610 (or its user) and the cloud service 640. The browser sends requests to the web server 642, and the web server 642 provides responses to the browser. This is done, for example, using the Hypertext Transfer Protocol (HTTP) or Secure Hypertext Transfer Protocol (HTTPS) protocol. The web server 642 provides a user interface to the browser, allowing the user to interact with the electronic trading platform. This user interface enables the display of market data and / or the issuance of trading orders. Alternatively, the trading device 610 may run an application that communicates with the web server 642 via an Application Programming Interface (API), etc., allowing the user to interact with the electronic trading platform. The application may provide a user interface that allows the user to interact with the electronic trading platform.
[0157] The trading device 610 can communicate with the PDS 644. The PDS 644 interfaces with the product DB 643. The product DB 643 stores definitions of products and user permissions for products. Specifically, the product DB 643 stores definitions of tradable objects and the permissions for users to place trading orders on tradable objects. This information may be provided to the trading device 610. This information is used by the user interface of the trading device 610 to determine which tradable objects a given user of the trading device 610 is permitted to place trading orders for.
[0158] The trading device 610 can communicate with the user configuration server 646. The user configuration server 646 interfaces with the user configuration database 645, which stores user settings, preferences, and other information related to the user's account. This information may be provided from the trading device 610 to the user configuration server 646 at the time of user registration or at certain times thereafter, and the user configuration server 646 may store this information in the user configuration database 645. This information may also be provided to the trading device 610. This information may be used by the user interface of the trading device 610 to determine which market data to display and in what format.
[0159] The transaction database 647 stores information about transactions executed using the electronic trading system 600. The transaction database 647 can store all trading orders submitted by users and all corresponding order execution reports provided by the exchange 630 when trading orders are executed. The transaction server 648 can query the transaction database 647 to generate an audit trail 649 for a specific user, for example. This audit trail 649 may be provided to the trading device 610 (or other device) to enable inspection and / or analysis of the trading activity of a specific user.
[0160] Data center 660 includes a multicast bus 661, a pricing server 662, an edge server 663, a risk server 664, a ledger uploader server 665, an order connector 666, and a strategy engine server 667. Various components within data center 660 communicate with each other via multicast bus 661. This enables efficient and scalable communication between components within data center 660. For example, information provided by one component may be received by several other components. By transmitting this information over multicast bus 661, which other components subscribe to, it becomes possible to transmit the information in a single message, regardless of the number of components receiving the information.
[0161] The price server 662 receives market data from the exchange 630. The price server 662 converts this information into a format and / or syntax associated with (e.g., used by) the electronic trading system 600. The price server 662 transmits the converted information as one or more multicast messages on the multicast bus 661. Specifically, the price server 662 multicasts this information on the first multicast bus A so that price clients can receive it. The edge server 663 and the strategy engine server 667 subscribe to the first multicast bus A and receive market data from the price server 662. The price server 662 can communicate with the cloud service 640. For example, the price server 662 can provide the PDS server 644 with information about products or tradable objects for use when the PDS server 644 defines tradable objects.
[0162] The edge server 663 communicates with the trading device 610. For example, the trading device 610 can communicate with the edge server 663 via a secure WebSocket or TCP connection. In some examples, the edge server 663 may be implemented as a server cluster. The number of servers in the cluster is determined and scaled as needed depending on usage. The edge server 663 receives market data via the first multicast bus A and routes that market data to the trading device 610. Users of the trading device 610 may decide to place trading orders based on the market data. The edge server 663 routes trading orders from the trading device 610 to the exchange 630. Specifically, when the edge server 663 receives an order message from the trading device 610, the edge server 663 multicasts the order message (or at least part of its contents) on the second multicast bus B so that it can be received by the order client. The risk server 664 subscribes to the second multicast bus B and receives order messages from the edge server 663.
[0163] Risk server 664 is used to determine the pre-trading risk for a given trading order contained in a given order message. For example, for a given trading order, risk server 664 can determine whether the user issuing the trading order is authorized to execute it. Risk server 664 determines whether the user is authorized to trade the quantity of tradable objects specified in the trading order. Risk server 664 prevents the issuance of fraudulent trading orders. Risk server 664 receives order messages from edge server 663 via the second multicast bus B and processes the order messages to determine the risk to the trading order in the message. If risk server 664 determines that a trading order should not be placed (for example, if the risk associated with the trading order exceeds a threshold), risk server 664 prevents the trading order from being placed. For example, in this case, risk server 664 may not send the order message to order connector 666, but instead send a message to the user notifying them that the trading order was not placed. If the risk server 664 determines that a trading order should be placed (for example, if the risk associated with the trading order falls below a threshold), the risk server 664 forwards the order message to the order connector 666. Specifically, the risk server 664 multicasts the order message over the second multicast bus B. The order connector 666 and the ledger uploader 665 subscribe to the second multicast bus B and receive the order message from the risk server 664.
[0164] The ledger uploader server 665 communicates with the transaction database 647 of the cloud service 640. The ledger uploader server 665 receives an order message from the risk server 664 and sends the order message to the transaction database 647. The transaction database 647 saves the order message (or at least part of its contents) to the ledger stored in the transaction database 647.
[0165] The order connector 666 communicates with the exchange 630. The order connector 666 receives order messages from the risk server 664, processes the order messages for transmission to the exchange 630, and sends the processed order messages to the exchange 630. Specifically, the processing includes converting the order messages into a data format that the exchange 630 can understand. If the trading order in the order message is executed by the exchange 630, the exchange 630 sends a corresponding execution report message to the order connector 666. The execution report message includes an execution report detailing the execution of the trading order. The order connector 666 applies processing to the execution report message. Specifically, this processing includes converting the execution report message into a data format that the electronic trading system and trading device 610 can understand. The order connector 666 multicasts the processed execution report message on the third multicast bus C so that execution report clients can receive it. The edge server 663 and the ledger uploader 665 subscribe to the third multicast bus C and receive the processed execution report message. The ledger uploader 665 communicates with the trading database 647 and updates the ledger with an execution report message (or at least part of its contents). The edge server 663 forwards the execution report message to the trading device 610. The trading device 610 can display information based on the execution report message to indicate that the trading order has been executed.
[0166] In some cases, order messages may be sent by the strategy engine server 667. For example, the strategy engine server 667 may implement one or more strategy engines using an algorithmic strategy engine and / or an autospreader strategy engine. The strategy engine 667 receives market data (received from the price server 662 via the first multicast bus A) and automatically generates order messages based on the market data and appropriately configured algorithms. The strategy engine server 667 may send order messages to the order connector 666 (via the risk server 664 and the second multicast bus B). The order connector 666 processes the order messages in the same manner as described above. Similarly, when the exchange 630 executes an order, the strategy engine 667 may receive a corresponding order execution report message from the order connector 666 (via the third multicast bus C). The order messages and execution report messages are sent to the ledger uploader 665 in the same manner as described above, which updates the ledger uploader 665, which is maintained by the trading database 647.
[0167] In some cases, trading orders sent by the trading device 610 may not be submitted by a human. For example, the trading device 610 may be a computing device implementing an algorithmic trading application. In these cases, the trading device 610 may not communicate with the web server 642, PDS 644, and / or user configuration server 646, and may not use a browser or user interface to submit trades. An application running on the trading device 610 may communicate with an adapter associated with the edge server 663. For example, the application and the adapter may communicate with each other using Financial Information Exchange (FIX) messages. In these cases, the adapter may be a FIX adapter. An application running on the trading device 610 may receive market data in FIX format (this market data is provided by the price server 662 and converted to FIX format by the FIX adapter associated with the edge server 663). The application running on the trading device 610 generates trading orders based on the received market data and sends order messages in FIX format to the FIX adapter associated with the edge server 663. The FIX adapter associated with edge server 663 may process order messages received in FIX format into a format that can be understood by components in data center 660.
[0168] It should be understood that the electronic trading system 600 is merely an example, and other electronic trading systems can also be used. For example, the electronic trading system 600 does not necessarily have to include the cloud service 640. Another example is that the data center 660 may have more or fewer components than those described above, with reference to Figure 6. Another example is that messaging other than multicast messaging may be used between the components of the data center 660.
[0169] In the examples, the data center 660 may be provided by or include the first DC210 or second DC220 of any example described above with reference to Figures 1 to 4. For example, the data center 660 may be configured to provide the same or similar functionality as the first DC210 or second DC220, according to any example described above with reference to Figures 1 to 4. Thus, exemplary, the first DC210 or second DC220 according to any example described above with reference to Figures 1 to 4 may be implemented by the data center 660. Exemplary, the exchange 630 may be provided by or include the second computing entity 260 according to any example described above with reference to Figures 1 to 4. Thus, the exchange 630 is an example of the second computing entity 260 described above with reference to Figures 1 to 4.
[0170] In an example where data center 660 is provided by the first DC 310, the strategy engine 667 is an example of the first application 314 described above with reference to Figure 3, the risk server 664 is another example of the first application 314 described above with reference to Figure 3, and the order connector 666 is an example of the second application 315 described above with reference to Figure 3. In an example where data center 660 is provided by the second DC 320, the strategy engine 667 is an example of the third application 314 described above with reference to Figure 3, the risk server 664 is another example of the third application 314 described above with reference to Figure 3, and the order connector 666 is an example of the fourth application 325 described above with reference to Figure 3.
[0171] Applications 667, 664, and 666 of data center 660 are controllable by one or more computing entities outside of data center 660. For example, one or more of the strategy engine 667, risk server 664, and order connector 666 are controllable by one or more computing entities outside of data center 660. For example, the strategy engine 667 can automatically generate order messages based on market data and appropriately configured algorithms. The strategy engine 667 is controllable by an external computing entity such as the trading device 610 or another device (not shown) that can start, stop, or modify the functions of the algorithm. The risk server 664 and / or order connector 666 are controllable by one or more computing entities outside of data center 660, for example, the computing devices of the data center 660 operator (not shown in Figure 6).
[0172] In the example where data center 660 is provided by the first DCs 210 and 310 as described above, if data center 660 detects that it has become unavailable to one or more computing entities outside of data center 660, applications 667, 664, and 666 will become uncontrollable by the external computing entities, and data center 660 can operate the applications in disaster recovery mode as described above. For example, the strategy engine 667 may stop processing market data, stop generating order messages, and / or shut down. This could prevent uncontrollable order messages from being sent from the strategy engine 667 to the exchange 530. Furthermore, this could be useful for an effective failover of the strategy engine 667 to another data center. For example, it may be beneficial for the strategy engine 667 not to be operating in two locations simultaneously, as this could result in the generation of distorted or other undesirable trading orders. Therefore, stopping data processing in the strategy engine 667 or shutting down the strategy engine 667 could potentially prevent trade orders from being issued twice to the exchange 630 when the strategy engine 667 fails over to another data center. As another example, the risk server 664 may stop routing order messages to the order connector 666. This could help prevent uncontrollable order messages from being sent to the exchange 630 and / or help provide an effective failover for the strategy engine 667 and / or the risk server 664. As yet another example, the order connector 666 may disconnect its communication connection to the exchange 630. This could help prevent uncontrollable order messages from being sent to the exchange 630. Furthermore, in the example, the exchange 630 may support or allow only one communication connection per trading account.Therefore, an order connector 666 that terminates its communication connection with the exchange 630 associated with a particular trading account can instead establish a communication connection for that particular trading account from the order connector in the failover data center to the exchange 630. This enables an effective failover of the order connector 666 to the failover data center.
[0173] For example, instances of the strategy engine 667, risk server 664, and / or order connector 666 may attempt to run after data center 660 becomes unavailable to one or more external computing entities. In this case, running the strategy engine 667 in disaster recovery mode may include preventing the processing of market data, preventing the generation of trading orders, and / or stopping the strategy engine 667 before it generates order messages. Another example is running the risk server 664 in disaster recovery mode, which may include preventing the risk server 664 from routing data, or shutting down the risk server 664 before order messages are routed to the order connector 666. Another example is running the order connector 666 in disaster recovery mode, which may include preventing a communication connection with the exchange 630 from being established, or stopping the order connector 666 before a communication connection with the exchange 630 is established. As above, this can prevent uncontrollable orders or order modifications / deletions from reaching the exchange 630 and may help effectively failover these applications to another data center.
[0174] In the example shown in Figure 2 or 3, where the second DC220, 330 provides the data center 660 and enables failover of the first DC210, if it is determined that an application in the first DC210, 310 has become uncontrollable by one or more external computing entities, the data center 660 can prompt the application to operate in failover mode in order to provide failover for the application in the first DC210, 310. For example, the strategy engine 667 can be operated to provide failover for the staging engine in the first DC210, the risk server 664 can be operated to provide failover for the risk server in the first DC210, and / or the order connector 666 can be operated to provide failover for the order connectors in the first DC210, 310.
[0175] For example, order connector 666 may receive instructions that it should operate in failover mode to provide failover for the order connectors of the first DCs 210 and 310. In response, order connector 666 can establish a communication connection with exchange 630 (in this case, exchange 630 was initially serviced by the first DCs 210 and 310). For example, order connector 666 can establish a communication connection with exchange 630 for each trading account that was active in the first DCs 210 and 310. For example, user configuration database 645 may store accounts for which the order connectors of the first DCs 210 and 310 had (or should have) established a connection with exchange 630. Order connector 666 can access user configuration DB 645 to identify accounts for which the order connectors of the first DCs 210 and 310 had (or should have) established a connection with exchange 630. Furthermore, the order connector 666 may determine information for each account (such as username and authentication information) that may be necessary to establish a connection with the exchange from the user configuration DB 645. The order connector 666 appropriately establishes a communication connection with the exchange 630, thereby providing failover for the order connectors of the first DC 210 and 310.
[0176] As another example, risk server 664 may receive instructions to operate in failover mode in order to provide failover for the risk servers of DC 210 and 310. In response, risk server 664 evaluates the risk of trade orders destined for exchange 630, and if the risk is within acceptable limits, routes the trade orders to order connector 666 and sends them to exchange 630. In some cases, trade orders destined for exchange 630 may be routed to risk server 664 first. Under normal operating conditions, risk server 664 evaluates the risk of these trade orders but does not route them to order connector 666. In these examples, operating risk server 664 in failover mode may include routing the orders (destined for exchange 630 and evaluated as having acceptable risk) to order connector 666. This may enable a rapid failover of the risk servers of DC 210 and 310.
[0177] As another example, the strategy engine 667 may receive instructions that it should operate in failover mode to provide failover for the strategy engines of the first DCs 210 and 310. In response, the strategy engine 667 can identify one or more algorithms that were being executed by the strategy engines of the first DCs 210 and 310. For example, the user configuration DB can store indications of algorithms executed by the first DCs 210 and 310 in relation to each user. The strategy engine 667 can access the user configuration DB 645 and identify one or more algorithms that were being executed by the strategy engines of the first DCs 210 and 310. The strategy engine 667 can then initialize the determined algorithms on its own. Furthermore, the strategy engine 667 can retrieve the latest state information of the algorithms being executed on the strategy engines of the first DCs 210 and 310. For example, the strategy engine 667 can access the trading database 647 to identify the most recent trading orders issued by specific algorithms of the strategy engines in the first DCs 210 and 310, and any execution reports received from the exchange regarding those trading orders. Using this information and market data about the exchange 630 from the price server 662, the strategy engine 667 can execute the algorithm from the point where the algorithm in the first DCs 210 and 310 stopped. The resulting trading orders related to the exchange 630 may be routed from the strategy engine 667 to the risk server 664 via multicast bus B. The risk server 664 may route those trading orders (only if they are of acceptable risk) to the order connector via multicast bus B. The order connector may then send these trading orders to the exchange 630 via the appropriate connection established with the exchange 630. The exchange 630 may return execution reports to the strategy engine 667 via the order connector 666 and multicast bus C, as described above. Therefore, data center 660 can provide effective failover for the first DC210 and 310.
[0178] In this specification, the terms “configured” and “adapted” also include the modification, arrangement, alteration, or transformation of an element, structure, or apparatus to perform a particular function or purpose.
[0179] Some of the figures described show exemplary block diagrams, systems, and / or flow diagrams that represent methods that may be used to implement all or part of a particular embodiment. One or more components, elements, blocks, and / or functionalities of the exemplary block diagrams, systems, and / or flow diagrams may be implemented, for example, alone or in combination, as hardware, firmware, discrete logic, computer-readable instructions stored on tangible computer-readable media, and / or any combination thereof. The exemplary block diagrams, systems, and / or flow diagrams may be implemented, for example, using application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable logic devices (FPLDs), discrete logic, hardware, and / or firmware in any combination.
[0180] The illustrated block diagrams, systems, and / or flow diagrams may be executed using, for example, one or more processors, controllers, and / or other processing units. For example, these examples may be implemented using coded instructions, such as computer-readable instructions stored on tangible computer-readable media. Tangible computer-readable media include, for example, various volatile and non-volatile storage media such as: random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), electrically rewritable read-only memory (EPROM), electrically erasable read-only memory (EEPROM), flash memory, hard disk drives, optical media, magnetic tape, file servers, and other tangible data storage devices, or any combination thereof. Tangible computer-readable media are non-transient.
[0181] Furthermore, while the above exemplary block diagrams, systems, and / or flow diagrams are illustrated with reference to the figures, other embodiments are also possible. For example, the execution order of components, elements, blocks, and / or functionalities can be changed, and / or some of the described components, elements, blocks, and / or functionalities can be modified, deleted, subdivided, or merged. In addition, all or some of the components, elements, blocks, and / or functionalities may be executed sequentially and / or in parallel, for example, by separate processing threads, processors, devices, discrete logic, and / or circuits.
[0182] While examples are disclosed, various modifications can be made, and equivalents can be substituted. Furthermore, many modifications can be made to adapt to specific situations and materials. Accordingly, the disclosed technology is not limited to the specific embodiments disclosed, but is intended to include all embodiments that fall within the scope of the appended claims.
[0183] IX. Clause 1. A method comprising: determining a first state of a first group of communication connections including one or more communication connections between a first data center and a second data center using a computing system; determining a second state of a second group of communication connections including one or more communication connections between a first data center and a third data center different from the second data center using a computing system; and determining, based on the first and second states, an indication of the availability of the first data center to one or more computing entities located outside the first data center using a computing system.
[0184] 2. A method according to paragraph 1, the method comprising determining a third state of a third group of communication connections, which includes one or more communication connections between a first data center and a fourth data center different from the second and third data centers, wherein the determination of an indication that the first data center is available is based on the third state.
[0185] 3. A method according to paragraph 1 or 2, wherein the determination of an indication of the availability of the first data center includes determining, for the state of each of a plurality of communication connection groups, that each of the plurality of communication connection groups is inoperable, that the first data center is unavailable to one or more computing entities located outside the first data center.
[0186] 4. A method according to any one of paragraphs 1 to 3, wherein, with respect to a predetermined group of communication connection groups, the determination of the state of the predetermined group of communication connection groups is based on the state of one or more communication connections within the predetermined group of communication connection groups.
[0187] 5. A method according to any one of paragraphs 1 to 4, wherein each of the groups of communication connections includes two or more communication connections.
[0188] 6. A method according to paragraph 5, wherein determining the state of each of a plurality of communication connection groups includes determining the state of each of two or more communication connections of a predetermined communication connection group, and determining that the predetermined communication connection group is inoperable in response to determining that each of two or more communication connections of the predetermined communication connection group is inoperable.
[0189] 7. A method according to any one of paragraphs 1 to 6, wherein each of a plurality of communication connection groups includes a first communication connection for receiving heartbeat messages from a predetermined data center, and determining the state of a predetermined communication connection group among the plurality of communication connection groups includes determining the state of the first communication connection of the predetermined communication connection group by monitoring the reception of heartbeat messages.
[0190] 8. The method described in paragraph 7, wherein the first communication connection includes a subscription to one or more multicast channels on which heartbeat messages are transmitted.
[0191] 9. A method according to any of paragraphs 1 to 8, wherein each of a plurality of communication connection groups includes a second communication connection that enables location inspection in a predetermined data center, and for a predetermined communication connection group among the plurality of communication connection groups, the determination of the state of the predetermined communication connection group includes determining whether the location in the predetermined data center can be successfully inspected using the second communication connection.
[0192] 10. The method described in paragraph 9, wherein the second communication connection includes a TCP connection.
[0193] 11. A method according to any of paragraphs 1 to 10, wherein the determination of the status of each of the multiple communication connection groups and the determination of the availability of the first data center are performed at the first data center.
[0194] 12. A method according to paragraph 11, comprising initiating one or more processes, including a shutdown process for shutting down one or more applications running in the first data center in response to the determination that the first data center is unavailable to one or more computing entities outside the first data center, and a disaster recovery process for running one or more applications running in the first data center in disaster recovery mode.
[0195] 13. A method according to any of paragraphs 1 to 10, wherein the determination of the status of each of the multiple communication connection groups and the determination of the availability of the first data center are performed in a data center different from the first data center.
[0196] 14. A method according to paragraph 13, wherein the determination of the status of each of the multiple communication connection groups and the determination of the availability of the first data center are performed at the third data center.
[0197] 15. A method according to paragraph 13 or 14, wherein, with respect to at least one predetermined group of communication connections among a plurality of groups of communication connections, the determination of the state of the predetermined group of communication connections includes obtaining information about one or more communication connections of the predetermined group of communication connections from a data center connected to the first data center by one or more communication connections of the predetermined group of communication connections, and determining the state of the predetermined group of communication connections based on the obtained information.
[0198] 16. A method according to any of paragraphs 13 to 15, comprising initiating a failover process to switch the functions of the first data center to a failover data center in response to the determination that the first data center is unavailable to one or more computing entities located outside the first data center.
[0199] 17. A method of the method described in paragraph 16, wherein initiating a failover process includes sending a message from a data center that has determined that the first data center is unavailable to a failover data center instructing it to failover the functions of the first data center.
[0200] 18. A method relating to paragraph 16 or lecture 17, wherein the failover data center is a second data center.
[0201] 19. A method according to any of paragraphs 16 to 18, wherein the failover process includes, for each of the one or more applications running in the first data center, running the application in the failover data center, and modifying the behavior of the existing application in the failover data center to include one or more functions of the application.
[0202] 20. A method according to any of paragraphs 13 to 19, wherein a data center different from the first data center performs a determination in response to the determination that the other data center is unable to perform the determination, wherein the different data center performs a determination of the status of each of a group of communication connections and a determination of the availability of the first data center.
[0203] 21. A computing system comprising memory and one or more processors, wherein one or more processors determine a first state of a first group of communication connections including one or more communication connections between a first data center and a second data center, determines a second state of a second group of communication connections including one or more communication connections between the first data center and a third data center different from the second data center, and determines an indication of the availability of the first data center to one or more computing entities located outside the first data center, based on the first and second states.
[0204] 22. A data center having a computing system, the computing system comprising memory and one or more processors, the processors determining a first state of a first group of communication connections including one or more communication connections between a first data center and a second data center, determining a second state of a second group of communication connections including one or more communication connections between the first data center and a third data center different from the second data center; and determining an indication of the availability of the first data center to one or more computing entities outside the first data center, based on the first and second states.
[0205] 23. A tangible computer-readable storage medium comprising instructions executed by one or more processors of a computing system, the instructions comprising: determining a first state of a first group of communication connections including one or more communication connections between a first data center and a second data center; determining a second state of a second group of communication connections including one or more communication connections between the first data center and a third data center different from the second data center; and determining, based on the first and second states, an indication of the availability of the first data center to one or more computing entities outside the first data center.
Claims
1. The computing system determines the first state of a first group of communication connections, which includes one or more communication connections between the first data center and the second data center. The computing system determines the second state of a second group of communication connections, which includes one or more communication connections between the first data center and a third data center different from the second data center, and A method comprising a computing system determining, based on a first state and a second state, an indication of the availability of a first data center to one or more computing entities located outside the first data center.
2. This includes determining the third state of a third group of communication connections, which includes one or more communication connections between the first data center and a fourth data center that is different from the second and third data centers. The method according to claim 1, wherein the determination of an indication that the first data center is available is based on a third state.
3. The method according to claim 1, wherein the determination of an indication of the availability of the first data center includes determining, for each of the multiple communication connection groups, that the first data center is unavailable to one or more computing entities located outside the first data center, in response to the status of each of the multiple communication connection groups indicating that each of the multiple communication connection groups is inoperable.
4. The method according to claim 1, wherein, for each of the predetermined communication connection groups among a plurality of communication connection groups, the determination of the state of the predetermined communication connection group is based on the state of one or more communication connections in the predetermined communication connection group.
5. The method according to claim 1, wherein each of the multiple groups of communication connections includes two or more communication connections.
6. Determining the state of a specific group of communication connections among multiple groups of communication connections is: Determine the status of each of two or more communication connections in a predetermined group of communication connections. The method according to claim 5, further comprising determining that a predetermined group of communication connections is inoperable in response to determining that each of two or more communication connections in a predetermined group of communication connections is inoperable.
7. Each of the multiple communication connection groups includes a first communication connection for receiving heartbeat messages from a designated data center. The method according to claim 1, wherein determining the state of a predetermined communication connection group for each of a plurality of communication connection groups includes determining the state of the first communication connection of the predetermined communication connection group by monitoring the reception of heartbeat messages.
8. The method according to claim 7, wherein the first communication connection includes a subscription to one or more multicast channels on which heartbeat messages are transmitted.
9. Each of the multiple communication connection groups includes a second communication connection that enables location inspection in a given data center, The method according to claim 1, wherein, for each of a predetermined group of communication connections among a plurality of communication connection groups, the determination of the state of the predetermined communication connection group includes determining whether the location in a predetermined data center can be properly inspected using a second communication connection.
10. The method according to claim 9, wherein the second communication connection includes a TCP connection.
11. The method according to claim 1, wherein the determination of the status of each of the Fukusu communication connection group and the determination of the availability of the first data center are performed at the first data center.
12. In response to the determination that the first data center is unavailable to one or more computing entities outside the first data center, A shutdown process for shutting down one or more applications running in the first data center. A disaster recovery process to run one or more applications running in the first data center in disaster recovery mode. The method according to claim 11, comprising initiating one or more of the processes.
13. The method according to claim 1, wherein the determination of the status of each of the multiple communication connection groups and the determination of the availability of the first data center are performed in a data center different from the first data center.
14. The method according to claim 13, wherein the determination of the status of each of the multiple communication connection groups and the determination of the availability of the first data center are performed in the third data center.
15. For at least one predetermined communication connection group among multiple communication connection groups, the determination of the state of the predetermined communication connection group is as follows: To obtain information about one or more communication connections in a predetermined communication connection group from a data center connected to the first data center by one or more communication connections in a predetermined communication connection group, and The method according to claim 13, further comprising determining the state of a predetermined group of communication connections based on acquired information.
16. The method according to claim 13, comprising initiating a failover process to switch the functions of the first data center to a failover data center in response to the determination that the first data center is unavailable to one or more computing entities located outside the first data center.
17. Initiating the failover process is The method according to claim 16, further comprising sending a message from a data center that has determined that the first data center is unavailable to a failover data center instructing it to failover the functions of the first data center.
18. The method according to claim 16, wherein the failover data center is a second data center.
19. The failover process applies to each of the one or more applications running in the primary data center: Running applications in a failover data center; and Modifying the behavior of an existing application in a failover data center to include one or more functions of the application in question. The method according to claim 16, comprising at least one of the above.
20. The method according to claim 13, wherein a data center different from the first data center performs a determination in response to the determination that the other data center is unable to perform the determination, the different data center performs a determination of the status of each of the multiple communication connection groups and a determination of the availability of the first data center.
21. A computing system, Memory and One or more processors, Equipped with, One or more processors, Determine the first state of the first group of communication connections, which includes one or more communication connections between the first data center and the second data center. Determine the second state of the second group of communication connections, which includes one or more communication connections between the first data center and a third data center different from the second data center. A computing system that determines an indication of the availability of the first data center to one or more computing entities located outside the first data center, based on the first and second states.
22. A data center having a computing system, wherein the computing system is Memory and One or more processors, Equipped with, The processor is Determine the first state of the first group of communication connections, which includes one or more communication connections between the first data center and the second data center. Determine the second state of the second group of communication connections, which includes one or more communication connections between the first data center and a third data center different from the second data center. Based on the first and second states, an indication is determined that shows the availability of the first data center to one or more computing entities located outside the first data center. A data center equipped with computing systems.
23. A tangible computer-readable storage medium that includes instructions executed by one or more processors of a computing system, The order is, To determine the first state of a first group of communication connections, which includes one or more communication connections between the first data center and the second data center. To determine the second state of a second group of communication connections, which includes one or more communication connections between the first data center and a third data center different from the second data center, and Based on the first and second states, determine an indication that the first data center is available to one or more computing entities located outside the first data center. A tangible, computer-readable storage medium, including [a specific type of storage medium].