Cross-cluster failover framework

The cross-cluster failover framework monitors the health status of the cluster and switches data flow to the remote cluster only when the local cluster is unhealthy, solving the traffic overhead and delay problems caused by data flow switching in the existing technology, realizing efficient data flow switching and data integrity management.

CN120569944APending Publication Date: 2025-08-29VISA INTERNATIONAL SERVICE ASSOCIATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380091972.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing failover frameworks add traffic overhead and latency to real-time data flow when switching data flows, and conventional frameworks cannot effectively manage data flow switching between local and remote clusters, resulting in unnecessary overhead and latency.

Method used

The Cross-cluster Failover Framework (CCFF) uses replicators to enable data mirroring and synchronization by monitoring the health status of local and remote clusters, only switches specific data flow categories to remote clusters when local clusters are unhealthy, maintains data integrity, and switches back to local clusters when local clusters recover health.

Benefits of technology

Reduces traffic overhead and latency of real-time data flows, improves the efficiency and availability of data flow switching, and ensures data integrity and high availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120569944A_ABST
    Figure CN120569944A_ABST
Patent Text Reader

Abstract

The present disclosure describes a cross-cluster failover framework configured to switch a synchronous or asynchronous connection from a first cluster data center to a second cluster data center based on a health assessment of the first and second clusters on an application level, data flow level, or topic level basis. The connection between the first cluster and the consumer application and the connection between the second cluster and the consumer application are maintained throughout the failover framework, allowing the failover process to be invisible at the consumer application side.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a cross-cluster failover framework configured to switch a synchronous or asynchronous connection from a first cluster data center to a second cluster data center based on a health assessment of the first cluster and the second cluster on a data flow class or topic level basis. Summary of the Invention

[0002] In one aspect, the present disclosure provides a distributed data processing system, the distributed data processing system comprising: a cluster management server, the cluster management server being configured to communicate with a cluster system, the cluster system comprising a first server cluster and a second server cluster in geographically different regions, the cluster management server comprising a processor and a memory coupled to the processor, the memory having machine executable instructions stored thereon, the machine executable instructions, when executed, causing the processor to: determine an initial health state of the first server cluster; establish a first push connection between a data source and the first server cluster based on determining that the first server cluster is healthy; monitor data being sent from the data source to the first server cluster, wherein the first server cluster publishes the data to a plurality of data stream categories corresponding to the data; and for the plurality of data stream categories associated with the data source, determining a first cluster health state of the first server cluster for each data flow category of the plurality of data flow categories associated with the data source; determining a second cluster health state of the second server cluster for each data flow category of the plurality of data flow categories associated with the data source; determining that the first server cluster is unhealthy for a first data flow category of the plurality of data flow categories associated with the data source; initiating a failover associated with the first data flow category for the data source, wherein the failover causes the processor to: establish a second push connection to the second server cluster based on determining that the second server cluster is healthy for the first data flow category associated with the data source; and switch the data source from sending to the first server cluster to sending to the second server cluster based on determining that the first server cluster is unhealthy for the data source and determining that the second server cluster is healthy for the data source.

[0003] In another aspect, the present disclosure provides a method for cross-cluster failover management, the method comprising: determining, by a cluster management server, an initial first cluster health state of a first server cluster and an initial second cluster health state of a second server cluster; establishing, by the cluster management server, a first push connection between a data source and the first server cluster based on determining that the initial first cluster health state of the first server cluster is healthy; monitoring, by the cluster management server, that data is sent from the data source to the first server cluster, wherein the first server cluster publishes the data to a plurality of data stream categories corresponding to the data; determining, by the cluster management server, a first cluster health state of the first server cluster for each of the plurality of data stream categories associated with the data source; determining, by the cluster management server, a second cluster health state of the second server cluster for each of the plurality of data stream categories associated with the data source; The cluster management server determines that the first server cluster is unhealthy for a first data flow category among the multiple data flow categories associated with the data source; the cluster management server initiates a failover associated with the first data flow category for the data source, wherein the failover also includes: the cluster management server establishes a second push connection to the second server cluster based on determining that the second server cluster is healthy for the first data flow category associated with the data source; the cluster management server switches the data source from sending to the first server cluster to sending to the second server cluster based on determining that the first server cluster is unhealthy for the data source and determining that the second server cluster is healthy for the data source; and the cluster management server maintains a first pull data connection between the first server cluster and a consumer application, and a second pull data connection between the second server cluster and the consumer application. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] In the description, for purposes of explanation rather than limitation, specific details are set forth, such as particular aspects, procedures, techniques, etc., to provide a thorough understanding of the technology of the present invention. However, it will be apparent to one skilled in the art that the technology of the present invention can be practiced in other aspects that depart from these specific details.

[0005] The accompanying drawings, together with the detailed description below, are incorporated into and form a part of the specification and serve to further illustrate aspects of the concepts including the claimed disclosure and to explain various principles and advantages of those aspects. In the accompanying drawings, like reference numerals refer to the same or functionally similar elements throughout the different views.

[0006] The cross-cluster failover framework devices, systems, and methods disclosed herein have been represented by conventional symbols in the accompanying drawings where appropriate, showing only those specific details relevant to understanding the various aspects of the disclosure so as not to obscure the disclosure with details that would be apparent to a person of ordinary skill in the art having the benefit of the description herein.

[0007] Figure 1 A block diagram illustrating a cross-cluster failover framework of a cluster according to at least one aspect of the present disclosure, the cluster including a local data center cluster and a remote data center cluster and controlled by a cluster manager.

[0008] Figure 2 is a block diagram of a cluster manager configured to continuously or periodically monitor the health of a remote cluster and a local cluster according to at least one aspect of the present disclosure.

[0009] Figure 3 Exemplary cluster configurations of node partitions and topics for a data center cluster in a remote cluster or a local cluster according to at least one aspect of the present disclosure are shown.

[0010] Figure 4 A first exemplary health check response evaluation by cluster manager 106 for a cluster configuration is shown in accordance with at least one aspect of the present disclosure.

[0011] Figure 5 A second exemplary health check response evaluation by a cluster manager for a cluster configuration is shown in accordance with at least one aspect of the present disclosure.

[0012] Figure 6 A third exemplary health check response evaluation by a cluster manager for a cluster configuration is illustrated in accordance with at least one aspect of the present disclosure.

[0013] Figure 7 is a flow chart of a failover process from a local cluster to a remote cluster according to a cross-cluster failover framework according to at least one aspect of the present disclosure.

[0014] Figure 8 is a block diagram of a cluster manager configured to buffer data from a data source according to at least one aspect of the present disclosure.

[0015] Figure 9 is a block diagram of a computer device having data processing subsystems or components according to at least one aspect of the present disclosure.

[0016] Figure 10 is a diagrammatic representation of an exemplary system including a host computer within which a set of instructions for performing any one or more of the methodologies discussed herein may be executed in accordance with at least one aspect of the present disclosure. DETAILED DESCRIPTION

[0017] The following disclosure may provide exemplary systems, devices, and methods for conducting financial transactions and related activities. While reference may be made to such financial transactions in the examples provided below, the various aspects are not limited thereto. That is, the systems, methods, and devices may be used for any suitable purpose.

[0018] Before discussing specific embodiments, aspects, or examples, some description of the terms used herein is provided below.

[0019] As used herein, the term "comprising" is not intended to be limiting, but rather may be a transitional term synonymous with "including," "containing," or "characterized by." Thus, the term "comprising" may be inclusive or open-ended, and does not exclude additional, unstated elements or method steps when used in a claim. For example, when describing a method, "comprising" indicates that the claim is open-ended and allows for additional steps. When describing an apparatus, "comprising" may mean that the named elements may be essential to an embodiment or aspect, but other elements may be added and still form a construction within the scope of the claim. In contrast, the transitional phrase "consisting of" excludes any element, step, or ingredient not specified in the claim. This is consistent with the usage of the term throughout this specification.

[0020] As used herein, the term "computing device" or "computer device" may refer to one or more electronic devices configured to communicate directly or indirectly with or on one or more networks. The computing device may be a mobile device, a desktop computer, or the like. As examples, the mobile device may include a cellular phone (e.g., a smartphone or a standard cellular phone), a portable computer, a wearable device (e.g., a watch, glasses, lenses, clothing, etc.), a personal digital assistant (PDA), and / or other similar devices. A computing device may not be a mobile device, such as a desktop computer. In addition, the term "computer" may refer to any computing device that includes the necessary components for sending, receiving, processing, and / or outputting data and typically includes a display device, a processor, a memory, an input device, and a network interface, and / or the like.

[0021] As used herein, references to "device," "server," "processor," and the like may refer to a previously stated device, server, or processor stated as performing a previously stated step or function, a different server or processor, and / or a combination of servers and / or processors. For example, as used in the specification and claims, a first server or first processor stated as performing a first step or a first function may refer to the same or a different server or the same or a different processor stated as performing a second step or a second function.

[0022] An "interface" may include any software module configured to handle communications. For example, an interface may be configured to receive, process, and respond to a specific entity in a specific communication format. Furthermore, a computer, device, and / or system may include any number of interfaces, depending on the functionality and capabilities of the computer, device, and / or system. In some embodiments or aspects, an interface may include an application programming interface (API) or other communication formats or protocols that may be provided to a third party or specific entity to allow communication with the device. Additionally, an interface may be designed based on functionality, a specified entity with which it is configured to communicate, or any other variable. For example, an interface may be configured to allow the system to respond to a specific request or may be configured to allow a specific entity to communicate with the system.

[0023] As used herein, the term "server" may include one or more computing devices, which may be individual stand-alone machines located at the same or different locations, may be owned or operated by the same or different entities, and may further be one or more clusters of distributed computers or "virtual" machines housed in a data center. Those skilled in the art will appreciate and understand that the functions performed by a "server" may be spread across multiple different computing devices for various reasons. As used herein, "server" is intended to refer to all such scenarios and should not be interpreted as or limited to a specific configuration. In addition, the server as described herein may, but need not, reside at (or be operated by) an agent of any one of a merchant, payment network, financial institution, medical provider, social media provider, government agency, or the aforementioned entity. The term "server" may also refer to or include one or more processors or computers, storage devices, or similar computer arrangements that operate or facilitate communication and processing by multiple parties in a network environment such as the Internet, but it should be understood that communication can be facilitated by one or more public or private network environments, and various other arrangements are possible. In addition, multiple computers (e.g., servers) or other computerized devices (e.g., point-of-sale devices) communicating directly or indirectly in a network environment may constitute a "system" (e.g., a merchant's point-of-sale system). As used herein, references to a "server" or "processor" may refer to a previously stated server and / or processor stated as performing a previous step or function, different servers and / or processors, and / or combinations of servers and / or processors. For example, as used in the specification and claims, a first server and / or first processor stated as performing a first step or function may refer to the same or a different server and / or processor stated as performing a second step or function.

[0024] A "server computer" may generally be a single powerful computer or cluster of computers. For example, a server computer may be a mainframe, a cluster of small computers, or a group of servers functioning as a unit. A server computer may be associated with an entity such as a payment processing network, a wallet provider, a merchant, an authentication cloud, an acquirer, or an issuer. In one example, a server computer may be a database server coupled to a web server. The server computer may be coupled to a database and may include any hardware, software, other logic, or combination of the foregoing for servicing requests from one or more client computers. The server computer may include one or more computing devices and may use any of a variety of computing structures, arrangements, and compilations for servicing requests from one or more client computers. In some embodiments or aspects, the server computer may provide and / or support payment network cloud services.

[0025] As used herein, the term "system" may refer to one or more computing devices or a combination of computing devices (eg, processors, servers, client devices, software applications, components of such computing devices, and / or the like).

[0026] A "user device" is an electronic device that can be transferred and / or operated by a user. A user device can provide telecommunication capabilities for a network. A user device can be configured to transmit data or communications to other devices and receive data or communications from other devices. In some embodiments or aspects, a user device can be portable. Examples of user devices can include mobile phones (e.g., smartphones, cellular phones, etc.), PDAs, portable media players, wearable electronic devices (e.g., smart watches, fitness bands, ankle bracelets, rings, earrings, etc.), e-reader devices, and portable computing devices (e.g., laptop computers, netbooks, ultrabooks, etc.). Examples of user devices can also include automobiles with telecommunication capabilities.

[0027] The present disclosure relates to a cluster computing platform for real-time data streams (e.g., Kafka, Hazelcast, etc.). The cluster computing platform includes a primary data center cluster (e.g., a local cluster) that receives data streams, such as messages or data records, from a data source (e.g., a producer application). The cluster computing platform publishes the data stream to a data stream category in the cluster so that it can be immediately accessed (e.g., consumed) by an application (e.g., a consumer application). The cluster is configured to operate in a redundant or parallel cluster with a primary cluster (e.g., a local cluster) and a secondary data center cluster (e.g., a remote cluster). The addition of a remote cluster can improve the throughput of reading and writing records from real-time data streams and / or create a high-availability data configuration for real-time data streams. In various aspects, if the local cluster is deemed unhealthy, data records (e.g., messages) from the data source will begin to fail (after some retries), and the data source will not be able to write data records to the local cluster until the local cluster is deemed healthy again. Instead of waiting for the local cluster to become healthy again, a conventional failover framework switches the data source to the remote cluster for all data stream categories (e.g., topics, maps, etc.). Conventional failover frameworks increase traffic overhead for real-time data flows and may create unnecessary failovers for the class of data flows that are written to healthy nodes within the cluster. Conventional failover frameworks ultimately create unnecessary overhead for the remote cluster and increased latency for otherwise healthy data flows.

[0028] This disclosure describes a Cross-Cluster Failover Framework (CCFF) to address the aforementioned issues. The CCFF solution adds the ability to write to a remote cluster when a specific data flow class is deemed unhealthy on the local cluster. Data sources (e.g., producer applications) will switch to the remote cluster only for the specific data flow class, while other data flow classes remain available and healthy on the local cluster. When the local cluster is deemed healthy again for the switched data flow class, the data source will switch back to the local cluster.

[0029] Throughout the failover process, instances of the consumer application can continue to run in both the local data center cluster and the remote data center cluster. The failover or switchover from the local cluster to the remote cluster is completely invisible to the consumer application because the consumer application runs multiple instances on both the remote cluster and the local cluster. The CCFF configuration is designed to maintain data integrity for one or more data flow categories (e.g., streams, topics, maps) that are replicated across multiple partitions and / or mirrored between the local and remote clusters.

[0030] Figure 1A block diagram of a cross-cluster failover framework 100 for a cluster 108, including a local data center cluster 110 and a remote data center cluster 112, is shown according to at least one aspect of the present disclosure. The cluster manager 106 is a server or computing system. The cluster 108 is configured to receive one or more data stream categories 122a-n, 124a-n (e.g., topics, maps, etc.) from one or more data sources 102a-n (e.g., producer applications). Data (e.g., messages, records, etc.) is written to or published by the cluster 108 and is immediately accessible to one or more consumer applications 104a-n (e.g., subscribers).

[0031] The cluster manager 106 provides a first connection between one or more data sources 102a-n and a local cluster 110 or a remote cluster 112. In one aspect, the local cluster can be set as the default cluster due to physical proximity and / or lower latency compared to the remote cluster 112. The data source 102a pushes data to the local cluster 110, where the data is associated with a particular data flow category and written (e.g., published) to the node 114a hosting the original data copy (P1) in the partition 118a-n.

[0032] Figure 1 Also shown is a non-exhaustive example of a local cluster 110 comprising seven nodes 114a-n (e.g., servers, proxies) configured to host original partitions 118a, 120a and replicated partitions 118b-n, 120b-n for three data flow classes 112a-n, 124a-n. In this example, data source 102a pushes data to node 114a, where the local cluster 110 publishes the original copy of the record in partition 118a. In various aspects, partition 118a is designated as the leader partition for data flow class 122a, and partitions 118b-n are slave partitions of 122a. In this example, the replication factor for data flow class 122a is four, and slave partitions 118b-d publish synchronized replicated (ISR) copies of the original data in the replicated partitions. The replication factor may be configured by the cluster manager 106 or an application (e.g., a client library).

[0033] Each node 114a-n, 116a-n in the cluster can be configured to host multiple data flow classes 122a-n, 124a-n through a partitioned configuration. For example, node 114b hosts the first ISR replica (P1R1) of data flow class 122a in partition 118b, and node 114b also hosts the original replica (P1) of topic 122n. The partitioned configuration improves data integrity and increases input / output throughput for data flow classes 122a-n, 124a-n by allowing synchronous reads by applications 104a-104n or synchronous / asynchronous writes by data sources 102a-n.

[0034] The cross-cluster failover framework relies on mirroring of use cases (e.g., topics) between the local cluster 110 and the remote cluster 112. In various aspects, multiple instances of the consumer applications 104a-n are simultaneously connected to the local cluster 110 and the remote cluster 112. Therefore, no replicator 126 is required to synchronize data sets between the local cluster 110 and the remote cluster 112. As long as a message is published on one healthy cluster, the consumer applications 104a-n can consume the data.

[0035] In one aspect, the replicator 126 (e.g., MirrorMaker) can be configured to mirror data between the local cluster 110 and the remote cluster 112. When the data source 102a connects to the remote cluster 112 and publishes to the remote cluster 112, the replicator 126 mirrors the data to the local cluster 110. Similarly, when the data source 102a connects to the local cluster 110 and publishes to the local cluster 110, the replicator 126 can mirror the data to the remote cluster 112. Thus, records in the partition 118a of the node 114a of the data stream class 122a at the local cluster 110 can be published in the partition 120a of the node 116a of the data stream class 124a at the remote cluster 112.

[0036] Thus, once a data record is published to a partition or published on any cluster, the data record is immediately available to consumer applications 104a-n (e.g., subscribers). The client application 128 or the cluster manager 106 may include a library of configuration parameters for the cross-cluster failover framework. In various aspects, these configuration parameters include permissions for data flow classes, the minimum number of synchronous replicas required for each data flow class, the replication factor, and firewall request parameters for connectivity between the cluster and the cluster manager 106.

[0037] Figure 2A block diagram of a cluster manager 106, according to at least one aspect of the present disclosure, is shown. The cluster manager is configured to continuously or periodically monitor the health of a remote cluster 112 and a local cluster 110. A client application 128 can configure the cluster manager 106 to perform health checks on the remote cluster 112 and the local cluster 110 at different time intervals. The cluster manager 106 maintains continuous connectivity with the local cluster 110 and the remote cluster 112 via a health check API to determine the health of each cluster. In one example, if the local cluster 110 is in use, the health check interval for the local cluster 110 can be more frequent than that for the remote cluster 112. The cluster manager 106 can check the health of each cluster by polling health at a specified health check interval via an application program interface (API), which specifies a canonical health check port for each node in the cluster. Thus, the cluster manager 106 determines the cluster status as locally healthy and remotely healthy based on the health of the cluster nodes, the minimum number of synchronized replicas required for each data stream, and the replication factor of the data stream partition. In one aspect, the local health and remote health can include granular status information for each node in the cluster. In another aspect, local health and remote health may include a binary determination of the health of data flows in the cluster based on a pass or fail of a health requirement.

[0038] Figure 3 An exemplary cluster configuration 200 is shown of nodes 114a-n, 116a-n, partitions 118a-n, 120a-n, and topics 122a-n, 124a-n for a data center cluster in a remote cluster 112 or a local cluster 110 according to at least one aspect of the present disclosure. Each column represents a node 114a-n in the cluster, and each row represents a data flow category or topic 122a-n hosted by multiple nodes 114a-n. Thus, the exemplary cluster configuration 200 may represent Figure 1 The cluster manager 126 polls each node 14a-n in the cluster to determine the health of the topics 122a-n, such as Figure 2 Each node 114a-n must respond to the health check message within a predetermined amount of time in order for the cluster manager 106 to determine that the node is healthy and to evaluate the health of each topic 112a-n.

[0039] Figure 4A first exemplary health check response evaluation performed by the cluster manager 106 for the cluster configuration 200 according to at least one aspect of the present disclosure is shown. In this example, the nodes 114a-c fail to respond to the health check message within a predetermined amount of time, and the cluster manager 106 determines that the nodes 114a-c are down. The health of the cluster is evaluated on a per-topic basis, where each topic is required to have a minimum number of ISRs active for a given topic. In various aspects, the local cluster 110 and the remote cluster 112 may require different minimum numbers of ISRs for a topic to be healthy. For example, one cluster may require a greater minimum number of ISRs to be active in order to maintain a higher degree of data integrity for the one cluster.

[0040] exist Figure 4 In the example, the minimum number of ISRs is 2, and at least 2 nodes hosting synchronous replicas of the topic's partition must be active. Figure 4 Topic T1 122a is shown as having one active node. Therefore, topic T1 122a does have at least two active ISRs. Node 114d is the only node that data source 102 can use to publish messages in the cluster for this topic 112a. Therefore, topic T1 122a is determined to be unhealthy. However, topics T2 112b and T3 122n have two or more active nodes and are therefore determined to be healthy.

[0041] Figure 5 122n 。 122a 。 122n 。 122b 。 122n ...

[0042] Figure 6122n 。 The cluster manager 106 performs a third exemplary health check response evaluation for the cluster configuration 200 according to at least one aspect of the present disclosure. In this example, the minimum number of ISRs is still set to 2, but the topic partition leader (e.g., the original record partition) must also be healthy. Here, the nodes 114b-c fail to respond to the health check message within a predetermined amount of time, and the cluster manager 106 determines that the nodes 114b-c are down. The cluster manager determines that topics T1 122a and T3 122n are healthy, and topic T2 122b is unhealthy. Although topic T2 122b has 3 active ISRs, the leader partition is on the down node 114c and therefore does not meet the topic health requirement.

[0043] After evaluating the health check responses, cluster manager 106 determines whether to initiate a failover process for the specific topic. In one aspect, once cluster manager 106 determines that the topic at local cluster 110 is unhealthy, cluster manager 106 initiates a failover process and switches the connection (e.g., synchronous or asynchronous connection) of data source 102 to remote cluster 112. If cluster manager 106 determines that local cluster 110 is healthy again, cluster manager 106 switches the connection of data source 102 back to local cluster 110.

[0044] In various aspects, the health check evaluation process may require a node to miss three consecutive health check responses before cluster manager 106 determines the node is unhealthy. This evaluation criterion takes into account that a node may be temporarily unhealthy for a few seconds due to various reasons, such as network connectivity issues. Once a node misses the first health check response, the allocated response time may be reduced for each subsequent negative health check result.

[0045] In the event that both local cluster 110 and remote cluster 112 are deemed unhealthy, data source 102 is not switched, and cluster manager 106 periodically checks both local cluster 110 and remote cluster 112. In addition, cluster manager 106 may send an error callback message to client application 128, log the error event, or trigger an alert to local cluster 110 and remote cluster 112. When the first cluster is determined to be healthy again, data source 102 switches to the healthy cluster or remains on the healthy cluster and begins publishing to the healthy cluster.

[0046] In one aspect, the failover process can be initiated based on an "all or nothing" switching approach. In the "all or nothing" switching approach, the client application 128 identifies a high-priority topic or a list of high-priority topics. The health of the cluster is checked for the high-priority topics, and if any of the high-priority topics are unhealthy in the current cluster 106 and all of the high-priority topics are healthy in another cluster, a failover process is initiated. Thus, the cluster manager 106 can initiate a failover for all topics of a data source, or not initiate a failover for any topic.

[0047] In another example of an "all-or-nothing" failover scenario, data source 102a asynchronously connects to local cluster 110 via a first connection and publishes messages for topics T1, T2, and T3 to local cluster 110. In this example, topics T1, T2, and T3 correspond to data source 102a. Client application 128 may identify topic T1 as a critical data flow category, such as a transactional data pipeline, while topics T2 and T3 may be identified as less critical data flow categories, such as reporting and logging pipelines. Due to the criticality of T1, client application 128 may specify failover priority at the topic level for T1. Therefore, if any health condition of topic T1 at local cluster 110 occurs, cluster manager 106 immediately fails over to remote cluster 112 by switching all three topics T1, T2, and T3 of data source 102a. However, if topics T2 and T3 are unhealthy and T1 is healthy, the current cluster can continue to operate for T1 without affecting the critical data of client application 128. In a system with a replicator 126, data for topics T2 and T3 that are not needed can be obtained in a synchronous manner from the remote cluster 112. Thus, the application is managed in such a way that the application is not affected by failover switching or down partitions.

[0048] Figure 7 is a flow diagram of a failover process 300 from a local cluster 110 to a remote cluster 112 according to a cross-cluster failover framework in accordance with at least one aspect of the present disclosure. At step 1, the cluster manager 106 polls 302 the health of each node 114a-n at the local cluster 110 and each node 116a-n at the remote cluster 112. The cluster manager 106 receives health check response messages from the nodes 114a-n, 116a-n and determines that the topic T1 at the local cluster 110 is unhealthy and the topic T1 at the remote cluster 112 is healthy. The cluster manager 106 initiates 304 a failover process from the local cluster 110 to the remote cluster 112 for only the first topic T1. Figure 4As shown, topics T2 and T3 remain healthy and operating on the local cluster 110, so only topic T1 is switched. As part of the failover process, cluster manager 106 may optionally begin buffering unpublished data records and / or new messages from data source 102. At step 3, cluster manager 106 closes the asynchronous data connection between data source 102 and local cluster 110. At step 4, cluster manager 106 receives confirmation that all asynchronous data connections between data source 102 and local cluster 110 have been closed. In alternative aspects, asynchronous data connections may be active at both the local cluster 110 and the remote cluster without closing the connections. At step 5, cluster manager 106 creates an instance of data source 102 for topic T1 at the remote cluster. At step 6, cluster manager 106 receives confirmation (e.g., an acknowledgment (ACK) message) that the instance of data source 102 at remote cluster 112 has been created. At step 7, if buffering is configured by client application 128, cluster manager 106 optionally flushes the buffer to remote cluster 112. Once the data in the buffer is published to the remote cluster 112, the buffer can be cleared for future use. At step 8, the cluster manager 106 continues to publish data to the remote cluster 112. The cluster manager 106 continues to poll 306 the health of each node 114a-n at the local cluster 110 and the remote cluster 112. If the cluster manager 106 determines that the remote cluster 112 is healthy and the local cluster 110 is healthy again, the cluster manager 106 will initiate a failover to switch back to the local cluster 110. The cluster manager 106 will initiate a switchback even when the remote cluster is healthy.

[0049] Figure 8 1 is a block diagram of a cluster manager 106 configured to buffer data from a data source according to at least one aspect of the present disclosure. After a failover is initiated, the data source 102a continues to push data (e.g., messages) to the cluster 108. The cluster manager 106 may include multiple buffers to temporarily store messages until the failover is complete. In one aspect, the cluster manager 106 includes a data source switch buffer 130 and an unpublished message buffer 132. The data source switch buffer 130 stores messages received by the data source 102a when the data source 102a attempts to publish a message after the failover is initiated but before the failover is complete. The unpublished message buffer 132 stores messages when the data source 102a attempts to publish a message but is unsuccessful because the local cluster 110 is unhealthy and a failover has not yet been initiated. Thus, the message buffers 130, 132 are configured to store received but unpublished messages to ensure zero message loss during the failover process. Once the failover process is complete, the buffers are flushed to the destination cluster of the failover process.

[0050] In various aspects, the buffers 130, 132 may include configurable size limits to prevent buffer overflow. Once the buffer limit is reached, the message may be routed to a secondary buffer or storage device until the failover is complete.

[0051] In one aspect, it may be desirable to fail unpublished messages quickly without using buffers 130, 132. Therefore, client application 128 can configure cluster manager 106 to set the buffer limit size to zero and receive fast failure notifications. After the cluster manager performs internal retries, all publish requests from the time the local cluster 110 becomes unhealthy until the failover completes will fail. For configurations without buffering, client application 128 must be configured to initiate a retry of the publish request after the failover completes.

[0052] In various aspects, the cluster manager 106 can test the cluster failover functionality without actually shutting down the cluster by using a switch integration feature. The cross-cluster failover framework can be integrated with a switch feature that allows cluster testing and dynamic control of cluster configuration. The test functionality of the switch integration allows for periodic checks of the failover framework to verify that the system is performing as expected. The dynamic control functionality of the switch integration allows the cluster manager 106 to dynamically disable the functionality of the failover framework without affecting or rolling back the release version of the client application 128. The switch integration feature can disable the failover framework functionality and restore the default settings of the cluster computing platform (e.g., Kafka, Hazelcast, etc.).

[0053] Figure 9 is a block diagram of a computer device 400 having data processing subsystems or components according to at least one aspect of the present disclosure. Figure 9 The subsystems shown in FIG4 are interconnected via a system bus 410. Additional subsystems are shown, such as a printer 418, a keyboard 426, a fixed disk 428 (or other memory including computer-readable media), and a monitor 422 coupled to a display adapter 420. Peripheral devices and input / output (I / O) devices coupled to an I / O controller 412 (which may be a processor or other suitable controller) may be connected to the computer system via any number of devices known in the art, such as a serial port 424. For example, a serial port 424 or an external interface 430 may be used to connect the computer device to a wide area network (e.g., the Internet), a mouse input device, or a scanner. Interconnection via the system bus allows the central processor 416 to communicate with each subsystem and control the execution of instructions from the system memory 414 or the fixed disk 428, as well as the exchange of information between the subsystems. The system memory 414 and / or the fixed disk 428 may contain computer-readable media.

[0054] Figure 10is a diagrammatic representation of an exemplary system 500 including a host computer 502 within which a set of instructions for performing any one or more of the methods discussed herein may be executed, in accordance with at least one aspect of the present disclosure. In various aspects, the host computer 502 operates as a standalone device or may be connected (e.g., via a network connection) to other machines. In a networked deployment, the host computer 502 may function as a server or client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The host computer 502 may be a computer or computing device, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a portable music player (e.g., a portable hard drive audio device such as a Moving Picture Experts Group Audio Layer 3 (MP3) player), a network appliance, a network router, a switch or bridge, or any machine capable of executing (sequentially or otherwise) a set of instructions specifying actions to be taken by the machine. Further, while a single machine is illustrated, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.

[0055] The exemplary system 500 includes a host 502, running a host operating system (OS) 504 on one or more processors / processor cores 506 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both) and various memory nodes 508. The host OS 504 may include a hypervisor 510 capable of controlling functions and / or communicating with virtual machines ("VMs") 512 running on machine-readable media. The VMs 512 may also include virtual CPUs or vCPUs 514. Memory nodes 508 may be linked or pinned to virtual memory nodes or vNodes 516. When a memory node 508 is linked or pinned to a corresponding vNode 516, data may then be mapped directly from the memory node 508 to its corresponding vNode 516.

[0056] All of the various components shown in host computer 502 can be connected to each other and can communicate with each other via a bus (not shown) or other coupling or communication channels or mechanisms. Host computer 502 may also include a video display, audio devices, or other peripherals 518 (e.g., a liquid crystal display (LCD), alphanumeric input devices including, for example, a keyboard, a cursor control device (e.g., a mouse), a voice recognition or biometric verification unit, an external drive, a signal generating device (e.g., a speaker)), a permanent storage device 520 (also known as a disk drive unit), and a network interface device 522. Host computer 502 may also include a data encryption module (not shown) for encrypting data. The components provided in host computer 502 are components typically present in computer systems suitable for use with various aspects of the present disclosure and are intended to represent a broad category of such computer components known in the art. Thus, system 500 may be a server, a minicomputer, a mainframe computer, or any other computer system. Computers may also include different bus configurations, networked platforms, multi-processor platforms, and the like. Various operating systems may be used, including UNIX, LINUX, WINDOWS, QNX ANDROID, IOS, CHROME, TIZEN, and other suitable operating systems.

[0057] The disk drive unit 524 may also be a solid state drive (SSD), a hard disk drive (HDD), or other drive including a computer or machine readable medium having stored thereon one or more instruction sets and data structures (e.g., data / instructions 526) embodying or utilizing any one or more of the methodologies or functionality described herein. The data / instructions 526 may also reside completely or at least partially within the main memory node 508 and / or within the processor(s) 506 during execution by the host 502. The data / instructions 526 may further be sent or received over the network 528 via the network interface device 522 utilizing any of several well-known transfer protocols (e.g., Hyper Text Transfer Protocol (HTTP)).

[0058] The processor(s) 506 and memory node 508 may also include machine-readable media. The term "computer-readable medium" or "machine-readable medium" should be understood to include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) that store one or more instruction sets. The term "computer-readable medium" should also be understood to include any medium capable of storing, encoding, or carrying an instruction set for execution by the host 502 and causing the host 502 to perform any one or more of the methods of the present application, or any medium capable of storing, encoding, or carrying data structures utilized by or associated with such an instruction set. Thus, the term "computer-readable medium" should be understood to include, but not be limited to, solid-state memory, optical and magnetic media, and carrier signals. Such media may also include, but are not limited to, hard disks, floppy disks, flash memory cards, digital video disks, random access memory (RAM), read-only memory (ROM), and the like. The exemplary aspects described herein may be implemented in an operating environment comprising software installed on a computer, in hardware, or in a combination of software and hardware.

[0059] Those skilled in the art will recognize that an Internet service can be configured to provide Internet access to one or more computing devices coupled to the Internet service, and that the computing devices can include one or more processors, buses, memory devices, display devices, input / output devices, etc. Furthermore, those skilled in the art will appreciate that an Internet service can be coupled to one or more databases, repositories, servers, etc., which can be used to implement any of the various aspects of the present disclosure as described herein.

[0060] The computer program instructions may also be loaded onto a computer, server, other programmable data processing device or other apparatus to cause a series of operational steps to be performed on the computer, other programmable device or other apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide a process for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0061] By way of example, a suitable network may include or be connected to any one or more of the following: a local intranet, a PAN (personal area network), a LAN (local area network), a WAN (wide area network), a MAN (metropolitan area network), a virtual private network (VPN), a storage area network (SAN), a frame relay connection, an advanced intelligent network (AIN) connection, a synchronous optical network (SONET) connection, a digital T1, T3, E1 or E3 line, a digital data service (DDS) connection, a DSL (digital subscriber line) connection, an Ethernet connection, an ISDN (integrated services digital network) line, a dial-up port (e.g., V.90, V.34 or V.34), a dual analog modem connection, a cable modem, an ATM (asynchronous transfer mode) connection, or an FDDI (fiber distributed data interface) or CDDI (copper distributed data interface) connection. In addition, communications may also include links to any of a variety of wireless networks, including WAP (Wireless Application Protocol), GPRS (General Packet Radio Service), GSM (Global System for Mobile Communications), CDMA (Code Division Multiple Access) or TDMA (Time Division Multiple Access), a cellular telephone network, GPS (Global Positioning System), CDPD (Cellular Digital Packet Data), RIM (Working Environment Photo) duplex paging network, Bluetooth radio, or IEEE 802.11 based radio frequency network. Network 530 may also include or interface with any one or more of the following: an RS-232 serial connection, an IEEE-1394 (FireWire) connection, a Fibre Channel connection, an IrDA (Infrared) port, a SCSI (Small Computer System Interface) connection, a USB (Universal Serial Bus) connection, or other wired or wireless, digital or analog interface or connection, mesh or Network connection.

[0062] Generally speaking, a cloud-based computing environment is a resource that typically combines the computing power of a large group of processors (e.g., within a web server) and / or the storage capacity of a large group of computer memory or storage devices. Systems that provide cloud-based resources can be employed solely by their owners, or such systems can be accessed by external users who deploy applications within the computing infrastructure to gain the benefits of large-scale computing or storage resources.

[0063] For example, a cloud is formed by a network of web servers comprising multiple computing devices (e.g., host 502), where each server 530 (or at least a plurality thereof) provides processor and / or storage resources. These servers manage workloads provided by multiple users (e.g., cloud resource clients or other users). Typically, the workload demands placed on the cloud by each user vary in real time, sometimes dramatically. The nature and extent of these changes typically depend on the type of business associated with the user.

[0064] It is noteworthy that any hardware platform suitable for performing the processing described herein is suitable for use with the technology. As used herein, the terms "computer-readable storage medium" and "computer-readable storage media" refer to any one or more media that participate in providing instructions to a CPU for execution. Such media can take a variety of forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks, such as fixed disks. Volatile media include dynamic memory, such as system RAM. Transmission media include coaxial cables, copper wires, and optical fibers, among others, including conductors that comprise one aspect of a bus. Transmission media can also take the form of sound waves or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media include, for example, a floppy disk, a hard disk, magnetic tape, any other magnetic medium, CD-ROM discs, digital video discs (DVDs), any other optical medium, any other physical medium with patterns of marks or holes, RAM, PROM, EPROM, EEPROM, FLASH EPROM, any other memory chip or data exchange adapter, a carrier wave, or any other medium from which a computer can read.

[0065] Various forms of computer-readable media can be used to carry one or more sequences of one or more instructions to the CPU for execution. The bus carries the data to the system RAM, from which the CPU retrieves the instructions and executes them. The instructions received by the system RAM can optionally be stored on a fixed disk before or after execution by the CPU.

[0066] The computer program code for performing operations on aspects of the present technology can be written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Java, Smalltalk, C++, etc.) and conventional procedural programming languages ​​(such as "C" programming language, Go, Python, or other programming languages ​​including assembly language). The program code can be executed entirely on the user's computer, partially on the user's computer; as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or an external computer can be connected (for example, using an Internet service provider to connect via the Internet).

[0067] Examples of methods according to various aspects of the present disclosure are provided below in the following numbered clauses. One aspect of the method may include any one or more than one and any combination of the numbered clauses described below.

[0068] Example 1. A distributed data processing system, the distributed data processing system comprising: a cluster management server, the cluster management server being configured to communicate with a cluster system, the cluster system comprising a first server cluster and a second server cluster in geographically different regions, the cluster management server comprising a processor and a memory coupled to the processor, the memory having machine-executable instructions stored thereon, the machine-executable instructions, when executed, causing the processor to: determine an initial health state of the first server cluster; establish a first push connection between a data source and the first server cluster based on determining that the first server cluster is healthy; monitor data being sent from the data source to the first server cluster, wherein the first server cluster publishes the data to a plurality of data stream categories corresponding to the data; for each data in the plurality of data stream categories associated with the data source The flow category determines a first cluster health state of the first server cluster; determines a second cluster health state of the second server cluster for each data flow category in the multiple data flow categories associated with the data source; determines that the first server cluster is unhealthy for a first data flow category in the multiple data flow categories associated with the data source; initiates a failover associated with the first data flow category for the data source, wherein the failover causes the processor to: establish a second push connection to the second server cluster based on determining that the second server cluster is healthy for the first data flow category associated with the data source; and switch the data source from sending to the first server cluster to sending to the second server cluster based on determining that the first server cluster is unhealthy for the data source and determining that the second server cluster is healthy for the data source.

[0069] Example 2. A distributed data processing system according to Example 1, wherein the processor is further configured to: determine the first cluster health state for each of the multiple data flow categories associated with the data source, wherein the first cluster health state for the first data flow category is healthy; determine the second cluster health state for each of the multiple data flow categories associated with the data source, wherein the second cluster health state for the first data flow category is healthy; initiate a failover associated with the first data flow category of the data source, wherein the failover causes the processor to: re-establish a first push connection to the first server cluster based on determining that the second server cluster is healthy for the first data flow category; and switch the data source from sending to the second server cluster back to sending to the first server cluster based on determining that the first server cluster is unhealthy for the first data flow category.

[0070] Example 3. A distributed data processing system according to Examples 1 and 2, wherein the processor is further configured to: determine the first cluster health state for each of the multiple data flow categories associated with the data source, wherein the first cluster health state for the first data flow category is unhealthy; determine the second cluster health state for each of the multiple data flow categories associated with the data source, wherein the second cluster health state for the first data flow category is unhealthy; maintain a second push connection between the data source and the second server cluster, and wait for a subsequent second cluster health state of the second server cluster and a subsequent first cluster health state of the first server cluster; and determine whether to continue sending data to the second server cluster or initiate a failover to the first server cluster based on the subsequent second cluster health state of the second server cluster and the subsequent first cluster health state of the first server cluster.

[0071] Example 4. A distributed data processing system according to Examples 1-3, wherein each data flow category in the multiple data flow categories is hosted by multiple nodes, wherein a first node in the multiple nodes includes a leader partition for the original copy of the data, and a second node in the multiple nodes includes a slave partition for the first replicated partition of the data.

[0072] Example 5. A distributed data processing system according to Example 4, wherein the cluster management server determines a node health status for each of the multiple nodes, wherein the node health status is determined by sending a health check poll to each of the multiple nodes, and wherein the active node responds to the health check poll within a predetermined time period.

[0073] Example 6. The distributed data processing system of example 5, wherein the cluster management server determines a down node based on a failure to respond to three consecutive health check polls within a predetermined time period.

[0074] Example 7. A distributed data processing system according to Example 6, wherein each time a health response fails to meet the predetermined time period for responding, the predetermined time period for responding to the health check poll is reduced, and wherein for the active node, the predetermined time period for responding restores a default response.

[0075] Example 8. The distributed data processing system of Examples 4-7, wherein the second cluster health state and the first cluster health state are determined based on a minimum number of synchronously replicated partitions available at a plurality of active nodes for each of the plurality of data flow categories.

[0076] Example 9. A distributed data processing system according to Example 8, wherein the processor is further configured to: determine the number of active nodes from the plurality of nodes for each data stream category in the plurality of data stream categories, wherein the number of active nodes for each data stream category in the plurality of data stream categories represents the number of synchronized replicas available for each data stream category in the plurality of data stream categories; and compare the number of synchronized replicas available for each data stream category in the plurality of data stream categories with the minimum number of synchronized replica partitions.

[0077] Example 10. The distributed data processing system of Example 9, wherein at least a second data flow category is determined to be unhealthy and at least a third data flow category is determined to be healthy.

[0078] Example 11. A distributed data processing system according to Example 10, wherein the second data flow category is determined to be unhealthy based on a failure to meet the minimum number of synchronized replicas; and wherein the third data flow category is determined to be healthy based on meeting the minimum number of synchronized replicas.

[0079] Example 12. A distributed data processing system according to Example 10, wherein the second data stream category is determined to be unhealthy based on a failure to meet the minimum number of synchronized replicas or a crash of the leader partition for the first data stream category; and wherein the third data stream category is determined to be healthy based on meeting the minimum number of synchronized replicas, and at least one of the synchronized replicas is the leader partition for the second data stream category.

[0080] Example 13. A distributed data processing system according to Examples 1-12, wherein the cluster management server includes a failover switching buffer, which is configured to store data after the failover is initiated but before the failover is completed, and wherein the data cannot be published to the cluster system before the failover between the first server cluster and the second server cluster is completed.

[0081] Example 14. A distributed data processing system according to Examples 1-13, wherein the cluster management server includes an unpublished data buffer, which is configured to store data that is not published by the cluster system before initiating a failover, and wherein the data fails to be published to the first server cluster or the second server cluster.

[0082] Example 15. A distributed data processing system according to Examples 4-14, wherein the cluster management server is configured to initiate a test failover from the first server cluster to the second server cluster, wherein the test failover does not shut down any of the multiple nodes in the cluster system.

[0083] Example 16. The distributed data processing system of example 1, wherein at least one data flow class of the plurality of data flow classes is identified as a high priority data flow class.

[0084] Example 17. A distributed data processing system according to Example 16, wherein the processor is further configured to: initiate a failover associated with all data flow categories of the multiple data flow categories based on determining that the first server cluster is unhealthy for any of the data flow categories of the data source and determining that the second server cluster is healthy for all of the data flow categories of the data source.

[0085] Example 18. A method for cross-cluster failover management, the method comprising: determining, by a cluster management server, an initial first cluster health state of a first server cluster and an initial second cluster health state of a second server cluster; establishing, by the cluster management server, a first push connection between a data source and the first server cluster based on determining that the initial first cluster health state of the first server cluster is healthy; monitoring, by the cluster management server, data being sent from the data source to the first server cluster, wherein the first server cluster publishes the data to a plurality of data stream categories corresponding to the data; determining, by the cluster management server, a first cluster health state of the first server cluster for each of the plurality of data stream categories associated with the data source; determining, by the cluster management server, a second cluster health state of the second server cluster for each of the plurality of data stream categories associated with the data source; The management server determines that the first server cluster is unhealthy for a first data flow category among the multiple data flow categories associated with the data source; the cluster management server initiates a failover associated with the first data flow category for the data source, wherein the failover also includes: the cluster management server establishes a second push connection to the second server cluster based on determining that the second server cluster is healthy for the first data flow category associated with the data source; the cluster management server switches the data source from sending to the first server cluster to sending to the second server cluster based on determining that the first server cluster is unhealthy for the data source and determining that the second server cluster is healthy for the data source; and the cluster management server maintains a first pull data connection between the first server cluster and a first application instance, and a second pull data connection between the second server cluster and a second application instance.

[0086] Example 19. The method according to Example 18 further includes: determining, by the cluster management server, the first cluster health state for each of the multiple data flow categories associated with the data source, wherein the first cluster health state for the first data flow category is healthy; determining, by the cluster management server, the second cluster health state for each of the multiple data flow categories associated with the data source, wherein the second cluster health state for the first data flow category is healthy; initiating, by the cluster management server, a failover associated with the first data flow category of the data source, wherein the failover further includes: reestablishing, by the cluster management server, a connection to the first server cluster based on determining that the second server cluster is healthy for the first data flow category; switching, by the cluster management server, the data source from sending to the second server cluster back to sending to the first server cluster based on determining that the first server cluster is unhealthy for the first data flow category; and maintaining, by the cluster management server, a first pull data connection between the first server cluster and a consumer application, and a second pull data connection between the second server cluster and the consumer application.

[0087] Example 20. The method according to Examples 18 and 19 further includes: the cluster management server determining the first cluster health state for each of the multiple data flow categories associated with the data source, wherein the first cluster health state for the first data flow category is unhealthy; the cluster management server determining the second cluster health state for each of the multiple data flow categories associated with the data source, wherein the second cluster health state for the first data flow category is unhealthy; the cluster management server maintaining a second push connection between the data source and the second server cluster, and waiting for the subsequent second cluster health state of the second server cluster and the subsequent first cluster health state of the first server cluster; and the cluster management server determining whether to continue sending data to the second server cluster or initiate a failover to the first server cluster based on the subsequent second cluster health state of the second server cluster and the subsequent first cluster health state of the first server cluster.

[0088] The foregoing detailed description has been described using block diagrams, flow charts, and / or examples to illustrate various forms of systems and / or processes. To the extent that such block diagrams, flow charts, and / or examples include one or more functions and / or operations, those skilled in the art will understand that each function and / or operation within such block diagrams, flow charts, and / or examples can be implemented individually and / or collectively by a wide range of hardware, software, firmware, or any combination thereof. Those skilled in the art will recognize that some aspects of the forms disclosed herein can be equivalently implemented in whole or in part in an integrated circuit as one or more computer programs running on one or more computers (e.g., as one or more programs running on one or more computer systems), as one or more programs running on one or more processors (e.g., as one or more programs running on one or more microprocessors), as firmware, or as any combination thereof, and recognize that designing circuit systems and / or writing code for software and / or firmware in accordance with the present disclosure will be well within the skill of those skilled in the art. Additionally, those skilled in the art will understand that the mechanisms of the subject matter described herein are capable of being distributed as one or more program products in various forms, and will understand that the illustrative forms of the subject matter described herein apply regardless of the particular type of signal-bearing medium used to actually perform the distribution.

[0089] Instructions for programming the logic to perform the various disclosed aspects may be stored in a memory in the system, such as a dynamic random access memory (DRAM), a cache, a flash memory, or other storage device. In addition, the instructions may be distributed via a network or with the aid of other computer-readable media. Thus, a machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), but is not limited to a floppy disk, an optical disk, a compact disc read-only memory (CD-ROM) and a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic or optical card, a flash memory, or a tangible machine-readable storage device for transmitting information via an electrical, optical, acoustic, or other form of propagation signal (e.g., a carrier wave, an infrared signal, a digital signal, etc.) via the Internet. Thus, a non-transitory computer-readable medium includes any type of tangible machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0090] Any software component or function described in this application can be implemented as a software code executed by a processor using any suitable computer language (e.g., Python, Java, C++, or Perl) using, for example, conventional or object-oriented techniques. The software code can be stored as a series of instructions or commands on a computer-readable medium such as RAM, ROM, magnetic media (e.g., hard disk or floppy disk), or optical media (e.g., CD-ROM). Any such computer-readable medium can reside on or within a single computing device and can exist on or within different computing devices within a system or network.

[0091] As used in any aspect herein, the term "logic" may refer to an application, software, firmware, and / or circuitry configured to perform any of the aforementioned operations. Software may be embodied as a software package, code, instructions, instruction sets, and / or data recorded on a non-transitory computer-readable storage medium. Firmware may be embodied as code, instructions, instruction sets, and / or data hard-coded (e.g., non-volatile) in a memory device.

[0092] As used in any aspect herein, the terms "component," "system," "module," and the like may refer to a computer-related entity, either hardware, a combination of hardware and software, software, or software in execution.

[0093] As used in any aspect herein, an "algorithm" refers to a self-consistent sequence of steps leading to a desired result, where a "step" refers to manipulations of physical quantities and / or logical states, which may, though not necessarily, take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. These signals are typically referred to as bits, values, elements, symbols, characters, terms, numbers, or the like. These and similar terms may be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities and / or states.

[0094] The network may include a packet-switched network. The communication devices may be able to communicate with each other using a selected packet-switched network communication protocol. An exemplary communication protocol may include an Ethernet communication protocol, which may allow communication using the Transmission Control Protocol / Internet Protocol (TCP / IP). The Ethernet protocol may conform to or be compatible with the Ethernet standard entitled "IEEE 802.3 Standard (IEEE 802.3 Standard)" and / or subsequent versions of this standard, published by the Institute of Electrical and Electronics Engineers (IEEE) in December 2008. Alternatively or in addition, the communication devices may be able to communicate with each other using the X.25 communication protocol. The X.25 communication protocol may conform to or be compatible with the standard promulgated by the International Telecommunication Union-Telecommunication Standardization Sector (ITU-T). Alternatively or in addition, the communication devices may be able to communicate with each other using the Frame Relay communication protocol. The frame relay communication protocol may conform to or be compatible with standards promulgated by the Consultative Committee for International Telegraph and Telephone (CCITT) and / or the American National Standards Institute (ANSI). Alternatively or additionally, the transceivers may communicate with each other using an asynchronous transfer mode (ATM) communication protocol. The ATM communication protocol may conform to or be compatible with the ATM standard entitled "ATM-MPLS Network Interworking 2.0," published by the ATM Forum in August 2001, and / or subsequent versions of such standard. Of course, different and / or later development connection-oriented network communication protocols are also contemplated herein.

[0095] Unless expressly indicated otherwise in the foregoing disclosure, it should be understood that throughout this disclosure, discussions using terms such as "process," "compute," "calculate," "determine," "display," or the like refer to the actions and processes of a computer system, or similar electronic computing device, that manipulates data represented as physical (electronic) quantities within the computer system's registers and memories and transforms it into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage, transmission, or display devices.

[0096] One or more components may be referred to herein as being "configured to," "configurable to," "operable / operable to," "suitable / adaptable to," "capable of," "compliant / compliant with," etc. Those skilled in the art will recognize that "configured to" may generally encompass active state components and / or inactive state components and / or standby state components, unless the context requires otherwise.

[0097] Those skilled in the art will recognize that, in general, the terms used herein, and particularly in the appended claims (e.g., the appended claim bodies), are generally intended as "open" terms (e.g., the term "including" should be interpreted as "including but not limited to," the term "having" should be interpreted as "having at least," the term "includes" should be interpreted as "includes but is not limited to," etc.). Those skilled in the art will further understand that if a specific number of an introduced claim recitation is intended, such intent will be expressly stated in the claim, and in the absence of such recitation, no such intent is present. For example, to aid understanding, the following appended claims may contain usage of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be construed as meaning that introducing a claim recitation by the indefinite article "a" or "an" limits any particular claim containing such introduced claim recitation to claims containing only one such recitation, even when the same claim includes the introductory phrase "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" should generally be construed to mean "at least one" or "one or more"); the same is true for the use of definite articles to introduce claim recitations.

[0098] In addition, even if a specific number of introduced claim statements is explicitly recited, those skilled in the art will recognize that such recitations should generally be interpreted to mean at least the recited number (e.g., simply reciting "two statements" without other modifiers generally means at least two statements, or two or more statements). Moreover, in those instances where a convention similar to "at least one of A, B, and C, etc." is used, generally, such construction is intended in the sense that one skilled in the art would understand the convention (e.g., "a system having at least one of A, B, and C" would include but is not limited to systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In those instances where a convention similar to "at least one of A, B, or C, etc." is used, generally, such construction is intended in the sense that one skilled in the art would understand the convention (e.g., "a system having at least one of A, B, or C" would include but is not limited to systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those skilled in the art that, whether in the specification, claims or drawings, alternative words and / or phrases generally presenting two or more alternative terms should be understood to include the possibility of one, any or both of the terms, unless the context dictates otherwise. For example, the phrase "A or B" will generally be understood to include the possibility of "A" or "B" or "A and B".

[0099] With respect to the appended claims, it will be understood by those skilled in the art that the operations described therein can generally be performed in any order. In addition, although the various operational flow charts are presented in (multiple) sequences, it will be understood that the various operations can be performed in other orders than those shown, or can be performed simultaneously. Examples of such alternative orderings may include overlapping, interleaved, interrupted, reordered, incremental, preliminary, supplementary, simultaneous, reversed, or other variations of ordering, unless the context dictates otherwise. In addition, terms such as "responsive to," "related to," or other past tense adjectives are generally not intended to exclude such variations, unless the context dictates otherwise.

[0100] It is important to note that any reference to "one aspect," "an aspect," "an example," "an example," and the like means that a particular feature, structure, or characteristic described in connection with the aspect is included in at least one aspect. Thus, the appearances of the phrases "in one aspect," "in an aspect," "in an example," and "in an example" in various places throughout this specification are not necessarily all referring to the same aspect. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more aspects.

[0101] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.

[0102] Any patent application, patent, non-patent publication, or other public material cited in this specification and / or listed in any application data sheet is incorporated herein by reference to the extent that the incorporated material is not inconsistent therewith. Thus, and to the extent necessary, the disclosure as expressly set forth herein supersedes any conflicting material incorporated by reference. Any material or portion thereof that is stated to be incorporated herein by reference but conflicts with existing definitions, statements, or other public material set forth herein will be incorporated only to the extent that there is no conflict between the incorporated material and the existing public material. No admission is made that they are prior art.

[0103] In summary, many benefits have been described that result from employing the concepts described herein. The foregoing description of one or more forms has been presented for purposes of illustration and description. It is not intended to be exhaustive or limited to the precise forms disclosed. Modifications and variations are possible in light of the above teachings. One or more forms have been selected and described to illustrate the principles and practical applications, thereby enabling one of ordinary skill in the art to utilize the various forms and make various modifications as appropriate for the specific application contemplated. The claims submitted herein are intended to define the overall scope.

Claims

1. A distributed data processing system comprising: A cluster management server configured to communicate with a cluster system including a first server cluster and a second server cluster in geographically different regions, the cluster management server comprising a processor and a memory coupled to the processor, the memory having stored thereon machine-executable instructions that, when executed, cause the processor to: determining an initial health status of the first server cluster; Based on determining that the first server cluster is healthy, establishing a first push connection between the data source and the first server cluster; Monitoring data is sent from the data source to the first server cluster, wherein the first server cluster publishes the data to a plurality of data flow categories corresponding to the data; determining a first cluster health state of the first server cluster for each data flow category of the plurality of data flow categories associated with the data source; determining a second cluster health state of the second server cluster for each data flow category of the plurality of data flow categories associated with the data source; determining that the first server cluster is unhealthy for a first data flow category of the plurality of data flow categories associated with the data source; Initiating a failover associated with a first data flow category for the data source, wherein the failover causes the processor to: establishing a second push connection to the second server cluster based on determining that the second server cluster is healthy for the first data flow category associated with the data source; as well as Based on a determination that the first server cluster is unhealthy for the data source and a determination that the second server cluster is healthy for the data source, the data source is switched from being sent to the first server cluster to being sent to the second server cluster.

2. The distributed data processing system of claim 1 , wherein the processor is further configured to: determining the first cluster health state for each of the plurality of data flow categories associated with the data source, wherein the first cluster health state for the first data flow category is healthy; determining the second cluster health state for each of the plurality of data flow categories associated with the data source, wherein the second cluster health state for the first data flow category is healthy; Initiating a failover associated with a first data flow category of the data source, wherein the failover causes the processor to: reestablishing a first push connection to the first server cluster based on determining that the second server cluster is healthy for the first data flow category; and Based on determining that the first server cluster is unhealthy for the first data flow category, the data source is switched from sending to the second server cluster back to sending to the first server cluster.

3. The distributed data processing system of claim 1 , wherein the processor is further configured to: determining the first cluster health state for each of the plurality of data flow categories associated with the data source, wherein the first cluster health state for the first data flow category is unhealthy; determining a second cluster health state for each of the plurality of data flow categories associated with the data source, wherein the second cluster health state for the first data flow category is unhealthy; maintaining the second push connection between the data source and the second server cluster, and waiting for a subsequent second cluster health status of the second server cluster and a subsequent first cluster health status of the first server cluster; as well as Based on the subsequent second cluster health status of the second server cluster and the subsequent first cluster health status of the first server cluster, it is determined whether to continue sending data to the second server cluster or to initiate a failover to the first server cluster.

4. A distributed data processing system according to claim 1, wherein each data flow category in the plurality of data flow categories is hosted by a plurality of nodes, wherein a first node in the plurality of nodes includes a leader partition for the original copy of the data, and a second node in the plurality of nodes includes a slave partition for the first replicated partition of the data.

5. A distributed data processing system according to claim 4, wherein the cluster management server determines a node health status for each of the multiple nodes, wherein the node health status is determined by sending a health check poll to each of the multiple nodes, and wherein the active node responds to the health check poll within a predetermined time period. 6 . The distributed data processing system of claim 5 , wherein the cluster management server determines a down node based on a failure to respond to three consecutive health check polls within a predetermined time period.

7. A distributed data processing system according to claim 6, wherein each time a health response fails to meet the predetermined time period for responding, the predetermined time period for responding to the health check poll is reduced, and wherein for the active node, the predetermined time period for responding restores a default response.

8. The distributed data processing system of claim 4, wherein the second cluster health state and the first cluster health state are determined based on a minimum number of synchronously replicated partitions available at a plurality of active nodes for each of the plurality of data flow categories.

9. The distributed data processing system of claim 8, wherein the processor is further configured to: determining, from the plurality of nodes, for each of the plurality of data flow categories, a number of active nodes, wherein the number of active nodes for each of the plurality of data flow categories represents a number of synchronized replicas available for each of the plurality of data flow categories; and The number of synchronous replicas available for each of the plurality of data flow classes is compared to a minimum number of synchronously replicated partitions.

10. The distributed data processing system of claim 9, wherein at least a second data flow category is determined to be unhealthy and at least a third data flow category is determined to be healthy.

11. A distributed data processing system according to claim 10, wherein the second data flow category is determined to be unhealthy based on failure to meet the minimum number of synchronized replicas; and wherein the third data flow category is determined to be healthy based on meeting the minimum number of synchronized replicas.

12. A distributed data processing system according to claim 10, wherein the second data stream category is determined to be unhealthy based on failure to meet the minimum number of synchronized replicas or the leader partition for the first data stream category is down; and wherein the third data stream category is determined to be healthy based on meeting the minimum number of synchronized replicas, and at least one of the synchronized replicas is the leader partition for the second data stream category.

13. A distributed data processing system according to claim 1, wherein the cluster management server includes a failover switching buffer, the failover switching buffer is configured to store data after the failover is initiated but before the failover is completed, and wherein the data cannot be published to the cluster system before the failover between the first server cluster and the second server cluster is completed.

14. A distributed data processing system according to claim 1, wherein the cluster management server includes an unpublished data buffer, the unpublished data buffer is configured to store data that is not published by the cluster system before initiating a failover, and wherein the data fails to be published to the first server cluster or the second server cluster.

15. The distributed data processing system of claim 4, wherein the cluster management server is configured to initiate a test failover from the first server cluster to the second server cluster, wherein the test failover does not shut down any of the plurality of nodes in the cluster system.

16. The distributed data processing system of claim 1, wherein at least one data flow class of the plurality of data flow classes is identified as a high priority data flow class.

17. The distributed data processing system of claim 16, wherein the processor is further configured to: Based on determining that the first server cluster is unhealthy for any of the data flow classes for the data source and determining that the second server cluster is healthy for all of the data flow classes for the data source, a failover associated with all of the plurality of data flow classes is initiated.

18. A method for cross-cluster failover management, the method comprising: determining, by the cluster management server, an initial first cluster health state of the first server cluster and an initial second cluster health state of the second server cluster; establishing, by the cluster management server, a first push connection between a data source and the first server cluster based on determining that an initial first cluster health state of the first server cluster is healthy; monitoring, by the cluster management server, data being sent from the data source to the first server cluster, wherein the first server cluster publishes the data to a plurality of data flow categories corresponding to the data; determining, by the cluster management server, a first cluster health status of the first server cluster for each of the plurality of data flow categories associated with the data source; determining, by the cluster management server, a second cluster health status of the second server cluster for each of the plurality of data flow categories associated with the data source; determining, by the cluster management server, that the first server cluster is unhealthy for a first data flow category of the plurality of data flow categories associated with the data source; Initiating, by the cluster management server, a failover associated with a first data flow category for the data source, wherein the failover further comprises: establishing, by the cluster management server, a second push connection to the second server cluster based on determining that the second server cluster is healthy for the first data flow category associated with the data source; Switching, by the cluster management server, sending the data source from the first server cluster to the second server cluster based on determining that the first server cluster is unhealthy for the data source and determining that the second server cluster is healthy for the data source; and A first pull data connection between the first server cluster and a first application instance and a second pull data connection between the second server cluster and a second application instance are maintained by the cluster management server.

19. The method according to claim 18, further comprising: determining, by the cluster management server, the first cluster health state for each of the plurality of data flow categories associated with the data source, wherein the first cluster health state for the first data flow category is healthy; determining, by the cluster management server, a second cluster health state for each of the plurality of data flow categories associated with the data source, wherein the second cluster health state for the first data flow category is healthy; Initiating, by the cluster management server, a failover associated with the first data flow category of the data source, wherein the failover further comprises: reestablishing, by the cluster management server, a first push connection to the first server cluster based on determining that the second server cluster is healthy for the first data flow category; The cluster management server switches the data source from being sent to the second server cluster back to being sent to the first server cluster based on determining that the first server cluster is unhealthy for the first data flow category; and A first pull data connection between the first server cluster and the first application instance and a second pull data connection between the second server cluster and the second application instance are maintained by the cluster management server.

20. The method of claim 18, further comprising: determining, by the cluster management server, a first cluster health state for each of the plurality of data flow categories associated with the data source, wherein the first cluster health state for the first data flow category is unhealthy; determining, by the cluster management server, a second cluster health state for each of the plurality of data flow categories associated with the data source, wherein the second cluster health state for the first data flow category is unhealthy; maintaining, by the cluster management server, a second push connection between the data source and the second server cluster, and waiting for a subsequent second cluster health status of the second server cluster and a subsequent first cluster health status of the first server cluster; as well as The cluster management server determines whether to continue sending data to the second server cluster or to initiate a failover to the first server cluster based on the subsequent second cluster health status of the second server cluster and the subsequent first cluster health status of the first server cluster.