Cluster Arbitration via Preemption Representatives
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In active-active data centers, the different arbitration mechanisms used by clusters can lead to inconsistent results when faults occur, resulting in probabilistic service interruptions as some clusters may survive in one data center while others survive in the other, causing inconsistent service access.
Innovation Solution
A cluster arbitration method where each cluster group determines a preemption representative to preempt an arbitration device, ensuring consistent arbitration results by having all sub-clusters in the successful cluster group continue service provision, with mechanisms in place for both clusters to attempt preemption if faults occur in both groups.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If each cluster uses its own arbitration mechanism independently, then each cluster can autonomously handle faults, but arbitration results become inconsistent across clusters leading to service interruptions
Solution Approach 1:
The patent merges multiple independent arbitration mechanisms into a unified arbitration system. A shared arbitration device is introduced that coordinates fault handling across all clusters, ensuring consistent arbitration results. The arbitration device receives fault information from different clusters and coordinates their arbitration processes, transforming independent operations into a coordinated unified action that maintains service continuity.
Solution Approach 2:
The patent introduces an arbitration device as an intermediary between clusters. This mediator receives fault notifications from various clusters, manages the arbitration process centrally, and coordinates the survival decisions across clusters. The intermediary ensures that arbitration results are consistent and prevents service interruptions by harmonizing the autonomous fault handling of different clusters.
2Reliability
If all clusters use a unified arbitration mechanism, then arbitration results are consistent across clusters, but system complexity increases
Solution Approach 1:
The patent segments the arbitration system into two distinct parts: a centralized arbitration device that manages coordination and consistency, and local cluster components that execute specific arbitration tasks. This segmentation allows the unified arbitration mechanism to maintain consistency while keeping the complexity localized and manageable. The arbitration device handles the complex coordination logic, while clusters follow standardized procedures.
Solution Approach 2:
The arbitration device is designed as a universal component that handles arbitration for multiple clusters simultaneously. It performs multiple functions including receiving fault information from different clusters, managing arbitration processes, and coordinating survival decisions. This multi-functionality reduces overall system complexity by consolidating arbitration capabilities into a single versatile device rather than requiring separate mechanisms for each cluster.
3Reliability
If one data center's cluster survives a fault, then service is maintained in that data center, but inconsistent arbitration may cause service interruption in other data centers
Solution Approach 1:
The patent implements a feedback mechanism where the arbitration device receives fault information from clusters across different data centers and uses this information to coordinate arbitration decisions. The arbitration device processes feedback from various clusters and adjusts its coordination to ensure consistent arbitration results. This feedback loop ensures that service availability in one data center does not compromise service access consistency in other data centers.
Solution Approach 2:
The patent creates equipotentiality in the arbitration process across all data centers by introducing a centralized arbitration device that applies the same arbitration logic uniformly. This ensures that all clusters, regardless of which data center they belong to, are treated equally during fault conditions. The arbitration device balances the arbitration process, preventing situations where one data center's cluster survival creates inconsistency in other data centers.
Data Source
Figure 1
Figure 2
AI summary
Embodiments of the present invention disclose a cluster arbitration method and a multi-cluster cooperation system. The method in the embodiments of the present invention includes: detecting whether a fault has occurred in a first cluster group or a second cluster group, where the first cluster group includes one portion of a first cluster and one portion of a second cluster, and the second cluster group includes another portion of the first cluster and another portion of the second cluster, and the first cluster and the second cluster cooperate with each other; when detecting that a fault has occurred, determining, by the first cluster group and the second cluster group, respective preemption representatives, where both the preemption representative of the first cluster group and the preemption representative of the second cluster group perform the following steps: determining whether a fault has occurred in the respective cluster group, and if no fault has occurred in the respective cluster group, attempting to preempt an arbitration device, where a cluster group whose preemption representative has successfully preempted the arbitration device according to a preset arbitration mechanism survives. The present invention can reduce a probability of interruption of service access.