Flexible Witness Service Architecture for Distributed System Arbitration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In distributed systems, maintaining a witness node in a datacenter is cumbersome due to the need for frequent bug fixing, security updates, configuration backups, and high administrative overhead, especially in scenarios like split-brain where equal network partitions require a third-site witness for decision-making.
Innovation Solution
A flexible witness service system architecture that allows the witness server to be deployed on multiple platforms, enabling it to run in the cloud, on local external devices, or embedded systems. This architecture includes a local witness service and a cloud witness service that perform identical arbitration functions, accessible through a uniform service interface, allowing for portability and deployment on various computing platforms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the witness node is placed in a datacenter, then the system can provide fault tolerance and continuous availability, but the administrative overhead and maintenance complexity increase significantly
Solution Approach 1:
The witness service is designed to be self-managing with automated health checks, self-healing capabilities, and automatic failover. The service monitors its own status and can recover from failures without manual intervention, eliminating the need for frequent bug fixing and security updates mentioned in the background.
Solution Approach 2:
The witness service is implemented as a multi-functional platform that can serve multiple clusters and partitions simultaneously. It provides not only arbitration functions but also health monitoring, logging, and automated recovery capabilities, reducing the need for separate infrastructure components.
2Reliability
If the witness node is placed in a datacenter, then the system can handle split-brain scenarios, but the physical infrastructure requirements and operational complexity increase
Solution Approach 1:
Instead of requiring physical witness hardware in datacenters, the invention uses virtualized witness instances that can be replicated across multiple locations. These virtual witnesses are software-based and can run on standard infrastructure, eliminating the need for specialized physical infrastructure while maintaining split-brain handling capabilities.
Solution Approach 2:
The witness service transitions from a single physical location model to a distributed multi-dimensional architecture where witness instances can be placed in cloud, on-premises, or hybrid environments. This dimensional flexibility allows the system to handle split-brain scenarios without being constrained by physical infrastructure limitations.
3Reliability
If the witness service requires frequent updates and maintenance, then the system can maintain security and stability, but the service availability and operational simplicity decrease
Solution Approach 1:
The witness service implements proactive security measures including pre-configured security policies, automated vulnerability scanning, and staged update deployment. Security updates are tested in isolated environments before being applied to production witnesses, ensuring security maintenance without service disruption.
Solution Approach 2:
The system employs periodic health checks, automated security scanning, and scheduled maintenance windows that minimize impact on service availability. Updates are deployed during predetermined maintenance periods, and the service automatically recovers if issues are detected, maintaining both security and availability.
Data Source
AI summary
A flexible witness service system architecture is provided that comprises one or more cluster sites each having at least two storage/compute nodes; at least one local external device associated with at least one of the one or more cluster sites, the at least one local external device configured to run a local witness service. A central cloud management platform is in communication with the one or more cluster sites, the central cloud management platform being configured to run a cloud witness service. The local witness service and the cloud witness service perform identical arbitration services if a storage/compute node in one of the one or more cluster sites fails or communication between storage/compute nodes in a cluster fails.


