Forcibly Completing Distributed Software Upgrades Amid Node Failures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
During a rolling upgrade of distributed software in a cluster, the presence of inaccessible nodes due to failures like hardware or software issues disrupts the process, requiring administrators to terminate the upgrade and downgrade the entire cluster, leading to significant downtime and productivity losses in business-critical environments.
Innovation Solution
A system that allows for the forced upgrade of a cluster by detecting inaccessible nodes and continuing the upgrade process on remaining nodes, enabling the installation and activation of a newer version of distributed software, while preventing inaccessible nodes from joining the cluster until they are upgraded, thus maintaining cluster operation and enabling new features and bug fixes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the administrator terminates the cluster upgrade process to handle inaccessible nodes, then the cluster reliability is maintained, but the productivity and upgrade completion time deteriorate due to full cluster downtime
Solution Approach 1:
The patent segments the cluster into accessible and inaccessible nodes, allowing the upgrade process to continue on accessible nodes while isolating failures to specific segments. This enables partial upgrade completion without requiring full cluster termination, resolving the contradiction between maintaining reliability and completing upgrades.
Solution Approach 2:
The patent enables continuous upgrade action on accessible nodes despite the presence of inaccessible nodes. The upgrade process continues uninterrupted on available nodes, maintaining productivity while reliability is preserved through proper handling of inaccessible nodes. This eliminates the need to terminate the entire upgrade process.
2Stability of the object's composition
If the administrator downgrades the entire cluster to handle inaccessible nodes, then the cluster stability is maintained, but the loss of time and productivity worsen due to manual intervention and full outage
Solution Approach 1:
The patent segments the cluster nodes into accessible and inaccessible groups, allowing stable operation to continue on accessible nodes while isolating issues to specific nodes. This prevents the need for full cluster downgrade and minimizes downtime, resolving the contradiction between stability and time loss.
Solution Approach 2:
The patent applies partial action by performing upgrades only on accessible nodes rather than requiring all nodes to be upgraded together. This partial completion maintains cluster stability while avoiding the time loss associated with full cluster downtime and manual intervention.
3Stability of the object's composition
If the cluster requires all nodes to be upgraded together, then the acting version consistency is maintained, but the productivity deteriorates due to the need to wait for all nodes including inaccessible ones
Solution Approach 1:
The patent segments the upgrade process into node-specific operations, allowing each accessible node to be upgraded independently. This maintains acting version consistency on upgraded nodes while eliminating the need to wait for inaccessible nodes, resolving the contradiction between consistency and upgrade speed.
Solution Approach 2:
The patent performs partial upgrades on accessible nodes without waiting for inaccessible nodes to become available. This partial action approach maintains version consistency on upgraded nodes while significantly improving upgrade speed by not being blocked by inaccessible nodes.
4Productivity
If inaccessible nodes are removed from the cluster during upgrade, then the upgrade process can complete, but the device complexity increases due to manual removal and readdition operations
Solution Approach 1:
The patent implements self-service by enabling inaccessible nodes to automatically rejoin the cluster after becoming accessible, without requiring manual removal and readdition operations. This reduces operation complexity while maintaining upgrade completion, resolving the contradiction between productivity and device complexity.
Data Source
AI summary
One embodiment of the present invention provides a system for facilitating an upgrade of a cluster of servers in the presence of one or more inaccessible nodes in the cluster. During operation, the system upgrades a version of a distributed software program on each of a plurality of nodes in the cluster. The system may detect that one or more nodes of the cluster are inaccessible. The system continues to upgrade nodes in the cluster other than the one or more nodes that were detected to be inaccessible, in which upgrading involves installing and activating a newer version of the distributed software on the nodes being upgraded. The system then upgrades an acting version of the cluster.


