Forcibly Completing Distributed Software Upgrades Amid Node Failures

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

During a rolling upgrade of distributed software in a cluster, the presence of inaccessible nodes due to failures like hardware or software issues disrupts the process, requiring administrators to terminate the upgrade and downgrade the entire cluster, leading to significant downtime and productivity losses in business-critical environments.

Innovation Solution

A system that allows for the forced upgrade of a cluster by detecting inaccessible nodes and continuing the upgrade process on remaining nodes, enabling the installation and activation of a newer version of distributed software, while preventing inaccessible nodes from joining the cluster until they are upgraded, thus maintaining cluster operation and enabling new features and bug fixes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the administrator terminates the cluster upgrade process to handle inaccessible nodes, then the cluster reliability is maintained, but the productivity and upgrade completion time deteriorate due to full cluster downtime

Engineering Contradiction:
Improvecluster reliabilityVSAvoidupgrade completion time
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the cluster into accessible and inaccessible nodes, allowing the upgrade process to continue on accessible nodes while isolating failures to specific segments. This enables partial upgrade completion without requiring full cluster termination, resolving the contradiction between maintaining reliability and completing upgrades.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent enables continuous upgrade action on accessible nodes despite the presence of inaccessible nodes. The upgrade process continues uninterrupted on available nodes, maintaining productivity while reliability is preserved through proper handling of inaccessible nodes. This eliminates the need to terminate the entire upgrade process.

Inventive Principle:
Principle #20Continuity of useful action

2Stability of the object's composition

If the administrator downgrades the entire cluster to handle inaccessible nodes, then the cluster stability is maintained, but the loss of time and productivity worsen due to manual intervention and full outage

Engineering Contradiction:
Improvecluster stabilityVSAvoiddowntime
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The patent segments the cluster nodes into accessible and inaccessible groups, allowing stable operation to continue on accessible nodes while isolating issues to specific nodes. This prevents the need for full cluster downgrade and minimizes downtime, resolving the contradiction between stability and time loss.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by performing upgrades only on accessible nodes rather than requiring all nodes to be upgraded together. This partial completion maintains cluster stability while avoiding the time loss associated with full cluster downtime and manual intervention.

Inventive Principle:
Principle #16Partial or excessive action

3Stability of the object's composition

If the cluster requires all nodes to be upgraded together, then the acting version consistency is maintained, but the productivity deteriorates due to the need to wait for all nodes including inaccessible ones

Engineering Contradiction:
Improveacting version consistencyVSAvoidupgrade speed
Core Design Contradiction:
Stability of the object's compositionVSProductivity

Solution Approach 1:

The patent segments the upgrade process into node-specific operations, allowing each accessible node to be upgraded independently. This maintains acting version consistency on upgraded nodes while eliminating the need to wait for inaccessible nodes, resolving the contradiction between consistency and upgrade speed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs partial upgrades on accessible nodes without waiting for inaccessible nodes to become available. This partial action approach maintains version consistency on upgraded nodes while significantly improving upgrade speed by not being blocked by inaccessible nodes.

Inventive Principle:
Principle #16Partial or excessive action

4Productivity

If inaccessible nodes are removed from the cluster during upgrade, then the upgrade process can complete, but the device complexity increases due to manual removal and readdition operations

Engineering Contradiction:
Improveupgrade completionVSAvoidoperation complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements self-service by enabling inaccessible nodes to automatically rejoin the cluster after becoming accessible, without requiring manual removal and readdition operations. This reduces operation complexity while maintaining upgrade completion, resolving the contradiction between productivity and device complexity.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10019250B2Forcibly completing upgrade of distributed software in presence of failures
Publication Date: 2018.07.10 ORACLE INT CORP
  • US10019250B2 patent drawing
  • US10019250B2 patent drawing
  • US10019250B2 patent drawing

AI summary

One embodiment of the present invention provides a system for facilitating an upgrade of a cluster of servers in the presence of one or more inaccessible nodes in the cluster. During operation, the system upgrades a version of a distributed software program on each of a plurality of nodes in the cluster. The system may detect that one or more nodes of the cluster are inaccessible. The system continues to upgrade nodes in the cluster other than the one or more nodes that were detected to be inaccessible, in which upgrading involves installing and activating a newer version of the distributed software on the nodes being upgraded. The system then upgrades an acting version of the cluster.