Clustered Node Software Upgrade via Site Control Transfer

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data protection systems face challenges such as system shutdown during backups, limited recovery points, and lengthy recovery processes, which result in potential data loss and downtime in the event of a disaster.

Innovation Solution

A method and system for upgrading software on nodes in a clustered environment, allowing seamless software upgrades on data protection appliances while maintaining site control, enabling independent upgrades and ensuring continuous operation by transferring site control between nodes, thus minimizing downtime and enabling recovery to any specified point in time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If software is upgraded on a node in a clustered environment, then the node can run newer software versions with improved functionality, but the node must shut down processes during upgrade causing downtime and loss of site control

Engineering Contradiction:
Improvesoftware versionVSAvoiddowntime
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system divides the clustered environment into multiple independent nodes, each capable of running software upgrades independently. The upgrade process is segmented across nodes rather than requiring system-wide shutdown, allowing continuous operation through node-level independence.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by designating a target node for upgrade before actual software installation. The target node is prepared and isolated from active site control duties beforehand, allowing upgrade processes to begin without forcing immediate shutdown of critical operations. Other nodes continue serving until the upgrade is complete.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 3:

A non-voting node serves as an intermediary during the upgrade process, taking over site control temporarily while the target node undergoes software upgrade. This intermediary node mediates between the need for continuous site control and the requirement for uninterrupted upgrade processes, preventing downtime.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If software is upgraded on a node, then the node can provide updated data protection capabilities, but the upgrade process may cause loss of site control and require system interruption

Engineering Contradiction:
Improvedata protection capabilityVSAvoidsite control continuity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system implements dynamic role assignment where nodes can transition between active and target statuses during upgrades. Site control is dynamically transferred to non-voting nodes when voting nodes are undergoing upgrades, and dynamically returned when upgrades complete. This dynamic flexibility maintains operational ease while enabling reliability improvements through updated software.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes operational parameters by allowing nodes to operate in different modes: active voting nodes, target nodes undergoing upgrade, and non-voting nodes providing temporary site control. These parameter changes enable the system to maintain data protection reliability through software updates without permanently compromising site control continuity.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If conventional backup systems are used, then data can be stored periodically, but the system must shut down during backup causing data unavailability

Engineering Contradiction:
Improvedata storage capacityVSAvoiddata availability
Core Design Contradiction:
Quantity of substanceVSSpeed

Solution Approach 1:

The system ensures continuous useful action by maintaining data protection capabilities throughout the upgrade process. While one node undergoes software upgrade, other nodes continue providing data protection services without interruption. This continuity eliminates the data unavailability period that plagues conventional backup systems requiring system shutdown.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system performs preliminary actions by pre-designating target nodes for upgrade and pre-arranging site control transfer to non-voting nodes. This preliminary preparation ensures that data protection operations can continue uninterrupted on other nodes while the upgrade proceeds, maintaining data availability throughout the process.

Inventive Principle:
Principle #10Preliminary action

4Reliability

If data replication is used for continuous protection, then recovery points can be maintained, but the recovery process still takes considerable time

Engineering Contradiction:
Improverecovery capabilityVSAvoidrecovery time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by maintaining ready-to-use backup nodes that are pre-configured and standing by during upgrade processes. When a node requires recovery, pre-positioned backup data and pre-configured recovery mechanisms enable immediate restoration, significantly reducing recovery time compared to conventional sequential backup restoration processes.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9286052B1Upgrading software on a pair of nodes in a clustered environment
Publication Date: 2016.03.15 EMC INT
  • US9286052B1 patent drawing
  • US9286052B1 patent drawing
  • US9286052B1 patent drawing

AI summary

In one aspect, a method to upgrade software on nodes in a clustered environment, includes terminating processes on a first node before upgrading the software on the first node, upgrading the software to a first version from a second version on the first node, running the processes on the first node after upgrading the software on the first node to the first version, determining whether a second node is about to upgrade to the first version of software, allowing transfer of site control from the second node to the first node, if the second node is about to upgrade to the first version of software and upgrading the software on the second node to the first version of software after the transferring of site control from the second node to the first node.