Appliance Switchover for SDDC Upgrade Downtime Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing software-defined data centers (SDDCs) face challenges in reducing downtime during application upgrades, managing high availability, and handling network partitions effectively.
Innovation Solution
The method involves deploying a second appliance with upgraded services, setting up a preemptive pair with fault domain management, performing a switchover to activate the upgraded services, and managing high availability and error handling during the upgrade process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If a traditional upgrade approach is used where the application is stopped and restarted with the new version, then the upgrade process is simple, but the downtime is long (hours)
Solution Approach 1:
The system is segmented into multiple appliances (first appliance with old version, second appliance with new version) that can operate independently. This allows the upgrade to proceed in stages without requiring a complete system shutdown, thereby reducing downtime while managing complexity through structured division of functions.
Solution Approach 2:
The second appliance with the new version is deployed and configured in advance before the actual switchover. This preliminary setup includes creating the preemptive pair relationship and configuring fault domain management, so that when the switchover occurs, it can happen quickly with minimal downtime.
2Reliability
If high availability is maintained during upgrade by having standby appliances, then availability is improved, but the system complexity increases due to additional components
Solution Approach 1:
The upgrade process and high availability mechanism are merged into a unified approach using preemptive pairs. The same appliance pair structure serves both as the HA backup mechanism and as the upgrade pathway, eliminating the need for separate upgrade infrastructure and reducing overall system complexity while maintaining reliability.
Solution Approach 2:
The preemptive pair structure serves multiple functions: it provides high availability through failover capability, enables seamless upgrades through controlled switchover, and maintains fault domain management. This multi-functionality reduces the need for additional specialized components, thereby managing complexity while improving reliability.
3Reliability
If fault domain management is used to protect against failures, then reliability is improved, but the complexity of managing failover scenarios increases
Solution Approach 1:
The fault domain management system automatically handles failover decisions and execution without requiring manual intervention. When a failure is detected, the system self-manages the recovery process by activating the appropriate appliance, thereby improving reliability while reducing the operational complexity of managing failover scenarios.
4Reliability
If a preemptive pair is set up with one protected and one unprotected appliance, then failover capability is improved, but the complexity of configuring and managing the pair increases
Solution Approach 1:
The protected/unprotected status of appliances in the preemptive pair is dynamic rather than static. The system can automatically adjust which appliance is protected based on operational status, version requirements, and failure conditions. This dynamic behavior simplifies configuration management while maintaining robust failover capability, as the system adapts automatically rather than requiring manual reconfiguration.
Data Source
AI summary
An example method of upgrading an application in a software-defined data center (SDDC) includes: deploying, by lifecycle management software executing in the SDDC, a second appliance, a first appliance executing services of the application at a first version, the second appliance having services of the application at a second version, the services in the first appliance being active and the services in the second appliance being inactive; setting, by the lifecycle management software, the first and second appliances as a preemptive pair, where the first appliance is protected and the second appliance is unprotected by fault domain management (FDM) software executing in the SDDC; performing, by the lifecycle management software, a switchover to stop the services of the first appliance and start the services of the second appliance; and setting, by the lifecycle management software, the first appliance as unprotected and the second appliance as protected by the FDM software.


