Cloud Service High Availability via Automatic VEE Restoration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional cloud systems lack integrated High Availability (HA) solutions, relying on manual service re-launches that cause significant delays and require third-party HA software, which is costly and difficult to implement, making it challenging to ensure continuous service availability in cloud-based infrastructure.
Innovation Solution
A method and computer program product that employs shared distributed storage with daemons monitoring node availability, using an exclusive file access mode and a master daemon to automatically restore Virtual Execution Environments (VEEs) on functional nodes, ensuring minimal service delays and easy implementation within the cloud system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual service re-launches are used in conventional cloud systems, then service availability can be restored after node failure, but significant delays occur and operational complexity increases
Solution Approach 1:
The patent implements automatic service re-launch mechanisms that are pre-configured and activated upon node failure detection. The system maintains service configuration metadata and automatically initiates re-launch procedures without waiting for manual intervention, thereby reducing service downtime while ensuring reliability restoration.
2Reliability
If third-party HA software is integrated into cloud systems, then high availability can be achieved, but device complexity and implementation difficulty increase
Solution Approach 1:
The patent merges the High Availability functionality directly into the cloud system's existing infrastructure components. The HA functions are integrated with the virtualization layer, storage system, and network routing devices, eliminating the need for separate third-party HA software and reducing overall system complexity.
Solution Approach 2:
The patent creates a universal HA solution that works across multiple cloud platforms and service types. The integrated HA mechanism provides multi-functional capabilities including service monitoring, automatic re-launch, IP failover, and coordination with various cloud infrastructure components, making it applicable to diverse cloud environments without requiring platform-specific third-party software.
3Reliability
If third-party HA software is used, then service monitoring and restoration can be performed, but cost-effectiveness decreases
Solution Approach 1:
The patent implements self-service HA capabilities where the cloud system automatically monitors its own services, detects failures, and performs restoration without external assistance. The system uses built-in monitoring agents, metadata storage, and automated re-launch mechanisms that eliminate the need for expensive third-party HA software licenses and external maintenance services.
4Reliability
If automatic service re-launch is implemented, then service availability is maintained, but network routing coordination complexity increases
Solution Approach 1:
The patent introduces an intermediary mechanism in the form of a coordination layer that manages communication between the service re-launch system and network routing devices. This intermediary handles IP failover coordination, updates routing tables, and manages network configuration changes, thereby simplifying the overall coordination complexity while maintaining continuous service availability.
Data Source
AI summary
A method and computer program product for providing High Availability (HA) of services in a cloud-based system. The services are the applications used by the end users of the cloud system. The system uses shared cloud distributed storage. The cloud data is distributed over several nodes and is duplicated in an on-line mode. Each user launches and runs his own Virtual Execution Environment VEE (VM or Container) used as a cloud service. If one of the hardware nodes becomes unavailable, the HA system restores all of the VEEs of the failed node on a functional hardware node. Each node has a daemon, which monitors the situation and, if one of the nodes crashes, a master daemon provides for the restoration of the services of this node on other available nodes of the cloud.


