Live Snapshotting Multiple Virtual Disks Failure Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtualized networked computer systems face challenges in creating live snapshots of virtual disks without shutting down virtual machines, and there is a lack of effective mechanisms to detect and manage snapshot failures across a networked environment.
Innovation Solution
A system and method that allow a first computer system to detect and manage the creation of live snapshots of virtual disks across a network, enabling the detection of snapshot failures and ensuring that all virtual disks are snapshotted correctly and at the same point in time, by issuing commands to destroy and deallocate snapshots if failures occur, without requiring the shutdown of virtual machines.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If live snapshots of multiple virtual disks are created simultaneously in a networked environment, then snapshot consistency across disks is improved, but the complexity of detecting and managing snapshot failures increases
Solution Approach 1:
The patent implements a feedback mechanism where the first computer system monitors the snapshot creation process on the second computer system and receives status information about snapshot successes and failures. This feedback loop enables the first system to detect failures and trigger appropriate recovery actions, resolving the contradiction by making failure detection manageable through structured information flow.
Solution Approach 2:
The first computer system acts as an intermediary that coordinates snapshot creation across the networked environment. It issues commands to the second computer system, monitors outcomes, and manages failure recovery centrally. This intermediary role simplifies failure detection and management complexity while ensuring snapshot consistency across multiple virtual disks.
2Productivity
If live snapshots are created without shutting down virtual machines, then system availability is improved, but the risk of snapshot failures increases
Solution Approach 1:
The system performs preliminary actions by issuing snapshot creation commands to the second computer system before actually needing the snapshots. The first computer system proactively monitors the process and prepares failure recovery mechanisms in advance. This preliminary approach allows live snapshotting without shutdown while maintaining reliability through pre-configured failure detection and management.
Solution Approach 2:
The patent implements beforehand cushioning by establishing failure detection and recovery mechanisms before snapshot failures can occur. The first computer system is positioned to detect failures and issue destroy commands to rollback unsuccessful snapshots, creating a safety cushion that protects against reliability issues while maintaining system availability during live snapshot operations.
3Reliability
If snapshot failure detection and management mechanisms are implemented across a network, then snapshot reliability is improved, but the communication overhead and system complexity increase
Solution Approach 1:
The patent segments the snapshot management functionality into distinct components: the first computer system handles command issuance and failure detection, while the second computer system executes snapshot creation locally. This segmentation distributes complexity across multiple systems rather than concentrating it in one place, making the overall networked system more manageable while improving failure detection reliability.
Data Source
AI summary
A system and method are disclosed for servicing requests to create live snapshots of a plurality of virtual disks in a virtualized environment. In accordance with one example, a first computer system detects that a second computer system has issued one or more commands to create a first snapshot of a first virtual disk of a virtual machine and a second snapshot of a second virtual disk of the virtual machine while the virtual machine is running on the second computer system. In response to a determination that the creating of the second snapshot failed, the first computer system issues one or more commands to destroy the first snapshot and deallocate an area of a storage device that stores the first snapshot.


