Storage Manager Snapshot Hierarchy for Multi-Role Application Root Cause Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Complex multi-role applications in distributed storage and computation systems face challenges in deployment, monitoring, and maintenance, particularly in ensuring prompt failure resolution and efficient resource management.

Innovation Solution

A system is developed that includes a storage manager coordinating snapshot creation and maintenance across compute and storage nodes, enabling efficient deployment, monitoring, and rollback of multi-role applications through snapshot hierarchy management and orchestration of resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual monitoring and maintenance methods are used for multi-role applications, then system complexity is reduced, but failure resolution time increases and reliability decreases

Engineering Contradiction:
Improvefailure resolution promptnessVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces a storage manager as an intermediary component that coordinates snapshot creation across multiple compute and storage nodes. This mediator automates the monitoring and maintenance tasks, enabling prompt failure detection and resolution without requiring complex manual intervention, thus improving reliability while managing system complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary actions by creating and maintaining snapshots of application states before failures occur. These pre-established snapshots enable rapid rollback and failure resolution, improving the promptness of failure response without adding operational complexity during critical moments

Inventive Principle:
Principle #10Preliminary action

2Difficulty of detecting and measuring

If comprehensive monitoring of all application roles is implemented, then failure detection capability improves, but system complexity and resource consumption increase

Engineering Contradiction:
Improvefailure detection capabilityVSAvoidmonitoring system complexity
Core Design Contradiction:
Difficulty of detecting and measuringVSDevice complexity

Solution Approach 1:

The patent creates snapshot copies of application states at specific points in time. These copies serve as monitorable representations of the actual running application, enabling failure detection through comparison without requiring direct intrusive monitoring of the live system, thus improving detection capability while limiting complexity increase

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The monitoring approach is segmented by focusing on snapshot points rather than continuous monitoring of all roles simultaneously. This segmentation allows failure detection capability to be improved through targeted snapshot creation and comparison, while reducing overall system complexity by avoiding comprehensive continuous monitoring

Inventive Principle:
Principle #1Segmentation

3Productivity

If snapshot hierarchy management is implemented across multiple nodes, then application deployment efficiency improves, but resource allocation complexity increases

Engineering Contradiction:
Improvedeployment efficiencyVSAvoidresource management complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The storage manager is designed as a universal component that handles snapshot creation, coordination, and management across multiple compute and storage nodes simultaneously. This multi-functional approach improves deployment efficiency by automating resource allocation across the distributed system, while the centralized management logic reduces the complexity burden on individual nodes

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Reliability

If rapid failure resolution through rollback is enabled, then system reliability improves, but storage resource consumption increases

Engineering Contradiction:
Improvefailure resolution speedVSAvoidstorage resource consumption
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system implements partial rollback by maintaining and utilizing only the necessary snapshot hierarchy required for failure resolution. Rather than preserving all possible application states indefinitely, the system keeps sufficient snapshots for rapid rollback while managing storage consumption, achieving reliable failure resolution with controlled resource usage

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11520650B2Performing root cause analysis in a multi-role application
Publication Date: 2022.12.06 RAKUTEN SYMPHONY INC
  • US11520650B2 patent drawing
  • US11520650B2 patent drawing
  • US11520650B2 patent drawing

AI summary

A new snapshot of a storage volume is created by instructing computing nodes to suppress write requests. A snapshot of the application may be created and used to rollback or clone the application. Clones snapshots of storage volumes may be gradually populated with data from prior snapshots to reduce loading on a primary snapshot. Components of cloned applications may communicate with one another using addresses of these components in the parent application. Jobs implementing a bundled application may be referenced with a simulated file system that generates reads to hosts only when the job log file is actually read. Job logs and a job hierarchy may be used to perform root cause analysis. Job logs may be for tasks such as creating the bundled application, cloning, rolling back, backing up, scaling out, scaling in, deleting, pruning unused application images, or the like.