Virtual Desktop Blast Radius Identification via Event Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In global virtual desktop systems, identifying the root cause of faults and determining the blast radius of affected users in near-real-time is challenging due to complex interactions among numerous dependent components, making it difficult to maintain high service availability and quickly respond to major incidents.

Innovation Solution

A diagnostic system that collects events from service components, merges them into a time-ordered stream, correlates events to generate a correlated event stream, and uses analysis modules to determine problem reports, identify spikes, and define the scope of major incidents, thereby determining the blast radius and recommending mitigating actions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a global virtual desktop system with numerous dependent components is deployed to provide comprehensive services, then service functionality and coverage are improved, but system complexity and difficulty of fault identification increase

Engineering Contradiction:
Improveservice functionalityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the complex global virtual desktop infrastructure into discrete, traceable components including datacenters, virtual desktop hosts, gateway hosts, cloud APIs, and client devices. Each component generates independent events that can be individually tracked and correlated, transforming the monolithic complex system into manageable segments for fault analysis.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary diagnostic system that acts as a mediator between the complex distributed components and the fault analysis process. This intermediary collects events from all components, merges them into a unified time-ordered stream, and presents a simplified view of system state, reducing the complexity burden on operators.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If comprehensive event collection from all service components is implemented, then fault detection capability is improved, but data processing complexity and resource consumption increase

Engineering Contradiction:
Improvefault detection capabilityVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges events from numerous distributed service components into a single time-ordered event stream. This consolidation approach maintains comprehensive fault detection capability by preserving all individual event data while simplifying processing through a unified stream structure that can be analyzed sequentially rather than managing multiple separate data sources.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The event stream processing system is designed with universal functionality to handle diverse event types from different components (virtual desktop hosts, gateway hosts, cloud APIs, client devices) through a single unified processing pipeline. This multi-functional approach reduces data processing complexity by applying consistent merging and correlation logic across all event sources.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of time

If real-time analysis of event streams is performed to identify root cause, then response time to incidents is improved, but computational resource requirements increase

Engineering Contradiction:
Improveresponse timeVSAvoidcomputational resource requirements
Core Design Contradiction:
Loss of timeVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary actions by continuously maintaining a time-ordered event stream with all events already collected and organized from service components. When a fault occurs, the root cause analysis can immediately query this pre-prepared stream without needing to collect and organize events in real-time, reducing the computational burden during incident response while maintaining fast detection capability.

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If detailed tracking of user scope, pool scope, and regional scope is implemented, then blast radius determination accuracy is improved, but system monitoring complexity increases

Engineering Contradiction:
Improveblast radius determination accuracyVSAvoidsystem monitoring complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements local quality by associating specific attributes with events based on their scope: user scope attributes track individual user impacts, pool scope attributes track virtual desktop pool impacts, and regional scope attributes track datacenter impacts. This localized attribution approach enables precise blast radius determination for each scope level while keeping monitoring complexity manageable through structured attribute tracking rather than holistic system monitoring.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20240362098A1Method and system for real-time identification of blast radius of a fault in a globally distributed virtual desktop fabric
Publication Date: 2024.10.31 WORKSPOT INC
  • US20240362098A1 patent drawing
  • US20240362098A1 patent drawing
  • US20240362098A1 patent drawing

AI summary

A system and method for determining a blast radius of a major incident occurring in a virtual desktop system is disclosed. The virtual desktop system has interconnected service components and provides access to virtual desktops by client devices. An event collection module collects events from the service components. An aggregation module merges the collected events in a time-ordered stream, provides context to the events in the time-ordered stream through relationships between the collected events, and generates a correlated event stream. An analysis module determines a stream of problem reports from the correlated event stream. The analysis module determines a spike in the stream of problem reports and determines the attributes of the problem reports in the spike to define the major incident. The analysis module determines a scope of the major incident and a corresponding attribute, to determine a blast radius associated with the major incident in the desktop system.