Cloud Outage Detection via Multi-Zone Web Agents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cloud environments face challenges in detecting and managing network connectivity outages across multiple availability zones, leading to service disruptions and downtime, as existing monitoring systems lack comprehensive and real-time health status assessment capabilities.

Innovation Solution

Implementing a system with internal and external web agents that perform network calls and collect response data to generate structured data on network connectivity status, allowing for the determination of health status, outage detection, and notification of affected entities, while also monitoring inbound, outbound, and internal connectivity across multiple availability zones.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If comprehensive network monitoring is implemented across multiple availability zones, then outage detection capability is improved, but system complexity increases

Engineering Contradiction:
Improveoutage detection capabilityVSAvoidmonitoring system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The monitoring system is segmented into multiple independent web agents deployed across different availability zones and network segments. Each agent independently monitors its local segment, and results are aggregated to provide comprehensive coverage. This segmentation allows precise outage detection in specific segments without requiring a monolithic complex system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Web agents serve as intermediaries between the monitoring system and the cloud infrastructure. These agents collect health status data from various components (applications, databases, services) and relay information to a central evaluation system. This intermediary layer simplifies the overall system architecture while maintaining comprehensive monitoring capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If real-time health status assessment is implemented, then service availability is improved, but data processing requirements increase

Engineering Contradiction:
Improveservice availabilityVSAvoiddata processing volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

Web agents continuously collect and pre-process health status data in real-time before outages occur. This preliminary action ensures that when an outage happens, the system already has current information available for immediate assessment and response, maintaining service availability without overwhelming processing requirements at the moment of failure.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The monitoring system uses automated web agents that independently collect, evaluate, and report health status data without requiring manual intervention. This self-service approach enables real-time assessment while reducing the processing burden on human operators and automated response systems.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If multiple web agents are deployed across network segments, then monitoring coverage is improved, but implementation complexity increases

Engineering Contradiction:
Improvemonitoring coverageVSAvoidsystem implementation ease
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The web agents are designed as universal, multi-functional components that can be deployed across different availability zones and network segments using a standardized implementation. Each agent performs the same core functions (collecting health data, evaluating status, reporting results) regardless of location, simplifying deployment and maintenance while achieving comprehensive monitoring coverage.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11888717B2Detecting outages in a multiple availability zone cloud environment
Publication Date: 2024.01.30 SAP SE
  • US11888717B2 patent drawing
  • US11888717B2 patent drawing
  • US11888717B2 patent drawing

AI summary

The present disclosure relates to computer-implemented methods, software, and systems for detecting outages in a multiple availability zone cloud environment. Structured data defining network connectivity statuses of network segments is received. Multiple availability zones of the first cloud platform are defined in a multiple availability zone cloud architecture. External structure data defining inbound connectivity statuses of the network segments correspondingly defined for the availability zones of the first cloud platform is iteratively collecting. The inbound connectivity statuses define availability for an entity running at an external cloud platform to the first cloud platform to connect to at least one entity running at the first cloud platform. In response to evaluating the internal and external structured data, determining a health status of the first cloud platform to be provided to platform services provided by the first cloud platform and/or applications running on the first cloud platform to support managing of lifecycle of entities running on the first cloud platform.