Event Notification Service Granular Resource Impact Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Computing resource service providers face challenges in providing granular information about operational issues affecting customers, leading to frustration and inefficient troubleshooting, as customers often spend resources determining the cause of malfunctions without sufficient detail on affected services.
Innovation Solution
A system event notification service that generates customized notifications and event data for customers, identifying impacted resources and preferences, allowing for targeted remedial actions and minimizing downtime by providing detailed event data and visualization tools.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If computing resource service providers publish general operational issue information, then customers receive notification of service problems, but customers cannot identify specific affected resources or take targeted remedial actions
Solution Approach 1:
The notification system segments operational issue information by customer, resource type, and event severity. Instead of publishing generic outage notices, the system divides notifications into specific categories (e.g., compute events, storage events, network events) and delivers tailored information to affected customers about their specific impacted resources, enabling precise identification and remediation without overwhelming complexity
Solution Approach 2:
The system applies local quality by customizing notification content and delivery based on individual customer needs and affected resource types. Each customer receives notifications tailored to their specific impacted resources (e.g., virtual machine instances, block storage volumes) rather than uniform general announcements, providing locally optimized information quality for each customer-resource pair
2Loss of time
If customers troubleshoot their own computing resources without sufficient event data, then they can identify root causes, but they spend significant time and resources determining the cause of malfunctions
Solution Approach 1:
The notification system performs preliminary action by proactively detecting operational issues and delivering detailed event data to customers before they need to troubleshoot. The system monitors resources, identifies events (such as storage volume failures or compute instance issues), and pushes notifications with specific event details (event type, affected resources, timestamps) to customers in advance, eliminating the need for time-consuming diagnostic investigations
Solution Approach 2:
The system implements feedback by continuously monitoring resource events and delivering real-time notifications to customers about their impacted resources. When an event occurs (e.g., storage I/O errors, compute resource failures), the system immediately feeds back specific event data including event classification, affected resource identifiers, and relevant details, enabling customers to take swift remedial actions without prolonged troubleshooting
Data Source
AI summary
A system event notification service detects that an event has occurred that impacts infrastructure of a computing resource service. In response to the event, the service identifies a customer account that is impacted by the event. The service generates, for the customer account, event data corresponding to a plurality of computing resources impacted by the event. The service provides the event data in accordance with one or more preferences specified in the customer account.


