Web-based Computing System Continuous Burn-in Testing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
As data centers scale, managing and provisioning physical computing resources becomes increasingly complex due to factors like aging equipment, software regression, and varying performance demands, necessitating continuous performance monitoring and ranking of computing devices to ensure optimal resource allocation.
Innovation Solution
Implementing a performance manager that continuously tests computing devices, updates ranking values based on performance metrics, and prioritizes testing based on time since last testing, while a placement manager allocates resource instances from higher-ranked, lightly loaded devices to ensure efficient resource distribution across the network.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If continuous performance testing is implemented to monitor computing device performance, then reliability and performance optimization improve, but device complexity and operational overhead increase
Solution Approach 1:
The computing devices perform self-testing by executing test code on themselves, eliminating the need for external testing infrastructure. Each device acts as both the tester and the tested object, reducing system complexity while maintaining continuous monitoring capability
Solution Approach 2:
The testing system uses universal test code that can evaluate multiple performance metrics (CPU performance, memory performance, storage performance, network performance) across different device types. This multi-functional approach simplifies the testing mechanism while comprehensively monitoring various performance aspects
2Reliability
If performance testing is performed frequently to detect degradation early, then performance optimization improves, but loss of time for actual computing tasks increases
Solution Approach 1:
Performance testing is conducted periodically at scheduled intervals rather than continuously, allowing the system to balance monitoring needs with productive workload execution. The periodic nature ensures timely detection of performance degradation while minimizing interference with computing tasks
Solution Approach 2:
The system performs partial testing by selecting specific performance metrics and sample workloads based on current system state and priorities. This selective approach provides sufficient monitoring coverage without the overhead of exhaustive continuous testing
3Productivity
If resource allocation is dynamically adjusted based on performance ranking, then productivity and efficiency improve, but device complexity and management overhead increase
Solution Approach 1:
The system pre-calculates performance rankings and maintains an updated hierarchy of computing devices before allocation decisions are needed. This preliminary ranking allows for rapid allocation decisions without complex real-time calculations, simplifying the management process
Solution Approach 2:
The system implements feedback loops where performance test results automatically update device rankings, which in turn influence resource allocation decisions. This closed-loop control automates the management process, reducing manual overhead while optimizing productivity through data-driven allocation
Data Source
AI summary
This document describes techniques for performance testing computing resources in a service provider network. In an example embodiment, a performance manager periodically tests the performance of computing devices in the service provider network using selected computing assets of each computing device, and updates, based on the performance, a ranking value that establishes precedence for allocation of resource instances of the computing device. A placement manager assigns resource instances from the computing devices based on the ranking values.


