Server Failure Diagnosis via Abnormal Fluctuation Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing complexity of network service systems makes manual failure diagnosis time-consuming and inefficient, leading to difficulties in timely and effective damage control.
Innovation Solution
A method and apparatus that analyze service monitoring indicators using an abnormality fluctuation detection algorithm, cluster servers with similar degrees of abnormal fluctuation, and determine the failure location, employing kernel density estimation and clustering algorithms like DBSCAN or hierarchical clustering to automate the diagnosis process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual analysis and troubleshooting are used for failure diagnosis, then the diagnosis process can be performed with simple tools, but it consumes a large amount of labor and time
Solution Approach 1:
The system performs automatic failure diagnosis by having the server itself generate and analyze monitoring indicators, eliminating the need for external manual intervention. The automated diagnosis system continuously monitors service health metrics and independently identifies failures, reducing both labor consumption and diagnosis time while maintaining high accuracy through algorithmic analysis of service monitoring indicators.
2Measurement precision
If comprehensive service monitoring indicators are analyzed for all servers, then the diagnosis accuracy is improved, but the complexity of the diagnosis system increases
Solution Approach 1:
The system transforms complex multi-dimensional monitoring data into a simplified one-dimensional abnormal fluctuation degree parameter. By calculating the degree of abnormal fluctuation for each service monitoring indicator and then computing a comprehensive abnormal fluctuation degree that aggregates these values, the system maintains high diagnosis accuracy while significantly reducing computational complexity and making the diagnosis results more interpretable.
3Reliability
If the scope of failure diagnosis is expanded to cover all service monitoring indicators, then the reliability of diagnosis is improved, but the time required for analysis increases
Solution Approach 1:
The system extracts and focuses on the most critical diagnostic information by calculating abnormal fluctuation degrees for individual indicators and then aggregating them into a comprehensive abnormal fluctuation degree. This extraction approach allows the system to maintain high diagnosis reliability by considering all monitoring indicators while achieving fast processing through efficient aggregation algorithms that identify the root cause server and service quickly.
Data Source
AI summary
A method, an apparatus and a storage medium for diagnosing failure based on a service monitoring indicator are provided. Service monitoring indicator of a server is analyzed to obtain a degree of abnormal fluctuation in the service monitoring indicator. Servers with similar degrees of abnormal fluctuation are clustered according to the degree of abnormal fluctuation of the service monitoring indicator, to obtain clustered results that include the servers and the service monitoring indicator. A location where the service fails is determined according to the clustered results.


