Cloud O&M Fault Troubleshooting Using Link Graph Root Cause Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing maintenance costs and labor costs associated with manually managing multiple deployment environments for Internet service platforms, particularly recommendation platforms, due to their complexity and widespread deployment.
Innovation Solution
An operation and maintenance platform that includes a debugging interface, proxy module, and multiple fault troubleshooting engines, which automatically receive and analyze operation and maintenance information to identify fault causes and generate troubleshooting reports, utilizing a fault troubleshooting link graph with prioritized branch sub-links and binary search to efficiently locate root causes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual operation and maintenance is performed by management personnel, then fault troubleshooting can be conducted, but maintenance costs and labor costs increase significantly
Solution Approach 1:
The system enables self-service fault troubleshooting through automated agents that independently detect faults, analyze link graphs, identify root causes, and execute repair actions without human intervention, thereby eliminating labor costs while maintaining troubleshooting capability
Solution Approach 2:
Manual mechanical troubleshooting operations are replaced by an automated digital system comprising fault troubleshooting engines, link graph analysis, and automated repair execution, substituting human labor with algorithmic processing
2Adaptability or versatility
If the service platform provides more services and has more deployment environments, then service coverage increases, but maintenance costs keep increasing
Solution Approach 1:
The fault troubleshooting engine is designed as a universal system that can handle multiple deployment environments and service types through a standardized link graph approach, allowing the same core technology to serve diverse services without proportional increases in maintenance complexity
Solution Approach 2:
The system segments fault troubleshooting into modular components: fault detection, link graph construction, root cause analysis, and repair execution. This segmentation allows the system to scale to multiple environments by reusing the same modular components rather than creating separate maintenance systems for each environment
3Productivity
If automated fault troubleshooting is implemented, then labor costs are reduced, but system complexity increases
Solution Approach 1:
The link graph serves as an intermediary data structure that bridges fault detection and root cause analysis. It automatically models service dependencies and data flows, enabling the system to handle complexity internally while presenting a simple automated troubleshooting interface to users
Data Source
AI summary
The present disclosure provides an operation and maintenance platform and a fault troubleshooting method. The operation and maintenance platform includes: a debugging interface, a proxy module, and multiple fault troubleshooting engines. The debugging interface is configured to receive operation and maintenance information and return a fault troubleshooting report. The proxy module is configured to determine a backend cloud environment based on environment information in the operation and maintenance information and submit the operation and maintenance information to a fault troubleshooting engine corresponding to the backend cloud environment. The fault troubleshooting engine is configured to determine a fault troubleshooting link graph based on the information on the problem description, perform fault troubleshooting on the maintenance object based on the fault troubleshooting link graph and an identity of the maintenance object, determine a root cause of a fault corresponding to the problem description, and generate the fault troubleshooting report.


