Cloud O&M Fault Troubleshooting Using Link Graph Root Cause Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing maintenance costs and labor costs associated with manually managing multiple deployment environments for Internet service platforms, particularly recommendation platforms, due to their complexity and widespread deployment.

Innovation Solution

An operation and maintenance platform that includes a debugging interface, proxy module, and multiple fault troubleshooting engines, which automatically receive and analyze operation and maintenance information to identify fault causes and generate troubleshooting reports, utilizing a fault troubleshooting link graph with prioritized branch sub-links and binary search to efficiently locate root causes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual operation and maintenance is performed by management personnel, then fault troubleshooting can be conducted, but maintenance costs and labor costs increase significantly

Engineering Contradiction:
Improvefault troubleshooting capabilityVSAvoidmaintenance cost
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system enables self-service fault troubleshooting through automated agents that independently detect faults, analyze link graphs, identify root causes, and execute repair actions without human intervention, thereby eliminating labor costs while maintaining troubleshooting capability

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical troubleshooting operations are replaced by an automated digital system comprising fault troubleshooting engines, link graph analysis, and automated repair execution, substituting human labor with algorithmic processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If the service platform provides more services and has more deployment environments, then service coverage increases, but maintenance costs keep increasing

Engineering Contradiction:
Improveservice coverageVSAvoidmaintenance cost
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The fault troubleshooting engine is designed as a universal system that can handle multiple deployment environments and service types through a standardized link graph approach, allowing the same core technology to serve diverse services without proportional increases in maintenance complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system segments fault troubleshooting into modular components: fault detection, link graph construction, root cause analysis, and repair execution. This segmentation allows the system to scale to multiple environments by reusing the same modular components rather than creating separate maintenance systems for each environment

Inventive Principle:
Principle #1Segmentation

3Productivity

If automated fault troubleshooting is implemented, then labor costs are reduced, but system complexity increases

Engineering Contradiction:
Improvemaintenance efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The link graph serves as an intermediary data structure that bridges fault detection and root cause analysis. It automatically models service dependencies and data flows, enabling the system to handle complexity internally while presenting a simple automated troubleshooting interface to users

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20260017139A1Operation and maintenance platform, fault troubleshooting method, and related device
Publication Date: 2026.01.15 BEIJING VOLCANO ENGINE TECH CO LTD
  • US20260017139A1 patent drawing
  • US20260017139A1 patent drawing
  • US20260017139A1 patent drawing

AI summary

The present disclosure provides an operation and maintenance platform and a fault troubleshooting method. The operation and maintenance platform includes: a debugging interface, a proxy module, and multiple fault troubleshooting engines. The debugging interface is configured to receive operation and maintenance information and return a fault troubleshooting report. The proxy module is configured to determine a backend cloud environment based on environment information in the operation and maintenance information and submit the operation and maintenance information to a fault troubleshooting engine corresponding to the backend cloud environment. The fault troubleshooting engine is configured to determine a fault troubleshooting link graph based on the information on the problem description, perform fault troubleshooting on the maintenance object based on the fault troubleshooting link graph and an identity of the maintenance object, determine a root cause of a fault corresponding to the problem description, and generate the fault troubleshooting report.