Cloud platform fault intelligent research and judgment system and method based on multi-dimensional evaluation and topology analysis
Through the cloud platform's intelligent fault analysis system with multi-dimensional evaluation and topological analysis, the scope of failure impact is predicted dynamically and the urgency is quantified, which solves the problems of inefficiency and misjudgment in cloud platform operation and maintenance, and achieves efficient fault response and handling.
Patent Information
- Application Number
- CN202510353789.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-18
AI Technical Summary
The cloud platform business system operation and maintenance fault analysis relies on human judgment, which is inefficient and prone to misjudgment and misjudgment, lacks multi-dimensional evaluation, dynamic prediction and quantitative evaluation, and the traditional weight allocation method does not consider the coupling degree of the system architecture.
The multi-dimensional evaluation matrix building module is adopted, including business impact, system health and time sensitivity calculation units, combined with the improved PageRank algorithm to build a service dependency map, dynamically predict the scope of the failure impact, and quantify the fault urgency through the emergency index calculation model, introduce fault knowledge graphs and reinforcement learning to optimize model parameters to improve accuracy.
It significantly improves the accuracy of fault grading, reduces the impact range prediction error rate and average repair time, reduces the need for excessive resource expansion, and achieves efficient fault response and processing.
Smart Images

Figure CN120336052A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud computing operation and maintenance, and particularly to an intelligent fault judgment system and method for a cloud platform based on multi-dimensional evaluation and topology analysis. Background Technique
[0002] With the rapid development of information technology, cloud platform business systems play an increasingly important role in information systems of various industries. However, the judgment of operation and maintenance faults in cloud platform business systems still heavily relies on manual judgment, resulting in many challenges in cloud platform operation and maintenance. It is not only inefficient but also prone to misjudgment and missed judgment. Therefore, there is an urgent need for an intelligent fault judgment method to improve the judgment ability and accurately grasp the impact of faults. This patent aims to introduce intelligent means to solve the following defect problems existing in the traditional operation and maintenance judgment of cloud platform systems:
[0003] 1. The fault evaluation dimension is single, lacking the linkage analysis between the business perspective and the system perspective;
[0004] 2. The judgment of the influence scope relies on manual experience and cannot dynamically predict the diffusion path;
[0005] 3. The emergency level division standard is fuzzy and a quantitative evaluation model has not been established;
[0006] 4. The traditional weight allocation method does not consider the system architecture coupling degree. Summary of the Invention
[0007] The purpose of the present invention is to provide an intelligent fault judgment system and method for a cloud platform based on multi-dimensional evaluation and topology analysis to solve the problems raised in the above background technique.
[0008] To achieve the above purpose, the present invention provides the following technical solution: An intelligent fault judgment system for a cloud platform based on multi-dimensional evaluation and topology analysis, a multi-dimensional evaluation matrix construction module for evaluating the emergency degree of operation and maintenance events of the cloud platform business system, specifically including:
[0009] A business impact degree BI calculation unit for calculating the business impact degree according to the number of affected customers and the business weight coefficient, where different weights are set for different levels of the number of affected customers, and business weights are set according to core business, important business, and ordinary business respectively;
[0010] A system health degree SH calculation unit for evaluating the system health status through the product of the resource occupancy rate and the service availability. The resource occupancy rate is calculated based on the weighted average of the CPU, memory, and disk usage rates, and the service availability is based on the product of the API success rate and the heartbeat detection normal rate;
[0011] A timeliness sensitivity TS calculation unit for considering the SLA remaining time coefficient and the repair complexity coefficient to evaluate the urgency of fault resolution.
[0012] Preferably, it further includes an impact diffusion model design module, which predicts the impact scope of faults through the following steps:
[0013] Construct a service dependency graph DAG, define node attributes and edge attributes. Node attributes include service type, redundancy, and traffic weight, and edge attributes cover call frequency, QPS, and timeout configuration;
[0014] Use an improved PageRank algorithm to calculate the impact weight of each service node, considering the weights of adjacent nodes and connection strength, and the connection strength is dynamically adjusted according to call frequency and timeout configuration;
[0015] Introduce a dynamic prediction mechanism, update node status in real time and adjust the impact diffusion path in combination with the fuse status, and predict the possible breadth of the impact of faults by calculating the impact scope formula.
[0016] Preferably, it further includes an emergency index calculation model, which comprehensively considers the business impact degree, the complement of the system health degree 1 - SH, and the timeliness sensitivity, and obtains the emergency index Urgency of the operation and maintenance event through weighted summation. And according to the emergency index, the operation and maintenance events are divided into five emergency levels from P0 to P4, and each level corresponds to different response requirements and time limits.
[0017] Preferably, the weighting coefficients in the emergency index calculation model are adjusted according to the actual operation situation and business requirements of the cloud platform business system, and the system has a continuous improvement mechanism, including establishing a fault knowledge graph, using reinforcement learning to optimize model parameters, and regularly reviewing the weight coefficients to ensure the accuracy and efficiency of fault judgment.
[0018] Preferably, the system also has an automated processing ability. For P0 - level emergency events, it automatically triggers the interruption of the maintenance process without manual approval, realizing fast response and efficient processing of operation and maintenance faults in the cloud platform business system.
[0019] A cloud platform fault intelligent judgment method for a cloud platform fault intelligent judgment system based on multi - dimensional evaluation and topology analysis includes:
[0020] Construct a multi - dimensional evaluation matrix, which includes three dimensions: business impact degree BI, system health degree SH, and timeliness sensitivity TS, where:
[0021] BI is calculated based on the number of affected customers and the business weight coefficient;
[0022] SH is evaluated by the product of resource occupancy rate and service availability;
[0023] TS considers the SLA remaining time coefficient and the repair complexity coefficient;
[0024] Use a multi-dimensional evaluation matrix to comprehensively evaluate the operation and maintenance events of the cloud platform business system, and output a five-level emergency index P0 - P4, where P0 represents the highest emergency level and P4 represents the lowest emergency level.
[0025] Preferably, it also includes designing an impact diffusion model, and the specific steps include:
[0026] Construct a service dependency graph DAG, and define node attributes and edge attributes, where node attributes include service type, redundancy, and traffic weight, and edge attributes include call frequency, QPS, and timeout configuration;
[0027] Use an improved PageRank algorithm to calculate the impact weight of service nodes. The impact weight considers the weights of adjacent nodes and connection strength, and the connection strength is dynamically adjusted according to call frequency and timeout configuration;
[0028] Introduce a dynamic prediction mechanism, update the node status in real time and adjust the impact diffusion path in combination with the fuse status to predict the scope of the fault impact.
[0029] Preferably, the calculation of the emergency index uses a comprehensive evaluation formula: Urgency = 0.5BI + 0.3(1 - SH) + 0.2*TS;
[0030] And classify the operation and maintenance events into five emergency levels from P0 to P4 according to the value of the emergency index. Each level corresponds to different response requirements and time limits. Among them, events at the P0 level need to automatically interrupt maintenance without manual approval.
[0031] Preferably, it also includes a continuous improvement mechanism, and the specific steps include:
[0032] Establish a fault knowledge graph to record historical fault cases and research and judgment results;
[0033] Use a reinforcement learning algorithm to optimize the parameters of the multi-dimensional evaluation matrix and the impact diffusion model to improve the accuracy of research and judgment;
[0034] Conduct a weight coefficient review regularly, and adjust the weights of evaluation dimensions according to the actual operation situation and business requirements of the cloud platform business system.
[0035] Preferably, the intelligent research and judgment method is integrated into the operation and maintenance management platform of the cloud platform business system to realize automatic monitoring, evaluation, and hierarchical response of operation and maintenance events, and improve the operation and maintenance efficiency and fault recovery speed of the cloud platform business system.
[0036] Compared with the prior art, the beneficial effects of the present invention are:
[0037] The intelligent fault judgment system and method for cloud platforms based on multi-dimensional evaluation and topological analysis proposed by the present invention can dynamically predict the impact scope of faults by constructing a service dependency graph and adopting an improved PageRank algorithm. In addition, the system also introduces an emergency index calculation model to quantify the emergency degree of faults and formulate corresponding response strategies accordingly. Through continuous improvement mechanisms, such as the establishment of a fault knowledge graph and the optimization of model parameters by reinforcement learning, the system can continuously improve the accuracy of fault judgment. This solution significantly improves the accuracy of fault grading, reduces the prediction error rate of the impact scope and the mean time to repair, and at the same time reduces the need for excessive resource expansion. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] In order to clearly and completely describe the objectives, technical solutions of the present invention and make the advantages more clearly understood, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are some but not all embodiments of the present invention, and are only used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0040] Embodiment 1, the present invention provides a technical solution: an intelligent fault judgment system for cloud platforms based on multi-dimensional evaluation and topological analysis, including:
[0041] 1. A multi-dimensional evaluation matrix for quantifying the impact of faults, including business impact degree, system health degree, and timeliness sensitivity; the multi-dimensional evaluation matrix includes:
[0042] 1) Business impact degree (BI), the calculation formula is BI = the number of affected customers × business weight coefficient, where the number of customers is classified as 1 - 100, 101 - 1000, > 1000, and the business weight is core business, important business, and ordinary business;
[0043] 2) System health degree (SH), the calculation formula is SH = (1 - resource occupancy rate) × service availability, where the resource occupancy rate is the weighted average of CPU / MEM / Disk, and the service availability is the API success rate × heartbeat detection normal rate;
[0044] Timeliness sensitivity (TS), the calculation formula is TS = SLA remaining time coefficient × repair complexity coefficient, where the SLA coefficient is set to 1.0, 0.7, 0.4 according to the remaining time < 1h, 1 - 4h, > 4h respectively, and the complexity is set to 0.3, 0.6, 1.0 for simple, medium, and complex respectively.
[0045] 2. Impact diffusion model, used to dynamically predict the impact scope of faults, including a service dependency graph and a method for calculating impact weights based on an improved PageRank algorithm; the impact diffusion model includes:
[0046] 1) Topological relationship modeling, constructing a service dependency graph (DAG), defining node attributes as service type, redundancy, and traffic weight, and edge attributes as call frequency, QPS, and timeout configuration;
[0047] Diffusion algorithm, using an improved PageRank algorithm to calculate impact weights, with the formula Impact = Σ(adjacent node weight × connection strength) + self-weight, where the connection strength is dynamically adjusted according to call frequency (QPS) and timeout configuration.
[0048] 3. Emergency index calculation model, used to quantify the emergency level of faults and formulate response strategies accordingly.
[0049] The emergency index calculation model includes:
[0050] Comprehensive evaluation formula, Urgency = 0.5×BI + 0.3×(1 - SH) + 0.2×TS; 2) Grading standard, dividing Urgency into five levels from P0 to P4, each level corresponding to different Urgency ranges and response requirements.
[0051] It also includes a continuous improvement mechanism, used to continuously improve the fault judgment mechanism by means of establishing a fault knowledge graph, optimizing model parameters through reinforcement learning, and quarterly review of weight coefficients.
[0052] The system can improve the fault grading accuracy to 93.2%, reduce the impact scope prediction error rate to <8%, reduce the MTTR (mean time to repair) by 30%, and reduce the excessive resource expansion demand by 60%.
[0053] Example 2, based on Example 1, proposes a cloud platform fault intelligent judgment method for a cloud platform fault intelligent judgment system based on multi-dimensional evaluation and topological analysis, including:
[0054] I. Construction of multi-dimensional evaluation matrix
[0055] 1. Business impact degree (BI)
[0056] BI = number of affected customers × business weight coefficient (0.1 - 1.0)
[0057] Customer number grading: 1 - 100 (0.3) | 101 - 1000 (0.6) | > 1000 (1.0)
[0058] Business weight: Core business (1.0) | Important business (0.7) | Ordinary business (0.4)
[0059] 2. System Health (SH)
[0060] SH = (1 - Resource occupancy rate) × Service availability
[0061] Resource occupancy rate: Weighted average of CPU / MEM / Disk
[0062] Service availability: API success rate × Heartbeat detection normal rate
[0063] 3. Timeliness Sensitivity (TS)
[0064] TS = SLA remaining time coefficient × Repair complexity coefficient
[0065] SLA coefficient: Remaining time < 1h (1.0) | 1 - 4h (0.7) | > 4h (0.4)
[0066] Complexity: Simple (0.3) | Medium (0.6) | Complex (1.0)
[0067] II. Impact Diffusion Model Design
[0068] 1. Topological Relationship Modeling
[0069] Build a service dependency graph (DAG)
[0070] Define node attributes: Service type, Redundancy, Traffic weight
[0071] Edge attributes: Call frequency, QPS, Timeout configuration
[0072] 2. Diffusion Algorithm
[0073] Use an improved PageRank algorithm to calculate the impact weight:
[0074] Impact = Σ(Adjacent node weight × Connection strength) + Self weight
[0075] Connection strength is dynamically adjusted according to call frequency (QPS) and timeout configuration
[0076] 3. Dynamic Prediction Mechanism
[0077] Real-time update the node status and adjust the propagation path in combination with the fuse status.
[0078] Predicted impact range formula: Impact_Scope = Σ(I(n)>θ) / N_total θ is the impact threshold (recommended 0.6, business period coefficient × 1.2 / non-peak × 0.8)
[0079] III. Emergency Index Calculation Model
[0080] 1. Comprehensive Evaluation Formula
[0081] Urgency=0.5BI+0.3(1-SH)+0.2*TS
[0082] 2. Grading Criteria:
[0083]
[0084] IV. Continuous Improvement Mechanism
[0085] The fault judgment mechanism is continuously improved by means of establishing a fault knowledge graph, optimizing model parameters through reinforcement learning, and quarterly review of weight coefficients.
[0086] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent fault judgment system for cloud platforms based on multi-dimensional evaluation and topological analysis, characterized in that: A multi-dimensional evaluation matrix construction module, which is used to evaluate the urgency of operation and maintenance events in the cloud platform business system, specifically including: A business impact degree BI calculation unit, which is used to calculate the business impact degree according to the number of affected customers and the business weight coefficient, where different weights are set for different levels of the number of affected customers, and the business weights are set according to core business, important business, and ordinary business respectively; A system health degree SH calculation unit, which evaluates the system health status through the product of the resource occupancy rate and the service availability. The resource occupancy rate is calculated based on the weighted average of the CPU, memory, and disk usage rates, and the service availability is based on the product of the API success rate and the heartbeat detection normal rate; A timeliness sensitivity TS calculation unit, which considers the SLA remaining time coefficient and the repair complexity coefficient to evaluate the urgency of fault resolution.
2. The intelligent fault judgment system for cloud platform based on multi-dimensional evaluation and topological analysis according to claim 1, characterized in that: It also includes an impact diffusion model design module, which realizes the prediction of the impact scope of faults through the following steps: Construct a service dependency graph DAG, define node attributes and edge attributes. Node attributes include service type, redundancy, and traffic weight, and edge attributes cover call frequency, QPS, and timeout configuration; Use an improved PageRank algorithm to calculate the impact weight of each service node, considering the weights of adjacent nodes and the connection strength, and the connection strength is dynamically adjusted according to the call frequency and timeout configuration; Introduce a dynamic prediction mechanism, update the node status in real time and adjust the impact diffusion path in combination with the fuse status, and predict the possible breadth of the impact of the fault by calculating the impact scope formula.
3. The intelligent fault judgment system for cloud platform based on multi-dimensional evaluation and topological analysis according to claim 2, characterized in that: It also includes an emergency index calculation model, which comprehensively considers the business impact degree, the complement of the system health degree 1 - SH, and the timeliness sensitivity, and obtains the emergency index Urgency of the operation and maintenance event through weighted summation. The operation and maintenance events are divided into five emergency levels from P0 to P4 according to the emergency index, and each level corresponds to different response requirements and time limits.
4. An intelligent fault judgment system for a cloud platform based on multi-dimensional evaluation and topological analysis according to claim 3, characterized in that: The weighted coefficients in the emergency index calculation model are adjusted according to the actual operation situation and business requirements of the cloud platform business system, and the system has a continuous improvement mechanism, including establishing a fault knowledge graph, using reinforcement learning to optimize model parameters, and regularly reviewing the weight coefficients to ensure the accuracy and efficiency of fault judgment.
5. An intelligent fault judgment system for a cloud platform based on multi-dimensional evaluation and topological analysis according to claim 1, characterized in that: The system also has an automated processing ability. For P0-level emergency events, it automatically triggers the interruption of the maintenance process without manual approval, realizing fast response and efficient processing of operation and maintenance faults in the cloud platform business system.
6. The cloud platform fault intelligent judgment method of the cloud platform fault intelligent judgment system based on multi-dimensional evaluation and topological analysis according to claim 5, characterized in that: Including: Construct a multi-dimensional evaluation matrix, which includes three dimensions: business impact degree BI, system health degree SH, and timeliness sensitivity TS, where: BI is calculated according to the number of affected customers and the business weight coefficient; SH is evaluated through the product of the resource occupancy rate and the service availability; TS considers the SLA remaining time coefficient and the repair complexity coefficient; Use the multi-dimensional evaluation matrix to comprehensively evaluate the operation and maintenance events of the cloud platform business system, and output a five-level emergency index P0 - P4, where P0 represents the highest emergency level and P4 represents the lowest emergency level.
7. The intelligent judgment method according to claim 6, wherein: It also includes the design of an impact diffusion model, and the specific steps include: Construct a service dependency graph DAG, define node attributes and edge attributes, where node attributes include service type, redundancy, and traffic weight, and edge attributes include call frequency, QPS, and timeout configuration; Use an improved PageRank algorithm to calculate the influence weight of service nodes. The influence weight considers the weights of adjacent nodes and connection strength, and the connection strength is dynamically adjusted according to call frequency and timeout configuration; Introduce a dynamic prediction mechanism to update node status in real time and adjust the influence diffusion path in combination with the fuse status to predict the scope of fault impact.
8. The intelligent judgment method according to claim 7, characterized in that: The calculation of the emergency index adopts a comprehensive evaluation formula: Urgency = 0.5BI + 0.3(1 - SH) + 0.2*TS; And classify the operation and maintenance events into five emergency levels from P0 to P4 according to the value of the emergency index. Each level corresponds to different response requirements and time limits. Among them, events at the P0 level need to automatically interrupt maintenance without manual approval.
9. The intelligent judgment method according to claim 8, characterized in that: It also includes a continuous improvement mechanism. The specific steps include: Build a fault knowledge graph to record historical fault cases and research results; Use reinforcement learning algorithms to optimize the parameters of the multi-dimensional evaluation matrix and the influence diffusion model to improve the accuracy of research and judgment; Conduct regular weight coefficient reviews and adjust the weights of evaluation dimensions according to the actual operation of the cloud platform business system and business requirements.
10. An intelligent judgment method according to claim 9, characterized in that: The intelligent research and judgment method is integrated into the operation and maintenance management platform of the cloud platform business system to realize automatic monitoring, evaluation, and hierarchical response of operation and maintenance events, improving the operation and maintenance efficiency and fault recovery speed of the cloud platform business system.
Citation Information
Cited By
Cloud computing platform intelligent operation and maintenance method and system based on knowledge graph
CN121711269A