A delimiting problem analysis method, system, electronic device and storage medium

By building a business topology diagram and a C4.5 decision tree model, the problems in the delimited cloud environment are solved, and the problem of relying on expert experience and data silos in the existing technology is improved, and operation and maintenance efficiency and system stability are improved.

CN119788496BActive Publication Date: 2025-07-04HANGZHOU HARMONYCLOUD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510279126.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-04
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

When monitoring and delimiting complex and heterogeneous cloud environment business problems, the existing technology relies on expert experience and lacks automation delimiting capabilities. Data island problems lead to low operation and maintenance efficiency. Traditional tools are unable to handle cross-system and cross-platform problems.

Method used

By collecting multi-source data, building a business topology diagram, setting thresholds to identify abnormal events, extracting abnormal nodes, and using the C4.5 algorithm to build a decision tree model to automatically delimit problems.

Benefits of technology

It has achieved rapid delimitation of application performance bottlenecks, network connection failures and other problems, improved operation and maintenance efficiency and system stability, and provided clear operation guidance for the operation and maintenance team.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119788496B_ABST
    Figure CN119788496B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, electronic device and storage medium for delimited problem analysis, including collecting data, where the data includes resource metrics, business application performance metrics, network status and network call metrics; constructing a business topology graph based on the call metrics; setting thresholds for each data, and events exceeding the thresholds are abnormal events; extracting abnormal nodes from the business topology graph based on the abnormal events; extracting abnormal features based on the abnormal nodes; extracting numerical features in the abnormal features and supplementing the missing numerical features; constructing a decision tree model for abnormal delimitation through the C4.5 algorithm based on the numerical features; after the decision tree model is trained, inputting the to-be-input features into the decision tree model to delimit the abnormal events. The present invention automates delimited problem analysis by integrating the decision tree algorithm and multi-source data collection technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of delimiter problem analysis, and particularly to a delimiter problem analysis method, system, electronic device and storage medium. Background Art

[0002] With the acceleration of enterprise digital transformation, numerous services are applied to multiple public and private cloud services. Monitoring and delimiting these complex and heterogeneous cloud environment business problems have become a severe challenge. The inherent differences between different cloud platforms not only easily lead to performance bottlenecks and scattered data links, but also result in the difficulty of unifying and interpreting monitoring data. Facing this complex situation, there is an urgent need for a method that can automatically analyze and delimit various problems and provide standardized troubleshooting logic for the operation and maintenance team to ensure the efficiency and consistency of cross-cloud platform business application delimitation problems.

[0003] Currently, problem delimitation mainly relies on experts' experience and a series of traditional tools. The Prometheus monitoring tool is used to monitor various performance metrics to ensure the capture of anomalies at the business application and machine metric levels. Log data is investigated and analyzed, and log collection and analysis tools such as the EFK stack (Fluentd, Elasticsearch, Kibana) are used to conduct detailed data mining at fixed query locations for preliminary fault delimitation. In addition, APM (Application Performance Management) tools are widely used to track the request links of business applications to achieve accurate positioning of fault links. For network-level problems, NPM (Network Performance Monitoring) tools are usually used for troubleshooting.

[0004] The current fault problem delimitation process relies on a series of traditional tools and technologies, which can meet the requirements to a certain extent but also have several limitations.

[0005] Reliance on expert experience

[0006] Although the Prometheus monitoring tool can effectively monitor various performance metrics and capture anomalies at the business application and machine metric levels, the final problem delimitation and solution still highly rely on experts' experience and judgment. This not only increases labor costs but may also lack flexibility when facing complex and changeable problems.

[0007] Fixed location query limitation

[0008] When using the EFK stack (Fluentd, Elasticsearch, Kibana) for log data mining, it is usually limited to fixed query locations. Although this approach helps with preliminary delimitation, it is insufficient when dealing with complex cross-system and cross-platform problems and is difficult to provide a comprehensive perspective.

[0009] Lack of automated delimitating ability

[0010] Although APM (Application Performance Management) tools can track the request links of business applications and achieve accurate positioning, for frequently occurring or complexly related problems, manual intervention is still required for further analysis. Similarly, NPM (Network Performance Monitoring) tools perform well in network problem troubleshooting, but their degree of automation is limited and they cannot completely replace manual decision-making.

[0011] Data silo problem

[0012] Data generated by different tools often exists in isolated environments and lacks an effective integration mechanism. This leads to problems such as scattered data, duplicate labor, and information lag, affecting the overall operation and maintenance efficiency. Summary of the Invention

[0013] In view of the deficiencies in the above problems, the present invention provides a method, system, electronic device, and storage medium for analyzing delimitating problems.

[0014] To achieve the above object, the present invention provides a method for analyzing delimitating problems, including:[[]]

[0015] Collecting data from different sources, the data including resource metrics, business application performance metrics, network status and network call metrics, and key events in log records;

[0016] Constructing a business topology graph based on the call metrics, mapping the resource metrics to corresponding nodes of the business topology graph, superimposing the business application performance metrics on the service links of the business topology graph, and marking the key times on relevant nodes of the business topology graph;

[0017] Setting a threshold for each piece of the data, and events exceeding the threshold are abnormal events;

[0018] Extracting abnormal nodes from the business topology graph based on the abnormal events;

[0019] Extracting abnormal features based on the abnormal nodes;

[0020] Extracting numerical features in the abnormal features and supplementing the missing numerical features;

[0021] Constructing a decision tree model for abnormal delimitating based on the numerical features through the C4.5 algorithm;

[0022] After the decision tree model is trained, inputting the to-be-input features into the decision tree model to delimit the abnormal events.

[0023] Preferably, the resource metrics include CPU usage rate, memory usage rate, and disk I / O metric usage rate; the business application performance metrics include response time and error rate; the network status and network call metrics include network traffic and latency.

[0024] Preferably, tags and types are set corresponding to the root cause of the exception event.

[0025] Preferably, extracting abnormal nodes from the business topology map includes:

[0026] Nodes where the request volume, error rate, or response time is monitored to exceed the threshold are set as the abnormal nodes;

[0027] Based on the tags and types, identify the types of the abnormal nodes;

[0028] Extract all the abnormal nodes from the business topology map and classify them to form a subset.

[0029] Preferably, extracting abnormal features based on the abnormal nodes includes:

[0030] Extract the data within the same time period of the data, calculate the mean value, maximum value, and minimum value within the time period, and count its features;

[0031] Analyze the occurrence frequency of specific events within a period of time;

[0032] Aggregate according to the service instance dimension and calculate the summary statistics under the service instance dimension.

[0033] Preferably, constructing a decision tree model for anomaly delimitation through the C4.5 algorithm includes:

[0034] The C4.5 algorithm uses the information gain ratio to select the optimal splitting attribute, and the formula is:

[0035] ;

[0036] In the formula: S is the current data set; A is the attribute to be evaluated; Gain(S, A) represents the information gain, which measures the importance of an attribute A in the classification task S; Intrinsic Value(A) represents the calculation of the intrinsic value, which calculates the intrinsic value of the attribute A and reflects the entropy of the different value distributions in the attribute A.

[0037] Preferably, the calculation formula of the information gain is:

[0038] ;

[0039] ;

[0040] Where: Values(A) represents all possible values of attribute A; Sv is the subset of the dataset where the value of attribute A is v; |S| and |Sv| represent the number of samples in dataset S and subset Sv respectively; Entropy(S) is the entropy of dataset S; P i is the probability of the i-th category;

[0041] The calculation formula for the intrinsic value is:

[0042] ;

[0043] This application also provides a delimiter problem analysis system, including:

[0044] An acquisition module, configured to acquire data from different sources, where the data includes resource metrics, business application performance metrics, network status and network call metrics, and key events in log records;

[0045] A construction module, configured to construct a business topology diagram based on the call metrics, map the resource metrics to the corresponding nodes of the business topology diagram, superimpose the business application performance metrics on the service links of the business topology diagram, and label the key times on the relevant nodes of the business topology diagram;

[0046] Set a threshold for each piece of the data, and events exceeding the threshold are abnormal events;

[0047] An extraction module, configured to extract abnormal nodes from the business topology diagram based on the abnormal events;

[0048] A feature module, configured to extract abnormal features based on the abnormal nodes;

[0049] A supplement module, configured to extract numerical features in the abnormal features and supplement the missing numerical features;

[0050] A decision tree module, configured to construct a decision tree model for abnormal delimitation based on the numerical features through the C4.5 algorithm;

[0051] A delimitation module, configured to, after the decision tree model is trained, input the to-be-input features into the decision tree model to delimit abnormal events.

[0052] This invention also provides an electronic device, including at least one processing unit and at least one storage unit, where the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the above method.

[0053] This invention also provides a storage medium, which stores a computer program executable by an electronic device, and when the program runs on the electronic device, the electronic device executes the above method.

[0054] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0055] By integrating the decision tree algorithm and multi-source data acquisition technology, the present invention automatically analyzes the problem of delimitation, quickly delimits specific problems when detecting anomalies, such as application performance bottlenecks, network connection failures, or host resource overloads, improves the operation and maintenance efficiency and system stability, and provides clear operation guidelines and automated auxiliary delimitation for the operation and maintenance team. Brief Description of the Drawings

[0056] Figure 1 It is a flowchart of the problem delimitation analysis method of the present invention. Detailed Embodiments

[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0058] Referring to Figure 1 , the present invention provides a problem delimitation analysis method, including:

[0059] Collect data from different sources, including resource metrics, business application performance metrics, network status and network call metrics, and key events in log records;

[0060] Based on the call metrics, construct a business topology graph, map the resource metrics to the corresponding nodes in the business topology graph, overlay the business application performance metrics on the service links of the business topology graph, and label the key times on the relevant nodes of the business topology graph;

[0061] Set a threshold for each data, and events exceeding the threshold are abnormal events;

[0062] Based on the abnormal events, extract abnormal nodes from the business topology graph;

[0063] Extract abnormal features based on the abnormal nodes;

[0064] Extract the numerical features in the abnormal features and supplement the missing numerical features;

[0065] Based on the numerical features, construct a decision tree model for abnormal delimitation through the C4.5 algorithm;

[0066] After the decision tree model is trained, input the features to be input into the decision tree model to delimit the abnormal events.

[0067] In this embodiment, the resource metrics include CPU usage rate, memory usage rate, and disk I / O metric usage rate; the business application performance metrics include response time and error rate; the network status and network call metrics include network traffic and latency.

[0068] Furthermore, each node in the business topology diagram represents a specific resource or service instance, while the edges represent the interactions and dependencies between them. According to the call topology diagram, multi-dimensional data is integrated:

[0069] A real-time updated business topology diagram is generated based on the network call data collected in multiple dimensions to construct the business topology diagram. For example, if business application A accesses business application B and then accesses business application C through the network within 1 hour, the dynamic business topology diagram A -> B -> C is obtained, covering all relevant services, components, and their call relationships through the topology diagram.

[0070] Metrics such as CPU usage rate, memory usage rate, and disk I / O metric usage rate (resource metrics) are mapped to the corresponding nodes to visually display the resource health status through color or icon changes.

[0071] Metrics such as response time and error rate (business application performance metrics) are then superimposed on the service link to identify performance bottlenecks. At the same time, metrics such as network traffic and latency (network status and network call metrics) are integrated into the link lines.

[0072] After the log data is parsed, the key events in the log records are directly marked on the relevant nodes to assist in the display of log information after fault problem demarcation.

[0073] After these data are standardized, collected, and processed, they are mapped to the corresponding topology diagram nodes to form a multi-dimensional monitoring view.

[0074] In this embodiment, a baseline within the normal operation range is defined for each data metric, and a reasonable threshold is set. An index event exceeding the threshold is considered an abnormal event. For example, the threshold for the memory usage rate of a business node is 90%. A business node with a memory usage rate exceeding 90% is considered an abnormal event. After preprocessing the collected abnormal data, the root cause of each abnormal event should be marked as the target variable, i.e., the "label" in the training set. The baseline threshold corresponding label lists for various metrics are set as shown in Table 1.

[0075] Table 1

[0076]

[0077] Furthermore, when an anomaly is detected, the system will automatically extract all nodes marked as anomalies from the topology graph to form an independent subset of sub - metric types for in - depth analysis. Different types of anomaly nodes (such as categories like application performance metrics, network metric data, and resource metric data) will be classified according to their data sources and labels, and data of the same metric will be stored in a subset.

[0078] The following are the specific steps for generating subsets of anomaly nodes for different types of data:

[0079] Identify anomaly nodes

[0080] Nodes with the request volume, error rate, or response time exceeding the set thresholds are set as anomaly nodes.

[0081] Label matching

[0082] According to the predefined types and labels (such as anomaly_request_volume, anomaly_error_rate, anomaly_response_time), identify the categories of these anomaly nodes.

[0083] Extract nodes

[0084] Extract nodes for classification. Extract all nodes marked as anomalies from the topology graph, and then classify the same metric to form a subset.

[0085] In this embodiment, extract features useful for anomaly judgment. For example, retain the calculated average response time, peak CPU utilization rate, memory usage, and average in - out traffic metrics. For different metric types, extract feature values in a specified manner. The specific methods are as follows:

[0086] Extraction method for resource metric feature values

[0087] Time - series analysis: Extract data of metrics such as CPU usage rate and memory usage rate (resource metrics) within the same time period, calculate the mean and extreme values within the time period, and statistically analyze its features. Such as the feature: the average CPU usage rate of the current node within 1 hour.

[0088] Identify event anomaly frequency: Analyze the occurrence frequency of specific events within a period of time (such as metric data where the CPU usage rate index exceeds the average CPU usage rate within 1 hour), and identify anomaly index information.

[0089] Aggregation and grouping: Aggregate according to the service instance dimension, and calculate the summary statistics under the service instance dimension. Example feature values: features such as the average CPU usage rate, average memory usage rate, and average utilization rate of disk I / O of anomaly nodes.

[0090] Method for Extracting Eigenvalues of Application Performance Metrics

[0091] Time series analysis: Extract data within the same time period of metrics such as response time and error rate (business application performance metrics), calculate the mean, maximum, and minimum values within the time period, and statistically analyze their characteristics. For example, the characteristic: the average response time of the current node within 1 hour.

[0092] Identifying event anomaly frequencies: Analyze the occurrence frequencies of specific events within a period of time (such as metric data where the single response time metric exceeds the average response time within 1 hour), and identify abnormal metric information.

[0093] Aggregation and grouping: Aggregate according to the service instance dimension, and calculate the summary statistics under the service instance dimension. Example eigenvalue: the average response time of the same instance.

[0094] Method for Extracting Eigenvalues of Network Condition Metrics

[0095] Time series analysis: Extract data within the same time period of metrics such as average round-trip time (network condition metrics), calculate the mean, maximum, and minimum values within the time period, and statistically analyze their characteristics. For example, the characteristic: the average round-trip time metric of the current node within 1 hour.

[0096] Identifying event anomaly frequencies: Analyze the occurrence frequencies of specific events within a period of time (such as metric data where the single round-trip time metric exceeds the average round-trip time within 1 hour), and identify abnormal metric information.

[0097] Aggregation and grouping: Aggregate according to the service instance dimension, and calculate the summary statistics under the service instance dimension. Example eigenvalues: the utilization rate of the average round-trip time of abnormal nodes, the TCP connection establishment success rate, and other eigenvalues.

[0098] In this embodiment, a decision tree for metric anomaly delimitation is constructed by the C4.5 algorithm, and the decision tree model is trained using the training set data to ensure that the model can accurately delimit abnormal problems based on the input features.

[0099] Algorithm implementation method:

[0100] Data preparation: Use the collected abnormal events to ensure that the dataset contains rich features, such as CPU utilization rate, memory usage, response time, etc.

[0101] Preprocessing: Complete the abnormal events with missing feature values and extract numerical features. Completing the missing feature values is an important step in data preprocessing, especially when dealing with abnormal events. For extracting numerical features and completing the missing values, the mean filling technique can be used.

[0102] The following are the detailed strategies and methods for completing the missing values:

[0103] Obtain the original sample data

[0104] # Sample data

[0105] data = {'cpu_usage': [None, 0.8, 0.9, 0.6, 0.5],

[0106] 'memory_usage': [0.7, 0.8, None, 0.5, 0.4]}

[0107] Calculate the mean of the original sample data

[0108] cpu_avg: 0.7

[0109] memory_avg: 0.6

[0110] Fill in the missing values of the original sample data

[0111] # Sample data after filling in the missing values

[0112] data = {'cpu_usage': [ 0.7, 0.8, 0.6, 0.9, 0.5],

[0113] 'memory_usage': [0.7, 0.8, 0.6, 0.5, 0.4]}

[0114] Select the optimal splitting attribute: The C4.5 algorithm uses the Gain Ratio to select the optimal splitting attribute. The Gain Ratio is the information gain divided by the intrinsic value of the attribute, which is used to penalize attributes with a large number of values, thus avoiding bias towards attributes with more possible values.

[0115] Calculate the information gain: Information Gain (IG) is a key concept used to measure the importance of an attribute in a classification task. It is based on the concept of entropy in information theory and is used to evaluate how much uncertainty can be reduced by introducing a certain attribute. Specifically, the information gain reflects how much the purity of the class distribution has been improved after using a certain attribute to partition the dataset.

[0116] ;

[0117] ;

[0118] Where: Values(A) represents all possible values of attribute A; Sv is the subset of the dataset where the value of attribute A is v; |S| and |Sv| represent the number of samples in dataset S and subset Sv respectively; Entropy(S) is the entropy of dataset S; P i is the probability of the i-th category;

[0119] Calculate the Intrinsic Value: Calculate the intrinsic value of attribute A, which reflects the entropy of the distribution of different values in attribute A.

[0120] ;

[0121] The C4.5 algorithm uses the information gain ratio to select the optimal splitting attribute, and the formula is:

[0122] ;

[0123] Where: S is the current dataset; A is the attribute to be evaluated; Gain(S, A) represents the information gain, which measures the importance of an attribute A in the classification task S; Intrinsic Value(A) represents the calculated intrinsic value, which calculates the intrinsic value of attribute A and reflects the entropy of the distribution of different values in attribute A. Example

[0124] Suppose we have a dataset that contains several log entries, each entry records the CPU usage rate and other relevant information, and is labeled as whether it is abnormal ("normal" or "abnormal"). The information gain of the C4.5 algorithm will be used to evaluate the importance of CPU usage rate as a splitting attribute.

[0125] The example dataset is shown in Table 2:

[0126] Table 2

[0127]

[0128] Calculate the Entropy of the original dataset

[0129] First, we need to calculate the entropy of the entire dataset. Suppose there are n samples in the dataset, where p are "normal" and q are "abnormal".

[0130] ;

[0131] For example, if there are 3 "normal" and 2 "abnormal" in the dataset:

[0132] ;

[0133] Select the splitting point and calculate the entropy of each subset

[0134] For continuous-valued attributes such as CPU usage, we need to select an optimal splitting point. Suppose we select a CPU usage of 0.7 as the splitting point and divide the dataset into two parts:

[0135] CPU usage ≤ 0.7: Includes samples with CPU usage of 0.7, 0.6, and 0.5.

[0136] CPU usage > 0.7: Includes samples with CPU usage of 0.8 and 0.9.

[0137] Subset 1 (CPU usage ≤ 0.7), see Table 3.

[0138] Table 3

[0139]

[0140] There are 3 samples in this subset, all of which are "normal". Therefore, its entropy is 0 (completely pure).

[0141] Entropy(S≤0.7)=0

[0142] Subset 2 (CPU usage > 0.7), see Table 4.

[0143] Table 4

[0144]

[0145] Entropy(S > 0.7)=0

[0146] Calculate the weighted average entropy

[0147] Calculate the weighted average entropy according to the proportion of each subset:

[0148] ;

[0149] Calculate the information gain

[0150] Finally, calculate the information gain:

[0151] Gain(S,CPU usage)=Entropy(S)−Weighted Entropy

[0152] Gain(S,CPU usage)=0.971−0=0.971

[0153] Calculate the split information

[0154] Based on the above splitting, we can calculate the intrinsic value SplitInfo(A):

[0155] ;

[0156] In this example, the attribute A (i.e., CPU usage rate) is divided into two subsets, and the total number of samples |S| = 5:

[0157] S ≤ 0.7 contains 3 samples;

[0158] S > 0.7 contains 2 samples;

[0159] ;

[0160] Calculate the information gain ratio

[0161] The information gain Gain(S,A) has been calculated to be 0.971, and the information gain ratio can be calculated:

[0162] ;

[0163] In this example, using the CPU usage rate as the splitting attribute, the information gain is 0.971, which indicates that dividing by the CPU usage rate significantly reduces the uncertainty and makes the subsets more pure. When using the CPU usage rate as the splitting attribute, the intrinsic value SplitInfo(A) is approximately 0.971. This indicates that although dividing by the CPU usage rate introduces a certain degree of complexity, since the information gain is also very high, the information gain ratio is close to 1, indicating that this is a very effective splitting attribute.

[0164] Handling continuous values: For continuous value attributes, C4.5 will try to find an optimal threshold to split the data. This process is called Binary Splitting. Traverse all possible threshold points and select the one that maximizes the information gain as the splitting point. The specific implementation of Binary Splitting is as follows:

[0165] Sorting: First, sort the continuous value attribute.

[0166] Traversal: Then traverse the sorted values to find the optimal splitting point. For each potential splitting point, calculate the information gain of the two subsets after splitting.

[0167] Select the optimal splitting point: Select the splitting point with the maximum information gain as the final splitting condition.

[0168] Construct the decision tree

[0169] Recursive splitting: Starting from the root node, select the best splitting attribute according to the above rules and create child nodes. For each child node, repeat this process until the stopping condition is met (such as all instances belonging to the same class, no more attributes available, etc.).

[0170] Pruning: To prevent overfitting, C4.5 introduces a post-pruning strategy. Pruning can simplify the tree structure by manually deleting unnecessary branches while maintaining the classification performance.

[0171] Model evaluation and optimization:

[0172] Evaluate the performance of the model using manual verification techniques, manually adjust the parameters to optimize the model performance, and further enhance the stability and accuracy of a single decision tree.

[0173] Apply the model for diagnosis:

[0174] When a new anomaly is detected, extract the corresponding features and input them into the trained decision tree model. The model will start from the root node along the tree structure, make judgments according to the feature values in turn, and finally reach a leaf node. The label carried by this leaf node is the predicted root cause of the anomaly.

[0175] Continuous improvement:

[0176] As new data arrives, continuously update the training set, retrain the model regularly to adapt to system changes, and maintain the effectiveness of the model.

[0177] Automated boundary determination of the decision tree model

[0178] When a new anomaly event is detected, extract the corresponding features and input them into the decision tree model. The model will start from the root node along the tree structure, make judgments according to the feature values in turn, and finally reach a leaf node. The label carried by this leaf node is the predicted root cause of the anomaly. The automated boundary determination process for anomaly problems in the decision tree model is as follows:

[0179] Preliminary inspection of business application resources:

[0180] The decision tree model starts by evaluating the key resource metrics of the business application, including CPU usage, memory usage, and disk I / O metrics. These resource - type metrics directly reflect the basic environment status during application runtime. If the detection results show that these resources are all within the normal range, exclude problems caused by resource shortages and continue in - depth analysis; conversely, if resource anomalies are found, immediately locate them as caused by resource bottlenecks.

[0181] Analyze network and application performance during slow - request periods:

[0182] Next, the decision tree model focuses on the network and application performance within 1 minute before and after the occurrence of slow requests. Specifically, it checks for connection establishment failures for backward requests (network - related metrics) or persistent slow requests (business application performance - related metrics). If significant network connection problems or performance degradation are detected during this period, it is defined as being caused by inefficient internal processing of the network or application; otherwise, continue with the next step of analysis.

[0183] Detect garbage collection activities:

[0184] The decision tree model further checks whether a full GC (Full Garbage Collection) occurred before and after the slow - request period. Full GC is a memory - management operation in Java applications, which may cause the application to pause execution and thus affect the response time. If a full GC event is confirmed, it prompts the user that there may be performance problems due to improper memory management; otherwise, continue to explore other possibilities.

[0185] Evaluate the resource status of the host where the application is located:

[0186] Finally, the decision tree model turns to check the overall resource situation of the host hosting the application, and reviews the usage of CPU, memory, and disk I / O again. If the host resources are abnormal, such as overloaded or saturated, it is attributed to problems at the host level; if the host resources are normal, it is finally determined that the problem stems from the logic or configuration defects of the application itself.

[0187] Through the above - mentioned hierarchical and progressive decision - making path, this model can not only accurately locate the causes of business application anomalies, but also provide clear operation guidelines for the operation and maintenance team, thereby accelerating the fault recovery speed and enhancing the stability and reliability of the system. This method ensures effective monitoring and rapid response to complex and changing application environments.

[0188] This application also provides a problem - delimiting analysis system, including:

[0189] A collection module, used to collect data from different sources. The data includes resource metrics, business application performance metrics, network status and network call metrics, and key events in log records;

[0190] A construction module, used to construct a business topology graph based on call metrics, map resource metrics to corresponding nodes in the business topology graph, overlay business application performance metrics on the service links of the business topology graph, and mark key times on relevant nodes of the business topology graph;

[0191] Set a threshold for each piece of data, and events exceeding the threshold are abnormal events;

[0192] An extraction module, used to extract abnormal nodes from the business topology graph based on abnormal events;

[0193] A feature module for extracting abnormal features based on abnormal nodes;

[0194] A supplementary module for extracting numerical features from the abnormal features and supplementing the missing numerical features;

[0195] A decision tree module for constructing a decision tree model for abnormal delimitation based on the numerical features through the C4.5 algorithm;

[0196] A delimitation module for delimiting abnormal events by inputting the to-be-input features into the decision tree model after the decision tree model training is completed.

[0197] The present invention also provides an electronic device, including at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the above method.

[0198] The present invention also provides a storage medium storing a computer program executable by an electronic device, and when the program runs on the electronic device, the electronic device executes the above method.

[0199] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for analyzing delimiter problems, characterized in that including: collecting data from different sources, where the data includes resource metrics, business application performance metrics, network conditions and network call metrics, and key events in log records; constructing a business topology map based on the call metrics, mapping the resource metrics to corresponding nodes of the business topology map, overlaying the business application performance metrics on the service links of the business topology map, and labeling the key times on relevant nodes of the business topology map; setting a threshold for each piece of the data, and events exceeding the threshold are abnormal events; extracting abnormal nodes from the business topology map based on the abnormal events; extracting abnormal features based on the abnormal nodes; extracting numerical features in the abnormal features and supplementing the missing numerical features; constructing a decision tree model for abnormal delimitation through the C4.5 algorithm based on the numerical features; after the decision tree model is trained, inputting the to-be-input features into the decision tree model to delimit abnormal events; wherein, constructing a decision tree model for abnormal delimitation through the C4.5 algorithm includes: the C4.5 algorithm uses the information gain ratio to select the optimal splitting attribute, and the formula is: ; In the formula: S is the current data set; A is the attribute to be evaluated; Gain(S, A) represents the information gain, which measures the importance of an attribute A in the classification task S; Intrinsic Value(A) represents calculating the intrinsic value, calculating the intrinsic value of attribute A, which reflects the entropy of the distribution of different values in attribute A; the calculation formula of the information gain is: ; ; Where: Values(A) represents all possible values of attribute A; Sv is the subset of the dataset where the value of attribute A is v; |S| and |Sv| represent the number of samples in dataset S and subset Sv respectively; Entropy(S) is the entropy of dataset S; P i is the probability of the i-th category; the calculation formula of the intrinsic value is: ; constructing a decision tree, starting from the root node, selecting the best splitting attribute, creating child nodes, and repeating this process for each child node until the stopping condition is met.

2. The delimiter problem analysis method according to claim 1, characterized in that the resource metrics include cpu usage rate, memory usage rate, and disk IO metric usage rate; the business application performance metrics include response time and error rate; the network conditions and network call metrics include network traffic and latency.

3. The delimitation problem analysis method according to claim 2, wherein setting corresponding labels and types according to the root cause of the abnormal events.

4. The delimitation problem analysis method according to claim 3, characterized in that extracting abnormal nodes from the business topology map includes: monitoring nodes where the request volume, error rate, or response time exceeds the threshold, and setting them as the abnormal nodes; identifying the types of the abnormal nodes based on the labels and types; extracting all the abnormal nodes from the business topology map and classifying them to form a subset.

5. The delimitation problem analysis method according to claim 4, characterized in that, extracting abnormal features based on the abnormal nodes includes: extracting data within the same time period of the data, calculating the mean value, maximum value, and minimum value within the time period, and statistically analyzing its features; analyzing the occurrence frequency of specific events within a period of time; aggregating according to the service instance dimension and calculating the summary statistics under the service instance dimension.

6. A delimited problem analysis system, characterized in that including: a collection module for collecting data from different sources, where the data includes resource metrics, business application performance metrics, network conditions and network call metrics, and key events in log records; A building block for constructing a business topology graph based on the call metrics, where the resource metrics are mapped to corresponding nodes of the business topology graph, the business application performance metrics are superimposed on the service links of the business topology graph, and the key times are marked on relevant nodes of the business topology graph; Set a threshold for each of the data, and an event exceeding the threshold is an abnormal event; An extraction module for extracting abnormal nodes from the business topology graph based on the abnormal events; A feature module for extracting abnormal features based on the abnormal nodes; A supplement module for extracting numerical features in the abnormal features and supplementing the missing numerical features; A decision tree module for constructing a decision tree model for abnormal delimitation based on the numerical features by using the C4.5 algorithm; A delimitation module for delimiting abnormal events by inputting the to-be-input features into the decision tree model after the decision tree model training is completed; Among them, constructing a decision tree model for abnormal delimitation by using the C4.5 algorithm includes: The C4.5 algorithm uses the information gain ratio to select the optimal splitting attribute, and the formula is: ; In the formula: S is the current data set; A is the attribute to be evaluated; Gain(S, A) represents the information gain, which measures the importance of an attribute A in the classification task S; Intrinsic Value(A) represents calculating the intrinsic value, calculating the intrinsic value of the attribute A, and reflecting the entropy of the distribution of different values in the attribute A; The calculation formula of the information gain is: ; ; Where: Values(A) represents all possible values of attribute A; Sv is the subset of the dataset where the value of attribute A is v; |S| and |Sv| represent the number of samples in dataset S and subset Sv respectively; Entropy(S) is the entropy of dataset S; P i is the probability of the i-th category; The calculation formula of the intrinsic value is: ; Construct a decision tree. Starting from the root node, select the best splitting attribute and create child nodes. For each child node, repeat this process until the stop condition is met.

7. An electronic device, characterized in that, It includes at least one processing unit and at least one storage unit. Among them, the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the method according to any one of claims 1 to 5.

8. A storage medium, characterized in that, It stores a computer program executable by an electronic device, and when the program runs on the electronic device, the electronic device executes the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Network anomaly positioning method and device

    CN114765574A

  • Micro-service fault diagnosis method and system

    CN115640159A