Service exception root cause analysis method and device, electronic equipment and storage medium
By collecting and correlating monitoring, logs, and tracking data from related services, the system automatically analyzes the root causes of anomalies, solving the problem of low troubleshooting efficiency caused by switching between multiple tools in operations and maintenance, and achieving efficient fault diagnosis and service stability assurance.
Patent Information
- Application Number
- CN202510834093.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-24
AI Technical Summary
In existing technologies, operation and maintenance personnel need to switch back and forth between multiple tools, resulting in low troubleshooting efficiency and inability to work effectively together.
Collect and correlate service monitoring data, log data, and tracking data. When monitoring data anomalies are detected, automatically correlate the log and tracking data for the corresponding time period, analyze and calculate the correlated data, determine the root cause of the anomaly, and generate an analysis report.
It enables intelligent linkage analysis of three major observable data points, quickly and accurately pinpointing the root cause of service anomalies, significantly improving the accuracy and efficiency of fault diagnosis, and reducing the technical threshold and workload of operation and maintenance personnel.
Smart Images

Figure CN120832259A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of information technology and operation and maintenance technology, and in particular, to a service exception root cause analysis method and device, electronic equipment and storage medium. BACKGROUND
[0002] In the field of modern information technology, the complexity and diversification of large systems have made the operation and maintenance task extremely heavy. The separate monitoring system, log retrieval platform and tracking tool often cannot work completely in coordination, resulting in the need for the operation and maintenance personnel to switch back and forth between multiple tools when facing problems, which reduces the troubleshooting efficiency. SUMMARY
[0003] The present disclosure provides a service exception root cause analysis method, device, electronic equipment and storage medium. The main purpose is to solve the problem that the operation and maintenance personnel need to switch back and forth between multiple tools when facing problems, which reduces the troubleshooting efficiency.
[0004] According to a first aspect of the present disclosure, a service exception root cause analysis method is provided, which comprises:
[0005] Collecting monitoring data, log data and tracking data of the service and establishing an association;
[0006] When detecting an abnormal monitoring data, automatically associating the log and tracking data of the corresponding time period;
[0007] Analyzing and calculating the associated data, determining the exception root cause and generating an analysis report.
[0008] Optionally, when detecting an abnormal monitoring data, automatically associating the log and tracking data of the corresponding time period further comprises:
[0009] When detecting an abnormal index data point, extracting the service identifier, cluster information and request characteristics corresponding to the time point of the abnormal index data point;
[0010] Constructing a log query condition according to the extracted service identifier, cluster information and request characteristics, and generating a log retrieval uniform resource locator with time range filtering;
[0011] Displaying a clickable link of the log retrieval uniform resource locator on the monitoring interface;
[0012] In response to a user click operation, jumping to the log system and automatically executing a query to display the matched detailed log record.
[0013] Optionally, in response to a user click operation, jumping to the log system and automatically executing a query to display the matched detailed log record further comprises:
[0014] When a request call link is received, an associated log query entry is generated for each interval element node;
[0015] In response to a user clicking on a query entry of an interval element node, the corresponding tracking identifier and span identifier are used to query and return a detailed log of the node; wherein the tracking identifier and span identifier are configured in advance.
[0016] Optionally, analyzing and calculating the associated data, determining the root cause of the anomaly, and generating an analysis report includes:
[0017] Build a topology analysis model that includes service nodes, dependent resources, and call relationships;
[0018] Analyze the performance indicators and error logs of each node along the call chain and calculate the anomaly probability score;
[0019] Generate a root cause analysis report containing the nodes with the highest abnormal probability and attach it to the alarm notification.
[0020] Optionally, when monitoring data anomalies are detected, automatically associating logs and tracking data for a corresponding time period includes:
[0021] Recording a tracking identifier of the current request for the abnormal indicator point of the abnormal monitoring data;
[0022] When displaying the indicator chart, add a tracking identifier to the abnormal indicator point;
[0023] In response to a user selection operation on the annotated data point, querying a distributed tracing system using the tracing identifier;
[0024] The complete request link associated with the abnormal indicator is displayed on the tracing system interface.
[0025] Optionally, after analyzing and calculating the associated data, determining the root cause of the anomaly, and generating an analysis report, the method further includes:
[0026] Regularly collect service monitoring indicators, alarm records, and tracking statistics;
[0027] Calculate resource utilization efficiency scores based on cost, failure rate scores based on stability, and vulnerability risk scores based on security.
[0028] The total service health score is calculated based on the weighted scores of each dimension. When the total health score is lower than a threshold, an optimization report containing specific improvement suggestions is generated.
[0029] According to a second aspect of the present disclosure, a device for analyzing root causes of service anomalies is provided, comprising:
[0030] The collection unit is configured to collect monitoring data, log data and tracking data of a service and establish an association;
[0031] The association unit is configured to automatically associate log data and tracking data of a corresponding time period when an abnormality in the monitoring data is detected.
[0032] The analysis and calculation unit is configured to analyze and calculate the associated data, determine an abnormality root cause and generate an analysis report.
[0033] Optionally, the association unit is further configured to:
[0034] extract a service identifier, cluster information and request characteristics corresponding to a time point of the abnormal indicator data point when the abnormal indicator data point is detected;
[0035] construct a log query condition according to the extracted service identifier, cluster information and request characteristics, and generate a log retrieval uniform resource locator with time range filtering;
[0036] display a clickable link of the log retrieval uniform resource locator on a monitoring interface;
[0037] in response to a user click operation, jump to a log system and automatically execute a query, and display matched detailed log records.
[0038] Optionally, the association unit is further configured to:
[0039] generate an associated log query entry for each interval element node when a request call link is received;
[0040] in response to a user click operation on the query entry of the interval element node, query and return detailed logs of the node using a corresponding tracking identifier and span identifier, wherein the tracking identifier and span identifier are obtained in advance.
[0041] Optionally, the analysis and calculation unit is further configured to:
[0042] build a topology analysis model containing service nodes, dependent resources and call relationships;
[0043] analyze performance indicators and error logs of each node along a call link, and calculate an abnormality probability score;
[0044] generate a root cause analysis report containing a node with the highest abnormality probability, and attach the report to an alarm notification for sending.
[0045] Optionally, the association unit is further configured to:
[0046] record a tracking identifier of a current request for the abnormal indicator point of the monitoring data abnormality;
[0047] When displaying the index chart, a tracking identifier label is added to the abnormal index point;
[0048] In response to a selection operation of a user on the labeled data point, the tracking identifier is used to query a distributed tracking system;
[0049] The complete request link associated with the abnormal index is displayed on the tracking system interface.
[0050] Optionally, the apparatus further comprises:
[0051] The acquisition unit is configured to periodically acquire the monitoring index, the alarm record and the tracking statistical data of the service after the analysis and calculation unit analyzes and calculates the associated data, determines the abnormal root cause and generates the analysis report;
[0052] The calculation unit is configured to calculate the resource use efficiency score according to the cost dimension, calculate the failure rate score according to the stability dimension, and calculate the vulnerability risk score according to the security dimension;
[0053] The generation unit is configured to calculate a total score of the service health degree based on the scores of each dimension, and generate an optimization report containing specific improvement suggestions when the total score of the health degree is lower than a threshold.
[0054] According to a third aspect of the present disclosure, an electronic device is provided, comprising:
[0055] at least one processor; and
[0056] a memory communicatively connected to the at least one processor; wherein
[0057] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the first aspect.
[0058] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method of the first aspect.
[0059] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method of the first aspect.
[0060] The service exception root cause analysis method, device, electronic equipment and storage medium provided by the present disclosure mainly include the following technical solutions: collecting monitoring data, log data and tracking data of a service and establishing an association; when detecting an abnormal monitoring data, automatically associating log data and tracking data of a corresponding time period; analyzing and calculating the associated data, determining an abnormal root cause and generating an analysis report. Compared with related technologies, the embodiments of the present disclosure collect monitoring data, log data and tracking data of a service and establish a multi-dimensional association, when detecting an abnormal monitoring, automatically associate log data and tracking data of a corresponding time period, and then cooperatively analyze and calculate the associated data, which can quickly and accurately locate the root cause of the service exception and generate a detailed analysis report. This technical solution effectively solves the problem of low fault positioning efficiency caused by the isolation of monitoring, log and tracking data in traditional operation and maintenance, realizes intelligent linkage analysis of the three major observability data, significantly improves the accuracy and efficiency of fault diagnosis, and greatly reduces the technical threshold and work burden of operation and maintenance personnel through automatic root cause analysis, thereby providing strong technical support for service stability guarantee. BRIEF DESCRIPTION OF DRAWINGS
[0061] The accompanying drawings are used to better understand the present scheme and do not constitute a limitation on the present disclosure. Among them:
[0062] Figure 1 A flowchart of a service exception root cause analysis method provided by the embodiments of the present disclosure;
[0063] Figure 2 A service exception root cause analysis architecture provided by the embodiments of the present disclosure;
[0064] Figure 3 A structural schematic diagram of a service exception root cause analysis device provided by the embodiments of the present disclosure;
[0065] Figure 4 A structural schematic diagram of another service exception root cause analysis device provided by the embodiments of the present disclosure;
[0066] Figure 5 A schematic block diagram of an example electronic device provided by the embodiments of the present disclosure. DETAILED DESCRIPTION
[0067] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to be clear and concise, the description below omits the description of well-known functions and structures.
[0068] A service exception root cause analysis method, device, electronic equipment and storage medium are described below with reference to the accompanying drawings.
[0069] Figure 1 A flowchart of a service exception root cause analysis method provided by an embodiment of the present disclosure.
[0070] As shown in the method comprises the following steps: Figure 1
[0071] Step 101, collecting monitoring data, log data and tracking data of the service and establishing correlation.
[0072] Please refer to Figure 2 , Figure 2 A service exception root cause analysis architecture provided by an embodiment of the present disclosure is shown in Figure 2 Monitoring data can reflect the running state of the service in real time, such as CPU usage, memory occupation, network traffic and other information of the server. Through the collection and analysis of these data, the performance bottleneck and potential problems of the service can be found in time. Log data records various events and operations in the service running process, such as user login behavior, request processing process, error occurrence record and the like, providing detailed clues and historical information for in-depth troubleshooting.
[0073] Tracking data can track the flow path of the request in the system, and clearly show the interaction relationship between each component, which is helpful for quickly locating the specific link where the problem is located. Establishing the correlation between data requires setting appropriate data collection mechanism on each key node and component of the service to ensure the accuracy and integrity of the data.
[0074] In some embodiments, for monitoring data, professional monitoring tools can be used for real-time collection; log data requires reasonable addition of log recording statements in the code to accurately record the occurrence time of events, related parameters and other information; the collection of tracking data needs the help of a distributed tracking system to add tracking identifiers at the initiation end of the request and each processing link.
[0075] After collecting these data, through the data processing and analysis platform, the monitoring data, log data and tracking data are correlated and integrated by using time stamp, request identifier and other key information. When the service has a problem, operation and maintenance personnel and developers can quickly and accurately locate the root cause of the problem by comprehensively analyzing these correlated data, take effective measures for repair and optimization, so as to ensure the stable and reliable operation of the service and improve the user experience.
[0076] Step 102, when detecting monitoring data exception, automatically correlating log and tracking data of the corresponding time period.
[0077] During service operation, the monitoring system monitors various indicators in real time. For example, a sudden increase in CPU load, reaching a critical value of memory usage, a significant increase in network latency, and the like are all manifestations of monitoring data anomalies. After detecting an anomaly, the system automatically associates the corresponding log and trace data for the time period. The log data records in detail the specific operations and events performed by the service during the time period, such as the processing flow of a certain request, read / write operations of a database, and the like, which can provide rich information for in-depth understanding of the specific state of the service when the anomaly occurs. The trace data clearly shows the transmission path of the request between various components of the system, helping to determine where the anomaly occurred.
[0078] To achieve automatic association, a perfect data association mechanism needs to be established during the system design phase. By adding uniform timestamps and unique request identifiers and the like to the monitoring data, log data, and trace data during data collection, relevant data for the corresponding time period can be quickly filtered based on these identifiers when an anomaly occurs. When monitoring data anomalies are triggered, operation and maintenance personnel or developers can quickly find the root cause of the anomaly by viewing the associated log and trace data, whether it is a code logic error, insufficient resources, or other external factors, and thus take appropriate measures to repair it in a timely manner, avoid greater service failures or performance degradation, and ensure service stability and reliability, and improve user satisfaction and trust in the service.
[0079] In some embodiments, when managing log and trace data, the following approach can also be used: collect microservice monitoring indicators, including general monitoring indicators required by all services and business-related custom monitoring indicators. There are mainly two types of microservice monitoring: one is general monitoring indicators required by all services, such as performance indicators such as interface time consumption, queries per second (QPS), success rate, and resource indicators such as CPU and memory usage; the other is business-related custom monitoring indicators, such as order volume and transaction amount.
[0080] After the collection of microservice monitoring indicators is complete, the effective convergence of the divergence dimension of some monitoring indicators is performed based on the collected microservice monitoring indicators.
[0081] Dimension divergence refers to the phenomenon that the number of label combinations grows exponentially due to the carrying of too many dynamic labels (such as random strings and meaningless parameters) by monitoring indicators. In order to ensure the comprehensiveness and accuracy of monitoring indicators, indicators will carry various dimension labels, such as service traffic monitoring, which generally carries cluster, instance, and interface information. In actual use, it is found that some monitoring indicators will have serious divergence, such as the use of RESTful (Representational State Transfer) style interfaces, which are prone to dimension divergence problems due to their stateless nature and dynamic URL structure, or custom business indicators containing temporary debugging labels, which will cause related monitoring indicators to have high base problems. This brings great stability risks to monitoring storage, and businesses also cannot obtain the real interface traffic distribution. Effective convergence is a process of reducing the number of time series through label rewriting, invalid data elimination, and indicator aggregation. Therefore, we design an effective convergence algorithm on the monitoring collection side, which can effectively converge the over-inflated monitoring indicators, including converging dynamic interface labels, static resource files, and meaningless strings such as random strings, to ensure that the monitoring indicator dimensions are within a controllable range.
[0082] For meaningless divergence dimensions, such as dynamic labels generated by RESTful interfaces and random strings, a real-time pattern recognition and regularization processing mechanism is established to automatically converge and aggregate. Effective convergence of some monitoring indicator divergence dimensions can not only reduce high base problems to ensure storage stability, but also improve query performance. According to business needs, monitoring indicators can be automatically aggregated, and users can directly observe real monitoring information without the need for secondary aggregation operations when querying, improving the usability of the monitoring system.
[0083] Step 103, analyze and calculate the associated data to determine the root cause of the anomaly and generate an analysis report.
[0084] For monitoring data, statistical analysis can be used to determine the trend of abnormal indicators, such as comparing CPU usage curves in normal and abnormal time periods to view their fluctuation amplitude and change rate to preliminarily determine the approximate direction of performance anomalies. Log data needs detailed text analysis to extract key events and error information, such as checking for database connection errors, file reading failures, and other prompts.
[0085] The tracking data helps to understand the flow path of the request in the entire system architecture, and if it is found that the processing time of a certain link is too long or the request interruption occurs, further focus can be made on the link for in-depth analysis. By comprehensively analyzing these associated data, combined with professional knowledge and experience, possible factors are gradually ruled out, and the root cause of the abnormality is finally determined, such as logic errors at the code level, insufficient system resources, faults of external dependent services, and various conditions. After determining the root cause of the abnormality, it is necessary to generate a detailed analysis report. The report should include the time of the abnormality, the specific monitoring index abnormality, the key information of the associated log and tracking data, the method and step adopted in the analysis process, the finally determined root cause of the abnormality, and the solution or improvement measure suggestion proposed for the root cause.
[0086] The analysis report can provide strong support for the solution of the current abnormal problem, and can also provide a reference for the optimization and improvement of subsequent services, help the technical team to better improve the stability and reliability of the service, reduce the probability of similar abnormal situations occurring in the future, and thus ensure the efficient operation of the entire system.
[0087] The service abnormal root cause analysis method provided by the present disclosure mainly includes the following technical solutions: collecting monitoring data, log data and tracking data of a service and establishing an association; when detecting monitoring data abnormality, automatically associating log and tracking data of the corresponding time period; analyzing and calculating the associated data to determine the root cause of the abnormality and generate an analysis report. Compared with related technologies, the embodiments of the present application can quickly and accurately locate the root cause of the service abnormality and generate a detailed analysis report by integrally collecting monitoring data, log data and tracking data of a service and establishing a multi-dimensional association, automatically associating log and tracking data of the corresponding time period when detecting monitoring abnormality, and then cooperatively analyzing and calculating the associated data. This technical solution effectively solves the problem of low fault positioning efficiency caused by the isolation of monitoring, log and tracking data in traditional operation and maintenance, realizes intelligent linkage analysis of the three major observability data, significantly improves the accuracy and efficiency of fault diagnosis, and greatly reduces the technical threshold and work burden of operation and maintenance personnel through automatic root cause analysis, thereby providing strong technical support for service stability guarantee.
[0088] In some embodiments, the automatically associating log and tracking data of the corresponding time period when detecting monitoring data abnormality further includes:
[0089] When detecting the abnormal index data point, extracting the service identifier, cluster information and request characteristics corresponding to the time point of the abnormal index data point;
[0090] Constructing a log query condition according to the extracted service identifier, cluster information and request characteristics, and generating a log retrieval uniform resource locator with time range filtering;
[0091] displaying a clickable link of the log retrieval uniform resource locator in the monitoring interface;
[0092] in response to a user click operation, jumping to the log system and automatically executing a query to display matching detailed log records.
[0093] When an abnormal indicator data point is detected, a series of key information corresponding to the time point of the data point is extracted, i.e., service identification, cluster information, and request characteristics. The service identification can accurately indicate which specific service has an abnormal condition, helping the operation and maintenance personnel to quickly locate the scope of the problem service; the cluster information includes the cluster environment where the service is located, including the configuration, load condition, etc. of the cluster; the request characteristics include detailed information such as the type of request, parameters, etc.
[0094] According to the extracted information, a log query condition is constructed, which can more accurately filter out the log records related to the anomaly by combining the service identification, cluster information, and request characteristics. For example, according to the service identification, the log of which service to query can be determined, combined with the cluster information, the interference of other clusters can be excluded, and the request characteristics can further refine the query range to ensure that the obtained log records are closely related to the request at the time of the anomaly. At the same time, a log retrieval uniform resource locator with time range filtering is generated, which contains accurate query conditions and time range, and can accurately point to the required log records.
[0095] The monitoring interface displays a clickable link of the log retrieval uniform resource locator, and the operation and maintenance personnel and developers can intuitively see and click on this link on the monitoring interface to quickly enter the log system for querying.
[0096] When the user clicks on the link, the system will respond to this operation, jump to the log system and automatically execute the query. The system will display matching detailed log records, which contain detailed information such as function calls, variable values, etc. of the service running at the time of the anomaly. More efficiently realize the association of monitoring data and log data, quickly locate the root cause of the abnormal problem, improve the operation and maintenance efficiency and troubleshooting ability of the system, and ensure the stable operation of the service.
[0097] In some embodiments, the response to the user click operation, jumping to the log system and automatically executing the query to display matching detailed log records further includes:
[0098] Upon receiving the request call link, an associated log query entry is generated for each interval element node;
[0099] In response to a user clicking on the query portal of the interval element node, the corresponding trace identifier and span identifier are used to query and return the detailed log of the node; wherein the trace identifier and span identifier are obtained by pre-configuration.
[0100] The request call link presents the flow path of the request between various components and links in the system, and the interval element node represents the key part on the path. Generating an associated log query portal for each such node is like setting a convenient "navigation point" for operation and maintenance personnel and developers on a complex system architecture map, so that they can more targetedly obtain relevant information when troubleshooting.
[0101] When the user clicks on the query portal of the interval element node, the system will respond and use the corresponding trace identifier and span identifier for querying. These trace identifiers and span identifiers are obtained by pre-configuration and are used for identification and positioning in the system. The trace identifier can uniquely identify the entire life cycle of a request in the system, from the initiation of the request to the end of the final response, and link all data related to the request; and the span identifier is used to identify the processing process of the request in a specific component or link, to help more accurately locate the specific location where the problem occurs.
[0102] Using these two identifiers for querying, the detailed log records corresponding to the node are filtered out and returned to the user. These detailed log records contain specific operation details of the node when processing the request, such as the calling order of functions, the passing of parameters, the throwing of exceptions, and the like. By viewing these detailed logs, the user can gain a deep understanding of the execution of the request at the node, and determine whether there are abnormalities or errors. This enhances the accuracy and efficiency of the system in the process of exception troubleshooting, so that operation and maintenance personnel and developers can more quickly locate the root cause of the problem and take appropriate measures to repair it, ensuring the stable operation of the system and the quality of the service. The entire process from monitoring to log querying to problem positioning is improved, providing strong support for fault diagnosis of complex systems.
[0103] In some embodiments, the analysis and calculation of the associated data, determination of the abnormal root cause, and generation of the analysis report include:
[0104] A topology analysis model containing service nodes, dependent resources, and call relationships is constructed;
[0105] The performance indicators and error logs of each node along the call link are analyzed, and an abnormal probability score is calculated;
[0106] A root cause analysis report containing the node with the highest abnormal probability is generated and attached to the alarm notification for sending.
[0107] In a complex system, there are numerous service nodes that cooperate with each other, rely on various resources such as databases, caches, etc., and have complex calling relationships. By constructing such a topology analysis model, the architecture of the system can be presented in a visual and structured manner, clearly showing the connection relationship between the service nodes, the distribution of dependent resources, and the calling path.
[0108] Along the calling link, the performance indicators and error logs of each node are analyzed to calculate the abnormal probability score. The calling link records the process of request passing from one node to another in the system. By analyzing the performance indicators of each node, such as response time, throughput, etc., it can be judged whether there is a performance bottleneck in the node. Error logs provide error information about the running process of the node, such as code exceptions, resource access failures, etc. By integrating these performance indicators and error log information, and using certain algorithms and rules to calculate the abnormal probability score of each node, the possibility of each node appearing abnormal can be quantified.
[0109] Generate a root cause analysis report containing the node with the highest abnormal probability and attach it to the alarm notification. After calculating the abnormal probability score of each node, find the node with the highest abnormal probability and generate a detailed root cause analysis report for it. The report should include the basic information of the node, the specific situation of the performance indicators, the detailed content of the error logs, and the analysis and reasoning process of these information, and finally conclude the possible causes of the abnormality. Attach this root cause analysis report to the alarm notification and send it to the relevant operation and development personnel, so that they can quickly understand the possible problem root when receiving the alarm, and take more targeted measures to troubleshoot and repair. This way improves the efficiency of problem handling, reduces the impact of faults on system operation, ensures the stability and reliability of the system, and also provides valuable reference for the continuous optimization of the system.
[0110] In some embodiments, when the monitoring data anomaly is detected, the corresponding time period of the log and the tracking data are automatically associated, including:
[0111] The abnormal indicator point of the monitoring data anomaly records the tracking identifier of the current request;
[0112] When displaying the indicator chart, add tracking identifier annotations to the abnormal indicator point;
[0113] In response to the user's selection operation on the annotated data point, the tracking identifier is used to query the distributed tracking system;
[0114] In the tracking system interface, the complete request link associated with the abnormal indicator is displayed.
[0115] When an abnormality in the monitoring data is detected and the abnormal indicator point is accurately identified, the tracking identifier of the current request is recorded. The tracking identifier is a key identifier throughout the entire system during the request flow process, which is like the "identity card" of the request, and can closely link all data related to the request, providing important clues for subsequent data association and analysis.
[0116] The indicator chart intuitively presents the change trend of the monitoring data, and the abnormal indicator point is often the focus of attention. By adding the tracking identifier label to these abnormal indicator points, the operation and maintenance personnel and the developers can clearly identify which data points are abnormal when viewing the chart, and quickly obtain the tracking identifier information related thereto.
[0117] When the user performs a selection operation on the labeled data point, the tracking identifier is used to query the distributed tracking system. The distributed tracking system records detailed flow information of the request in each component and link of the system, and by querying the system through the tracking identifier, the request data related to the abnormal indicator can be accurately filtered.
[0118] The complete request link associated with the abnormal indicator is displayed on the tracking system interface. The complete request link details the entire process of the request from the initiation end, through each intermediate component, to the response end. By viewing this link, the user can clearly understand the processing of the request at each link, including processing time, transmitted parameters, and other information. The implementation of this mechanism enables the monitoring data, tracking data, and subsequent analysis to be closely combined, providing strong support for quickly locating and solving system abnormalities, effectively improving the system operation efficiency and fault troubleshooting ability.
[0119] In some embodiments, after analyzing and calculating the associated data, determining the abnormal root cause, and generating an analysis report, the method further comprises:
[0120] Periodically collecting monitoring indicators, alarm records, and tracking statistical data of the service;
[0121] Calculating the resource use efficiency score according to the cost dimension, the failure rate score according to the stability dimension, and the vulnerability risk score according to the security dimension;
[0122] Based on the weighted calculation of the scores of each dimension, the total score of the service health degree is calculated, and when the total score of the health degree is lower than a threshold, an optimization report containing specific improvement suggestions is generated.
[0123] The monitoring indicators can reflect the running state of the service in real time, such as CPU usage, memory occupation, network bandwidth, etc. Through long-term collection and analysis of these indicators, the changing trend of service performance can be understood. The alarm record records the abnormal conditions occurring in the service running process, providing a basis for analyzing the causes and frequency of faults. The tracking statistical data can show the flow path and processing time of the request in the system. By regularly collecting these data, the running status of the service can be comprehensively and dynamically mastered.
[0124] The service is scored from different dimensions, and the resource use efficiency score is calculated according to the cost dimension, which helps to evaluate the benefits of the service in resource utilization. For example, by analyzing the use of CPU and memory, it can be determined whether there is a problem of resource waste or excessive configuration. If the resource use efficiency is low, it may mean that the service needs to be optimized to reduce operating costs. The failure rate score is calculated according to the stability dimension, which reflects the reliability of the service. A high failure rate indicates that the service has many faults, which may affect user experience and require in-depth investigation and improvement of the system. The vulnerability risk score is calculated according to the security dimension. Security is an important guarantee for the service. By evaluating the security vulnerabilities and risks in the system, timely measures are taken to repair and prevent, avoiding data leakage and system attacks.
[0125] The service health score is calculated based on the weighted scores of each dimension, and this comprehensive evaluation method can comprehensively reflect the overall condition of the service. The scores of different dimensions are given different weights according to their importance, and finally an objective service health score is obtained. When the health score is lower than the threshold, it means that the service has more serious problems and needs to be optimized. At this time, an optimization report containing specific improvement suggestions is generated, which will propose targeted improvement measures according to the analysis results of each dimension. For example, in terms of resource use efficiency, it is suggested to reasonably allocate and optimize the configuration of resources; in terms of stability, it is suggested to perform code review and performance optimization on modules with frequent faults; in terms of security, it is suggested to update security patches and strengthen access control in a timely manner.
[0126] Corresponding to the above-mentioned service exception root cause analysis method, the present application also provides a service exception root cause analysis device. Since the device embodiments of the present application correspond to the above-mentioned method embodiments, the details not disclosed in the device embodiments can be referred to the above-mentioned method embodiments, which will not be described in detail in the present application.
[0127] Figure 3 A structural schematic diagram of a service exception root cause analysis device provided by the embodiments of the present application is shown in Figure 3 as shown, comprising:
[0128] The acquisition unit 21 is configured to acquire monitoring data, log data and tracking data of the service and establish an association.
[0129] The association unit 22 is configured to automatically associate the log and the trace data of the corresponding time period when the monitoring data exception is detected;
[0130] The analysis and calculation unit 23 is configured to analyze and calculate the associated data, determine the exception root cause, and generate an analysis report.
[0131] Further, in a possible implementation manner of the embodiment of the present disclosure, the association unit 22 is further configured to:
[0132] extract the service identifier, the cluster information, and the request feature of the time point corresponding to the abnormal index data point when the abnormal index data point is detected;
[0133] construct a log query condition according to the extracted service identifier, the cluster information, and the request feature, generate a log retrieval uniform resource locator with time range filtering;
[0134] display a clickable link of the log retrieval uniform resource locator on the monitoring interface;
[0135] in response to a user click operation, jump to a log system and automatically execute a query, and display the matched detailed log record.
[0136] Further, in a possible implementation manner of the embodiment of the present disclosure, the association unit 22 is further configured to:
[0137] generate an associated log query entry for each interval element node when the request call link is received;
[0138] in response to a user click operation on the query entry of the interval element node, query and return the detailed log of the node using the corresponding trace identifier and span identifier, wherein the trace identifier and the span identifier are obtained in advance.
[0139] Further, in a possible implementation manner of the embodiment of the present disclosure, the analysis and calculation unit 23 is further configured to:
[0140] build a topology analysis model containing service nodes, dependent resources, and call relations;
[0141] analyze the performance indicators and error logs of each node along the call link, and calculate an exception probability score;
[0142] generate a root cause analysis report containing the node with the highest exception probability, and attach the report to an alarm notification for sending.
[0143] Further, in a possible implementation manner of the embodiment of the present disclosure, the association unit 22 is further configured to:
[0144] record the trace identifier of the current request for the abnormal index point of the monitoring data exception.
[0145] In displaying the index chart, a tracking identifier label is added to the abnormal index point;
[0146] In response to a selection operation of a user on the labeled data point, the tracking identifier is used to query a distributed tracking system;
[0147] In the tracking system interface, a complete request link associated with the abnormal index is displayed.
[0148] Further, in a possible implementation manner of the embodiment of the present disclosure, as shown in Figure 4 The apparatus further includes:
[0149] The collection unit 24 is configured to periodically collect the monitoring index, the alarm record and the tracking statistical data of the service after the analysis and calculation unit 23 analyzes and calculates the associated data, determines the abnormal root cause and generates the analysis report;
[0150] The calculation unit 25 is configured to calculate the resource use efficiency score according to the cost dimension, calculate the failure rate score according to the stability dimension, and calculate the vulnerability risk score according to the security dimension;
[0151] The generation unit 26 is configured to calculate a total score of the service health degree based on the dimension scores, and generate an optimization report containing specific improvement suggestions when the total score of the health degree is lower than a threshold.
[0152] It should be noted that the foregoing explanation and description of the method embodiments are also applicable to the apparatus of the embodiment of the present disclosure, and the principle is the same, which is not limited in the embodiment of the present disclosure.
[0153] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0154] Figure 5 A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.
[0155] As Figure 5As shown, the device 400 includes a computing unit 401 that can perform various appropriate actions and processes in accordance with a computer program stored in a ROM (Read-Only Memory) 402 or a computer program loaded into a RAM (Random Access Memory) 403 from the storage unit 408. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An I / O (Input / Output) interface 405 is also connected to the bus 404.
[0156] A plurality of components in the device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, and the like; an output unit 407, such as various types of displays, speakers, and the like; a storage unit 408, such as a magnetic disk, an optical disk, and the like; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 409 allows the device 400 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0157] The computing unit 401 can be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, and the like. The computing unit 401 performs various methods and processes described above, such as the service anomaly root cause analysis method. For example, in some embodiments, the service anomaly root cause analysis method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to perform the aforementioned service anomaly root cause analysis method by any other appropriate means, such as by means of firmware.
[0158] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application-Specific Integrated Circuit (ASIC), an Application Specific Standard Product (ASSP), a System on a Chip (SOC), a Complex Programmable Logic Device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0159] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general or special purpose computer, such that the program code, when executed by the processor or controller, causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be implemented in a wholly in machine language, in partially in machine language, in partially in a high level language, and other combinations thereof. The program code can execute entirely on the machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0160] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a linearly-programmed electrical connection, a portable computer diskette, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory), or flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0161] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0162] The systems and techniques described here can be implemented in a computing system that includes a back-end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front-end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.
[0163] The computer system can include clients and servers. This relationship can be between a client and a server that are typically remote from each other and typically interact through a communication network. The relationship between client and server exists by virtue of computer programs running on the respective computer systems and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS (Virtual Private Server, or VPS for short) services. The server can also be a server of a distributed system, or a server combined with a blockchain.
[0164] It should be noted that artificial intelligence is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors of people (such as learning, reasoning, thinking, planning, etc.), both hardware and software technologies. Artificial intelligence hardware technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing, etc.; artificial intelligence software technology mainly includes computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, etc. several major directions.
[0165] It should be understood that the various forms of the flow shown above can be used to reorder, add or delete steps. For example, each step described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which is not limited herein.
[0166] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A service abnormality root cause analysis method characterized by comprising: The method comprises: Collecting monitoring data, log data and tracking data of a service and establishing association; When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; Analyzing and calculating the associated data to determine the root cause of the abnormality and generate an analysis report.
2. The method of claim 1, wherein, The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; 3. The method of claim 2, wherein, The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises:
4. The method of claim 1, wherein, When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises:
5. The method of claim 3, wherein, When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; 6. The method according to any one of claims 1-5, characterized in that, The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; 7. A service abnormality root cause analysis apparatus characterized by comprising: The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; 8. An electronic device, comprising: The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period; The method further comprises: When detecting an abnormality in the monitoring data, automatically associating log and tracking data of a corresponding time period The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
9. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, The computer instructions are for causing the computer to perform the method of any one of claims 1-6.
10. A computer program product, characterised in that, A computer program comprising instructions which, when executed by a processor, implement the method of any one of claims 1-6.
Citation Information
Cited By
Abnormal monitoring method for data transmission and electronic equipment
CN121012731A
Server operation and maintenance management method and device, electronic equipment and medium
CN122064555A