Problem root cause positioning method and device for application

By importing probe components into the computer system to capture abnormal information and perform correlation analysis, the root cause positioning complexity and uncertainty problems applied in the computer system are solved, real-time and accurate root cause positioning and system optimization are achieved.

CN119988066APending Publication Date: 2025-05-13BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311510059.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The problems in applications in computer systems are due to complex positioning and uncertainty, the performance overhead of existing technology, high requirements for operators, and relying on third-party services and libraries.

Method used

By importing preset probe components when the target application is started, abnormal information is captured, and operation data is collected, and correlation analysis is used to determine the degree of correlation between abnormal information and performance indicators, thereby determining the root cause information that affects the performance of the target application.

Benefits of technology

Real-time and accurate root cause positioning is achieved, abnormal recurrence is reduced, the system architecture is optimized, and the system availability and stability are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988066A_ABST
    Figure CN119988066A_ABST
Patent Text Reader

Abstract

The invention discloses a problem root cause positioning method and device of an application, and relates to the technical field of computers. A specific embodiment of the method comprises the following steps: when a target application is started, importing a preset probe assembly into the target application, and capturing abnormal information of the target application by utilizing the probe assembly; collecting operation data of the target application, and determining index data of the target application in at least one performance index according to the operation data; and obtaining a correlation degree between the abnormal information and the index data in the same statistical period, and determining root cause information influencing the performance of the target application according to the correlation degree. According to the embodiment, the problem root cause can be accurately positioned in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and device for locating the root cause of an application problem. Background Art

[0002] Locating the root cause of problems in computer system applications (e.g., various Java applications and services, such as gateway log storage, traffic monitoring, etc.) is a complex task that requires the use of multiple technologies and tools to achieve accurate and efficient results. In the related technologies, there are root cause locating methods such as distributed tracing systems, performance analysis tools, breakpoint debugging and debuggers, fault injection and stress testing, code review and static analysis, version control and rollback strategies. These methods have varying degrees of complexity and uncertainty, high performance overhead, high requirements for operators, and reliance on third-party services and libraries. Summary of the invention

[0003] In view of this, an embodiment of the present invention provides a method and device for locating the root cause of an application problem, which implements a root cause locating solution based on probes and correlation analysis, and can accurately locate the root cause of the problem in real time.

[0004] To achieve the above objective, according to one aspect of the present invention, a method for locating the root cause of an application problem is provided.

[0005] The method for locating the root cause of an application problem in an embodiment of the present invention includes: when a target application is started, importing a preset probe component into the target application, and using the probe component to capture abnormal information of the target application; collecting operating data of the target application, and determining indicator data of at least one performance indicator of the target application based on the operating data; obtaining the degree of correlation between the abnormal information and the indicator data in the same statistical period, and determining the root cause information affecting the performance of the target application based on the degree of correlation.

[0006] Optionally, the exception information includes the exception occurrence time and exception type, and also includes at least one of the following data: exception interface, exception calling method, exception description information, and exception call stack information.

[0007] Optionally, the degree of association is indicated by a correlation coefficient; and obtaining the degree of association between the abnormal information and the indicator data in the same statistical period includes: generating an abnormal sequence of a preset aggregation dimension based on the abnormal information in the statistical period; wherein the abnormal sequence characterizes the changing trend of the number of abnormal occurrences in the aggregation dimension over time; generating an indicator sequence of the performance indicator based on the indicator data of any performance indicator in the statistical period, wherein the indicator sequence characterizes the changing trend of the indicator data over time; and determining the correlation coefficient between the abnormal sequence and the indicator sequence.

[0008] Optionally, the aggregation dimensions include: exception type, exception interface or exception calling method; and the determining of the root cause information affecting the performance of the target application according to the degree of correlation includes: for any performance indicator in the indicator sequence of the statistical period, respectively determining the correlation coefficient between the indicator sequence and the exception sequence of each exception type, and determining the exception type corresponding to the maximum correlation coefficient as the root cause exception type of the performance indicator; for any performance indicator in the indicator sequence of the statistical period, respectively determining the correlation coefficient between the indicator sequence and the exception sequence of each exception interface, and determining the exception interface corresponding to the maximum correlation coefficient as the root cause interface of the performance indicator; for any performance indicator in the indicator sequence of the statistical period, respectively determining the correlation coefficient between the indicator sequence and the exception sequence of each exception calling method, and determining the exception calling method corresponding to the maximum correlation coefficient as the root cause calling method of the performance indicator.

[0009] Optionally, the method further comprises: storing the exception information captured by the probe component in an exception log file of the target application, and storing the exception log file in a preset database.

[0010] Optionally, the performance indicators include: response time, availability or access volume; and the determining of the root cause information affecting the performance of the target application based on the degree of correlation includes: for the indicator sequence of the response time indicator in the statistical period: respectively determining the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal type, and determining the abnormal type corresponding to the maximum correlation coefficient as the root cause abnormal type of the response time indicator of the target application; respectively determining the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal interface, and determining the abnormal interface corresponding to the maximum correlation coefficient as the root cause interface of the response time indicator of the target application; respectively determining the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal calling method, and determining the abnormal calling method corresponding to the maximum correlation coefficient as the root cause calling method of the response time indicator of the target application; for the indicator sequence of the availability indicator in the statistical period: respectively determining the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal type, and determining the abnormal type corresponding to the maximum correlation coefficient as the availability indicator of the target application. The root cause abnormal type; respectively determine the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal interface, and determine the abnormal interface corresponding to the maximum correlation coefficient as the root cause interface of the availability indicator of the target application; respectively determine the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal call method, and determine the abnormal call method corresponding to the maximum correlation coefficient as the root cause call method of the availability indicator of the target application; for the indicator sequence of the access volume indicator in the statistical period: respectively determine the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal type, and determine the abnormal type corresponding to the maximum correlation coefficient as the root cause abnormal type of the access volume indicator of the target application; respectively determine the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal interface, and determine the abnormal interface corresponding to the maximum correlation coefficient as the root cause interface of the access volume indicator of the target application; respectively determine the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal call method, and determine the abnormal call method corresponding to the maximum correlation coefficient as the root cause call method of the access volume indicator of the target application.

[0011] To achieve the above objective, according to another aspect of the present invention, a device for locating the root cause of an application problem is provided.

[0012] The device for locating the root cause of the problem of an application in an embodiment of the present invention includes: an exception capture unit, which is used to import a preset probe component into the target application when the target application is started, and use the probe component to capture exception information of the target application; a performance monitoring unit, which is used to collect operating data of the target application, and determine indicator data of at least one performance indicator of the target application based on the operating data; and a correlation calculation unit, which is used to obtain the degree of correlation between the exception information and the indicator data in the same statistical period, and determine the root cause information affecting the performance of the target application based on the degree of correlation.

[0013] Optionally, the exception information includes the exception occurrence time and exception type, and also includes at least one of the following data: exception interface, exception calling method, exception description information, and exception call stack information.

[0014] To achieve the above objective, according to another aspect of the present invention, an electronic device is provided.

[0015] An electronic device of the present invention includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method for locating the root cause of the application problem provided by the present invention.

[0016] To achieve the above objective, according to yet another aspect of the present invention, a computer-readable storage medium is provided.

[0017] A computer-readable storage medium of the present invention stores a computer program, which, when executed by a processor, implements the method for locating the root cause of an application problem provided by the present invention.

[0018] According to the technical solution of the present invention, the embodiments of the above invention have the following advantages or beneficial effects:

[0019] On the one hand, when the target application is started, the preset probe component is imported into the target application, and the probe component is used to capture the abnormal information of the target application. On the other hand, the operation data of the target application is collected, and the indicator data of the target application in at least one performance indicator is determined based on the operation data. Thereafter, the degree of correlation between the abnormal information and the indicator data in the same statistical period is calculated, and the root cause information affecting the performance of the target application is determined based on the degree of correlation. In this way, accurate and real-time root cause location of the problem is achieved based on probe technology and correlation analysis, which helps to optimize the system architecture and reduce the recurrence of abnormalities.

[0020] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with the specific implementation manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention.

[0022] Figure 1 Schematic diagram of main steps of the problem root cause location method used in an embodiment of the present invention;

[0023] Figure 2 Schematic diagram of the architecture of the problem root cause location method used in an embodiment of the present invention;

[0024] Figure 3 It is a schematic diagram of components of a problem root cause location device used in an embodiment of the present invention;

[0025] Figure 4 is an exemplary system architecture diagram in which embodiments of the present invention may be applied;

[0026] Figure 5 It is a schematic diagram of the structure of an electronic device used to implement the problem root cause locating method applied in an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.

[0028] It should be pointed out that the embodiments of the present invention and the technical features therein may be combined with each other without conflict.

[0029] Figure 1 1 is a schematic diagram of main steps of a method for locating the root cause of a problem applied in an embodiment of the present invention.

[0030] like Figure 1 As shown, the method for locating the root cause of the application problem in the embodiment of the present invention can be specifically performed according to the following steps:

[0031] Step S101: when a target application is started, a preset probe component is imported into the target application, and abnormal information of the target application is captured by using the probe component.

[0032] In this step, the probe component can be a Java probe. Optionally, when the target application is started, the javaagent technology can be used to import the probe component into the target application to complete the implantation of the tracking point. After capturing the exception information, the exception information captured by the probe component can be stored in the exception log file of the target application, and the exception log file can be stored in a preset database, such as Mysql. Preferably, the above exception information may include the time when the exception occurs and the type of exception, and also include at least one of the following data: exception interface, exception call method, exception description information, exception call stack information. Among them, the time when the exception occurs can be represented by a timestamp in seconds or minutes; the exception interface refers to the interface where the exception occurs, and the exception call method refers to the call method where the exception occurs. Generally, the exception interface can be indicated by the name of the interface where the exception occurs, and the exception call method can be indicated by the class name where the exception occurs and the corresponding method name.

[0033] In actual applications, each exception has an exception occurrence time and exception type. The exception occurrence time can help form the exception sequence to be described later, and the exception type can provide basic correlation analysis between application exceptions and application performance. In addition, each exception can also have an exception interface, an exception call method, exception description information or exception call stack information. The exception interface and exception call method can provide a more fine-grained correlation analysis based on the exception type correlation analysis, while the exception description information and exception call stack information can provide a basis for subsequent exception handling and problem solving.

[0034] Step S102: collecting operation data of the target application, and determining indicator data of at least one performance indicator of the target application according to the operation data.

[0035] In the embodiment of the present invention, the performance index may include: response time, availability or access volume. In this step, a third-party tool may be used to monitor the running state of the target application and obtain the above index data.

[0036] Step S103: Obtain the correlation degree between the abnormal information and the indicator data in the same statistical period, and determine the root cause information affecting the target application performance according to the correlation degree.

[0037] Preferably, the above correlation degree can be indicated by a correlation coefficient (such as Pearson correlation coefficient, cosine similarity, etc.). Specifically, the following steps can be performed to calculate the correlation coefficient. First, an abnormal sequence of a preset aggregation dimension is generated based on the abnormal information in the statistical period, wherein the abnormal sequence characterizes the changing trend of the number of abnormal occurrences of the corresponding aggregation dimension over time; then, an indicator sequence of the performance indicator is generated based on the indicator data of any performance indicator in the statistical period, and the indicator sequence characterizes the changing trend of the indicator data over time; thereafter, the correlation coefficient between the abnormal sequence and the indicator sequence is determined.

[0038] In practical applications, the above aggregation dimensions may include: exception type, exception interface or exception calling method. Accordingly, exception sequences may be formed in the exception type dimension, exception interface dimension or exception calling method dimension respectively. The exception sequence of any exception type represents the change of the number of exceptions of this type over time. The exception sequence of any exception interface represents the change of the number of exceptions when calling the interface over time. The exception sequence of any exception calling method represents the change of the number of exceptions when calling the method over time.

[0039] As a preferred solution, the root cause information affecting the performance of the target application can be determined by the following steps. For the indicator sequence of any performance indicator in the statistical period, the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal type is determined respectively, and the abnormal type corresponding to the maximum correlation coefficient is determined as the root cause abnormal type of the performance indicator. For the indicator sequence of any performance indicator in the statistical period, the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal interface is determined respectively, and the abnormal interface corresponding to the maximum correlation coefficient is determined as the root cause interface of the performance indicator. For the indicator sequence of any performance indicator in the statistical period, the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal calling method is determined respectively, and the abnormal calling method corresponding to the maximum correlation coefficient is determined as the root cause calling method of the performance indicator.

[0040] That is to say, for the indicator sequence of the response time indicator in the statistical period: the correlation coefficients between the indicator sequence and the exception sequences of each exception type can be determined respectively, and the exception type corresponding to the maximum correlation coefficient is determined as the root cause exception type of the response time indicator of the target application; the correlation coefficients between the indicator sequence and the exception sequences of each exception interface are determined respectively, and the exception interface corresponding to the maximum correlation coefficient is determined as the root cause interface of the response time indicator of the target application; the correlation coefficients between the indicator sequence and the exception sequences of each exception calling method are determined respectively, and the exception calling method corresponding to the maximum correlation coefficient is determined as the root cause calling method of the response time indicator of the target application.

[0041] For the indicator sequence of the availability indicator in the statistical period: the correlation coefficients between the indicator sequence and the exception sequences of each exception type can be determined respectively, and the exception type corresponding to the maximum correlation coefficient is determined as the root cause exception type of the availability indicator of the target application; the correlation coefficients between the indicator sequence and the exception sequences of each exception interface are determined respectively, and the exception interface corresponding to the maximum correlation coefficient is determined as the root cause interface of the availability indicator of the target application; the correlation coefficients between the indicator sequence and the exception sequences of each exception calling method are determined respectively, and the exception calling method corresponding to the maximum correlation coefficient is determined as the root cause calling method of the availability indicator of the target application.

[0042] For the indicator sequence of the access volume indicator in the statistical period: the correlation coefficients between the indicator sequence and the exception sequences of each exception type can be determined respectively, and the exception type corresponding to the maximum correlation coefficient is determined as the root cause exception type of the access volume indicator of the target application; the correlation coefficients between the indicator sequence and the exception sequences of each exception interface are determined respectively, and the exception interface corresponding to the maximum correlation coefficient is determined as the root cause interface of the access volume indicator of the target application; the correlation coefficients between the indicator sequence and the exception sequences of each exception calling method are determined respectively, and the exception calling method corresponding to the maximum correlation coefficient is determined as the root cause calling method of the access volume indicator of the target application.

[0043] In the technical solution of the embodiment of the present invention, on the one hand, when the target application is started, the preset probe component is imported into the target application, and the probe component is used to capture the abnormal information of the target application. On the other hand, the operation data of the target application is collected, and the indicator data of the target application in at least one performance indicator is determined based on the operation data. Thereafter, the degree of correlation between the abnormal information and the indicator data in the same statistical period is calculated, and the root cause information affecting the performance of the target application is determined based on the degree of correlation. In this way, accurate and real-time root cause location of the problem is achieved based on probe technology and correlation analysis, which helps to optimize the system architecture and reduce the recurrence of abnormalities.

[0044] A specific embodiment of the present invention is described below.

[0045] Locating the root cause of a problem is a complex task that requires the use of a variety of technologies and tools to achieve accurate and efficient results. The following are some commonly used root cause location techniques. Log analysis and monitoring tools: Use log analysis tools (such as ELK Stack, Splunk, Logstash, etc.) and monitoring tools (such as Prometheus, Grafana, etc.) to collect, analyze and visualize system logs and indicator data to discover potential problems and anomalies. Distributed tracing system: Use distributed tracing systems (such as Zipkin, Jaeger, etc.) to track and analyze the request and response links between different components in distributed applications to locate possible delays or errors. Performance analysis tools: Use performance analysis tools (such as Profiler, Flame Graphs, etc.) to identify application performance bottlenecks and hot spots, and help locate potential performance problems and optimization opportunities. Breakpoint debugging and debugger: Use debuggers (such as GDB, Xcode debugger, etc.) and breakpoint debugging techniques to locate problems in specific code segments during the development phase or in the production environment, and perform operations such as variable viewing and stack tracing. Fault injection and stress testing: Fault injection techniques (such as Chaos Engineering) and stress testing tools (such as JMeter, Gatling, etc.) can be used to simulate system failures and high-load scenarios to identify system weaknesses and stability issues. Code review and static analysis: Code review and static analysis tools (such as SonarQube, CodeClimate, etc.) can be used to identify potential code quality issues, security vulnerabilities, and incorrect usage patterns, thereby reducing the frequency of potential problems. Version control and rollback strategy: Use version control tools (such as Git, SVN, etc.) to track code changes and establish a rollback strategy so that you can quickly roll back to a stable version when problems occur. These technologies complement each other, and the appropriate combination of technologies should be selected according to the specific situation to locate the root cause of the problem. At the same time, continuous learning and practice are also the key to improving the ability to locate the root cause of the problem.

[0046] Although existing root cause location technologies provide developers and engineers with valuable tools and methods, they still have some shortcomings and challenges:

[0047] 1. Complexity: Modern applications are often complex distributed systems consisting of multiple components and microservices. This increases the complexity of problem location because it may involve interactions and collaborations between multiple components. Understanding and tracing the behavior and data flow of the entire system may take a lot of time and effort.

[0048] 2. Uncertainty: Root cause location usually involves tracking and analyzing a large amount of logs, metrics, and data. However, sometimes this information may not be clear enough, may be ambiguous, or lack sufficient context. This leads to a certain degree of uncertainty in root cause location, requiring developers to perform more analysis and reasoning.

[0049] 3. Performance overhead: Some root cause location techniques themselves have an impact on system performance. For example, collecting and analyzing large amounts of log and metric data may cause performance degradation, adding additional network traffic and storage requirements. This may make root cause location more difficult because developers need to balance performance and debugging requirements.

[0050] 4. Human resource requirements: Effective root cause location usually requires professionals with certain domain knowledge and skills. This can be a challenge for organizations because they need to have sufficient human resources to handle and solve complex problems. For small teams or organizations with limited resources, professional support may not be immediately available.

[0051] 5. Dependency and integration issues: Modern applications often rely on multiple third-party services and libraries. When problems arise with these dependencies, it becomes more difficult to locate the root cause. In addition to your own code and systems, you also need to consider external dependency issues and integration.

[0052] This embodiment aims to solve the technical problem of using probe technology to capture anomalies in real time, generate real-time data of anomalies, and perform correlation analysis with performance indicators such as availability, so as to quickly locate the root cause of anomalies. The traditional root cause location process of anomalies may have the following problems:

[0053] 1. Delayed anomaly discovery: Traditional methods usually require waiting for users to report anomalies or checking logs regularly, which leads to delayed discovery of anomalies and may cause more serious impacts.

[0054] 2. Difficult anomaly correlation analysis: In complex systems, anomalies may be related to specific business scenarios or reduced availability. However, it is difficult and time-consuming to manually analyze the correlation between abnormal data and business anomalies.

[0055] 3. Complexity of manually locating the root cause: Traditional methods usually require developers to locate the root cause based on limited information and experience, which can be very complex and require a lot of time and effort.

[0056] In response to the above problems, this embodiment provides a solution using probe technology, which captures anomalies in real time and generates real-time anomaly data, combined with correlation analysis of performance indicators such as availability, making it possible to quickly locate the root cause of the anomaly. This can speed up the anomaly troubleshooting, reduce the impact of system failures on the business, and improve the availability and stability of the system.

[0057] Probe technology can be used to capture anomalies in real time and generate real-time data on anomalies. By performing correlation analysis with business anomalies or availability, the root cause of anomalies can be quickly located. The main steps of this solution are as follows. Figure 2 .

[0058] 1. Install the Java probe into the application. Specifically, first, install the custom Java probe tool into the application. When the client program generates an exception, the exception information is captured and written into the log file. For example, in the Exception <init>Insert code at the place to capture all exception information.

[0059] 2. Capture exception information and save it to the exception log file: In the application, use the exception capture function provided by the probe tool to monitor and capture exceptions. Once an exception is found, the exception information (such as exception type, stack trace, etc.) can be recorded in the exception log file. This can be achieved by using a logging library, an exception handler, or a custom logging function.

[0060] 3. Store the exception log in the database: Parse the data in the exception log file and store it in the database for subsequent query and analysis. You can use a relational database or a dedicated log storage system to save the exception log data.

[0061] 4. Reporting business errors or availability data: In addition to capturing and storing exceptions, you need to use probe tools or third-party tools to monitor and report business performance indicators such as response time and availability. By defining appropriate indicators and collecting relevant data, performance indicators such as response time and availability can be sent to specified targets (such as monitoring systems, log collectors, or other data processing tools) on a regular or real-time basis.

[0062] 5. Perform correlation analysis between business performance indicators and captured anomalies: Use the stored business indicator data and captured anomaly log data to perform positive and negative correlation analysis. You can use data analysis tools or programming languages ​​(such as Python, R, etc.) to analyze the relationship between data, and analyze the relationship between business indicators and captured anomalies to determine the impact of anomalies on indicators such as availability.

[0063] 6. Determine the root cause of the problem based on correlation analysis: Based on the results of correlation analysis, the root cause of the anomaly can be inferred. By observing the associations between strongly correlated anomalies, the potential root cause of the problem can be determined.

[0064] Compared with conventional abnormal root cause location, this solution has the following advantages:

[0065] 1. Real-time: Probe technology can capture and generate real-time data of exceptions in real time, enabling developers to discover and locate exceptions more quickly and reduce the negative impact of failures on the system.

[0066] 2. Correlation analysis: By correlating the abnormal data with business abnormalities and availability data, the relationship between the abnormality and the business scenario or availability degradation can be more accurately determined, helping developers locate the root cause more quickly.

[0067] 3. Accurate positioning: Detailed context, stack trace and other information in the exception data can help developers locate the root cause of the exception more accurately and speed up problem solving.

[0068] 4. Data-driven decision making: Real-time generated exception data enables more in-depth data-driven decision making. The trend, correlation, and impact of exceptions can be analyzed to optimize system architecture and reduce the occurrence of repeated exceptions.

[0069] In summary, this embodiment combines probe technology, real-time data generation and correlation analysis, and rapid root cause location to provide an innovative technical solution that can locate the root cause of anomalies more quickly and accurately, thereby improving the availability and stability of the system.

[0070] It should be noted that the collection, collection, update, analysis, processing, use, transmission, storage and other aspects of user personal information that may be involved in the technical solution of the present invention are in compliance with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken for user personal information to prevent illegal access to user personal information data and maintain the security of user personal information, network security and national security.

[0071] For the above-mentioned method embodiments, for the convenience of description, they are expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the order of the actions described, and some steps can actually be performed in other orders or simultaneously. In addition, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily necessary to implement the present invention.

[0072] In order to better implement the above-mentioned solution of the embodiment of the present invention, relevant devices for implementing the above-mentioned solution are also provided below.

[0073] See also Figure 3 As shown, the device 300 for locating the root cause of an application problem provided by an embodiment of the present invention may include: an exception capturing unit 301 , a performance monitoring unit 302 , and a correlation calculating unit 303 .

[0074] Among them, the exception capture unit 301 is used to import the preset probe component into the target application when the target application is started, and use the probe component to capture the exception information of the target application; the performance monitoring unit 302 is used to collect the operation data of the target application, and determine the indicator data of the target application in at least one performance indicator according to the operation data; the correlation calculation unit 303 is used to obtain the correlation degree between the exception information and the indicator data in the same statistical period, and determine the root cause information affecting the performance of the target application according to the correlation degree.

[0075] In an embodiment of the present invention, the exception information includes the exception occurrence time and exception type, and also includes at least one of the following data: exception interface, exception calling method, exception description information, and exception call stack information.

[0076] In a specific application, the degree of association is indicated by a correlation coefficient; and the association calculation unit 303 can be further used to: generate an exception sequence of a preset aggregation dimension based on the exception information in the statistical period; wherein the exception sequence represents the changing trend of the number of exception occurrences in the aggregation dimension over time; generate an indicator sequence of the performance indicator based on the indicator data of any performance indicator in the statistical period, wherein the indicator sequence represents the changing trend of the indicator data over time; and determine the correlation coefficient between the exception sequence and the indicator sequence.

[0077] In practical applications, the aggregation dimensions include: exception type, exception interface or exception calling method; and the association calculation unit 303 can be further used to: for any performance indicator in the indicator sequence of the statistical period, respectively determine the correlation coefficient between the indicator sequence and the exception sequence of each exception type, and determine the exception type corresponding to the maximum correlation coefficient as the root cause exception type of the performance indicator; for any performance indicator in the indicator sequence of the statistical period, respectively determine the correlation coefficient between the indicator sequence and the exception sequence of each exception interface, and determine the exception interface corresponding to the maximum correlation coefficient as the root cause interface of the performance indicator; for any performance indicator in the indicator sequence of the statistical period, respectively determine the correlation coefficient between the indicator sequence and the exception sequence of each exception calling method, and determine the exception calling method corresponding to the maximum correlation coefficient as the root cause calling method of the performance indicator.

[0078] Preferably, the exception capturing unit 301 may be further configured to: store the exception information captured by the probe component in an exception log file of the target application, and store the exception log file in a preset database.

[0079] In addition, in an embodiment of the present invention, the performance indicators include: response time, availability or access volume. The association calculation unit 303 can be further used for: for the indicator sequence of the response time indicator in the statistical period: respectively determine the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal type, and determine the abnormal type corresponding to the maximum correlation coefficient as the root cause abnormal type of the response time indicator of the target application; respectively determine the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal interface, and determine the abnormal interface corresponding to the maximum correlation coefficient as the root cause interface of the response time indicator of the target application; respectively determine the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal call method, and determine the abnormal call method corresponding to the maximum correlation coefficient as the root cause call method of the response time indicator of the target application; for the indicator sequence of the availability indicator in the statistical period: respectively determine the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal type, and determine the abnormal type corresponding to the maximum correlation coefficient as the root cause abnormal type of the availability indicator of the target application; respectively determine the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal interface, and determine the abnormal interface corresponding to the maximum correlation coefficient as the root cause abnormal type of the response time indicator of the target application. The correlation coefficients between the abnormal sequences of interfaces, and the abnormal interface corresponding to the maximum correlation coefficient is determined as the root cause interface of the availability index of the target application; the correlation coefficients between the indicator sequence and the abnormal sequence of each abnormal calling method are determined respectively, and the abnormal calling method corresponding to the maximum correlation coefficient is determined as the root cause calling method of the availability index of the target application; for the indicator sequence of the access volume index in the statistical period: the correlation coefficients between the indicator sequence and the abnormal sequence of each abnormal type are determined respectively, and the abnormal type corresponding to the maximum correlation coefficient is determined as the root cause abnormal type of the access volume index of the target application; the correlation coefficients between the indicator sequence and the abnormal sequence of each abnormal interface are determined respectively, and the abnormal interface corresponding to the maximum correlation coefficient is determined as the root cause interface of the access volume index of the target application; the correlation coefficients between the indicator sequence and the abnormal sequence of each abnormal calling method are determined respectively, and the abnormal calling method corresponding to the maximum correlation coefficient is determined as the root cause calling method of the access volume index of the target application.

[0080] According to the technical solution of the embodiment of the present invention, on the one hand, when the target application is started, the preset probe component is imported into the target application, and the probe component is used to capture the abnormal information of the target application. On the other hand, the operation data of the target application is collected, and the indicator data of the target application in at least one performance indicator is determined based on the operation data. Thereafter, the degree of correlation between the abnormal information and the indicator data in the same statistical period is calculated, and the root cause information affecting the performance of the target application is determined based on the degree of correlation. In this way, accurate and real-time root cause location of the problem is achieved based on probe technology and correlation analysis, which helps to optimize the system architecture and reduce the recurrence of abnormalities.

[0081] Figure 4 An exemplary system architecture 400 is shown to which the method for locating the root cause of an application problem or the apparatus for locating the root cause of an application problem according to an embodiment of the present invention can be applied.

[0082] like Figure 4 As shown, the system architecture 400 may include terminal devices 401, 402, 403, a network 404 and a server 405 (this architecture is only an example, and the components included in the specific architecture may be adjusted according to the specific application). The network 404 is used to provide a medium for a communication link between the terminal devices 401, 402, 403 and the server 405. The network 404 may include various connection types, such as wired, wireless communication links or optical fiber cables.

[0083] The user can use the terminal devices 401, 402, 403 to interact with the server 405 through the network 404 to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 401, 402, 403, such as a root cause positioning application (only as an example).

[0084] The terminal devices 401 , 402 , and 403 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0085] The server 405 may be a server that provides various services, such as a background server (only an example) that provides support for the root cause location application operated by the user using the terminal devices 401, 402, 403. The background server may process the received root cause location request, etc., and feed back the processing result (such as the root cause of the problem - only an example) to the terminal devices 401, 402, 403.

[0086] It should be noted that the method for locating the root cause of an application problem provided in the embodiment of the present invention is generally executed by the server 405 , and accordingly, the device for locating the root cause of an application problem is generally disposed in the server 405 .

[0087] It should be understood that Figure 4 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0088] The present invention also provides an electronic device. The electronic device of an embodiment of the present invention includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method for locating the root cause of the application problem provided by the present invention.

[0089] Reference below Figure 5 , which shows a schematic diagram of the structure of a computer system 500 of an electronic device suitable for implementing an embodiment of the present invention. Figure 5 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0090] like Figure 5 As shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage part 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the computer system 500 are also stored. The CPU 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0091] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that a computer program read therefrom is installed into the storage section 508 as needed.

[0092] In particular, according to the embodiments disclosed in the present invention, the process described in the main step diagram above can be implemented as a computer software program. For example, the embodiments of the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the main step diagram. In the above embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit 501, the above functions defined in the system of the present invention are executed.

[0093] It should be noted that the computer-readable medium shown in the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0094] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0095] The units involved in the embodiments of the present invention may be implemented by software or hardware. The units described may also be arranged in a processor, for example, it may be described as: a processor includes an exception capture unit, a performance monitoring unit, and an associated calculation unit. The names of these units do not constitute a limitation on the units themselves under certain circumstances, for example, the exception capture unit may also be described as a "unit that provides exception information to the associated calculation unit".

[0096] As another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiment; or it may exist independently and not be assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the device, the steps executed by the device include: when the target application is started, importing the preset probe component into the target application, and using the probe component to capture the abnormal information of the target application; collecting the operation data of the target application, and determining the index data of the target application in at least one performance indicator according to the operation data; obtaining the correlation degree between the abnormal information and the index data in the same statistical period, and determining the root cause information affecting the performance of the target application according to the correlation degree.

[0097] In the technical solution of the embodiment of the present invention, on the one hand, when the target application is started, the preset probe component is imported into the target application, and the probe component is used to capture the abnormal information of the target application. On the other hand, the operation data of the target application is collected, and the indicator data of the target application in at least one performance indicator is determined based on the operation data. Thereafter, the degree of correlation between the abnormal information and the indicator data in the same statistical period is calculated, and the root cause information affecting the performance of the target application is determined based on the degree of correlation. In this way, accurate and real-time root cause location of the problem is achieved based on probe technology and correlation analysis, which helps to optimize the system architecture and reduce the recurrence of abnormalities.

[0098] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.< / init>

Claims

1. A method for locating the root cause of an application problem, characterized in that: include: When the target application is started, the preset probe component is imported into the target application, and the abnormal information of the target application is captured by using the probe component; Collecting operation data of the target application, and determining indicator data of at least one performance indicator of the target application according to the operation data; The degree of correlation between the abnormal information and the indicator data in the same statistical period is obtained, and the root cause information affecting the performance of the target application is determined according to the degree of correlation.

2. The method according to claim 1, characterized in that The exception information includes the exception occurrence time and exception type, and also includes at least one of the following data: exception interface, exception calling method, exception description information, and exception calling stack information.

3. The method according to claim 2, characterized in that The degree of said association is indicated by the correlation coefficient; And, obtaining the correlation between the abnormal information and the indicator data in the same statistical period includes: Generate an abnormal sequence of a preset aggregation dimension according to the abnormal information in the statistical period; wherein the abnormal sequence represents the changing trend of the number of abnormal occurrences of the aggregation dimension over time; Generate an indicator sequence of the performance indicator according to the indicator data of any performance indicator in the statistical period, wherein the indicator sequence represents the change trend of the indicator data over time; A correlation coefficient between the abnormal sequence and the indicator sequence is determined.

4. The method according to claim 3, characterized in that The aggregation dimension includes: exception type, exception interface or exception calling method; and the root cause information affecting the performance of the target application is determined according to the association degree, including: For the indicator sequence of any performance indicator in the statistical period, respectively determine the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal type, and determine the abnormal type corresponding to the maximum correlation coefficient as the root cause abnormal type of the performance indicator; For the indicator sequence of any performance indicator in the statistical period, respectively determine the correlation coefficient between the indicator sequence and the abnormal sequence of each abnormal interface, and determine the abnormal interface corresponding to the maximum correlation coefficient as the root cause interface of the performance indicator; For the indicator sequence of any performance indicator in the statistical period, the correlation coefficients between the indicator sequence and the abnormal sequences of each abnormal calling method are determined respectively, and the abnormal calling method corresponding to the maximum correlation coefficient is determined as the root cause calling method of the performance indicator.

5. The method according to claim 1, characterized in that The method further comprises: The exception information captured by the probe component is stored in an exception log file of the target application, and the exception log file is stored in a preset database.

6. The method according to claim 4, characterized in that The performance indicators include: response time, availability or access volume; and the root cause information affecting the performance of the target application determined according to the correlation degree includes: For the indicator sequence of the response time indicator in the statistical period: respectively determine the correlation coefficients between the indicator sequence and the abnormal sequence of each abnormal type, and determine the abnormal type corresponding to the maximum correlation coefficient as the root cause abnormal type of the response time indicator of the target application; respectively determine the correlation coefficients between the indicator sequence and the abnormal sequence of each abnormal interface, and determine the abnormal interface corresponding to the maximum correlation coefficient as the root cause interface of the response time indicator of the target application; respectively determine the correlation coefficients between the indicator sequence and the abnormal sequence of each abnormal calling method, and determine the abnormal calling method corresponding to the maximum correlation coefficient as the root cause calling method of the response time indicator of the target application; For the indicator sequence of the availability index in the statistical period: respectively determine the correlation coefficients between the indicator sequence and the abnormal sequence of each abnormal type, and determine the abnormal type corresponding to the maximum correlation coefficient as the root cause abnormal type of the availability index of the target application; respectively determine the correlation coefficients between the indicator sequence and the abnormal sequence of each abnormal interface, and determine the abnormal interface corresponding to the maximum correlation coefficient as the root cause interface of the availability index of the target application; respectively determine the correlation coefficients between the indicator sequence and the abnormal sequence of each abnormal calling method, and determine the abnormal calling method corresponding to the maximum correlation coefficient as the root cause calling method of the availability index of the target application; For the indicator sequence of the access volume indicator in the statistical period: respectively determine the correlation coefficients between the indicator sequence and the abnormal sequences of each abnormal type, and determine the abnormal type corresponding to the maximum correlation coefficient as the root cause abnormal type of the access volume indicator of the target application; respectively determine the correlation coefficients between the indicator sequence and the abnormal sequences of each abnormal interface, and determine the abnormal interface corresponding to the maximum correlation coefficient as the root cause interface of the access volume indicator of the target application; respectively determine the correlation coefficients between the indicator sequence and the abnormal sequences of each abnormal calling method, and determine the abnormal calling method corresponding to the maximum correlation coefficient as the root cause calling method of the access volume indicator of the target application.

7. A device for locating the root cause of an application problem, characterized in that: include: An exception capture unit is used to import a preset probe component into the target application when the target application is started, and use the probe component to capture the exception information of the target application; A performance monitoring unit, configured to collect operating data of the target application and determine indicator data of at least one performance indicator of the target application according to the operating data; The correlation calculation unit is used to obtain the correlation degree between the abnormal information and the indicator data in the same statistical period, and determine the root cause information affecting the performance of the target application according to the correlation degree.

8. The device according to claim 7, characterized in that The exception information includes the exception occurrence time and exception type, and also includes at least one of the following data: exception interface, exception calling method, exception description information, and exception calling stack information.

9. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.