A service fault diagnosis method and system

CN116107780BActive Publication Date: 2026-10-09YGSOFT INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111320809.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-09
Publication Date
2026-10-09
Estimated Expiration
2041-11-09

AI Technical Summary

Technical Problem

[0002]当被执行服务故障诊断的设备发生故障时,故障难以定位,无法实时了解设备的实时jvm信息

Benefits of technology

[0038] The service fault diagnosis method provided in this embodiment of the invention receives a trace command sent by a service fault diagnosis server. The trace command carries a class name and the name of the method corresponding to the class name to be diagnosed. An interception method is added to the interface method. This interception method is used to obtain method-related indicator information from the method and its sub-methods. When it is detected that the method and its sub-methods corresponding to the method name are about to be called or about to finish execution, the interface method with the added interception method is automatically called. The interception method then obtains the method-related indicator information from the method and its sub-methods and sends this information to the service fault diagnosis server. The service fault diagnosis server then performs service fault diagnosis based on the method-related indicator information to obtain the service fault diagnosis result. This method facilitates service fault diagnosis and has the beneficial effects of minimal performance impact and strong real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116107780B_ABST
    Figure CN116107780B_ABST
Patent Text Reader

Abstract

The application discloses a service fault diagnosis method and system, belongs to the technical field of Java virtual machines, facilitates service fault diagnosis, and has beneficial effects of small program performance loss and strong real-time performance. The method comprises the following steps: receiving a trace command sent by a service fault diagnosis server; the trace command carries a class name and a method name corresponding to the class name to be subjected to service fault diagnosis; an interception method is added in an interface method; the interception method is used to acquire method-related index information from a method and a sub-method; when it is found that the method corresponding to the method name and the sub-method will be called or execution ends, the interface method after the interception method is automatically called, the method-related index information is acquired from the method and the sub-method through the interception method, and the method-related index information is sent to the service fault diagnosis server, so that the service fault diagnosis server performs service fault diagnosis according to the method-related index information, and service fault diagnosis results are obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Java Virtual Machine technology, and in particular to a service fault diagnosis method and system. Background Technology

[0002] When a device undergoing service fault diagnosis malfunctions, the fault is difficult to locate, and real-time JVM information about the device is unavailable. Obtaining information such as method return values, input parameters, and execution time is impossible without restarting the service application, and the cause of the service error remains unknown. Existing methods typically involve adding logs, repackaging, deploying, and then reproducing the scenario. However, such methods often lead to a series of unknown and serious problems. Summary of the Invention

[0003] Based on the above analysis, the embodiments of the present invention aim to provide a service fault diagnosis method and system, which facilitates service fault diagnosis and has the beneficial effects of low program performance loss and strong real-time performance.

[0004] This invention discloses a service fault diagnosis method, comprising:

[0005] Receive a trace command sent by the service fault diagnosis server; the trace command carries the class name and the method name corresponding to the class name to be diagnosed in the service fault diagnosis.

[0006] Add an interception method to the interface method; the interception method is used to obtain method-related indicator information from the method and its sub-methods;

[0007] When the interception method is detected as being about to be called or when the method and its sub-methods corresponding to the method name are about to finish execution, the interface method with the added interception method is automatically called. The interception method is used to obtain the relevant indicator information of the method and its sub-methods, and the relevant indicator information is sent to the service fault diagnosis server so that the service fault diagnosis server can perform service fault diagnosis based on the relevant indicator information and obtain the service fault diagnosis result.

[0008] Furthermore, the addition of an interception method to the interface method includes:

[0009] Add a first interception method to the interface method corresponding to the execution of the method and its sub-methods, for obtaining method-related indicator information before the execution of the method and its sub-methods;

[0010] A second interception method is added to the interface method corresponding to the execution of the method and its sub-methods, for obtaining method-related indicator information after the execution of the method and its sub-methods.

[0011] Furthermore, the step of automatically invoking the interface method after adding the interception method when it is detected that the method and sub-methods corresponding to the method name are about to be called or have finished executing, obtaining the method-related indicator information from the method and sub-methods through the interception method, and sending the method-related indicator information to the service fault diagnosis server includes:

[0012] When it is detected that the method and its sub-methods are about to be called, the corresponding interface method before the execution of the method and its sub-methods is automatically called.

[0013] When the execution of the method and its sub-methods is detected to have ended, the corresponding interface method after the execution of the method and its sub-methods is automatically invoked.

[0014] In the interface method corresponding to the execution of the method and its sub-methods, the listener object is called through the methodEnd method to output the relevant indicator information of the method; the methodEnd method is used to output all the relevant indicator information of the method when the execution of the method and its sub-methods ends, and send all the relevant indicator information of the method to the service fault diagnosis server.

[0015] Furthermore, the service fault diagnosis method also includes:

[0016] Receive device detection commands sent by the service fault diagnosis server;

[0017] The local device is detected according to the device detection command to obtain device-related indicator information for service fault diagnosis;

[0018] The device-related indicator information and the method-related indicator information are sent to the service fault diagnosis server so that the service fault diagnosis server can perform service fault diagnosis based on the method-related indicator information and the device-related indicator information to obtain the service fault diagnosis result.

[0019] This invention discloses a service fault diagnosis method, comprising:

[0020] A trace command is sent to the device performing the service fault diagnosis; the trace command carries the class name and the method name corresponding to the class name for which the service fault diagnosis is to be performed;

[0021] The system receives method-related indicator information returned by the device based on the method name corresponding to the class name, and performs service fault diagnosis based on the method-related indicator information to obtain service fault diagnosis results.

[0022] Furthermore, the method-related indicator information includes method input parameters, return values, and method execution time; correspondingly, the step of performing service fault diagnosis based on the method-related indicator information to obtain service fault diagnosis results includes:

[0023] The class name and method name in the input parameters of the method are normalized to obtain the method name of the class;

[0024] The execution time of the method corresponding to the method name of the class is determined, and the service fault diagnosis result is obtained based on the comparison result of the method execution time with the preset method execution time threshold and the return value.

[0025] Further, obtaining the service fault diagnosis result based on the comparison result of the method execution time and the preset method time threshold, and the return value, includes:

[0026] If the execution time of the method is less than the preset method time threshold, and the return value is normal, then the service fault diagnosis result is determined to be normal.

[0027] If the execution time of the method is greater than or equal to the preset method time threshold, and / or the return value is an abnormal value, then the service fault diagnosis result is determined to be abnormal.

[0028] Furthermore, the service fault diagnosis method also includes:

[0029] Receive device-related indicator information and method-related indicator information sent by the device; and perform service fault diagnosis based on the method-related indicator information and the device-related indicator information to obtain service fault diagnosis results.

[0030] Further, the device-related indicator information includes memory-related indicator information, thread-related indicator information, and stack-related indicator information; correspondingly, the step of performing service fault diagnosis based on the method-related indicator information and the device-related indicator information to obtain service fault diagnosis results includes:

[0031] The service fault diagnosis result is determined based on the comparison between the memory-related indicator information and the preset memory-related indicator thresholds;

[0032] The service fault diagnosis result is determined based on the comparison between the thread-related indicator information and the preset thread-related indicator threshold.

[0033] The service fault diagnosis result is determined based on the comparison between the stack-related indicator information and the preset stack-related indicator thresholds.

[0034] The service fault diagnosis result is determined based on the comparison result between the execution time of the method and the preset method time threshold, and the return value.

[0035] This invention discloses a service fault diagnosis system, comprising: a device for performing service fault diagnosis and a service fault diagnosis server;

[0036] The device and the service fault diagnosis server respectively execute the corresponding service fault diagnosis method.

[0037] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:

[0038] The service fault diagnosis method provided in this embodiment of the invention receives a trace command sent by a service fault diagnosis server. The trace command carries a class name and the name of the method corresponding to the class name to be diagnosed. An interception method is added to the interface method. This interception method is used to obtain method-related indicator information from the method and its sub-methods. When it is detected that the method and its sub-methods corresponding to the method name are about to be called or about to finish execution, the interface method with the added interception method is automatically called. The interception method then obtains the method-related indicator information from the method and its sub-methods and sends this information to the service fault diagnosis server. The service fault diagnosis server then performs service fault diagnosis based on the method-related indicator information to obtain the service fault diagnosis result. This method facilitates service fault diagnosis and has the beneficial effects of minimal performance impact and strong real-time performance.

[0039] The service fault diagnosis system provided in this embodiment of the invention includes a device performing service fault diagnosis and a service fault diagnosis server. The method executed by the device includes: receiving a trace command sent by the service fault diagnosis server; the trace command carries a class name and a method name corresponding to the class name to be diagnosed; adding an interception method to the interface method; the interception method is used to obtain method-related indicator information from the method and its sub-methods; when it is detected that the method and its sub-methods corresponding to the method name are about to be called or about to finish execution, the interface method with the added interception method is automatically called, the method-related indicator information is obtained from the method and its sub-methods through the interception method, and the method-related indicator information is sent to the service fault diagnosis server so that the service fault diagnosis server can perform service fault diagnosis based on the method-related indicator information to obtain a service fault diagnosis result. This system facilitates service fault diagnosis and has the beneficial effects of low performance loss and strong real-time performance.

[0040] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description

[0041] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.

[0042] Figure 1 This is a flowchart of the service fault diagnosis method in an embodiment of the present invention;

[0043] Figure 2 This is a flowchart of the service fault diagnosis method in an embodiment of the present invention;

[0044] Figure 3 This is a schematic diagram of the service fault diagnosis system in an embodiment of the present invention. Detailed Implementation

[0045] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.

[0046] A specific embodiment of the present invention discloses a service fault diagnosis method, the flowchart of which is shown below. Figure 1 As shown, it includes the following steps:

[0047] Step S1: Receive a trace command sent by the service fault diagnosis server; the trace command carries the class name and the method name corresponding to the class name to be diagnosed in the service fault diagnosis.

[0048] Step S2: Add an interception method to the interface method; the interception method is used to obtain method-related indicator information from the method and its sub-methods.

[0049] Step S3: When it is detected that the method and sub-methods corresponding to the method name are about to be called or have finished executing, the interface method after adding the interception method is automatically called. The interception method is used to obtain the method-related indicator information from the method and sub-methods, and the method-related indicator information is sent to the service fault diagnosis server so that the service fault diagnosis server can perform service fault diagnosis based on the method-related indicator information and obtain the service fault diagnosis result.

[0050] In step S1 above, the device undergoing service fault diagnosis receives a trace command sent by the service fault diagnosis server. The trace command carries the class name and the method name corresponding to that class name, for which service fault diagnosis is to be performed. The device undergoing service fault diagnosis is hereinafter referred to as "the device". Probes can be pre-deployed in the device. These probes can locate methods within the Java process using the complete class name and method name, thereby enabling dynamic code enhancement when the method is invoked.

[0051] Dynamic code enhancement processing refers to processing that does not actually modify the code used to run the program. Instead, it collects some method-related metrics before and after method execution for service fault diagnosis. These metrics include class name, method name, method description, and execution command. Ultimately, it obtains the method input parameters, method execution time, and return value, and uses these three types of information to determine whether the method execution is abnormal.

[0052] In step S2 above, the device performing service fault diagnosis adds an interception method to the interface method; the interception method is used to obtain method-related indicator information from the method and sub-methods. Further, dynamic code enhancement processing can be performed on the method and sub-methods using ASM to enable the device to add an interception method to the interface method; the interception method is used to obtain method-related indicator information from the method and sub-methods.

[0053] ASM is a general-purpose Java bytecode manipulation and analysis framework. It can be used to modify existing classes or dynamically generate classes directly in binary form, and provides some common bytecode transformation and analysis algorithms from which you can build custom, complex transformation and code analysis tools.

[0054] This approach is based on the ASM framework and extends its methods such as onMethodEnter, onMethodExit, visitMaxs, and visitMethodInsn. These overloaded methods are dynamically modified during the actual runtime of the method being detected and its sub-methods, allowing for customization of the method-related metrics information to be obtained before and after method execution.

[0055] The following is an example illustrating dynamic code enhancement for a method and its sub-methods:

[0056] void function1(int parameter){

[0057] subFunction(parameter);

[0058] }

[0059] After dynamic code enhancement, method-related metrics such as method name, parameter values, class name, client ID, and start execution time are recorded before entering the function1 method. Similarly, corresponding method-related metrics are recorded before entering the subFunction method of function1. After the subFunction and function1 methods are executed, method-related metrics such as method input parameters, return value, and execution time are recorded. At the same time, the statistics of method-related metrics are completed and sent to the service fault diagnosis server.

[0060] Furthermore, the addition of an interception method to the interface method includes:

[0061] Add a first interception method to the interface method corresponding to the execution of the method and sub-methods, which is used to obtain method-related indicator information before the execution of the method and sub-methods; referring to the above example, the principle of the method and sub-method is the same, and only the method is used as an example here, for example, the class name is A1 and the method name is abc.

[0062] Add a first interceptor method "ON_BEFORE_METHOD" to the interface method "onMethodEnter" that corresponds to the execution of the aforementioned method and its sub-methods. This interceptor method "abc" will output the relevant metric information for the method named "abc" through the "methodStart" method. This method-related metric information is the metric information for the method before its execution. The relevant program code is as follows, and details can be found in the text comments following the code:

[0063]

[0064]

[0065] ASM adds an "ON_BEFORE_METHOD" method to the original interface method, and passes M parameters to this method. These parameters can include the class name to be intercepted, method name, method description, method parameters, ClassLoader, and execution command client ID. The "ON_BEFORE_METHOD" method ultimately executes the "methodStart" method. The relevant code is shown below, and the relevant content can be found in the text comments after the code:

[0066]

[0067]

[0068] The "methodStart" method above will save relevant information about the first intercepted method "ON_BEFORE_METHOD", such as class name, method name, method description, traceId, method ID, parent method ID, and client ID.

[0069] A second interception method is added to the interface method corresponding to the execution of the method and sub-methods to obtain method-related indicator information after the execution of the method and sub-methods; referring to the above example, the principle of the method and sub-methods is the same, and only the method is used as an example here.

[0070] Add a second interceptor method "ON_RETURN_METHOD" to the interface method "onMethodExit" to intercept "abc". The "methodEnd" method will output the relevant metrics information for the method named "abc". This method-related metric information is the metric information after the method execution. The relevant code is as follows; see the text comments after the code for details:

[0071]

[0072]

[0073] ASM adds an "ON_RETURN_METHOD" method to the original interface method, and passes two parameters to this method: the method execution time and the return value (the return value can be a normal value or an exception value). The "ON_RETURN_METHOD" method ultimately executes the "methodEnd" method, as follows:

[0074] private static void methodEnd(Object result, boolean isThrow) / / The method after execution

[0075]

[0076]

[0077] In this way, the "methodEnd method" above will save relevant information about the second interception method "ON_RETURN_METHOD": method execution time and return value. Together with the information saved in the "methodStart method", they constitute the method-related indicator information that needs to be collected.

[0078] In step S2 above, the device performing service fault diagnosis adds an interception method to the interface method; the interception method is used to obtain method-related indicator information from the method and sub-methods.

[0079] In step S3 above, when it is detected that the method and sub-methods corresponding to the method name are about to be called or about to finish execution, the interface method after adding the interception method is automatically called. The interception method is used to obtain the method-related indicator information from the method and the sub-methods, and the method-related indicator information is sent to the service fault diagnosis server so that the service fault diagnosis server can perform service fault diagnosis based on the method-related indicator information and obtain the service fault diagnosis result.

[0080] Furthermore, the step of automatically invoking the interface method after adding the interception method when it is detected that the method and sub-methods corresponding to the method name are about to be called or have finished executing, obtaining the method-related indicator information from the method and sub-methods through the interception method, and sending the method-related indicator information to the service fault diagnosis server includes:

[0081] When it is detected that the method and its sub-methods are about to be called, the corresponding interface method before the execution of the method and its sub-methods is automatically called.

[0082] When the execution of the method and its sub-methods is detected to have ended, the corresponding interface method after the execution of the method and its sub-methods is automatically invoked.

[0083] Specifically, taking a method as an example, when ASM detects that a method is about to be called, it automatically calls the "onMethodEnter" interface method. At this time, the monitored method will be paused and the interface method will be executed instead, thereby executing the first interception method "ON_BEFORE_METHOD". The relevant indicator information before the method is executed is obtained through the first interception method.

[0084] Once the first interception method finishes executing, the suspended monitored method will continue to execute.

[0085] When the monitored method finishes execution, ASM detects the end of the monitored method's execution and automatically calls the "onMethodExit" interface method, thereby executing the second intercept method "ON_RETURN_METHOD" to obtain relevant indicator information after the method execution.

[0086] Furthermore, in the interface method corresponding to the execution of the method and its sub-methods, the listener object is called through the methodEnd method to output the method-related indicator information; the methodEnd method is used to output all method-related indicator information when the execution of the method and its sub-methods ends, and to send all method-related indicator information to the service fault diagnosis server.

[0087] Specifically, a listener object DListener listener is defined in the methodEnd method. The listener listener.afterMethodReturn method is called in the methodEnd method. This method will eventually execute the methodEnd method as follows. This method outputs the method-related indicator information collected earlier. In this way, the process of outputting all method-related indicator information will be completed when the methodEnd method finishes execution.

[0088] The code is as follows:

[0089]

[0090]

[0091]

[0092] In the code above, the variable costTime stores the execution time of the method (in milliseconds), the variable result stores the return value of the method, where "RESULT" indicates a normal return value and "ERROR" indicates an abnormal return value, the variable args stores the parameters passed to the method, and the variable sb integrates the above three variable names and their corresponding values, and adds the class name and method name as the final output.

[0093] Then, send the relevant indicator information of the method to the service fault diagnosis server.

[0094] Dynamic code enhancement can be stopped, reverted, or canceled by pressing Ctrl+C. This is extremely helpful for locating problems, and the pluggable dynamic code enhancement feature has no performance penalty whatsoever; enhancements are applied immediately upon use.

[0095] Furthermore, the service fault diagnosis method also includes:

[0096] The device receives a device probe command sent by the service fault diagnosis server. The device probe command can be sent to the device before, after, or simultaneously with the trace command. Probes can be pre-deployed in the device, and the probes perform corresponding probe actions based on the device probe command.

[0097] The device probes the local device according to the device probe command to obtain device-related indicator information for service fault diagnosis. The device-related indicator information obtained from the local device probe may include memory-related indicator information, thread-related indicator information, and stack-related indicator information. Among them, the memory-related indicator information may include the Heap area, Metaspace area, and PS old Gen area of ​​memory. The device probe command may include `memory 20 3000`, indicating that JVM memory information is output every three seconds, for a total of 20 times, to obtain data from various memory regions of the host program's JVM, including the Heap area, Metaspace area, and PS old Gen area.

[0098] Memory-related metrics can also include gc, which is short for garbage collection. Device detection commands can include gc 20 3000, which means outputting GC information every 3 seconds for 20 times to obtain the GC data of the host program's JVM.

[0099] Memory-related metrics can also include jmap, which is used to obtain memory information via MemoryMXBean after interaction with the probe.

[0100] By collecting jmap information from the host program using probes, such as jmap-dump, and dumping the program's memory snapshot to the service fault diagnosis server, these operations may cause the program to pause briefly. However, they are of great help in obtaining memory information and then optimizing the program.

[0101] It may also include:

[0102] jmap-heap: Used to generate a snapshot of the memory details file and upload it to the service troubleshooting server.

[0103] jmap-histo: Used to generate a snapshot of statistics files of objects in the heap and upload it to the service fault diagnosis server.

[0104] jmap –permstat: This is used to snapshot the statistics file of the class loader that generates the heap permanent generation to the service fault diagnosis server.

[0105] `jmap --finalizerinfo`: This command is used to capture a snapshot of the object information file generated in the F-Queue waiting for the Finalizer thread to execute the `finalize` method, and upload it to the service fault diagnosis server. The corresponding device detection commands are not detailed here.

[0106] Thread-related metrics include: thread, which obtains real-time Java thread information of the host program through probes. By viewing the status of each thread and the CPU resources it occupies, thread information can be obtained, including the total number of threads, the total number of active threads, and the target thread that occupies more than a preset percentage of CPU resources.

[0107] Stack-related metrics include: detailed information about each thread obtained through the stack, specifically the number of blocked threads.

[0108] The device sends its relevant indicator information and the method-related indicator information to the service fault diagnosis server, so that the service fault diagnosis server can perform service fault diagnosis based on the method-related indicator information and the device-related indicator information to obtain a service fault diagnosis result. The device can either package the device-related indicator information and the method-related indicator information and send them to the service fault diagnosis server synchronously, or it can sort the device-related indicator information and the method-related indicator information in order and send them to the service fault diagnosis server asynchronously.

[0109] A specific embodiment of the present invention discloses a service fault diagnosis method, the flowchart of which is shown below. Figure 2 As shown, it includes the following steps:

[0110] Step R1: Send a trace command to the device performing the service fault diagnosis; the trace command carries the class name and the method name corresponding to the class name for which the service fault diagnosis is to be performed.

[0111] Step R2: Receive method-related indicator information returned by the device based on the method name corresponding to the class name, and perform service fault diagnosis based on the method-related indicator information to obtain service fault diagnosis results.

[0112] In step R1 above, the service fault diagnosis server sends a trace command to the device to which service fault diagnosis is performed; the trace command carries the class name and the method name corresponding to the class name to be diagnosed.

[0113] In step R2 above, the service fault diagnosis server receives method-related indicator information returned by the device based on the method name corresponding to the class name, and performs service fault diagnosis based on the method-related indicator information to obtain the service fault diagnosis result.

[0114] Furthermore, the method-related indicator information includes method input parameters, return value, and method execution time; the method input parameters correspond to the above-mentioned class name and the method name corresponding to the class name; the return value includes whether the method execution is normal or abnormal; the method execution time is the time interval from the start to the end of the method execution.

[0115] Accordingly, the service fault diagnosis is performed based on the relevant indicator information of the method to obtain the service fault diagnosis result:

[0116] The class name and method name in the input parameters of the method are normalized to obtain the method name of the class. The method name of the class can be represented in the form of class name + method name. For example, if the class name is A1 and the method name is abc, the method name of the class can be A1 + abc. The same class can correspond to multiple method names. Therefore, the method name of the class can conveniently determine the method-related indicator information corresponding to different methods, as well as the class to which different methods belong.

[0117] The execution time of the method corresponding to the method name of the class is determined, and the service fault diagnosis result is obtained based on the comparison result of the method execution time with a preset method execution time threshold and the return value. The preset method execution time threshold can be set independently according to the actual situation, for example, 3 seconds.

[0118] Further, obtaining the service fault diagnosis result based on the comparison result of the method execution time and the preset method time threshold, and the return value, includes:

[0119] If the execution time of the method is less than the preset method time threshold and the return value is normal, then the service fault diagnosis result is determined to be normal; that is, when both the method execution time and the return value meet the conditions, it indicates that the service fault diagnosis result is normal.

[0120] If the execution time of the method is greater than or equal to the preset method execution time threshold, and / or the return value is an abnormal value, then the service fault diagnosis result is determined to be abnormal. That is, if at least one of the method execution time and return value does not meet the conditions, the service fault diagnosis result is considered abnormal.

[0121] Furthermore, the service fault diagnosis method also includes:

[0122] Receive device-related indicator information and method-related indicator information sent by the device; and perform service fault diagnosis based on the method-related indicator information and the device-related indicator information to obtain service fault diagnosis results.

[0123] Further, the device-related indicator information includes memory-related indicator information, thread-related indicator information, and stack-related indicator information; correspondingly, the step of performing service fault diagnosis based on the method-related indicator information and the device-related indicator information to obtain service fault diagnosis results includes:

[0124] The service fault diagnosis result is determined based on the comparison between the memory-related indicator information and the preset memory-related indicator threshold. Referring to the above example, the memory-related indicator information may include memory, gc, and jmap. For memory, the memory-related indicator information corresponding to memory includes the memory usage rate corresponding to the Heap area, Metaspace area, and PS old Gen area, respectively. The corresponding preset memory-related indicator threshold can be set independently according to the actual situation and can be selected as 90%.

[0125] Taking the Heap area as an example, obtain the memory usage rate of the Heap area over N days, determine the average memory usage rate of the Heap area over N days, and compare the real-time memory usage rate of the Heap area with the average value. If the real-time memory usage rate of the Heap area is greater than the average value and the duration is longer than the preset duration, then the memory of the Heap area is determined to be abnormal. The specific values ​​of N and the preset duration can be set according to the actual situation. N can be selected as 30, and the preset duration can be selected as 5 minutes.

[0126] The Metaspace and PS old Gen areas can be referenced from the Heap area description, and will not be repeated here. If at least one memory area is abnormal, the service fault diagnosis result is abnormal.

[0127] For GC, obtain the real-time average duration corresponding to GC. The corresponding preset time threshold can be selected as 500ms. Execute X times consecutively. When the real-time average duration exceeds the preset time threshold at least once, the service fault diagnosis result is determined to be abnormal. X can be set independently according to the actual situation and can be selected as 2.

[0128] For jmap, retrieve the memory usage of the class instance corresponding to jmap. The default memory usage can be selected as 5% or 200MB. A class instance is a class object, such as java.lang.Thread. If the memory usage of the class instance corresponding to jmap is greater than 5% or 200MB, the service fault diagnosis result is determined to be abnormal.

[0129] Memory-related metrics can be used alone as a basis for service fault diagnosis, or they can be combined with other metrics to serve as a basis for service fault diagnosis.

[0130] For memory, garbage collection (GC), and jmap, if at least one memory-related parameter is abnormal, the service fault diagnosis result is considered abnormal. If memory-related metrics are abnormal, the method can be terminated. If all memory-related metrics are normal, the service fault diagnosis result is determined by comparing the thread-related metrics with preset thread-related metric thresholds. This step specifically includes:

[0131] Thread-related metrics include the total number of threads. The corresponding preset thresholds for thread-related metrics can be set according to actual conditions, with an option of 400.

[0132] Thread-related metrics include the total number of active threads. The corresponding preset thresholds for thread-related metrics can be set independently according to actual conditions, and can be selected as 300.

[0133] Taking the total number of threads as an example, the average number of threads over N days is used as a reference value. When the real-time value of the total number of threads is greater than the average and the duration is greater than the preset duration, it is determined that there is an anomaly in the threads, and the service fault diagnosis result is determined to be abnormal, so the method can be terminated and continue execution. The specific values ​​of N and the preset duration can be set according to the actual situation. N can be selected as 30 and the preset duration can be selected as 5 minutes.

[0134] Thread-related metrics can be used alone as a basis for service fault diagnosis, or they can be combined with other metrics to serve as a basis for service fault diagnosis.

[0135] If either the total number of active threads or the total number of threads is abnormal, then the service fault diagnosis result is abnormal.

[0136] If all threads are normal, the service fault diagnosis result is determined by comparing the stack-related indicator information with preset stack-related indicator thresholds. This step specifically includes:

[0137] Stack-related metrics include the number of blocked threads. The corresponding preset stack-related metric thresholds can be set independently according to the actual situation, and can be selected as 1.

[0138] The average number of blocked threads over N days is used as a reference value. When the real-time number of blocked threads exceeds the average and also exceeds the preset stack trace threshold, an anomaly is identified in the stack trace, thus confirming an abnormal service fault diagnosis result, and the method can be terminated. The specific value of N can be set according to the actual situation; N can be selected as 30.

[0139] Stack-related metrics can be used alone as a basis for service fault diagnosis, or they can be combined with other metrics to serve as a basis for service fault diagnosis.

[0140] If the stack is normal, the service fault diagnosis result is determined based on the comparison result of the method execution time and the preset method time threshold, as well as the return value. This can be referred to the above embodiment for explanation, and will not be repeated here.

[0141] In summary, by systematically diagnosing service faults based on memory-related, thread-related, stack-related, and method-related metrics, the overall efficiency of service fault diagnosis for the device can be improved.

[0142] Compared with existing technologies, the service fault diagnosis method provided in this embodiment of the invention receives a trace command sent by a service fault diagnosis server. The trace command carries a class name and the name of the method corresponding to the class name to be diagnosed. An interception method is added to the interface method. The interception method is used to obtain method-related indicator information from the method and its sub-methods. When it is detected that the method and its sub-methods corresponding to the method name are about to be called or about to finish execution, the interface method with the added interception method is automatically called. The interception method obtains the method-related indicator information from the method and its sub-methods and sends the method-related indicator information to the service fault diagnosis server, so that the service fault diagnosis server can perform service fault diagnosis based on the method-related indicator information and obtain the service fault diagnosis result. This method facilitates service fault diagnosis and has the beneficial effects of low program performance loss and strong real-time performance.

[0143] A specific embodiment of the present invention discloses a service fault diagnosis system, the structural diagram of which is shown below. Figure 3 As shown, the system includes a device for performing service fault diagnosis and a service fault diagnosis server. The device for performing service fault diagnosis has a built-in probe. The methods performed by the device for performing service fault diagnosis and the service fault diagnosis server can be referred to the above method embodiments for description, and will not be repeated here.

[0144] The service fault diagnosis system may also include a client. After the user sends a command to the service fault diagnosis server through the client's browser, the service fault diagnosis server will send the command information to the corresponding service probe according to the service address and port of the device it carries. The probe is responsible for parsing the corresponding command and returning the detection results.

[0145] Compared with existing technologies, the service fault diagnosis system provided in this embodiment of the invention includes a device performing service fault diagnosis and a service fault diagnosis server. The method executed by the device includes: receiving a trace command sent by the service fault diagnosis server; the trace command carries a class name and a method name corresponding to the class name to be diagnosed; adding an interception method to the interface method; the interception method is used to obtain method-related indicator information from the method and its sub-methods; when it is detected that the method and its sub-methods corresponding to the method name are about to be called or about to finish execution, the interface method with the added interception method is automatically called, the method-related indicator information is obtained from the method and its sub-methods through the interception method, and the method-related indicator information is sent to the service fault diagnosis server so that the service fault diagnosis server can perform service fault diagnosis based on the method-related indicator information to obtain the service fault diagnosis result. This facilitates service fault diagnosis and has the beneficial effects of low program performance loss and strong real-time performance.

[0146] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.

[0147] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A service fault diagnosis method, characterized in that, include: Receive the trace command sent by the service fault diagnosis server; The trace command carries the class name and the method name corresponding to the class name for which service fault diagnosis is to be performed; Add an interception method to the interface method; The interception method is used to obtain method-related indicator information from the method and its sub-methods; The method-related metrics include: class name, method name, method description, and execution command, ultimately obtaining the method input parameters, method execution time, and return value. The addition of an interception method to the interface method includes: adding a first interception method to the interface method corresponding to the execution of the method and sub-methods before execution, for obtaining method-related indicator information before the execution of the method and sub-methods; and adding a second interception method to the interface method corresponding to the execution of the method and sub-methods after execution, for obtaining method-related indicator information after the execution of the method and sub-methods. When it is detected that the method and its sub-methods corresponding to the method name are about to be called or have finished executing, the interface method with the added interception method is automatically called. The interception method retrieves relevant indicator information from the method and its sub-methods and sends this information to the service fault diagnosis server. The service fault diagnosis server then performs service fault diagnosis based on this information, obtaining the service fault diagnosis result, including: When the listener detects that the method and its sub-methods are about to be called, it automatically calls the interface method corresponding to the method and its sub-methods before execution; when the listener detects that the method and its sub-methods have finished execution, it automatically calls the interface method corresponding to the method and its sub-methods after execution; in the interface method corresponding to the method and its sub-methods after execution, the listener object is called through the methodEnd method to output the relevant indicator information of the method; the methodEnd method is used to output all relevant indicator information of the method when the method and its sub-methods have finished execution, and to send all relevant indicator information of the method to the service fault diagnosis server.

2. The service fault diagnosis method according to claim 1, characterized in that, The service fault diagnosis method also includes: Receive device detection commands sent by the service fault diagnosis server; The local device is detected according to the device detection command to obtain device-related indicator information for service fault diagnosis; The device-related indicator information and the method-related indicator information are sent to the service fault diagnosis server so that the service fault diagnosis server can perform service fault diagnosis based on the method-related indicator information and the device-related indicator information to obtain the service fault diagnosis result.

3. A service fault diagnosis method, characterized in that, include: Send the trace command to the device performing the service fault diagnosis; The trace command carries the class name and the method name corresponding to the class name for which service fault diagnosis is to be performed; The device adds an interception method to the interface method, including: adding a first interception method to the interface method corresponding to the execution of the method and sub-methods before execution, for obtaining method-related indicator information before the execution of the method and sub-methods; and adding a second interception method to the interface method corresponding to the execution of the method and sub-methods after execution, for obtaining method-related indicator information after the execution of the method and sub-methods. When the device detects that the method and sub-methods corresponding to the method name are about to be called or have finished executing, it automatically calls the interface method after adding the interception method, obtains the method-related indicator information from the method and sub-methods through the interception method, and sends the method-related indicator information to the service fault diagnosis server. The service fault diagnosis server receives method-related indicator information returned by the device based on the method name corresponding to the class name, and performs service fault diagnosis based on the method-related indicator information to obtain a service fault diagnosis result. The method-related indicator information includes method input parameters, return value, and method execution time. The method input parameters correspond to the aforementioned class name and the method name corresponding to the class name. The return value includes whether the method execution is normal or abnormal. The method execution time is the time interval from the start to the end of the method execution.

4. The service fault diagnosis method according to claim 3, characterized in that, The method-related indicator information includes method input parameters, return values, and method execution time; correspondingly, the step of performing service fault diagnosis based on the method-related indicator information to obtain service fault diagnosis results includes: The class name and method name in the input parameters of the method are normalized to obtain the method name of the class; The execution time of the method corresponding to the method name of the class is determined, and the service fault diagnosis result is obtained based on the comparison result of the method execution time with the preset method execution time threshold and the return value.

5. The service fault diagnosis method according to claim 4, characterized in that, The process of obtaining the service fault diagnosis result based on the comparison result of the execution time of the method with the preset method time threshold and the return value includes: If the execution time of the method is less than the preset method time threshold, and the return value is normal, then the service fault diagnosis result is determined to be normal. If the execution time of the method is greater than or equal to the preset method time threshold, and / or the return value is an abnormal value, then the service fault diagnosis result is determined to be abnormal.

6. The service fault diagnosis method according to claim 4, characterized in that, The service fault diagnosis method also includes: Receive device-related indicator information and method-related indicator information sent by the device; and perform service fault diagnosis based on the method-related indicator information and the device-related indicator information to obtain service fault diagnosis results.

7. The service fault diagnosis method according to claim 6, characterized in that, The device-related metrics include memory-related metrics, thread-related metrics, and stack-related metrics; correspondingly, the step of performing service fault diagnosis based on the method-related metrics and the device-related metrics to obtain service fault diagnosis results includes: The service fault diagnosis result is determined based on the comparison between the memory-related indicator information and the preset memory-related indicator thresholds; The service fault diagnosis result is determined based on the comparison between the thread-related indicator information and the preset thread-related indicator threshold. The service fault diagnosis result is determined based on the comparison between the stack-related indicator information and the preset stack-related indicator thresholds. The service fault diagnosis result is determined based on the comparison result between the execution time of the method and the preset method time threshold, and the return value.

8. A service fault diagnosis system, characterized in that, include: The equipment and service fault diagnosis server that performs the service fault diagnosis; The device performs the service fault diagnosis method as described in any one of claims 1 to 2; The service fault diagnosis server executes the service fault diagnosis method as described in any one of claims 3 to 7.

Citation Information

Patent Citations

  • Method and interceptor system facing to tangent plane programming

    CN101276271A

  • WEB service distributed intelligent monitoring method and device, computer equipment and storage medium

    CN111385123A