Fault analysis methods, apparatus, computer equipment and computer-readable storage media
By acquiring and filtering events from performance counters, accurate output logs are generated, solving the problem of low efficiency in application system fault analysis in the fintech field, and achieving fast and accurate fault location and efficient fault analysis.
Patent Information
- Application Number
- CN202211551965.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-05
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-12-05
AI Technical Summary
Existing technologies in the fintech field suffer from low efficiency in application system fault analysis, especially when performance failures occur within a single application system. They cannot be analyzed quickly and accurately, and log recording consumes a large amount of computing and storage resources, resulting in low analysis efficiency.
By acquiring all events recorded by the performance counter, filtering start and end events, generating record information for each event, and generating output logs for the performance counter based on this information, and generating fault analysis information if a performance fault is detected.
It improves the efficiency of fault analysis, reduces the consumption of computing and storage resources, quickly locates the fault range, and improves the accuracy and efficiency of fault analysis.
Smart Images

Figure CN115904905B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fault analysis, and more particularly to a fault analysis method, apparatus, computer equipment, and computer-readable storage medium that can be used in financial technology or other related fields. Background Technology
[0002] In the fintech field, the operational performance of computer application systems is a crucial indicator of their quality. Ensuring good application system performance and rapidly analyzing performance failures are important research topics for guaranteeing application system quality. Typically, external systems with distributed tracing can be deployed to provide APIs (Application Programming Interfaces) to access the application system's development code or collect and analyze information from the application system's network layer for performance failure analysis. However, external systems with distributed tracing are usually used for fault analysis across multiple application systems and cannot accurately analyze performance failures in a single application system.
[0003] When a performance failure occurs within a single application system, current technology typically records the application system's processing flow as logs and analyzes the performance failure by printing out these logs. However, application systems generate a large number of logs during operation, and printing these logs consumes significant computing and storage resources and takes a considerable amount of time, leading to inefficient application system failure analysis. Furthermore, the logs contain much information irrelevant to performance failure analysis, requiring time to filter out irrelevant information, further contributing to the inefficiency of application system failure analysis. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to overcome the shortcomings of the prior art and provide a fault analysis method, apparatus, computer equipment and computer-readable storage medium that can be used in financial technology or other related fields to solve the problem of low efficiency in fault analysis of application systems.
[0005] In a first aspect, this application provides a fault analysis method applied to a computer device including a performance counter, the method comprising:
[0006] Retrieve all events recorded by the performance counter, and filter the start and end events from all the events;
[0007] Based on the processing information and description information of each event, record information for each event is generated, wherein the processing information is used to determine whether the event was processed successfully;
[0008] Based on the start event, the end event, and the recorded information for each event, the output log of the performance counter is obtained;
[0009] If a performance failure is detected, failure analysis information is generated based on the output logs.
[0010] In conjunction with the first aspect, in a first possible implementation, obtaining the output log of the performance counter based on the start event, the end event, and the recording information of each event includes:
[0011] Based on the title information of the performance counter, determine the component processing flow corresponding to the performance counter;
[0012] In response to the received process completion information of the component processing flow, the output log of the performance counter is obtained based on the start event, the end event, and the record information of each event.
[0013] In conjunction with the first possible implementation of the first aspect, in the second possible implementation, the processing information includes processing completion information and processing stage information. The processing completion information is used to describe whether the event was successfully processed, and the processing stage information is used to describe the processing stage of the component processing flow corresponding to the event.
[0014] In conjunction with the first aspect, in a third possible implementation, generating the record information for each event based on the processing information and description information of each event includes:
[0015] Obtain the time interval between the current event and the previous event, and calculate the duration of the current event based on the time interval;
[0016] Based on the duration, processing information, and description of the current event, generate record information for the current event.
[0017] In conjunction with the first aspect, in a fourth possible implementation, the method further includes:
[0018] Obtain the running memory threshold of the performance counter, and release the memory resources of the performance counter according to the running memory threshold.
[0019] In conjunction with the first aspect, in the fifth possible implementation, obtaining the output log of the performance counter based on the start event, the end event, and the recording information of each event includes:
[0020] The start time of the performance counter is determined based on the start event, and the end time of the performance counter is determined based on the end event.
[0021] Based on the start time, the end time, and the recorded information for each event, the output log of the performance counter is obtained.
[0022] In conjunction with the first aspect, in the sixth possible implementation, obtaining the output log of the performance counter based on the start event, the end event, and the recording information of each event includes:
[0023] If all events include external call events, then obtain the external time consumption of the external call events;
[0024] The performance counter's output log is generated based on the startup event, the termination event, the recorded information for each event, and the external time consumption.
[0025] Secondly, this application provides a fault analysis apparatus applied to a computer device including a performance counter, the apparatus comprising:
[0026] The event acquisition module is used to acquire all events recorded by the performance counter and filter the start events and end events among all events;
[0027] The record information generation module is used to generate record information for each event based on the processing information and description information of each event, wherein the processing information is used to determine whether the event was processed successfully;
[0028] The log output module is used to obtain the output log of the performance counter based on the start event, the end event, and the recording information of each event;
[0029] The fault analysis module is used to generate fault analysis information based on the output logs if a performance fault is detected.
[0030] Thirdly, this application provides a computer device, which includes a performance counter, a memory, and a processor. The memory stores a computer program, which, when executed by the processor, implements the fault analysis method as described in the first aspect.
[0031] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the fault analysis method as described in the first aspect.
[0032] This application provides a fault analysis method applied to a computer device including a performance counter. The method includes: acquiring all events recorded by the performance counter and filtering start events and end events from all events; generating record information for each event based on processing information and description information; obtaining an output log of the performance counter based on the start event, the end event, and the record information for each event; and generating fault analysis information based on the output log if a performance fault is detected. By accurately recording events of the computer device using event processing and description information, and summarizing them to generate an output log including a small amount of key information, the method eliminates the need to output large amounts of system process logs, thus improving fault analysis efficiency. Attached Figure Description
[0033] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope of protection of the present invention. In the various drawings, similar components are numbered similarly.
[0034] Figure 1 A flowchart of the first fault analysis method provided by an embodiment of the present invention is shown;
[0035] Figure 2 A flowchart of the second fault analysis method provided by an embodiment of the present invention is shown;
[0036] Figure 3 An example diagram of the output log provided in an embodiment of the present invention is shown;
[0037] Figure 4 A schematic diagram of the fault analysis device provided in an embodiment of the present invention is shown. Detailed Implementation
[0038] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0039] The components of the embodiments of the invention described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0040] In the following, the terms “comprising,” “having,” and their cognates, which may be used in various embodiments of the invention, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as excluding, firstly, the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more features, numbers, steps, operations, elements, components, or combinations thereof.
[0041] Furthermore, the terms "first," "second," and "third" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0042] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of the invention pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be interpreted as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of the invention.
[0043] Example 1
[0044] Please see Figure 1 , Figure 1 A flowchart of the first fault analysis method provided by an embodiment of the present invention is shown. Figure 1 The fault analysis methods described in the text are applied to computer devices that include performance counters, including:
[0045] S110, obtain all events recorded by the performance counter, and filter the start events and end events among all events.
[0046] A performance counter, also known as a performance monitor, collects and analyzes performance events of applications, services, drivers, etc., within a system in real time. It then analyzes system bottlenecks and monitors components based on the performance data from these events. This application applies to computer devices that include performance counters. The performance counters record all events occurring during the completion of the processing flow of server-side components within the computer device, as well as the time consumption information for each step in the processing flow. Start-up and end-up events are filtered from all events to mark the start and end events of the performance counter, which are then used as reference time points for subsequent performance recordings by the performance counter.
[0047] S120 generates record information for each event based on the processing information and description information of each event.
[0048] In computer devices, an event typically refers to a response operation, such as clicking, double-clicking, or moving the human-computer interaction device. Each event's description is used to provide specific details and annotations; it can be any type of character information, without any restrictions. The description information identifies each event and distinguishes its type.
[0049] The processing information for each event is used to record key information such as whether the event was successfully processed. By combining the processing and descriptive information for each event, a record is generated for each event. This provides the ability to record detailed performance events through performance counters, and the recorded event information can be used for system fault analysis of computer equipment.
[0050] As an example, generating the record information for each event based on the processing information and description information of each event includes:
[0051] Obtain the time interval between the current event and the previous event, and calculate the duration of the current event based on the time interval;
[0052] Based on the duration, processing information, and description of the current event, generate record information for the current event.
[0053] During the process of acquiring the processing and description information of the current event, the time interval between the current event and the previous event is also acquired, and the processing time of the current event is obtained based on the time interval between the current event and the previous event. Based on the processing time, processing information, and description information of the current event, record information for the current event is generated. During the recording of performance event details, the time interval between the current event and the previous event is recorded to obtain the processing time of the current event, facilitating system fault analysis of computer equipment based on the time consumption.
[0054] S130, based on the start event, the end event, and the record information of each event, obtain the output log of the performance counter.
[0055] The system retrieves the performance recording time interval between the start and end events, and collects all events within that interval, summarizing all events between the start and end events. Based on the recorded information of all events between the start and end events, the performance counter output log is obtained. The performance counter output log summarizes the details of all performance events that occurred from the start to the end of the process. This precise recording of information during the processing, along with the aggregated output of each event's record, avoids scattering event information across different locations, facilitating system fault analysis of the computer equipment.
[0056] As an example, the output log of the performance counter, obtained based on the start event, the end event, and the recording information of each event, includes:
[0057] Based on the title information of the performance counter, determine the component processing flow corresponding to the performance counter;
[0058] In response to the received process completion information of the component processing flow, the output log of the performance counter is obtained based on the start event, the end event, and the record information of each event.
[0059] It's important to understand that each server-side component corresponds to a different performance counter. Computer devices can provide the ability to record the titles of these performance counters, thus indicating their purpose. Based on the title information, the processing flow of the component corresponding to that performance counter is determined. For example, if the title of a performance counter is "HTTP Server," then the performance counter is used to record the processing flow of an HTTP (Hypertext Transfer Protocol) server-side component.
[0060] Once an HTTP server-side processing flow is complete, the performance counter output log is generated based on the start events, end events, and information recorded by the performance counters. The performance counter output log is printed all at once after the server-side component completes its processing, resulting in high processing efficiency and providing convenient performance analysis capabilities for computer devices.
[0061] In an optional example, the processing information includes processing completion information and processing stage information. The processing completion information describes whether the event was processed successfully, and the processing stage information describes the processing stage of the component processing flow corresponding to the event.
[0062] Both completion information and processing stage information are key pieces of information for performance events recorded by performance counters. Completion information describes whether the event was successfully processed, while processing stage information describes the processing stage of the component's processing flow corresponding to the event. Completion information can be used for system fault analysis of computer equipment. Simultaneously, processing stage information can be used to determine the corresponding processing stage of the event, thereby quickly locating the fault.
[0063] The system accurately records events during processing based on their duration, processing information, and description. Once the server-side component completes its processing, the logs are printed all at once, resulting in high efficiency for fault analysis and preventing log information from being scattered across different locations. Only key information such as processing and descriptions needs to be output as logs, eliminating the need for extensive process logs and improving the efficiency of application system fault analysis.
[0064] As an example, the output log of the performance counter, obtained based on the start event, the end event, and the recording information of each event, includes:
[0065] The start time of the performance counter is determined based on the start event, and the end time of the performance counter is determined based on the end event.
[0066] Based on the start time, the end time, and the recorded information for each event, the output log of the performance counter is obtained.
[0067] The system retrieves the time of the performance counter's startup event and marks it as the performance counter's startup time. It also retrieves the time of the performance counter's end event and marks it as the performance counter's end time. Based on the performance counter's startup and end times, it obtains the recording time interval for performance events. Based on the recorded information for each event within the time interval, it generates the performance counter's output log. By setting the startup and end times as the time reference point for the output log, it outputs a log containing key information such as processing and description information. This eliminates the need for large amounts of process logs, improving the efficiency of fault analysis in the application system.
[0068] Please see Figure 2 , Figure 2 A flowchart of a second fault analysis method provided by an embodiment of the present invention is shown.
[0069] As an example, the output log of the performance counter, obtained based on the start event, the end event, and the recording information of each event, includes:
[0070] S131, if all events include external call events, then obtain the external time consumption of the external call events.
[0071] During the system processing of computer equipment, external calls may occur. When an internal performance failure of the computer equipment system is detected, it is necessary to provide the ability to mark external calls with performance counters to exclude the external call period. Specifically, if all events in the service component process of the computer equipment system include external call events, then the external call events are marked, and the external call time is obtained.
[0072] S132, based on the start event, the end event, the record information of each event, and the external time consumption, generate the output log of the performance counter.
[0073] Based on the start and end events recorded by the performance counters, along with the recording information and external execution time for each event, an output log for the performance counters is generated. This output log allows for quick determination of the external call time and whether the external call event was successfully processed.
[0074] S140, if a performance failure is detected, generate failure analysis information based on the output log.
[0075] Please see Figure 3 , Figure 3 An example diagram of the output log provided in an embodiment of the present invention is shown.
[0076] As shown in the figure, in this embodiment, a performance counter output log is generated based on the start event, end event, the recorded information of each event, the time consumed, and the external time consumed. When a performance failure of a computer device is detected, the generated output log quickly determines the processing stage, time consumed, and whether the event was successfully processed for each event. Simultaneously, if the event is an external call event, the external time consumed by the external call event can also be quickly determined, generating fault analysis information. This fault analysis information allows for rapid localization of the fault scope, identifying events that cause performance failures in the system, improving fault analysis efficiency, and helping developers handle performance failures promptly.
[0077] By utilizing performance counters included in computer equipment, and without relying on complex external systems, it independently completes fine-grained, easily analyzable performance recording within the application process of the computer equipment system, thus aiding performance analysis. When performance failures of computer equipment are detected, situations where performance data cannot be collected on-site due to a lack of timely tool loading are avoided, providing direct on-site capture of performance failures.
[0078] As an example, the method also includes:
[0079] Obtain the running memory threshold of the performance counter, and release the memory resources of the performance counter according to the running memory threshold.
[0080] Because performance counters require computer memory resources to record performance events, a memory cleanup mechanism is developed to obtain the performance counter's memory threshold. When the memory resources used by the performance counter exceed the threshold, the memory resources of the performance counter are released to provide memory cleanup capabilities for the performance counter.
[0081] This application provides a fault analysis method applied to a computer device including a performance counter. The method includes: acquiring all events recorded by the performance counter and filtering start events and end events from all events; generating record information for each event based on processing information and description information; obtaining an output log of the performance counter based on the start event, the end event, and the record information for each event; and generating fault analysis information based on the output log if a performance fault is detected. By accurately recording events of the computer device using event processing and description information, and summarizing them to generate an output log including a small amount of key information, the method eliminates the need to output large amounts of system process logs, thus improving fault analysis efficiency.
[0082] Example 2
[0083] Please see Figure 4 , Figure 4 A schematic diagram of the fault analysis device provided in an embodiment of the present invention is shown. Figure 4 The fault analysis device 200 is applied to a computer device including a performance counter, comprising:
[0084] The event acquisition module 210 is used to acquire all events recorded by the performance counter and filter the start events and end events among all events;
[0085] The record information generation module 220 is used to generate record information for each event based on the processing information and description information of each event, wherein the processing information is used to determine whether the event was successfully processed;
[0086] The log output module 230 is used to obtain the output log of the performance counter based on the start event, the end event, and the recording information of each event;
[0087] The fault analysis module 240 is used to generate fault analysis information based on the output log if a performance fault is detected.
[0088] As an example, the log output module 230 includes:
[0089] The processing flow determination submodule is used to determine the component processing flow corresponding to the performance counter based on the title information of the performance counter;
[0090] The output log submodule is used to output the process completion information of the component processing flow received in response, and to obtain the output log of the performance counter based on the start event, the end event and the record information of each event.
[0091] In an optional example, the processing information includes processing completion information and processing stage information. The processing completion information describes whether the event was processed successfully, and the processing stage information describes the processing stage of the component processing flow corresponding to the event.
[0092] As an example, the record information generation module 220 includes:
[0093] The time consumption submodule is used to obtain the time interval between the current event and the previous event, and to obtain the time consumption of the current event based on the time interval;
[0094] The recording information submodule is used to generate recording information for the current event based on the duration, processing information, and description information of the current event.
[0095] As an example, in the fault analysis apparatus 200, the method further includes:
[0096] The memory resource release module is used to obtain the running memory threshold of the performance counter and release the memory resources of the performance counter according to the running memory threshold.
[0097] As an example, the log output module 230 includes:
[0098] The time determination submodule is used to determine the start time of the performance counter based on the start event, and to determine the end time of the performance counter based on the end event;
[0099] The logging submodule is used to obtain the output log of the performance counter based on the start time, the end time, and the recording information of each event.
[0100] As an example, the log output module 230 includes:
[0101] The external time consumption submodule is used to obtain the external time consumption of the external call event if the events include an external call event.
[0102] A log generation submodule is used to generate the output log of the performance counter based on the start event, the end event, the record information of each event, and the external time consumption.
[0103] The fault analysis device 200 is used to execute the corresponding steps in the fault analysis method described above. The specific implementation of each function will not be described in detail here. In addition, the optional examples in Embodiment 1 are also applicable to the fault analysis device 200 in Embodiment 2.
[0104] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the fault analysis method as described in Embodiment 1.
[0105] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the fault analysis method as described in Embodiment 1.
[0106] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, as an alternative implementation, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0107] In addition, the functional modules or units in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0108] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0109] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A fault analysis method, characterized in that, Applied to a computer device including a performance counter, the method includes: Retrieve all events recorded by the performance counter, and filter the start and end events from all the events; Based on the processing information and description information of each event, record information for each event is generated. The processing information is used to determine whether the event was processed successfully, and the description information for each event is used to provide specific descriptions and notes for the event. Each event is identified through the description information to distinguish the type of the event. Based on the start event, the end event, and the recorded information for each event, the output log of the performance counter is obtained; If a performance failure is detected, failure analysis information is generated based on the output logs.
2. The fault analysis method according to claim 1, characterized in that, The output log of the performance counter is obtained based on the start event, the end event, and the record information of each event, including: Based on the title information of the performance counter, determine the component processing flow corresponding to the performance counter; In response to the received process completion information of the component processing flow, the output log of the performance counter is obtained based on the start event, the end event, and the record information of each event.
3. The fault analysis method according to claim 2, characterized in that, The processing information includes processing completion information and processing stage information. The processing completion information describes whether the event was processed successfully, and the processing stage information describes the processing stage of the component processing flow corresponding to the event.
4. The fault analysis method according to claim 1, characterized in that, The step of generating the record information for each event based on the processing information and description information of each event includes: Obtain the time interval between the current event and the previous event, and calculate the duration of the current event based on the time interval; Based on the duration, processing information, and description of the current event, generate record information for the current event.
5. The fault analysis method according to claim 1, characterized in that, The method further includes: Obtain the running memory threshold of the performance counter, and release the memory resources of the performance counter according to the running memory threshold.
6. The fault analysis method according to claim 1, characterized in that, The output log of the performance counter is obtained based on the start event, the end event, and the record information of each event, including: The start time of the performance counter is determined based on the start event, and the end time of the performance counter is determined based on the end event. Based on the start time, the end time, and the recorded information for each event, the output log of the performance counter is obtained.
7. The fault analysis method according to claim 1, characterized in that, The output log of the performance counter is obtained based on the start event, the end event, and the record information of each event, including: If all events include external call events, then obtain the external time consumption of the external call events; The performance counter's output log is generated based on the startup event, the termination event, the recorded information for each event, and the external time consumption.
8. A fault analysis device, characterized in that, An apparatus applicable to a computer device including a performance counter, the apparatus comprising: The event acquisition module is used to acquire all events recorded by the performance counter and filter the start events and end events among all events; The record information generation module is used to generate record information for each event based on the processing information and description information of each event. The processing information is used to determine whether the event was processed successfully, and the description information of each event is used to make specific descriptions and remarks about the event. Each event is identified by the description information to distinguish the type of the event. The log output module is used to obtain the output log of the performance counter based on the start event, the end event, and the recording information of each event; The fault analysis module is used to generate fault analysis information based on the output logs if a performance fault is detected.
9. A computer device, characterized in that, The computer device includes a performance counter, a memory, and a processor. The memory stores a computer program, which, when executed by the processor, implements the fault analysis method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the fault analysis method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Apache service error processing method and system based on Linux
CN106209427A
Log information analysis method and related device
CN111597550A