Application program monitoring method and system
By building a time series model and log detection model, the problem that the existing technology cannot effectively detect application abnormalities is solved, and the stability and user experience of the application are improved.
Patent Information
- Application Number
- CN202510091650.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The existing application monitoring system cannot effectively detect abnormal situations and failure points of the application, resulting in unstable application and reducing user experience.
By building a time series model and log detection model, obtain the time and log information of the application request response server, and perform exception detection to discover performance bottlenecks and failure points.
It realizes detailed monitoring of application metrics and logs, quickly locates performance bottlenecks and failure points, and improves application stability and user experience.
Smart Images

Figure CN120011175A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to an application monitoring method; in addition, the present invention also relates to an application monitoring system. Background Art
[0002] With the popularity of various intelligent terminal devices and their increasingly powerful processing capabilities, the number of installed applications continues to increase. During the operation of applications, some abnormal situations often occur, such as network abnormalities, database abnormalities, file operation abnormalities, etc. However, the existing application monitoring system simply monitors the logs and cannot better detect abnormal situations and fault points, which makes the application unstable and reduces the user experience. Summary of the invention
[0003] In order to solve the problems existing in the prior art, at least one embodiment of the present invention provides an application monitoring method, which can better monitor and detect application indicator anomalies and log anomalies, find performance bottlenecks, thereby improving application stability and user experience. To this end, at least one embodiment of the present invention also provides an application monitoring system.
[0004] In a first aspect, an embodiment of the present invention provides an application monitoring method, comprising:
[0005] Get the time it takes for an application to request and respond to a server, and build a time series model;
[0006] Anomaly detection on application metrics through time series models;
[0007] Obtain application log information and build a log detection model;
[0008] Detect abnormal logs through log detection models.
[0009] In some embodiments, the present invention provides an application monitoring method, wherein the consumed time includes network delay time, client delay time, routing time, database processing time, cache processing time, external service call time, and response generation return time.
[0010] In some embodiments, the present invention provides an application monitoring method, wherein application indicators include throughput, response time, request error rate, CPU usage, and memory usage that vary with consumption time.
[0011] In some embodiments, the present invention provides an application monitoring method, and constructing a time series model includes:
[0012] The time series model is constructed by the exponential moving average method. The time series model is expressed by the following equations 1 and 2:
[0013] S(t)=aX(t)+(1-a)S(t-1) Formula 1;
[0014]
[0015] Among them, S(t) is the smoothed value, a is the weighting coefficient, X(t) is the time series, and S(t-1) is the original smoothed value.
[0016] In some embodiments, the present invention provides an application monitoring method, and building a log detection model includes:
[0017] The log detection model is constructed by clustering algorithm. The log detection model is expressed by the following formula 3:
[0018] d(x,c)>threshold Equation 3;
[0019] Where x is the log item, c is the cluster center, d(x,c) is the distance between x and c, and threshold is the threshold determined based on the clustering result.
[0020] In some embodiments, the present invention provides an application monitoring method, the method comprising:
[0021] When a single request from an application spreads responses across multiple services, the performance bottleneck of the service is expressed by the following equation 4:
[0022]
[0023] Among them, T i For Service i The processing time, C i→j For Service i Calling Service S j Time, C i→j For Service i With Service j The calling relationship between them.
[0024] In a second aspect, an embodiment of the present invention further provides an application monitoring system, including:
[0025] The consumption time acquisition module is used to obtain the consumption time of the application request response server;
[0026] Time series model building module, used to build time series models;
[0027] Anomaly indicator detection module, used to detect anomalies of application indicators through time series models;
[0028] Log acquisition module, used to obtain application log information;
[0029] Log detection model building module, used to build log detection model;
[0030] The abnormal log detection module is used to detect abnormal logs through the log detection model.
[0031] In a third aspect, an embodiment of the present invention further provides an application monitoring device, comprising at least one processor; a memory coupled to the at least one processor, the memory storing executable instructions, which when executed by the at least one processor implement the steps of any one of the methods of the first aspect above.
[0032] In a fourth aspect, an embodiment of the present invention further provides a chip for executing the steps of the method in the first aspect. Specifically, the chip includes: a processor for calling and running a computer program from a memory, so that a device equipped with the chip is used to execute the steps of the method in the first aspect.
[0033] In a fifth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the methods in the first aspect above are implemented.
[0034] It can be seen that an application monitoring method and system in an embodiment of the present invention, through the constructed time series model and log detection model, tracks the processing flow of each request response in detail and detects its anomalies, which helps to quickly locate performance bottlenecks and fault points in the application, and realizes the monitoring guarantee of users based on their own business goals and their own platform stability, thereby improving the stability of the application and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0036] Figure 1 Shown is a flow chart of an application monitoring method according to an embodiment of the present invention
[0037] Figure 2 Shown is a schematic diagram of the framework of an application monitoring system in an embodiment of the present invention. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0039] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. In this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0040] [Example 1]
[0041] In the prior art, the existing application monitoring system simply monitors the logs and cannot better detect abnormal situations and fault points, thereby making the application unstable and reducing the user experience. Embodiment 1 of the present invention provides the following solution:
[0042] like Figure 1 As shown, this embodiment provides an application monitoring method, the method comprising the following steps:
[0043] Step 101, obtain the time consumed by the application program requesting and responding to the server, and build a time series model.
[0044] In some embodiments, when a user operates an application, the application request response server goes through multiple stages of processing, the main processing process is as follows:
[0045] S1: Client request initiated (user request);
[0046] S2: The request enters the entry layer (such as a web server or API gateway);
[0047] S3: Authentication and Authorization;
[0048] S4: Request routing to the business logic layer;
[0049] S5: database operations (query, insert, etc.);
[0050] S6: Cache operation;
[0051] S7: external service calls (e.g. third-party APIs);
[0052] S8: Response generation and return.
[0053] The total time taken for the request is T total ,The total time consumption refers to the time consumed in the entire life cycle of the request from initiation to the final return. The time consumption of each stage is summed up:
[0054] T total =T1+T2+T3+T4+T5+T6+T7+T8;
[0055] The time consumed in each stage includes:
[0056] T1: Network delay time, including the network delay of the request and the delay of the client.
[0057] T2: Client latency, the time it takes for a request to reach an entry point, such as a web server or API gateway, from the client.
[0058] T3: The time for authentication and authorization, which involves calling the authentication service.
[0059] T4: Routing time, request routing and business logic layer processing time.
[0060] T5: Database processing time, such as query, insert, update, etc.
[0061] T6: Cache processing time, such as the time it takes for a request to pass through the cache layer, including cache query time and cache invalidation query time.
[0062] T7: The time of external service call, such as API call or third-party service request time.
[0063] T8: Response generation return time, including data processing, formatting and response sending.
[0064] In some embodiments, in a distributed system, transaction tracking can be used to associate the processing of each component through Trace ID and Span ID. Assume that a request spans multiple microservices, each microservice will generate a Span, and finally form a complete Trace. For a request spanning multiple services, the transaction tracking time is expressed by the following formula 5:
[0065]
[0066] Among them, Ti The processing time consumed for each service.
[0067] It should be noted that Trace ID is a globally unique identifier used to represent the entire tracking process of a single user request or transaction in a distributed system. When a user request enters the system, the system generates a globally unique Trace ID at the first layer of the RPC call network. In distributed systems and microservice architectures, Span is the basic unit for tracking request flows. When each service processes a request, it generates a Span to represent the processing process of the service. Span includes:
[0068] Span ID, which is used to uniquely identify the Span;
[0069] Parent Span ID, used if the span is caused by another span (for example, a request in a call chain), it will have a parent span ID;
[0070] The start time and end time are used to record the duration of the span, expressed in the form of a timestamp;
[0071] Tags are used to mark some key context information, such as service name, request parameters, status code, etc.
[0072] Events, used to record important events that occur during processing, usually as a time series log;
[0073] Status, which indicates the success or failure status of the span.
[0074] The processing time of each service can be further broken down into the time spent in each processing stage, ultimately forming a complete processing chain. By monitoring and tracking the time spent in each link, performance bottlenecks can be discovered. If a link takes too long, system performance can be improved through optimization solutions, such as index optimization, caching mechanism, asynchronous processing, etc.
[0075] In some embodiments, application metrics include throughput, response time, request error rate, CPU usage, and memory usage as a function of elapsed time.
[0076] It should be noted that as the time consumed by the application's request responses at each stage changes, the throughput, response time, request error rate, CPU usage, and memory usage will also change accordingly.
[0077] Throughput is the number of requests processed per unit time, indicating the load capacity of the system, and is expressed by the following formula 6:
[0078]
[0079] Among them, N requests is the number of requests processed in the time window, T window is the time window length, and the average throughput is expressed by the following equation 7:
[0080]
[0081] Where m is the number of time windows for detection.
[0082] Response time refers to the time required for the system to process a request, which is expressed by the following formula 8:
[0083] T requests =T end -T start Formula 8;
[0084] Among them, T end The timestamp of request processing completion, T start is the timestamp of request arrival, and the average response time is expressed by the following formula 9:
[0085]
[0086] The request error rate is the proportion of failed requests, which is expressed by the following formula 10:
[0087]
[0088] Among them, N errors is the number of error requests, N requests The total number of requests.
[0089] CPU usage is the percentage of CPU resources consumed by the application to the total available CPU resources of the system, which is expressed by the following formula 11:
[0090]
[0091] Among them, T used is the CPU usage time, T total Total CPU usage time.
[0092] Memory usage is the percentage of total system memory consumed by the application, expressed as follows:
[0093]
[0094] Among them, M used is the current memory usage, M total The total system memory.
[0095] The overall performance can be expressed by weighting each indicator, as shown in the following formula 13:
[0096]
[0097] Among them, ω1, ω2, ω3, ω4 and ω5 are weights and can be adjusted according to the application scenario.
[0098] In some embodiments, a time series model is constructed by an exponential moving average method, and the time series model is represented by the following formula 1 and the following formula 2:
[0099] S(t)=aX(t)+(1-a)S(t-1) Formula 1;
[0100]
[0101] Among them, S(t) is the smoothed value, a is the weighting coefficient, X(t) is the time series, and S(t-1) is the original smoothed value.
[0102] It should be noted that the exponential smoothing method performs a weighted average of historical data, giving higher weights to recent data and lower weights to more distant data, so that the exponential smoothing method can more accurately capture the changing trends in the data.
[0103] Step 102: Perform anomaly detection on application indicators using a time series model.
[0104] In some embodiments, anomaly detection of application metrics may be performed based on setting a threshold for a time series model, as represented by the following formula 14:
[0105]
[0106] Among them, a and b are the set thresholds.
[0107] In some other embodiments, anomaly detection of application metrics may be detected by standard deviation, which is expressed by the following formula 15:
[0108]
[0109] Among them, μ is the mean, σ is the standard deviation, and k is the adjustment parameter.
[0110] In some embodiments, the capacity utilization rate can be modeled and predicted by an ARIMA model or a Prophet model, and then a threshold is set to trigger an early warning of the need for capacity expansion. For example, the early warning threshold is set to 0.8.
[0111] Step 103, obtain the log information of the application and build an abnormal log detection model.
[0112] It should be noted that the log information mainly includes log level, timestamp, message content, error code, and request ID.
[0113] In some embodiments, the abnormal log detection model is constructed by a clustering algorithm, and the abnormal log detection model is represented by the following formula 3:
[0114] d(x,c)>threshold Equation 3;
[0115] Where x is the log item, c is the cluster center, d(x,c) is the distance between x and c, and threshold is the threshold determined based on the clustering result.
[0116] In some other embodiments, anomaly detection is performed on log data by statistical analysis methods, such as detection based on standard deviation and detection based on moving average. Suppose there are n log samples L = {l1, l2, ..., l n}, the response time of each sample is x i , the mean value μ and standard deviation σ of the log are expressed by the following equations 16 and 17 respectively:
[0117]
[0118] Based on statistics, set an anomaly detection rule. For example, if the response time of a log item is x i If it exceeds μ+kσ or is lower than μ-kσ, the log is considered abnormal.
[0119] Step 104: Detect abnormal logs using an abnormal log detection model.
[0120] In some embodiments, the relationship between the error log and the exception log is expressed by calculating the correlation between them. Let E(t) be the number of error logs occurring at time t, A(t) be the number of exception logs, and the correlation d(E, A) between the error log and the exception log is expressed by the following formula 18:
[0121]
[0122] Among them, σ E is the standard deviation of the error log, σ A is the standard deviation of the abnormal log. If d(E,A) is close to 1, it means that there is a strong positive correlation between them, indicating that most error logs may be accompanied by abnormalities.
[0123] In some embodiments, an error log trend graph, an abnormal log trend graph, and a log density distribution may also be displayed on the interface of the application or on the server side. The error log trend graph shows the number of error logs in different time periods, the abnormal log trend graph shows the change in the number of abnormal logs, and the log density distribution: shows the density distribution of different types of logs (such as error logs, abnormal logs) in different time periods.
[0124] In some embodiments, when a single request of an application propagates responses between multiple services, distributed tracing usually identifies a single request through a unique Trace ID. The request may span multiple services and generate a Span in each service node. Each Span records the processing time of the request in a service, the call relationship between services, and other key information. Each request is processed by multiple services, and each service generates a Span. The request starts from the client, passes through multiple microservices, and finally the response is returned to the client. The propagation path of the request is: Trace = {S1, S2, ..., S n}, S1 is the first service that the request enters, S n It is the last service. Each S1 generates a Span to record the processing time and other related information of the service.
[0125] For example, suppose there are three microservices, A, B, and C. The processing path of the client request through the three services is expressed by the following formula 19:
[0126]
[0127] Where T1 is the processing time of the request in service A, and T2 is the processing time of the request in service B. Let the time of requesting service A be T A , the service time of B is T B , the service time of C is T C , the total time of the entire request is expressed by the following formula 20:
[0128] T total =T A +T B +T C Formula 20.
[0129] In some embodiments, the performance bottleneck of the service is expressed by the following formula 4:
[0130]
[0131] Among them, T i For Service i The processing time, C i→j For Service i Calling Service S j Time, C i→j For Service i With Service j The calling relationship between them is reduced by i and C i→j Based on the above bottleneck analysis, the following optimization measures can be adopted:
[0132] Optimize slow services: Optimize the performance of services with long response times, such as optimizing database queries and reducing unnecessary calculations.
[0133] Introduce a cache mechanism: For data that is frequently queried and whose results do not change much, use a cache to reduce the response time of requests, such as Redis and Memcached.
[0134] Asynchronous processing: If some services take a long time to process and do not affect the main business process, change them to asynchronous processing to reduce the impact on overall performance;
[0135] Service parallelization: Parallelize the serial dependencies between multiple services to reduce the overall response time.
[0136] [Example 2]
[0137] like Figure 2 As shown, this embodiment provides an application monitoring system, including:
[0138] The consumption time acquisition module 201 is used to acquire the consumption time of the application program requesting and responding to the server;
[0139] A time series model building module 202 is used to build a time series model;
[0140] An abnormal indicator detection module 203, used to perform abnormality detection on application indicators through a time series model;
[0141] The log acquisition module 204 is used to acquire the log information of the application program;
[0142] A log detection model building module 205 is used to build a log detection model;
[0143] The abnormal log detection module 206 is used to detect abnormal logs through a log detection model.
[0144] [Example 3]
[0145] This embodiment provides an application monitoring device, including:
[0146] At least one processor; a memory coupled to the at least one processor, the memory storing executable instructions, wherein the executable instructions, when executed by the at least one processor, implement the method steps of embodiment 1 of the present invention.
[0147] In an application monitoring device provided by an embodiment of the present invention, a processor and a memory may be provided separately or integrated together.
[0148] For example, the memory may include a random access memory, a flash memory, a read-only memory, a programmable read-only memory, a non-volatile memory or a register, etc. The processor may be a central processing unit (CPU), etc. or a graphic processing unit (GPU). The memory may store executable instructions. The processor may execute the executable instructions stored in the memory to implement the various processes described herein.
[0149] It can be understood that the memory in this embodiment can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a ROM (Read-Only Memory), a PROM (Programmable ROM), an EPROM (Erasable PROM), an EEPROM (Electrically EPROM), or a flash memory. The volatile memory can be a RAM (Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as SRAM (Static RAM), DRAM (Dynamic RAM), SDRAM (Synchronous DRAM), DDR SDRAM (Double Data Rate SDRAM), ESDRAM (Enhanced SDRAM), SLDRAM (Synchlink DRAM), and DRRAM (Direct Rambus RAM). The memories described herein are intended to include, but are not limited to, these and any other suitable types of memories.
[0150] In some embodiments, the memory stores the following elements, upgrade packages, executable units or data structures, or a subset thereof, or an extended set thereof: an operating system and an application program.
[0151] The operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic services and process hardware-based tasks. The application program includes various application programs, which are used to implement various application services. The program for implementing the method of the embodiment of the present invention can be included in the application program.
[0152] In the embodiment of the present invention, the processor calls a program or instruction stored in the memory, specifically, a program or instruction stored in an application program, and the processor is used to execute the method steps of the embodiment 1 of the present invention.
[0153] [Example 4]
[0154] This embodiment provides a chip for executing the method of the above-mentioned embodiment 1 of the present invention. Specifically, the chip includes: a processor for calling and running a computer program from a memory, so that a device equipped with the chip is used to execute the method of the embodiment 1 of the present invention.
[0155] [Example 5]
[0156] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method of Embodiment 1 of the present invention are implemented.
[0157] For example, machine-readable storage media may include, but are not limited to, various known and unknown types of non-volatile memory.
[0158] In summary, embodiments 1-5 of the present invention provide an application monitoring method and system, which constructs a time series model and a log detection model to track the processing flow of each request response in detail and detect its anomalies, which helps to quickly locate performance bottlenecks and fault points in the application, and enables users to monitor and ensure the stability of their own business goals and their own platforms, thereby improving application stability and user experience.
[0159] It will be appreciated by those skilled in the art that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0160] In the embodiments of the present application, the disclosed systems, devices and methods can be implemented in other ways. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system. In addition, the coupling between the various units can be direct coupling or indirect coupling. In addition, the various functional units in the embodiments of the present application can be integrated into a processing unit, or can be separate physical existences, etc.
[0161] It should be understood that in the various embodiments of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0162] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a machine-readable storage medium. Therefore, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a machine-readable storage medium, which can include several instructions to enable an electronic device to perform all or part of the technical solution described in the embodiments of the present application. The above-mentioned storage medium may include various media that can store program codes, such as ROM, RAM, removable disk, hard disk, magnetic disk or optical disk.
[0163] The above contents are only specific implementation methods of the present application, and the protection scope of the present application is not limited thereto. Those skilled in the art may make changes or substitutions within the technical scope disclosed in the present application, and these changes or substitutions shall be within the protection scope of the present application.
Claims
1. An application monitoring method, characterized in that: include: Get the time it takes for an application to request and respond to a server, and build a time series model; Perform anomaly detection on application metrics through the time series model; Obtain application log information and build a log detection model; Abnormal logs are detected by the log detection model.
2. The application monitoring method according to claim 1, characterized in that: The consumption time includes network delay time, client delay time, routing time, database processing time, cache processing time, external service call time and response generation and return time.
3. The application monitoring method according to claim 1, characterized in that: The application metrics include throughput, response time, request error rate, CPU usage, and memory usage as a function of elapsed time.
4. The application monitoring method according to claim 1, characterized in that: The constructing of the time series model comprises: The time series model is constructed by the exponential moving average method, and the time series model is represented by the following formula 1 and the following formula 2: S(t)=aX(t)+(1-a)S(t-1) Formula 1; Among them, S(t) is the smoothed value, a is the weighting coefficient, X(t) is the time series, and S(t-1) is the original smoothed value.
5. The application monitoring method according to claim 1, characterized in that: The log detection model construction includes: The log detection model is constructed by a clustering algorithm, and the log detection model is represented by the following formula 3: d(x,c)>threshold Equation 3; Where x is the log item, c is the cluster center, d(x,c) is the distance between x and c, and threshold is the threshold determined based on the clustering result.
6. The application monitoring method according to claim 1, characterized in that: The method comprises: When a single request from an application spreads responses across multiple services, the performance bottleneck of the service is expressed by the following equation 4: Among them, T i For Service i The processing time, C i→j For Service i Calling Service S j Time, C i→j For Service i With Service j The calling relationship between them.
7. An application monitoring system, characterized in that: include: The consumption time acquisition module is used to obtain the consumption time of the application request response server; Time series model building module, used to build time series models; An abnormal indicator detection module, used to perform abnormality detection on application indicators through the time series model; Log acquisition module, used to obtain application log information; Log detection model building module, used to build a log detection model; The abnormal log detection module is used to detect abnormal logs through the log detection model.
8. An application monitoring device, comprising at least one processor; a memory coupled to the at least one processor, the memory storing executable instructions, characterized in that: The executable instructions, when executed by the at least one processor, enable the steps of the method according to any one of claims 1 to 6 to be implemented.
9. A chip, characterized in that: It comprises a processor, which is used to call and run a computer program from a memory, so that a device equipped with the chip executes the steps of the method as claimed in any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.