Log monitoring method and device, storage medium and electronic equipment
By extracting the log data of the target application feature and determining the dynamic threshold interval, using the 3σ principle of statistical process control and the SPCC algorithm, the problem of low log monitoring accuracy is solved, real-time and accurate log monitoring is achieved, and the system stability and operation and maintenance efficiency is improved.
Patent Information
- Application Number
- CN202510442118.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
AI Technical Summary
The log monitoring methods in the prior art have problems with low accuracy, especially in Internet trading systems, where the fixed-configured alarm threshold cannot fully cover various abnormal situations in the system, resulting in false alarms, missed alarms or lagged alarms.
By obtaining the current log data of the target application, performing feature extraction, dynamically determine the feature threshold interval based on historical feature data, and using the 3σ principle of statistical process control and the SPCC algorithm to determine whether the current feature data is within the threshold interval, realizing abnormal monitoring.
Real-time monitoring of log data abnormalities is achieved, the accuracy and efficiency of log monitoring is improved, false alarms and missed reports are reduced, and the stability and reliability of the system are enhanced.
Smart Images

Figure CN120371638A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a log monitoring method, device, storage medium and electronic device. Background Art
[0002] In the related art, the log alarm method generally monitors based on dimensions such as log keywords and the success rate of database transaction orders. When key indicators such as the number of keywords and the success rate value exceed the alarm threshold, the monitoring system triggers an alarm. The alarm method in the related art configures alarm indicators (success rate, keywords), scans based on a timed task, and configures the scan period of the timed task (generally in minutes, hours, or days according to the business dimension) and the monitoring threshold. Since the transactions of the Internet trading system occur in real time and are affected by various factors such as time, business, and banks, the fixed configuration of the alarm threshold cannot fully cover all abnormal situations of the system, and problems such as false alarms, missed alarms, or delayed alarms often occur.
[0003] In view of the problem of low accuracy of log monitoring in the log monitoring method in the above related art, no effective solution has been proposed yet. Summary of the Invention
[0004] Embodiments of the present invention provide a log monitoring method, device, storage medium and electronic device to at least solve the technical problem of low accuracy of log monitoring in the log monitoring method in the related art.
[0005] According to an aspect of an embodiment of the present invention, a log monitoring method is provided, including: obtaining current log data of a target application at the current moment; extracting features from the current log data to obtain current feature data of the target application; determining a feature threshold interval based on historical feature data of the target application in a target historical period, where the historical feature data is obtained by extracting features from historical log data of the target application in the target historical period, and the target historical period is a period before the current moment; determining whether the current feature data is within the feature threshold interval to obtain a determination result; and determining a monitoring result of the current log data according to the determination result, where the monitoring result is used to indicate whether the current log data is abnormal.
[0006] According to another aspect of the embodiments of the present invention, there is also provided a log monitoring device, including: a data acquisition module, configured to acquire current log data of a target application at the current moment; a feature extraction module, configured to perform feature extraction on the current log data to obtain current feature data of the target application; an interval determination module, configured to determine a feature threshold interval based on historical feature data of the target application in a target historical period, where the historical feature data is obtained by performing feature extraction on historical log data of the target application in the target historical period, and the target historical period is a period before the current moment; a judgment module, configured to judge whether the current feature data is within the feature threshold interval to obtain a judgment result; and a log monitoring module, configured to determine a monitoring result of the current log data according to the judgment result, where the monitoring result is used to indicate whether the current log data is abnormal.
[0007] According to another aspect of the embodiments of the present invention, there is also provided a non-volatile storage medium storing multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform any one of the log monitoring methods.
[0008] According to another aspect of the embodiments of the present invention, there is also provided an electronic device including one or more processors and a memory, where the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement any one of the log monitoring methods.
[0009] In the embodiments of the present invention, by acquiring current log data of a target application at the current moment; performing feature extraction on the current log data to obtain current feature data of the target application; determining a feature threshold interval based on historical feature data of the target application in a target historical period, where the historical feature data is obtained by performing feature extraction on historical log data of the target application in the target historical period, and the target historical period is a period before the current moment; judging whether the current feature data is within the feature threshold interval to obtain a judgment result; and determining a monitoring result of the current log data according to the judgment result, where the monitoring result is used to indicate whether the current log data is abnormal, the purpose of automatically collecting and analyzing current and historical log data of the target application, accurately performing log monitoring by using feature extraction and dynamically calculating the feature threshold interval is achieved, thereby realizing the technical effect of real-time monitoring of abnormal log data, improving the log monitoring efficiency and monitoring accuracy, and further solving the technical problem of low log monitoring accuracy existing in the log monitoring methods in the related art. Description of the Drawings
[0010] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0011] Figure 1 is a flowchart of a log monitoring method according to an embodiment of the present invention;
[0012] Figure 2 is a flowchart of an alternative log monitoring method according to an embodiment of the present invention;
[0013] Figure 3 is an alternative system framework diagram according to an embodiment of the present invention;
[0014] Figure 4 is another alternative system framework diagram according to an embodiment of the present invention;
[0015] Figure 5 is a schematic diagram of a log monitoring device according to an embodiment of the present invention. Detailed Embodiments
[0016] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0017] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0018] First, for the convenience of understanding the embodiments of the present invention, some terms or nouns involved in the present invention will be explained below:
[0019] Kafka middleware is a distributed message queue service based on the publish-subscribe model. It was originally designed to handle real-time data streams and is now widely used in various scenarios such as data pipelines, stream processing, log collection, and website activity tracking.
[0020] According to an embodiment of the present invention, there is provided an embodiment of a log monitoring method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0021] Figure 1 It is a flowchart of the log monitoring method according to an embodiment of the present invention. As Figure 1 shown, the method includes the following steps:
[0022] Step S102, obtain the current log data of the target application at the current moment;
[0023] Optionally, the target application can be any business application in an Internet application system (such as in a payment business system). The current log data can be collected based on a message middleware. The message middleware can be but is not limited to Kafka middleware. The message middleware is used to collect log data from each application in the Internet application system and transmit it to the log processing system. The Internet application system has a distributed architecture, and data generation points are distributed on different servers. The message middleware provides a unified data collection entry to ensure the integrity and timeliness of log data.
[0024] Step S104, perform feature extraction on the current log data to obtain the current feature data of the target application;
[0025] Optionally, feature extraction is used to extract log keywords (i.e., key features) in the current log data, so as to subsequently perform targeted log monitoring on the target application based on the feature data of the target application.
[0026] In an alternative embodiment, feature extraction is performed on the current log data to obtain the current feature data of the target application, including: extracting key features from the current log data according to a preset log cleaning rule to obtain the current feature data; wherein the key features include at least one of the following: interface success rate, response time, request frequency; the interface success rate is used to indicate the ratio of the number of requests successfully responded by the application programming interface of the application to the total number of requests within a predetermined time; the response time is used to indicate the time from when a request is sent from the user side to when the application system completes processing and returns a response; the request frequency is used to indicate the number of requests sent by the user side to the application within a unit time; different key features correspond to different feature threshold intervals.
[0027] Optionally, before extracting the log data, preset log cleaning rules are adopted to ensure that the extracted feature data is clean, standardized, and valuable for analysis. These rules can include, but are not limited to, removing irrelevant noise information, formatting log times, separating request and response data, etc., so as to provide high-quality data input for subsequent feature extraction. The interface success rate is used to measure the ratio of the number of requests successfully responded to by the application interface to the total number of requests within a predetermined time. It can be calculated by counting the number of successes and failures in interface calls. It is crucial for monitoring the health status of the application, especially in a high-traffic Internet payment system, where fluctuations in the interface success rate can quickly reflect potential problems in the system. The response time indicates the time from when a request is sent from the user side to when the application system finishes processing and returns a response. The response time is an important indicator of system performance. Especially when processing a large number of real-time transactions, an increase in the response time may mean that the system is overloaded or there are potential faults. The request frequency is used to indicate the number of requests sent from the user side to the application within a unit time. High-frequency requests can lead to excessive consumption of system resources or reflect abnormal user behavior. Different key features correspond to different feature threshold intervals, which means that the system can set corresponding monitoring rules and anomaly detection thresholds according to different types of feature data (interface success rate, response time, request frequency). For example, the feature threshold interval for response time may focus on identifying performance degradation, while the feature threshold interval for request frequency is more concerned with detecting abnormal traffic patterns. By setting independent feature threshold intervals for different key features, it is possible to more flexibly adapt to different business scenarios and monitoring requirements. For example, during peak payment periods, the threshold interval for the interface success rate may need to be relaxed to avoid misjudging normal high-load situations as anomalies; while during off-peak hours, the threshold interval for request frequency may need to be more stringent to quickly identify potential attack behaviors. The extraction and analysis of key features are based on actual log data, which enables the system to make anomaly detections based on the real performance of the data, rather than relying on static rules or thresholds. By observing changes in key feature data, system operation and maintenance can more scientifically and accurately judge the running state of the application and promptly discover and handle abnormal situations.
[0028] In the above way, key feature data is extracted by presetting log cleaning rules to more meticulously and comprehensively monitor the running state of the application. Different key features correspond to different feature threshold intervals, and this mechanism can ensure the dynamism and adaptability of log monitoring rules, accurately detect anomalies in various business scenarios, and improve the stability and security of Internet application systems.
[0029] Step S106: Determine a characteristic threshold interval based on the historical characteristic data of the target application during the target historical period, where the historical characteristic data is obtained by extracting characteristics from the historical log data of the target application during the target historical period, and the target historical period is the period before the current moment.
[0030] It should be noted that the monitoring systems in related technologies often rely on fixed thresholds, which may not accurately reflect the dynamic changes of the business. By analyzing the characteristics of historical log data, a dynamic threshold interval based on historical performance can be obtained. This interval can adjust itself according to the fluctuations of the business. For example, during the peak business period, the threshold interval may be widened, while during the low business period, it may be stricter, so as to more accurately reflect the normal operating state of the system at different time points. Moreover, the dynamic characteristic threshold interval is based on the statistical analysis of historical data, which can more intelligently identify real anomalies, reduce false alarms and missed alarms, and improve the accuracy and reliability of alarms. In addition, the Internet business scenario is complex and changeable, and fixed rules are difficult to cover all abnormal situations. By extracting historical characteristic data and dynamically adjusting the threshold interval, it can better adapt to different types and scales of business scenarios, and improve the universality and flexibility of the log monitoring solution.
[0031] In an optional embodiment, when the historical log data includes the characteristic data corresponding to multiple historical sampling moments within the target historical period, determining the characteristic threshold interval based on the historical characteristic data of the target application during the target historical period includes: determining the statistical parameters of the characteristic data corresponding to the multiple historical sampling moments respectively, where the statistical parameters include the mean and the standard deviation; determining the characteristic threshold interval based on the mean and the standard deviation.
[0032] Optionally, the characteristic data corresponding to the multiple historical sampling moments respectively is obtained by extracting characteristics from the historical log data corresponding to the multiple sampling moments respectively, and the specific implementation method is the same as the method for obtaining the current characteristic data, which will not be elaborated here.
[0033] Optionally, for the characteristic data of each historical sampling moment, calculate its mean and standard deviation, and these means and standard deviations together constitute the statistical profile of the historical characteristic data. The mean (μ) is the arithmetic average of the characteristic data of all historical sampling moments within the target historical period, which can reflect the average level of the characteristic data of the target application during the target historical period. The standard deviation (σ) is an index to measure the degree of dispersion of the distribution of historical characteristic data, which can reflect the average difference between the data points and the mean, and more intuitively represent the fluctuation size of the data. By combining the mean and variance to dynamically determine the characteristic threshold interval, the dynamically determined characteristic threshold interval can more accurately reflect the range of the normal operation of the system compared with the static threshold, thereby reducing false alarms and missed alarms.
[0034] In an alternative embodiment, a feature threshold interval is determined based on the mean and the standard deviation, including: determining the feature threshold interval based on the mean and the standard deviation as: [μ - 3σ, μ + 3σ], where μ represents the mean and σ represents the standard deviation.
[0035] Optionally, it is possible but not limited to determining the feature threshold interval by using the Statistical Process Control Chart with Confidence Intervals (SPCC) algorithm based on the statistical process control, for the feature data corresponding to multiple historical sampling moments within the target historical period. This process involves the calculation of the mean μ and the standard deviation σ, and the determination of the feature threshold interval is based on the determined mean μ and standard deviation σ. The feature threshold interval is determined based on the 3σ principle. The 3σ principle means that for a dataset with a normal distribution, approximately 99.73% of the data points will fall within three standard deviations (3σ) of the mean μ. The probability distribution of the 3σ interval and the probability distribution of normal data within it are as follows:
[0036] The probability of 1σ, P(μ - σ < X ≤ μ + σ) = 68.26%;
[0037] The probability of 2σ, P(μ - 2σ < X ≤ μ + 2σ) = 95.44%;
[0038] The probability of 3σ, P(μ - 3σ < X ≤ μ + 3σ) = 99.74%.
[0039] As long as the sample is large enough, the interval calculated according to [μ - 3σ, μ + 3σ] can cover 99.74% of the normal data. When the data is not within this interval, it can be considered abnormal data, thus triggering an alarm.
[0040] It should be noted that any characteristic data falling outside the interval [μ - 3σ, μ + 3σ] has a relatively high probability of being considered abnormal, and this interval then becomes the characteristic threshold interval. Based on the mean μ and the standard deviation σ, the interval [μ - 3σ, μ + 3σ] can dynamically reflect the range of characteristic data when the target application is operating normally during the target historical period. With the addition of new historical data, the mean and the standard deviation will also be updated accordingly, thereby dynamically adjusting the characteristic threshold interval to more accurately reflect the current state of the system. Through the interval [μ - 3σ, μ + 3σ], the current characteristic data can be quickly compared with the historical characteristic data to determine whether it is abnormal. This interval not only covers almost all normal data points but also can sensitively capture abnormal fluctuations, thus improving the efficiency of anomaly detection while ensuring a low false alarm rate. Since the characteristic threshold interval is based on the statistical analysis of historical data, it can better adapt to the fluctuations of the business. For example, during the peak business period, the mean and the standard deviation of the characteristic data may be higher than usual, and the dynamically adjusted threshold interval can prevent misreporting normal high-load situations as anomalies. Similarly, during the low period, the threshold interval can be more stringent to reduce missed reports. By automatically calculating the characteristic threshold interval, system operation and maintenance can greatly simplify the alarm configuration process, reducing the workload and potential errors caused by manually setting the threshold. This automated configuration mechanism can improve the overall operation and maintenance efficiency and reliability of the system.
[0041] In an alternative embodiment, before determining the characteristic threshold interval based on the historical characteristic data of the target application during the target historical period, the method further includes: obtaining the characteristic information corresponding to the target application, where the characteristic information includes at least one of the following: the data volume and data quality of the characteristic data of the target application stored in the log database, the business cycle and business scenario corresponding to the target application; determining the target historical period based on the characteristic information; and obtaining the historical log data of the target application during the target historical period from the log database.
[0042] Optionally, each application may exhibit different normal operation characteristics due to differences in its business cycle, business scenario, data volume, and data quality. By obtaining this characteristic information, personalized monitoring strategies can be developed for each application, ensuring that the monitoring rules are closer to its specific operating conditions and improving the accuracy of anomaly detection. Selecting an appropriate target historical period is crucial for determining a reasonable characteristic threshold range. For example, the business cycle may mean that at certain time points or dates, there are expected fluctuations in the interface success rate, response time, and request frequency of the application. By considering these periodic changes, a historical period that represents the normal operating state of the application can be selected, thereby more accurately calculating the characteristic threshold range. The size of the data volume directly affects the reliability and effectiveness of statistical analysis. A sufficiently large data volume can reduce the impact of random fluctuations on the statistical results; while high data quality means the accuracy and integrity of the data, which is crucial for data-based anomaly detection. By evaluating the data volume and data quality of the characteristic data stored in the log database, it can be ensured that the data used to calculate the characteristic threshold range is sufficient and reliable. By comprehensively considering the business cycle, business scenario, data volume, and data quality of the target application, anomalies can be identified more accurately, avoiding false positives and false negatives caused by insufficient or poor-quality data. At the same time, the target historical period determined based on the characteristic information can ensure that the calculation of the threshold range is more scientific, improve the robustness of the algorithm, and enable it to operate effectively in different business environments. In the above ways, by considering the characteristic information of the target application to optimize the monitoring strategy and select the most suitable analysis period, the accuracy and sensitivity of the log monitoring method can be improved.
[0043] Optionally, data quality can be determined by the number of missing values or the proportion of missing values. If there are abnormal fluctuations or missing values in the data within the past 30 days, a longer time period can be traced back, such as the past 60 days or 90 days, to avoid the impact of abnormal data on statistical analysis. Alternatively, abnormal days can be automatically filtered out, and only the data of normal days can be used to calculate the statistic. For businesses with periodic characteristics, such as certain payment businesses having significant traffic changes on specific dates (such as paydays, weekends, holidays), the target historical period should be selected as the time period with the same business cycle as the current moment. For example, if it is Friday currently, the target historical period may select the Friday data of the past 4 weeks to obtain a threshold range reflecting the business cycle. Different business scenarios may require different monitoring strategies. For example, in special scenarios such as the launch of a new product or a promotional activity, the target historical period may need to be adjusted to the most similar past period to more accurately reflect the normal operation range of the current business.
[0044] In an alternative embodiment, when the feature information includes data volume, data quality, business cycle, and business scenario, based on the feature information, determining a target historical period includes: determining a first weight value corresponding to the data volume, a second weight value corresponding to the data quality, a third weight value corresponding to the business cycle, and a fourth weight value corresponding to the business scenario; performing a weighted calculation based on the data volume, data quality, business cycle, business scenario, first weight value, second weight value, third weight value, and fourth weight value to obtain a target score value; determining the target score interval to which the target score value belongs; based on the correspondence between the score interval and the duration, determining the target duration corresponding to the target score interval; and determining the target historical period based on the target duration.
[0045] Optionally, by determining the weight values of the data volume, data quality, business cycle, and business scenario, a quantitative assessment of the importance of different features can be made. This enables a comprehensive consideration of the impact of various factors on data representativeness when selecting the target historical period, rather than just a single dimension (such as time length). A weighted calculation is performed based on the above four features and their corresponding weight values to generate a target score value. This score value can comprehensively reflect the overall health status of the target application in historical data and is a key basis for selecting the target historical period. The target score value is divided into different score intervals, and each score interval corresponds to a preset target duration. The setting of this correspondence relationship takes into account the characteristics of various application programs in different business scenarios and can automatically select the optimal target historical period length according to the score interval for calculating the feature threshold interval. By determining the target historical period based on the score interval, a historical data set that best matches the current business state can be selected for analysis, thereby improving the accuracy of calculating the feature threshold interval and reducing false positives and false negatives in anomaly detection. The weighted scoring mechanism enables the system to flexibly adapt to changes in different business cycles and scenarios. Even during peak business periods or in special scenarios, the most appropriate historical period can be selected for analysis based on the comprehensive quality of the data to ensure the adaptability and effectiveness of the log monitoring rules.
[0046] Step S108, determining whether the current feature data is within the feature threshold interval to obtain a determination result;
[0047] Optionally, by comparing the current feature data with the feature threshold interval calculated based on historical data, it can be quickly identified whether the current data deviates from the normal range. If the current feature data falls outside the threshold interval, it is regarded as an anomaly, which can promptly reflect possible problems with the target application, such as performance degradation, increased errors, etc., thereby triggering the subsequent anomaly handling process.
[0048] Step S110, determine the monitoring result of the current log data according to the judgment result, where the monitoring result is used to indicate whether there is an abnormality in the current log data.
[0049] Optionally, based on the judgment result, it can be immediately confirmed whether the current log data is indeed abnormal. If the current feature data exceeds the pre-defined feature threshold range, this will trigger an abnormal monitoring result, and the abnormal information can be promptly communicated to the operation and maintenance personnel or relevant responsible persons in the form of emails, text messages, application notifications, etc., to ensure that they can respond quickly and conduct fault troubleshooting and system repair. The monitoring result is an important basis for the operation and maintenance personnel to make decisions. It not only indicates that there is an abnormality in the log data, but also provides the specific nature of the abnormality (such as performance degradation, frequent errors, etc.) and the severity of the abnormality, helping the operation and maintenance personnel quickly judge the priority and decide whether to intervene immediately or can be processed later.
[0050] Optionally, the abnormal monitoring result will be recorded and become the log analysis and fault tracing materials. These records can be used for subsequent fault analysis, helping to understand the reasons and patterns of the abnormality occurrence, and providing reference for preventing and handling similar problems in the future. Through the feedback of the monitoring result, continuous self-adjustment and optimization can be carried out. For example, if it is found that certain specific types of log abnormalities occur frequently, the feature extraction strategy can be adjusted or the feature threshold range can be further refined to more accurately identify and handle these abnormalities.
[0051] In an optional embodiment, determining the monitoring result of the current log data according to the judgment result includes: when the judgment result indicates that the current log data is within the feature threshold range, determining the monitoring result as: there is no abnormality in the current log data; or when the judgment result indicates that the current log data is not within the feature threshold range, determining the monitoring result as: there is an abnormality in the current log data.
[0052] Optionally, the characteristic threshold interval calculated by the SPCC algorithm can be used to determine whether the current log data deviates from the historical normal range. When the data is within the interval of [μ - 3σ, μ + 3σ], it is determined that there is no abnormality in the current log data. This is a judgment based on statistical principles, which can greatly reduce false alarms and ensure the accuracy of anomaly detection. If the current log data falls outside the characteristic threshold interval, it is determined that the current log data is abnormal, and the corresponding monitoring result is triggered, that is, there is an anomaly. This immediate response can prevent small problems from evolving into major failures, thereby reducing the negative impacts that may be caused by anomalies not being processed in a timely manner. Through automated judgment, the dependence on manual monitoring can be reduced, the workload of operation and maintenance personnel and the monitoring cost can be lowered. Operation and maintenance personnel can concentrate on dealing with truly abnormal situations rather than needlessly checking a large amount of normal data, improving work efficiency. Accurate anomaly detection can also timely discover and handle potential problems in the system, avoid the spread of failures, and thus enhance the stability of the entire system. Especially for critical services such as payment systems, this can greatly improve the user experience and reduce user complaints and trust crises caused by system failures.
[0053] Through the above steps S102 to S110, the purpose of automatically collecting and analyzing the current and historical log data of the target application program, using feature extraction and dynamically calculating the characteristic threshold interval, and accurately performing log monitoring can be achieved, so as to realize real-time monitoring of log data anomalies, improve the efficiency and accuracy of log monitoring, and further solve the technical problem of low log monitoring accuracy existing in the related log monitoring methods.
[0054] Based on the above embodiments and optional embodiments, the present invention proposes an optional implementation manner. Figure 2 It is a flowchart of an optional log monitoring method according to an embodiment of the present invention. Figure 3 It is an optional system framework diagram according to an embodiment of the present invention. Figure 4 It is another optional system framework diagram according to an embodiment of the present invention. This method can be applied to system frameworks such as Figure 3 and Figure 4 As shown in Figure 2 The method includes:
[0055] S1. Standardize the log output format, formulate the "Monitoring Log Specification", and require each application system (i.e., application program) to strictly output according to the log specification when printing application logs. The standardized log format is the basis for subsequent execution of the confidence interval monitoring algorithm SPCC algorithm for calculating the alarm threshold based on statistical process control. Among them:
[0056] The interface request message format is as follows:
[0057] [Date and Time][Log Level][Log Print Code Location][Thread ID]traceLogid:[traceLogid]dstTraceId:[dstID]call[Class Name][Method Name]PARAMETER:[Interface Request Message];
[0058] Among them, the interface request message format is a standardized format used to describe and record detailed information during the invocation of application program interfaces. The following are the explanations for each part in the format:
[0059] [Date and Time] represents the specific time when the interface request occurred, usually in the format of year - month - day hour:minute:second, and is used to trace and locate specific requests or problems.
[0060] [Log Level] represents the importance or type of this log message. Common log levels include DEBUG (debugging), INFO (information), WARNING (warning), ERROR (error), FATAL (fatal error), etc. This helps filter and prioritize log messages.
[0061] [Log Print Code Location] represents the specific code file and line number, identifying the code location where this log message is generated. This is very helpful for developers to debug problems and can quickly find the relevant code for inspection.
[0062] [Thread ID] is used to identify the specific thread handling this interface request. In a multi - threaded environment, this is key information for tracing the request processing flow.
[0063] The traceLog id is a globally unique identifier used to trace the entire lifecycle of the request, from request sending to response receiving. This is particularly important for tracing cross - service request flows in a distributed system.
[0064] dstTraceId:[dstID] identifies the trace ID of the destination service or component. In a microservices architecture, a request may pass through multiple services, and the dstTraceId helps trace the transfer of the request between different services.
[0065] call[Class Name][Method Name] represents the class and specific method name being called, providing context information for request processing and facilitating understanding of the request processing logic and location.
[0066] PARAMETER:[Interface Request Message] contains the specific parameter information of the interface request. This is the core content of the interface call and is crucial for analyzing the type, purpose of the request, and whether it is processed as expected.
[0067] The interface return message format is as follows:
[0068] [Date and Time][Log Level][Log Print Code Location][Thread ID]traceLogid:[traceLogid]dstTraceId:[dstID]call[Class Name][Method Name][Service Call Duration][Interface Call Status][Response Code]RESPONSE:Result{Interface Return Message};
[0069] Among them, RESPONSE is a keyword used to identify that the following content is the response information of the interface call. It is relative to the request, indicating that the system is feeding back or returning the processing result of the previously received interface request. Result means "result", and in the context of the log, it refers to the specific output or result data generated by the interface call.
[0070] S2, the application system pushes it to the log processing system through the message middleware according to the log specification, and the log processing system saves the log interface data to the log database according to the log collection and log cleaning rules.
[0071] S3, the SPCC algorithm takes the log collection records in the log database for analysis. According to the data collected in step S3, applying the SPCC algorithm, calculate the average value μ of the success rate of a certain API interface in the past 30 days. The accuracy rate that the current success rate of the API interface falls within the interval [μ - 3σ, μ + 3σ] is as high as 99.74%. When the current success rate of the API of the application system is not within this interval range, it can be considered that the API system corresponding to the application has an abnormality and an alarm is issued.
[0072] This embodiment can achieve at least one of the following effects: 1) Realize intelligent alarm triggering through the algorithm, completely eliminating situations such as inaccurate and omitted business configuration rules caused by human intervention. 2) Greatly save personnel costs. Configuration personnel do not need to master the business logic for configuration, and the monitoring configuration function can be uniformly centralized to the monitoring department without the need for each business department to configure. 3) Fast alarm response. Based on the Kafka-based log collection system, it can achieve quasi-real-time log collection to the greatest extent. After actual combat testing, the log collection delay is within 1 minute. 4) High accuracy. Through mathematical calculation, the accuracy rate of the alarm triggered based on the SPCC algorithm is 99.74%. 5) The ability to predict faults. The SPCC algorithm can predict possible faults in the system. Since the SPCC algorithm is a calculation method based on historical data, it realizes intelligent prediction of faults through it. For example, when the data calculated by the SPCC for a certain API interface falls between [0.4, 0.8], if the current success rate of this interface is 0.38, it is already a precursor to possible problems in the system. If the success rate continues to deteriorate, it means that this API interface is already at the margin of being unable to handle business. At this time, issuing an alarm has the ability to predict faults.
[0073] In this embodiment, a log monitoring device is further provided. This device is used to implement the above-mentioned embodiment and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, the terms "module" and "device" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0074] According to an embodiment of the present invention, an apparatus embodiment for implementing the above-mentioned log monitoring method is further provided. Figure 5 It is a schematic structural diagram of a log monitoring device according to an embodiment of the present invention. As Figure 5 shown, the above-mentioned log monitoring device includes: a data acquisition module 500, a feature extraction module 502, an interval determination module 504, a judgment module 506, and a log monitoring module 508, where:
[0075] The data acquisition module 500 is used to acquire the current log data of the target application at the current moment.
[0076] The feature extraction module 502 is connected to the data acquisition module 500 and is used to extract features from the current log data to obtain the current feature data of the target application.
[0077] The interval determination module 504 is connected to the feature extraction module 502 and is used to determine a feature threshold interval based on the historical feature data of the target application in the target historical period, where the historical feature data is obtained by extracting features from the historical log data of the target application in the target historical period, and the target historical period is the period before the current moment.
[0078] The judgment module 506 is connected to the interval determination module 504 and is used to judge whether the current feature data is within the feature threshold interval to obtain a judgment result.
[0079] The log monitoring module 508 is connected to the judgment module 506 and is used to determine the monitoring result of the current log data according to the judgment result, where the monitoring result is used to indicate whether there is an abnormality in the current log data.
[0080] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following manner: the above-mentioned various modules can be located in the same processor; or, the above-mentioned various modules are located in different processors in any combination.
[0081] It should be noted that the above data acquisition module 500, feature extraction module 502, interval determination module 504, judgment module 506, and log monitoring module 508 correspond to steps S102 to S110 in the embodiment. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above modules can run in a computer terminal as part of the device.
[0082] It should be noted that the optional or preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiment, and will not be elaborated here.
[0083] The above log monitoring device may further include a processor and a memory. The above data acquisition module 500, feature extraction module 502, interval determination module 504, judgment module 506, log monitoring module 508, etc. are all stored in the memory as program modules, and the processor executes the above program modules stored in the memory to implement corresponding functions.
[0084] The processor contains a kernel, and the kernel retrieves the corresponding program module from the memory. One or more kernels can be set. The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0085] According to an embodiment of the present application, an embodiment of a non-volatile storage medium is also provided. Optionally, in this embodiment, the above non-volatile storage medium includes a stored program, wherein when the above program runs, it controls the device where the non-volatile storage medium is located to execute any of the above log monitoring methods.
[0086] Optionally, in this embodiment, the above non-volatile storage medium can be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group. The above non-volatile storage medium includes a stored program.
[0087] Optionally, when the program is running, control the device where the non-volatile storage medium is located to perform the following functions: obtain the current log data of the target application at the current moment; extract features from the current log data to obtain the current feature data of the target application; determine the feature threshold interval based on the historical feature data of the target application in the target historical period, where the historical feature data is obtained by extracting features from the historical log data of the target application in the target historical period, and the target historical period is the period before the current moment; determine whether the current feature data is within the feature threshold interval to obtain a judgment result; determine the monitoring result of the current log data according to the judgment result, where the monitoring result is used to indicate whether the current log data is abnormal.
[0088] According to an embodiment of the present application, an embodiment of a processor is further provided. Optionally, in this embodiment, the above-mentioned processor is used to run a program, where the above-mentioned program performs any one of the above-mentioned log monitoring methods when running.
[0089] According to an embodiment of the present application, an embodiment of a computer program product is further provided. When executed on a data processing device, it is adapted to execute a program initialized with the steps of any one of the above-mentioned log monitoring methods.
[0090] Optionally, when the above-mentioned computer program product is executed on a data processing device, it is adapted to execute a program initialized with the following method steps: obtain the current log data of the target application at the current moment; extract features from the current log data to obtain the current feature data of the target application; determine the feature threshold interval based on the historical feature data of the target application in the target historical period, where the historical feature data is obtained by extracting features from the historical log data of the target application in the target historical period, and the target historical period is the period before the current moment; determine whether the current feature data is within the feature threshold interval to obtain a judgment result; determine the monitoring result of the current log data according to the judgment result, where the monitoring result is used to indicate whether the current log data is abnormal.
[0091] An embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, the following steps are implemented: obtaining current log data of a target application at the current moment; performing feature extraction on the current log data to obtain current feature data of the target application; determining a feature threshold interval based on historical feature data of the target application in a target historical period, where the historical feature data is obtained by performing feature extraction on historical log data of the target application in the target historical period, and the target historical period is a period before the current moment; determining whether the current feature data is within the feature threshold interval to obtain a determination result; and determining a monitoring result of the current log data according to the determination result, where the monitoring result is used to indicate whether there is an abnormality in the current log data.
[0092] The order of the above embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments.
[0093] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0094] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the above module division can be a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of modules or modules can be in an electrical or other form.
[0095] The modules described above as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0096] In addition, in each embodiment of the present invention, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0097] If the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable non-volatile storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a non-volatile storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned non-volatile storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.
[0098] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A log monitoring method, characterized in that, Including: Obtain the current log data of the target application at the current moment; Extract features from the current log data to obtain the current feature data of the target application; Based on the historical feature data of the target application in the target historical period, determine the feature threshold interval, where the historical feature data is obtained by extracting features from the historical log data of the target application in the target historical period, and the target historical period is the period before the current moment; Judge whether the current feature data is within the feature threshold interval to obtain a judgment result; Determine the monitoring result of the current log data according to the judgment result, where the monitoring result is used to indicate whether there is an abnormality in the current log data.
2. The method according to claim 1, characterized in that When the historical log data includes the feature data corresponding to multiple historical sampling moments within the target historical period, the determining the feature threshold interval based on the historical feature data of the target application in the target historical period includes: Determine the statistical parameters of the feature data corresponding to the multiple historical sampling moments, where the statistical parameters include the mean and the standard deviation; Based on the mean and the standard deviation, determine the feature threshold interval.
3. The method according to claim 2, wherein The determining the feature threshold interval based on the mean and the standard deviation includes: Based on the mean and the standard deviation, determine the feature threshold interval as: [μ - 3σ, μ + 3σ], where μ represents the mean and σ represents the standard deviation.
4. The method according to claim 1, characterized in that, The extracting features from the current log data to obtain the current feature data of the target application includes: Extract the key features in the current log data according to the preset log cleaning rules to obtain the current feature data; Wherein, the key features include at least one of the following: interface success rate, response time, request frequency; the interface success rate is used to indicate the ratio of the number of requests successfully responded by the application interface of the application to the total number of requests within a predetermined time; the response time is used to indicate the time from when the request is sent from the user side to when the application system completes processing and returns a response; the request frequency is used to indicate the number of requests sent by the user side to the application within a unit time; different key features correspond to different feature threshold intervals.
5. The method according to claim 1, characterized in that Before the determining the feature threshold interval based on the historical feature data of the target application in the target historical period, the method further includes: Obtain the feature information corresponding to the target application, where the feature information includes at least one of the following: the data volume and data quality of the feature data of the target application stored in the log database, the business cycle and business scenario corresponding to the target application; Based on the feature information, determine the target historical period; Obtain the historical log data of the target application in the target historical period from the log database.
6. The method according to claim 5, wherein When the feature information includes the data volume, the data quality, the business cycle and the business scenario, the determining the target historical period based on the feature information includes: Determine a first weight value corresponding to the data volume, a second weight value corresponding to the data quality, a third weight value corresponding to the business cycle, and a fourth weight value corresponding to the business scenario; Perform weighted calculation based on the data volume, the data quality, the business cycle, the business scenario, the first weight value, the second weight value, the third weight value, and the fourth weight value to obtain a target score value; Determine the target score interval to which the target score value belongs; Based on the correspondence between the score interval and the duration, determine a target duration corresponding to the target score interval; Based on the target duration, determine the target historical period.
7. The method according to any one of claims 1 to 6, characterized in that The determining the monitoring result of the current log data according to the judgment result includes: In the case where the judgment result indicates that the current log data is within the feature threshold interval, determining that the monitoring result is: the current log data has no anomaly; or In the case where the judgment result indicates that the current log data is not within the feature threshold interval, determining that the monitoring result is: the current log data has an anomaly.
8. A log monitoring device, characterized in that, Including: A data acquisition module, configured to acquire current log data of a target application at the current moment; A feature extraction module, configured to perform feature extraction on the current log data to obtain current feature data of the target application; An interval determination module, configured to determine a feature threshold interval based on historical feature data of the target application in a target historical period, where the historical feature data is obtained by performing feature extraction on historical log data of the target application in the target historical period, and the target historical period is a period before the current moment; A judgment module, configured to judge whether the current feature data is within the feature threshold interval to obtain a judgment result; A log monitoring module, configured to determine the monitoring result of the current log data according to the judgment result, where the monitoring result is used to indicate whether the current log data has an anomaly.
9. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform the log monitoring method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Including one or more processors and a memory, the memory is configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the log monitoring method according to any one of claims 1 to 7.
Citation Information
Cited By
Large packet anomaly detection method and device based on dynamic baseline, medium and product
CN121309161A
Code quality access control method and system based on statistical learning in AI coding scene
CN122064378A