Log data analysis method and system based on AI

Through AI-based log data analysis methods, error patterns are identified and theoretical markers for stack traces are defined. Abnormal stack trace analysis is screened out by combining time deviation and asynchronous keywords. This solves the problem of incomplete stack traces caused by the misordering of asynchronous log timestamps in traditional stack analysis, and improves the accuracy and efficiency of analysis.

CN120670252AInactive Publication Date: 2025-09-19SHANGHAI FENGSHEN NEW ENERGY TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510857939.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In a distributed microservices architecture, traditional stack trace analysis suffers from incomplete stack traces in cross-service call scenarios due to the misordering of asynchronous log timestamps. This leads to the inability to identify asynchronous events and the loss of exception stack information.

Method used

Through AI-based log data analysis methods, we obtain log content feature data, identify error patterns, and define the theoretical start and end markers of stack traces. We combine time deviation and asynchronous keyword recognition to filter out asynchronous log entries in the vicinity of abnormal stack traces. We then optimize the accuracy of the end point timestamp of stack trace analysis by setting up an abnormal optimization detection mechanism.

Benefits of technology

It achieves the accuracy of stack trace analysis and the integrity of exception information, reduces misjudgments caused by asynchronous log interference, and improves the accuracy and efficiency of fault root cause analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670252A_ABST
    Figure CN120670252A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of log data analysis, and provides an AI-based log data analysis method and system, and the method comprises the following steps: obtaining log content feature data, and judging whether to trigger stack tracking analysis or not based on error pattern recognition which comprises high-frequency error recognition and new error type recognition; if the stack tracking analysis is triggered, a corresponding error log entry is obtained, a theoretical starting mark and a theoretical ending mark of the stack tracking analysis are defined, an actual starting point and an actual ending point for executing the stack tracking analysis are obtained, consistency comparison is conducted on the theoretical ending mark and the actual ending point, and whether stack tracking is abnormal or not is judged; and if stack tracking is abnormal, identifying and screening out asynchronous log entries in an abnormal stack tracking adjacent area based on time deviation and asynchronous keywords, analyzing a correlation between stack tracking abnormity and the asynchronous log entries, and identifying a correlation degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of log data analysis, and specifically relates to an AI-based log data analysis method and system. Background Art

[0002] Log data refers to structured or unstructured text data automatically generated by information systems during operation, recording system behaviors and events. It is aggregated through multi-source collection mechanisms (such as application servers, databases, and middleware) and contains information such as operation records, state changes, and abnormal events. It typically includes key fields such as log level, timestamp, error type, and request identifier, and is stored in formats such as JSON, text, or Syslog. It is used to support core scenarios such as system operation and maintenance monitoring, fault root cause analysis, performance bottleneck identification, and security audits.

[0003] In a distributed microservices architecture, stack traces, as key diagnostic information that records program execution paths in complex distributed architectures, are the core basis for locating the root cause of exceptions. However, when log entries contain compound events in highly concurrency scenarios involving cross-service calls, traditional stack trace analysis techniques face significant challenges. When log entries contain asynchronous log fragments across service calls, stack trace analysis suffers from the following issues:

[0004] In cross-service call scenarios, asynchronous logs (such as those generated by message queues and delayed tasks) often have timestamp out-of-order due to network latency or batch processing mechanisms. Traditional stack trace analysis relies solely on theoretical timestamps to mark event boundaries. When asynchronous log timestamps are involved, it can easily trigger the identification of "pseudo-end markers," causing the stack trace to terminate prematurely and resulting in missing exception stack information. For example, in a microservice call chain, the log timestamp of an asynchronous task may be earlier than the theoretical end time. Traditional methods cannot identify such asynchronous events, resulting in incomplete stack traces.

[0005] To this end, the present invention provides an AI-based log data analysis method and system. Summary of the Invention

[0006] In order to make up for the deficiencies of the prior art, at least one technical problem raised in the background technology is solved.

[0007] The technical solution adopted by the present invention to solve the technical problem is: an AI-based log data analysis method, comprising the following steps:

[0008] Obtain log content feature data and determine whether to trigger stack trace analysis based on error pattern recognition, where error pattern recognition includes high-frequency error recognition and new error type recognition;

[0009] If stack trace analysis is triggered, obtain the corresponding error log entry, define the theoretical start and end markers for the stack trace analysis, obtain the actual start and end points for executing the stack trace analysis, compare the theoretical end marker with the actual end point for consistency, and determine whether the stack trace is abnormal;

[0010] If the stack trace is abnormal, asynchronous log entries in the vicinity of the abnormal stack trace are filtered out based on time deviation and asynchronous keyword recognition. The correlation between the stack trace abnormality and the occurrence of asynchronous log entries is analyzed to identify the degree of correlation.

[0011] If the correlation is strong, set up an exception optimization detection mechanism to achieve accurate detection of the end point timestamp through stack trace analysis;

[0012] The exception optimization detection mechanism includes setting a fixed time window and filtering the end point timestamp of the exception stack trace based on the timestamp distribution of the log entries of the historical normal stack trace.

[0013] As a further solution of the present invention: the process of error pattern recognition and judgment is:

[0014] Set a time window, use a counter to count the number of times the same error type occurs within the window, and calculate the ratio of the number of occurrences to the length of the time window to obtain the frequency of occurrence;

[0015] Extract confirmed error types from historical logs, build a historical whitelist database, and use Bloom filters to quickly determine whether the obtained standardized error types exist in the historical whitelist database.

[0016] When the occurrence frequency is greater than or equal to the frequency threshold or the error type does not exist in the historical whitelist library, stack trace analysis is triggered.

[0017] As a further solution of the present invention: the process of determining whether the stack trace is normal is:

[0018] Extract error log entries containing more than one event as log entries to be analyzed;

[0019] Define the theoretical start mark and theoretical end mark of each log entry to be analyzed. The theoretical start mark is the start timestamp of the abnormal event propagation chain of the log entry to be analyzed, and the theoretical end mark is the end timestamp of the abnormal event propagation chain of the log entry to be analyzed.

[0020] Perform stack tracing on each log entry to be analyzed, and obtain the actual start and end timestamps of the stack trace;

[0021] If the end timestamp defined in the log entry to be analyzed is inconsistent with the end timestamp of the stack trace, an exception occurs in the stack trace.

[0022] As a further solution of the present invention, the process of filtering out asynchronous log entries in the adjacent area of ​​the exception stack trace is as follows:

[0023] Get candidate log entries and scan whether they contain asynchronous processing keywords;

[0024] Set the regular expression basic mode; perform regular expression matching on each candidate log entry, and record the matching asynchronous processing keywords and their occurrence counts;

[0025] For each candidate log entry, calculate the sum of the occurrences of all matching asynchronous processing keywords as the keyword matching density;

[0026] If the keyword matching density is greater than the keyword matching density limit, it is marked as an asynchronous log entry.

[0027] As a further solution of the present invention: the process of obtaining the candidate log entries is as follows:

[0028] Filter log entries containing the same request ID within the maximum time difference from the timestamp of the exception stack trace end point as candidate log entries.

[0029] As a further solution of the present invention: the process of identifying the degree of correlation is:

[0030] Count the number of abnormal log entries with asynchronous log entries and calculate the ratio with the total number of abnormal log entries to obtain the percentage value;

[0031] If the quantity proportion value is greater than the quantity proportion threshold, the correlation is strong.

[0032] As a further solution of the present invention: the abnormal optimization detection mechanism includes:

[0033] Define a fixed time window, using the theoretical start timestamp and theoretical end timestamp of the current log entry to be analyzed as the fixed time window;

[0034] If there are multiple timestamps within the fixed time window, the timestamp filtering process is performed.

[0035] As a further solution of the present invention: the timestamp screening process includes:

[0036] Count several historical normal stack trace log entries, and extract the timestamp sequence and adjacent time interval of each historical normal stack trace log entry;

[0037] Calculate the historical adjacent interval average and historical adjacent interval standard deviation of all historical normal stack trace log entries; count the total time consumed by all historical normal stack trace log entries and perform mean calculation to obtain the historical average total time consumed;

[0038] Get multiple timestamps within a fixed window during the current stack trace analysis process, calculate all adjacent time interval values, and use them as the current adjacent time interval value;

[0039] Calculate the Euclidean distance between the current adjacent time interval value and the average value of the historical adjacent intervals, and convert the Euclidean distance value into a similarity value;

[0040] Calculate the deviation rate between the current total time and the historical average total time;

[0041] The value obtained by subtracting the deviation rate from 1 is weighted and fused with the similarity value to output a comprehensive score;

[0042] The timestamp with the highest comprehensive score is extracted as the timestamp of the end point.

[0043] As a further solution of the present invention: the calculation process of the similarity value and the deviation rate is:

[0044] Calculate the ratio of the Euclidean distance value to the standard deviation of the historical adjacent intervals, and then divide the ratio by 1 to obtain the similarity value;

[0045] The difference between the current total time and the historical average total time is calculated and the absolute value is taken. The difference is then compared with the historical average total time to obtain the deviation rate.

[0046] An AI-based log data analysis system, comprising:

[0047] Stack trace trigger module: This module obtains log content feature data and determines whether to trigger stack trace analysis based on error pattern recognition. Error pattern recognition includes high-frequency error recognition and new error type recognition.

[0048] Stack trace judgment module: If stack trace analysis is triggered, obtain the corresponding error log entry, define the theoretical start mark and theoretical end mark of the stack trace analysis, obtain the actual start point and actual end point of the stack trace analysis, compare the theoretical end mark and the actual end point for consistency, and determine whether the stack trace is normal;

[0049] Asynchronous log analysis module: If a stack trace is abnormal, it filters out asynchronous log entries in the vicinity of the abnormal stack trace based on time deviation and asynchronous keyword recognition. It then counts the number of asynchronous log entries, analyzes the correlation between the stack trace abnormality and the occurrence of asynchronous log entries, and identifies the degree of correlation.

[0050] Abnormal optimization module: If the correlation is strong, set up abnormal optimization detection mechanism to optimize the accuracy of stack trace analysis for end point timestamp detection;

[0051] The exception optimization detection mechanism includes setting a fixed time window and filtering the end point timestamp of the exception stack trace based on the timestamp distribution of the log entries of the historical normal stack trace.

[0052] The beneficial effects of the present invention are as follows:

[0053] 1. This invention collects logs from multiple sources and extracts key fields. It uses Redis counters to count the frequency of errors within a time window to identify high-frequency errors. It then uses Bloom filters to match new error types with a historical whitelist database. When trigger conditions are met, it performs stack trace analysis, reducing ineffective analysis of normal logs and improving the pertinence and efficiency of problem location.

[0054] 2. This invention defines the theoretical start and end markers of the log entry to be analyzed and compares them with the start and end points of the actual stack trace to accurately determine whether the stack trace is normal. This solves the problem of premature termination of stack traces due to the inability to identify event boundaries, increases the integrity of exception stack information, and provides comprehensive and accurate data support for fault root cause analysis.

[0055] 3. The present invention filters asynchronous log entries based on time deviation and asynchronous keyword recognition, quantifies the correlation strength between asynchronous operations and stack exceptions by the number ratio, defines a fixed time window when the correlation is strong, and filters the end point timestamp based on the timestamp distribution of historical normal stack traces. This eliminates asynchronous log interference, optimizes the problem of end point misjudgment, improves the accuracy of stack trace analysis, and enables the complete capture of abnormal stack information. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The present invention will be further described below with reference to the accompanying drawings.

[0057] Figure 1 This is a flowchart of the steps of an AI-based log data analysis method of the present invention;

[0058] Figure 2 This is a flowchart of step 1 in an AI-based log data analysis method of the present invention;

[0059] Figure 3 This is an architecture diagram of an AI-based log data analysis system of the present invention. DETAILED DESCRIPTION

[0060] In order to make the technical means, creative features, objectives and effects achieved by the present invention easier to understand, the present invention is further described below in conjunction with specific implementation methods.

[0061] Example 1

[0062] See also Figure 1 and Figure 2 As shown, an AI-based log data analysis method according to an embodiment of the present invention includes the following steps:

[0063] Step 1: Obtain log content feature data and determine whether to trigger stack trace analysis based on error pattern recognition. Error pattern recognition includes high-frequency error recognition and new error type recognition.

[0064] Specifically, we first need to obtain log content feature data. The process is as follows:

[0065] Collect logs from multiple sources including application servers, databases, and middleware through the Agent, supporting at least JSON, text, and Syslog formats;

[0066] Log content feature data includes at least: log level, error type, error message, timestamp, request ID, log source, and module ID;

[0067] For content feature extraction, we input log content and use log parsing tools to extract key fields. We then map error types to unified identifiers for easy subsequent analysis.

[0068] For example, the original error type (such as "java.lang.NullPointerException") is mapped to a unified identifier (such as NULL_PTR_EXCEPTION);

[0069] Key fields are extracted through Grok pattern matching. Key fields include at least: log level, error type, error message, timestamp, and request ID.

[0070] Specifically, high-frequency errors are identified by counting errors per unit time. The process is as follows:

[0071] Set a time window. The time window is set by those skilled in the art based on business volume and experience. Use a counter to count the number of times the same error type occurs within the window. Ratio the number of occurrences to the time window length to calculate the occurrence frequency.

[0072] Among them, the counter can use Redis counter;

[0073] Identify new error types by matching with the historical error database. The process is as follows:

[0074] Extract confirmed error types from historical logs, build a historical whitelist library, convert error types to lowercase, remove leading and trailing spaces, and reduce matching failures caused by format differences.

[0075] Among them, an update mechanism is set up for the historical whitelist library, and newly confirmed error types are synchronized from the log library to the historical whitelist library every day;

[0076] The obtained standardized error type is used to quickly determine whether the error type exists in the historical whitelist database through a Bloom filter.

[0077] When the occurrence frequency is greater than or equal to the frequency threshold or the error type does not exist in the historical whitelist library, stack trace analysis is triggered; otherwise, stack trace analysis is not triggered;

[0078] The frequency threshold is set by those skilled in the art based on industry standards and experience. For example, in accordance with ITIL operation and maintenance specifications, "high frequency" is generally defined as exceeding the average error rate by 2-3 times, or the number of errors within a specific time window (such as 5 minutes or 1 hour) is significantly higher than the daily average.

[0079] For example, a scenario where an e-commerce API reports a "ConnectionTimeout" error between 10:00 and 10:05.

[0080] Sliding window count: If the error occurs 12 times within 5 minutes, it is considered a high frequency error.

[0081] Historical database comparison: This error type is not recorded in the historical whitelist database and is a new error;

[0082] Trigger analysis: When the trigger condition is met, stack trace analysis is performed;

[0083] This step has at least the following benefits: by collecting logs from multiple sources and extracting key fields, combined with mechanisms for identifying high-frequency errors and new error types, log entries requiring stack trace analysis can be quickly identified. For example, Redis counters are used to count the frequency of errors within a time window, and Bloom filters are used to match new error types against a historical whitelist. Analysis is triggered when the frequency is greater than or equal to a threshold or the error type is new. This reduces ineffective analysis of normal logs and improves the pertinence and efficiency of problem location.

[0084] Step 2: If stack trace analysis is triggered, obtain the corresponding error log entry, define the theoretical start and end markers for the stack trace analysis, obtain the actual start and end points for executing the stack trace analysis, compare the theoretical end marker with the actual end point for consistency, and determine whether the stack trace is normal;

[0085] Since this embodiment analyzes the situation where a log entry contains multiple compound events, it is necessary to first filter the error log entries;

[0086] Specifically, error log entries containing more than one event are extracted as log entries to be analyzed;

[0087] For each log entry to be analyzed;

[0088] Define the theoretical start mark and theoretical end mark of the log entry to be analyzed. The theoretical start mark is the start timestamp of the abnormal event propagation chain of the log entry to be analyzed, and the theoretical end mark is the end timestamp of the abnormal event propagation chain of the log entry to be analyzed.

[0089] Execute stack trace and obtain the actual start point timestamp and actual end point timestamp of the stack trace;

[0090] Compare the end timestamp of the log entry to be analyzed with the end timestamp of the stack trace.

[0091] The purpose of consistency comparison is to determine whether the stack trace is complete and whether the stack trace analysis is interrupted due to misjudgment of the end timestamp during the analysis process;

[0092] If the end timestamp defined by the log entry to be analyzed is consistent with the end timestamp of the stack trace, it is a normal stack trace;

[0093] If the end timestamp defined in the log entry to be analyzed is inconsistent with the end timestamp of the stack trace, the stack trace will be abnormal;

[0094] This step has at least the following effects: It defines the theoretical start and end markers of the log entry to be analyzed. By comparing the theoretical start and end points with the actual stack trace, it can accurately determine whether the stack trace is normal. This improves the problem of premature termination of stack traces due to the inability to identify event boundaries, ensures the integrity of the exception stack information, and provides comprehensive and accurate data support for subsequent root cause analysis.

[0095] Step 3: If the stack trace is abnormal, we filter out asynchronous log entries in the vicinity of the abnormal stack trace based on time deviation and asynchronous keyword identification. We count the number of asynchronous log entries, analyze the correlation between the abnormal stack trace and the occurrence of asynchronous log entries, and identify the degree of correlation.

[0096] In asynchronous logging scenarios, log entries from different service nodes may have timestamp misordering due to network latency or batch processing, triggering pseudo-end marker recognition and causing premature termination of the stack trace. Therefore, this step analyzes whether the inconsistency between the end point timestamp of the exception stack trace and the defined end timestamp is due to a misjudgment of the end point timestamp caused by asynchronous log entries.

[0097] Specifically, based on the stack trace exception, the corresponding log entry to be analyzed is obtained as the abnormal log entry, and the abnormal log entry is intelligently identified. The process is as follows:

[0098] Quickly filter candidate logs based on time difference thresholds and asynchronous keywords;

[0099] Filter log entries containing the same request ID within the maximum time difference from the timestamp of the exception stack trace end point as candidate log entries;

[0100] The maximum time difference is a time deviation threshold set by those skilled in the art to filter out log entries with timestamps close to the end point of the stack trace;

[0101] Scan candidate log entries to see if they contain the asynchronous processing keyword;

[0102] Asynchronous processing keywords are text patterns in the log content that directly or indirectly indicate that the log was generated by an asynchronous task, delayed processing, or event-driven asynchronous execution mechanism, including but not limited to:

[0103] Task scheduling: such as async, defer, schedule, queue; message queues: such as MQ, RabbitMQ, Kafka, consumer; batch processing: such as batch, bulk, chunk; retry mechanisms: such as retry, RETRY_SUCCESS, MAX_RETRY; delayed execution: such as delay, timeout, deferred; distributed coordination: such as @Async (Spring annotation) and Celery (Python task queue);

[0104] Use regular expression matching method to determine the matching degree of abnormal keywords;

[0105] Set the regular expression basic mode: \b(async|defer|retry|queue|batch|delay)\b;

[0106] \b: word boundary matching to avoid partial matching (e.g. asynchronous does not trigger async);

[0107] |: logical OR operator, matches any keyword;

[0108] i modifier: ignore case (such as ASYNC and async are equivalent);

[0109] Remove special characters from the log and convert them to lowercase (regular expression case is irrelevant). Special characters include but are not limited to the [ERROR] prefix.

[0110] Perform regular expression matching on each candidate log entry and record the matching asynchronous processing keywords and their occurrence counts.

[0111] For each candidate log entry, calculate the sum of the occurrences of all matching asynchronous processing keywords as the keyword matching density;

[0112] Compare the keyword matching density with the keyword matching density limit. If the keyword matching density is greater than the keyword matching density limit, mark it as an asynchronous log entry; if the keyword matching density is less than or equal to the keyword matching density limit, mark it as a non-asynchronous log entry.

[0113] The keyword matching density limit is set by those skilled in the art based on industry experience and industry standards.

[0114] Since the extraction condition for candidate log entries is to meet the maximum time difference of the exception stack trace end point timestamp, the exception stack trace analysis end point timestamp is close to the timestamp of the asynchronous log entry;

[0115] Count the number of abnormal log entries with asynchronous log entries and calculate the ratio with the total number of abnormal log entries to obtain the percentage value;

[0116] Among them, for the quantity ratio value, the quantity ratio value quantifies the correlation strength between asynchronous operations and stack exceptions.

[0117] Indicates the percentage of exceptions caused by asynchronous log entries in stack trace exception events. Given a known stack trace exception, the higher the frequency of asynchronous log entries, the stronger the correlation.

[0118] Compare the quantity ratio value with the quantity ratio threshold. If the quantity ratio value is less than or equal to the quantity ratio threshold, the correlation is weak. If the quantity ratio value is greater than the quantity ratio threshold, the correlation is strong, and the error in detecting the timestamp of the end point of the exception stack trace is caused by the occurrence of asynchronous log entries.

[0119] The quantity percentage threshold is set by those skilled in the art based on industry experience and business tolerance.

[0120] This step has at least the following effects: Based on the time deviation and asynchronous keyword identification indicators, it filters out asynchronous log entries in the vicinity of the abnormal stack trace. It also quantifies the correlation strength between asynchronous operations and stack exceptions by their percentage. This allows us to determine whether the stack trace exception is caused by an asynchronous log entry. This provides a clear direction for subsequent optimization, reduces misjudgments caused by asynchronous log interference, and improves the reliability of analysis results.

[0121] Step 4: If the correlation is strong, set up an exception optimization detection mechanism to optimize the accuracy of stack trace analysis for end point timestamp detection;

[0122] Specifically, as a preferred embodiment of this embodiment, a fixed time window is defined, and the theoretical start timestamp and the theoretical end timestamp of the current log entry to be analyzed are used as the fixed time window; when defining the fixed time window, all asynchronous log timestamps outside the window can be excluded;

[0123] For example, the theoretical time range of the abnormal event chain recorded in the log entry to be analyzed is [10:00:00, 10:00:05];

[0124] If there are multiple timestamps within the fixed time window, the timestamp filtering process is performed:

[0125] Count several historical normal stack trace log entries, and extract the timestamp sequence and adjacent time interval of each historical normal stack trace log entry;

[0126] Calculate the historical adjacent interval average and historical adjacent interval standard deviation of all historical normal stack trace log entries; count the total time consumed by all historical normal stack trace log entries and perform mean calculation to obtain the historical average total time consumed;

[0127] Count the time interval distribution in different business scenarios and establish a time pattern library for each scenario;

[0128] Get multiple timestamps within a fixed window during the current stack trace analysis process, calculate all adjacent time interval values, and use them as the current adjacent time interval value;

[0129] Calculate the Euclidean distance between the current adjacent time interval and the average value of the historical adjacent intervals, and convert the Euclidean distance value into a similarity value by dividing 1 by (the ratio of the Euclidean distance value to the standard deviation of the historical adjacent intervals). The calculation formula is: ; Where X is the similarity value, D is the Euclidean distance, and Z is the standard deviation of the historical adjacent intervals;

[0130] It is worth noting that if the standard deviation of historical adjacent intervals is zero, it indicates that there is no variation in the historical data and all adjacent interval values ​​are equal to the average value. If the Euclidean distance value D between the current adjacent interval value and the historical average value is 0, it means that the current interval is completely consistent with the historical pattern, and the similarity value X is directly assigned to 1. If the Euclidean distance value D≠0, the denominator of the original formula is zero due to Z=0. In this case, it can be considered that there is an absolute difference between the current interval and the historical pattern. The similarity value X is assigned to 0 (minimum value), or a minimum value (such as 0.01) can be set according to business logic to avoid program exceptions. Add a conditional judgment in the algorithm. When Z=0 is detected, skip the division operation and directly determine the extreme value of X by whether D is zero, which increases the robustness of the calculation process.

[0131] Calculate the deviation rate between the current total time and the historical average total time. That is, calculate the difference between the current total time and the historical average total time, take the absolute value, and then calculate the ratio of the difference to the historical average total time to obtain the deviation rate.

[0132] The value obtained by subtracting the deviation rate from 1 is weighted and fused with the similarity value to output a comprehensive score;

[0133] It should be noted that the similarity values ​​need to be normalized before calculating the comprehensive score;

[0134] For the calculation of the comprehensive score, the comprehensive score maps the current timestamp sequence into a probability value belonging to a normal stack trace by fusing the time interval similarity (matching degree with historical normal patterns) and the total time deviation rate;

[0135] The weight coefficient is set by those skilled in the art based on business-sensitive scenarios;

[0136] Extract the timestamp with the highest comprehensive score as the timestamp of the end point;

[0137] It will be understood by those skilled in the art that a minimum judgment limit for the comprehensive score is also set. If the highest total score is less than the minimum judgment limit for the comprehensive score, a re-detection mechanism is triggered to re-acquire the timestamp;

[0138] This step has at least the following effects: When it is determined that the stack trace exception is highly correlated with an asynchronous log entry, a fixed time window is defined to lock the detection of the end point timestamp to the current log entry, eliminating interference from asynchronous log timestamps outside the window. At the same time, based on the time distribution pattern of historical normal stack traces, multiple timestamps within the window are quantitatively screened and the timestamp with the highest comprehensive score is selected as the end point. This reduces the problem of misjudgment of the end point due to asynchronous log fragments, improves the accuracy of stack trace analysis, and ensures that the exception stack information is fully captured.

[0139] This instance collects logs from multiple sources and extracts key fields, then triggers stack trace analysis by combining high-frequency error recognition and new error type identification mechanisms. It defines theoretical start and end markers and compares them with the start and end points of the actual stack trace to determine whether the stack trace is normal. It filters asynchronous log entries in neighboring areas based on time deviation and asynchronous keyword recognition, and quantifies the correlation strength by the percentage of the number. When the correlation is strong, it defines a fixed time window and filters the end point timestamp based on the timestamp distribution of historical normal stack traces. This optimizes traditional stack trace analysis in high-concurrency scenarios across cross-service calls, addressing issues such as "pseudo-end marker" recognition, premature stack trace termination, and missing exception stack information caused by the misordering of asynchronous log timestamps. This improves the accuracy of stack trace analysis and the completeness of exception information capture.

[0140] Example 2

[0141] Based on the same inventive concept as the AI-based log data analysis method in the aforementioned embodiment, Figure 3 As shown, the present application provides an AI-based log data analysis system, wherein the system specifically includes:

[0142] Stack trace trigger module: This module obtains log content feature data and determines whether to trigger stack trace analysis based on error pattern recognition. Error pattern recognition includes high-frequency error recognition and new error type recognition.

[0143] Specific implementation: Collect log data from multiple sources through Agent (supporting JSON, text, and Syslog formats), extract key fields such as log level and error type, map the error type to a unified identifier, and then use Redis counters to count the frequency of errors per unit time to identify high-frequency errors. Use Bloom filters to match the historical whitelist library to identify new error types. When the error frequency is ≥ the threshold or it is a new error type, trigger stack trace analysis, otherwise do not trigger.

[0144] Stack trace judgment module: If stack trace analysis is triggered, obtain the corresponding error log entry, define the theoretical start mark and theoretical end mark of the stack trace analysis, obtain the actual start point and actual end point of the stack trace analysis, compare the theoretical end mark and the actual end point for consistency, and determine whether the stack trace is normal;

[0145] Specific implementation: If analysis is triggered, select error log entries containing multiple compound events as the objects to be analyzed, define their theoretical start / end markers (timestamps of the abnormal event propagation chain), perform stack tracing to obtain the actual start / end point timestamps, and compare the consistency of the theoretical and actual end markers: if they are consistent, it is a normal stack trace; if they are inconsistent, it is an abnormal stack trace;

[0146] Asynchronous log analysis module: If a stack trace is abnormal, it filters out asynchronous log entries in the vicinity of the abnormal stack trace based on time deviation and asynchronous keyword recognition. It then counts the number of asynchronous log entries, analyzes the correlation between the stack trace abnormality and the occurrence of asynchronous log entries, and identifies the degree of correlation.

[0147] Specific implementation: For exception stack traces, candidate log entries with timestamps close to the end point are screened based on the time difference threshold. Regular expressions are used to match asynchronous keywords, and the keyword matching density is calculated and compared with the threshold. Asynchronous log entries are marked, and the percentage of asynchronous log entries in the exception log is calculated and compared with the threshold. If the percentage exceeds the threshold, the stack exception is strongly correlated with the asynchronous log; otherwise, the correlation is weak.

[0148] Abnormal optimization module: If the correlation is strong, set up abnormal optimization detection mechanism to optimize the accuracy of stack trace analysis for end point timestamp detection;

[0149] Specific implementation: If the correlation is strong, define a fixed time window with the theoretical start / end timestamps, exclude the interference of asynchronous logs outside the window, calculate the average time interval, standard deviation, and average total time consumption of historical normal stack traces, establish a time pattern library, calculate the time interval similarity and total time consumption deviation rate of the current stack trace, and perform weighted fusion to obtain a comprehensive score. Extract the timestamp with the highest score as the end point. If the score is less than the minimum judgment limit, trigger re-detection.

[0150] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. An AI-based log data analysis method, characterized by: The following steps are involved: Obtain log content feature data and determine whether to trigger stack trace analysis based on error pattern recognition, where error pattern recognition includes high-frequency error recognition and new error type recognition; If stack trace analysis is triggered, obtain the corresponding error log entry, define the theoretical start and end markers for the stack trace analysis, obtain the actual start and end points for executing the stack trace analysis, compare the theoretical end marker with the actual end point for consistency, and determine whether the stack trace is abnormal; If the stack trace is abnormal, asynchronous log entries in the vicinity of the abnormal stack trace are filtered out based on time deviation and asynchronous keyword recognition. The correlation between the stack trace abnormality and the occurrence of asynchronous log entries is analyzed to identify the degree of correlation. If the correlation is strong, set up an exception optimization detection mechanism to achieve accurate detection of the end point timestamp through stack trace analysis; The exception optimization detection mechanism includes setting a fixed time window and filtering the end point timestamp of the exception stack trace based on the timestamp distribution of the log entries of the historical normal stack trace.

2. The AI-based log data analysis method according to claim 1, characterized in that: The process of error pattern recognition and judgment is as follows: Set a time window, use a counter to count the number of times the same error type occurs within the window, and calculate the ratio of the number of occurrences to the length of the time window to obtain the frequency of occurrence; Extract confirmed error types from historical logs, build a historical whitelist database, and use Bloom filters to quickly determine whether the obtained standardized error types exist in the historical whitelist database. When the occurrence frequency is greater than or equal to the frequency threshold or the error type does not exist in the historical whitelist library, stack trace analysis is triggered.

3. The AI-based log data analysis method according to claim 1, characterized in that: The process of determining whether the stack trace is normal is as follows: Extract error log entries containing more than one event as log entries to be analyzed; Define the theoretical start mark and theoretical end mark of each log entry to be analyzed. The theoretical start mark is the start timestamp of the abnormal event propagation chain of the log entry to be analyzed, and the theoretical end mark is the end timestamp of the abnormal event propagation chain of the log entry to be analyzed. Perform stack tracing on each log entry to be analyzed, and obtain the actual start and end timestamps of the stack trace; If the end timestamp defined in the log entry to be analyzed is inconsistent with the end timestamp of the stack trace, an exception occurs in the stack trace.

4. The AI-based log data analysis method according to claim 1, characterized in that: The process of filtering out asynchronous log entries in the area adjacent to the exception stack trace is as follows: Get candidate log entries and scan whether they contain asynchronous processing keywords; Set the regular expression basic mode; perform regular expression matching on each candidate log entry, and record the matching asynchronous processing keywords and their occurrence counts; For each candidate log entry, calculate the sum of the occurrences of all matching asynchronous processing keywords as the keyword matching density; If the keyword matching density is greater than the keyword matching density limit, it is marked as an asynchronous log entry.

5. The AI-based log data analysis method according to claim 4, characterized in that: The process of obtaining the candidate log entries is as follows: Filter log entries containing the same request ID within the maximum time difference from the timestamp of the exception stack trace end point as candidate log entries.

6. The AI-based log data analysis method according to claim 1, characterized in that: The process of identifying the degree of relevance is: Count the number of abnormal log entries with asynchronous log entries and calculate the ratio with the total number of abnormal log entries to obtain the percentage value; If the quantity proportion value is greater than the quantity proportion threshold, the correlation is strong.

7. The AI-based log data analysis method according to claim 1, characterized in that: The abnormal optimization detection mechanism includes: Define a fixed time window, using the theoretical start timestamp and theoretical end timestamp of the current log entry to be analyzed as the fixed time window; If there are multiple timestamps within the fixed time window, the timestamp filtering process is performed.

8. The AI-based log data analysis method according to claim 7, characterized in that: The timestamp screening process includes: Count several historical normal stack trace log entries, and extract the timestamp sequence and adjacent time interval of each historical normal stack trace log entry; Calculate the historical adjacent interval average and historical adjacent interval standard deviation of all historical normal stack trace log entries; count the total time consumed by all historical normal stack trace log entries and perform mean calculation to obtain the historical average total time consumed; Get multiple timestamps within a fixed window during the current stack trace analysis process, calculate all adjacent time interval values, and use them as the current adjacent time interval value; Calculate the Euclidean distance between the current adjacent time interval value and the average value of the historical adjacent intervals, and convert the Euclidean distance value into a similarity value; Calculate the deviation rate between the current total time and the historical average total time; The value obtained by subtracting the deviation rate from 1 is weighted and fused with the similarity value to output a comprehensive score; The timestamp with the highest comprehensive score is extracted as the timestamp of the end point.

9. The AI-based log data analysis method according to claim 8, characterized in that: The calculation process of the similarity value and the deviation rate is: Calculate the ratio of the Euclidean distance value to the standard deviation of the historical adjacent intervals, and then divide the ratio by 1 to obtain the similarity value; The difference between the current total time and the historical average total time is calculated and the absolute value is taken. The difference is then compared with the historical average total time to obtain the deviation rate.

10. An AI-based log data analysis system, characterized in that: The system is used to execute the method according to any one of claims 1 to 9, and the system comprises: Stack trace trigger module: This module obtains log content feature data and determines whether to trigger stack trace analysis based on error pattern recognition. Error pattern recognition includes high-frequency error recognition and new error type recognition. Stack trace judgment module: If stack trace analysis is triggered, obtain the corresponding error log entry, define the theoretical start mark and theoretical end mark of the stack trace analysis, obtain the actual start point and actual end point of the stack trace analysis, compare the theoretical end mark and the actual end point for consistency, and determine whether the stack trace is normal; Asynchronous log analysis module: If a stack trace is abnormal, it filters out asynchronous log entries in the vicinity of the abnormal stack trace based on time deviation and asynchronous keyword recognition. It then counts the number of asynchronous log entries, analyzes the correlation between the stack trace abnormality and the occurrence of asynchronous log entries, and identifies the degree of correlation. Abnormal optimization module: If the correlation is strong, set up abnormal optimization detection mechanism to optimize the accuracy of stack trace analysis for end point timestamp detection; The exception optimization detection mechanism includes setting a fixed time window and filtering the end point timestamp of the exception stack trace based on the timestamp distribution of the log entries of the historical normal stack trace.