Online Hybrid Mining System for Multi-Source Logs' Log Templates
By summarizing log features into frequent items and variable-length variables, and using segmented mining methods to adjust the strategy in time, the accuracy and efficiency problems in multi-source log template mining are solved, adaptive log template extraction is achieved, and the accuracy and efficiency of log templates are improved.
Patent Information
- Application Number
- CN202310489871.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-04-28
AI Technical Summary
When faced with multi-source logging, existing log template mining algorithms are difficult to adapt to log changes, resulting in reduced accuracy and efficiency. Especially when log formats are diverse or have different lengths, they are prone to error extraction or inefficiency.
A log template online hybrid mining system for multi-source logs is adopted. By summarizing log features into frequent items and variable-length variables, the mining strategy is timely adjusted using segmented mining methods, including log preprocessing, template mining and evaluation optimization subsystems, identifying log sources and adaptively adjusting mining strategies.
On the basis of ensuring accuracy and efficiency, the degree of personnel participating in log preprocessing is reduced, adaptive mining of multi-source logs is realized, and the accuracy and efficiency of log template extraction is improved.
Smart Images

Figure CN116521628B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of log template mining, and particularly to an online hybrid mining system for log templates for multi-source logs. Background Art
[0002] Document 1 [He P, Zhu J, Zheng Z, et al. Drain: An online log parsing approach with fixed depth tree [C] / / 2017 IEEE international conference on web services (ICWS). IEEE, 2017: 33-40.] is an online log parsing based on a fixed-depth tree. Its main idea is to distinguish and cluster different log types by the feature of length, calculate the text similarity through a word list, and thus mine the template. However, in the case of diverse or uneven log formats, it may lead to an explosion of the tree structure branches and the efficiency is affected. In our opinion, it should be able to adapt to the diversity of log sequence types.
[0003] Document 2 [M. Du and F. Li, “Spell: Streaming parsing of system event logs,” in Proc. IEEE International Conference on Data Mining (ICDM). IEEE, 2016, pp. 859–864.] is a method for parsing log sequences based on the longest common subsequence. Its main idea is to obtain the correlation between logs by comparing the common sequences existing between logs, obtain the constants existing in the log sequences, and match the log templates. However, too many parameter values in the log sequences may be judged as constants, and at the same time, this algorithm needs to traverse the entire log sequence, and the efficiency is low in establishing rules for large-scale log sequences.
[0004] Reference 3 [Shima K. Length Matters: Clustering System Log Messages using Length of Words:, 10.48550 / arXiv.1611.03213[P]. 2016.] introduced positional similarity based on the number of shared words at the same position in the log parsing method of classifying logs based on length. Its main purpose is to solve the problem that simply using length to parse logs cannot guarantee the correctness of the results, and to judge whether two log sequences are similar according to the number of shared words at the same position. It solves the drawback that the conventional clustering algorithm needs to traverse twice, improving the running efficiency. However, it may cause two different templates to be integrated into the same class.
[0005] Existing log template mining algorithms often have good effects only for logs with one feature. However, once the object changes, they cannot adapt well. Reference 1 is sensitive to the length of logs, and uncertain variable outputs may cause the same log template to be partitioned into different lengths, resulting in incorrect extraction of log templates. Reference 2 can effectively handle the problems that occur in Reference 1, but it requires the existence of a common subsequence in the log that exceeds the threshold requirement. The occurrence of log position offsets and a large number of variables will reduce the effectiveness of the matching. The shared word similarity strategy in Reference 3 may classify two different log template types into one category due to a small positional similarity index for words with low frequencies, bringing uncertainty to the accuracy of the results.
[0006] In summary, the existing algorithms are too targeted. Once the logs change due to software upgrades or other reasons, the algorithm results will be unsatisfactory. The present invention hopes to overcome this problem as much as possible through the working mode of "trial mining → evaluation → adjustment → mining". Summary of the Invention
[0007] In order to better adapt to the changes in logs, the present invention proposes an online hybrid mining system for log templates for multi-source logs. The present invention classifies log features into two features: frequent items and variable-length variables, formulates corresponding mining strategies for these two types of features, and discovers the features existing in the logs in a timely manner through segmented mining, and can adjust the subsequent mining strategies in a timely manner. The ultimate goal of the present invention is: given a log, this system can identify the log source according to its features and adaptively adjust the mining strategy, and on the basis of ensuring accuracy and efficiency, minimize the degree of human participation in log preprocessing.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] An online hybrid mining system for multi-source logs, which consists of a log preprocessing subsystem, a log template mining subsystem, and an evaluation and optimization subsystem:
[0010] The log preprocessing subsystem is used to identify the log type, strip the relatively fixed components from the original log from the log items, and submit the processed log to the log template mining subsystem;
[0011] The log template mining subsystem is used to divide the logs into several consecutive windows according to the quantity, and implement mining by using two basic methods of frequent item mining and similarity clustering;
[0012] The evaluation and optimization subsystem is used to evaluate the mining results within the window, adjust the mining parameters according to the evaluation results, and output the log template.
[0013] Furthermore, the log preprocessing subsystem includes a log recognition module, a log element extraction module, and a log message processing module; the log recognition module is used to identify the source and type of the log; the log element extraction module is used to extract log elements according to the log type based on the pre-written format and rules; the log message processing module is used to reprocess the log messages.
[0014] Furthermore, the log template mining subsystem includes a log template filtering module, a frequent item mining module, and a similarity clustering module; the log template filtering module is used to classify some log messages into known log templates according to the already mined log templates to reduce the subsequent mining workload; the frequent item mining module is used to extract keywords from the logs by using the method of frequent vocabulary mining to generate log templates; the similarity clustering module is used to cluster the log records without frequent words by using the method of combining length and vocabulary similarity.
[0015] Furthermore, the evaluation and optimization subsystem includes a mining result evaluation module, a mining parameter maintenance module, and a log template generation module; the mining result evaluation module is used to evaluate the mining results of the previous window, including examining the proportion of frequent items in the window and the existence of variable-length variables; the mining parameter maintenance module is used to adjust the mining strategy according to the evaluation results; the log template generation module is used to periodically generate log templates and output the log templates after all log mining is completed.
[0016] Furthermore, the log recognition module is specifically used for:
[0017] Extract the header of the log item, decompose the target log item header into multiple dimensions, and construct the log feature code x of the log item header; read all log feature codes y in the log feature set i , 1 < i < n, where n is the number of log feature information contained in the log feature set; detect each dimension x and yi Similarity; According to the detection result, construct vector v, by default, the feature code vector in the log feature set is w, and calculate the cosine similarity between v and w; Screen the log feature code y corresponding to the maximum cosine similarity and meeting the minimum similarity threshold. d ; If y d does not exist, return that the query is empty; If it exists, in the log feature set, query the log type corresponding to y d and return it.
[0018] Furthermore, the log template filtering module is specifically used for:
[0019] Perform preprocessing according to the log templates already available in the system, filter the log messages belonging to the log templates, and no longer perform subsequent mining: For the log template S_c, its corresponding maximum common subsequence of the log messages is LCS. If the log message msg also has LCS, and Len(msg)-Len(LCS)<=2, then msg belongs to the log template S_c, where Len(msg) represents the length of msg, and Len(LCS) represents the length of LCS.
[0020] Furthermore, the frequent item mining module is specifically used for:
[0021] Scan whether the log messages contain constants. During the scanning, group the log messages containing constants according to the contained constants, cluster the log messages in the same group, and retain the clusters whose number of the same type reaches the minimum clustering threshold: Take any two log messages in the same group, calculate the similarity. If it exceeds the similarity threshold, form a cluster; Take the word with the longest length in the cluster as the identification message, and calculate the similarity between any log message entering the cluster and it. Only after meeting the similarity threshold can it enter the cluster; When all log messages are processed, or the number of times of calculating similarity reaches 2n, where n refers to the total number of log messages in the same group, stop clustering; Among the clusters, the log messages whose quantity meets the minimum clustering threshold are officially output as template clusters, and the remaining log messages are left for subsequent mining.
[0022] Furthermore, the similarity clustering module is specifically used for:
[0023] Classify the message sequence set according to the length, and store the node information sequentially;
[0024] Classify the node message sequences according to the anchor points and complete the following operations:
[0025] Check whether the first word y of the current log message m is an anchor word;
[0026] According to the above results, determine whether m is a child node without an anchor point or an anchor point (y);
[0027] Check if there is an anchor (y) node in the query tree. If not, create the node;
[0028] Perform word vector feature classification on the descendant nodes of the anchor (y) node. The operations are as follows:
[0029] Calculate the average word length of the logs in the node;
[0030] Filter the log entries with an average word length within a word vector feature threshold range and regard them as a classification;
[0031] Perform the following operations in the similar clustering layer of the classification tree:
[0032] Construct the word length vector of the log entries;
[0033] Calculate the cosine similarity Sc pairwise. If both log entries contain the anchor word, divide them into two parts, x1 and x2, and y1 and y2, on the left and right of the anchor respectively, and calculate the similarity between x1 and y1, and between x2 and y2. Finally, the similarity Sc is the average of the two values;
[0034] Filter the log entries with calculation results exceeding the similarity threshold and group them into a cluster;
[0035] Perform cross-cluster clustering on the log entries that have not been clustered through word length similarity comparison. The operations are as follows:
[0036] Obtain the set of free log messages x_Msg and the set of clusters C;
[0037] Take any cluster classification C in C i and any log message c_m in it i and compare the similarity with any log message x_m in x_Msg i Perform similarity comparison;
[0038] Incorporate the log messages with similarity higher than the similarity threshold into the cluster. If it is lower than all cluster classifications, regard x_m i as a new class and incorporate it into C;
[0039] Return the obtained set of clusters C after processing.
[0040] Furthermore, the mining parameter maintenance module is specifically used for:
[0041] When the system is initialized, the constant word set and the anchor word set are empty. Start the constant word and anchor word generation process. The method for obtaining the constant set and the anchor set is as follows: First, traverse the log messages in the window, count the occurrence times of each vocabulary in the log messages, and compare them with the constant threshold and the anchor threshold to obtain the constant set and the anchor set;
[0042] After each window is processed, the system automatically analyzes the hit count of each word and executes the elimination process for constant words and anchor words. If the number of constant words and anchor words does not meet the minimum capacity requirement, during the next window processing, the constant word and anchor word generation process will start first;
[0043] After each round of window processing, the constant words and anchor words are sorted according to the hit count of each word, sorted in descending order of the hit count value, so that the most frequently used words can be compared first during the next window processing; after each window processing, the bubble sort method is used to sort the constant words and anchor words to reduce the computational complexity of sorting.
[0044] Further, the log template generation module is specifically used for:
[0045] Set the log template review period. After each log template review period is completed, mine the log template in the following manner: If the number of log messages included in the current classification exceeds the log template quantity threshold, continue with the subsequent process; otherwise, consider this classification not a log template; count the length of the common subsequence Len(LCS) of the log messages. If Len(x)-Len(LCS)<=2, where x represents any log message, then consider this classification a log template;
[0046] If all log messages have been processed, then output all classifications as log templates in the following manner: Extract the longest common subsequence of all log messages in the same classification as constants, and the other parts as variables, and output.
[0047] Compared with the prior art, the beneficial effects of the present invention are:
[0048] In order to adapt to the adaptive preprocessing of multi-type logs, the present invention proposes an online hybrid mining system for log templates for multi-source logs. Due to the diversity of logs, it is difficult to adopt a single strategy to suit the mining tasks of multiple types of logs. Based on the identification of logs, the present invention adopts an online progressive mining method, processes the log content in stages, thereby understanding the characteristics presented by recent logs, and adjusts the corresponding mining strategy according to the characteristics, so that the log template mining can achieve a better balance between accuracy and efficiency.
[0049] The present invention summarizes the log features into two features: frequent items and variable-length variables, formulates corresponding mining strategies for these two types of features, discovers the features existing in the logs in a timely manner through the method of segmented mining, and can adjust the subsequent mining strategy in a timely manner. For different types of logs, this system can identify the log source according to its features and adaptively adjust the mining strategy, and on the basis of ensuring accuracy and efficiency, minimize the degree of human participation in log preprocessing as much as possible. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 Schematic diagram of the mining principle of an online hybrid mining system for log templates for multi-source logs according to an embodiment of the present invention;
[0051] Figure 2 Schematic diagram of the architecture of an online hybrid mining system for log templates for multi-source logs according to an embodiment of the present invention;
[0052] Figure 3 Schematic diagram of the architecture of the log preprocessing subsystem according to an embodiment of the present invention;
[0053] Figure 4 Schematic diagram of the log feature code dimension according to an embodiment of the present invention;
[0054] Figure 5 Example diagram of Hadoop logs according to an embodiment of the present invention;
[0055] Figure 6 Flow chart of frequent item mining according to an embodiment of the present invention;
[0056] Figure 7 Schematic diagram of the clustering mining method according to an embodiment of the present invention. Detailed implementation manners
[0057] For the convenience of understanding, the following explanations are given for some terms that appear in the detailed implementation manners of the present invention:
[0058] Constants and variables: Except for log preprocessing items, log content is generally divided into constants and variables. A constant is a part of the log content that remains unchanged, while a variable is a part that changes. For example, in the logs IPC Server listener on 62260:starting and IPC Server listener on 62270:starting, the words IPC, Server, listener, on, and starting are included, which belong to logs in the same pattern. The words IPC, Server, listener, on, and starting can be regarded as constants in the log content. 62260 and 62270 are in the same position in the log content and are two changes in this position, which can be regarded as variables in the log content.
[0059] Constant words and anchor words: In logs, both constant words and anchor words are words that appear more frequently within a window. However, in terms of frequency, the frequency of constant words is higher than that of anchor words. For example, if the total number of words in a log is 10,000, and the constant word threshold and anchor word threshold are defined as 1% and 0.05% respectively, and the words for and admin appear 200 times and 70 times respectively, then for is a constant word and admin is an anchor word.
[0060] The present invention will be further explained and illustrated below with reference to the accompanying drawings and specific embodiments:
[0061] In order to adapt to the adaptive preprocessing of multiple types of logs, the present invention proposes a log template online hybrid mining system for multi-source logs. The log template mining principle of this system is as Figure 1 shown. Due to the diversity of logs, it is difficult to adopt a single strategy to suit the mining tasks of multiple types of logs. Based on the identification of logs, the present invention adopts an online progressive mining method. By processing the log content in stages, it can understand the characteristics presented by recent logs and adjust the corresponding mining strategy according to the characteristics, so that the log template mining can achieve a better balance between accuracy and efficiency.
[0062] The key points of the present invention are: log identification, online segmented processing, and multi-strategy hybrid collaboration. Log identification refers to analyzing log data to preliminarily determine the source and type of logs. Online segmented processing refers to dividing logs into multiple consecutive parts (referred to as windows in the present invention) according to time, and estimating the possible characteristics of subsequent logs based on the processing results of each window to provide support for the processing of subsequent windows. Multi-strategy hybrid collaboration refers to integrating mining strategies suitable for logs with different characteristics, and adaptively adjusting the window processing strategy according to the window characteristics, so as to achieve a balance between the efficiency and accuracy of log mining within the window.
[0063] The idea of the present invention can be summarized as: performing preprocessing according to the log source, and adjusting the log mining strategy for the latter part based on the mining results of the previous part of the logs. For example, when we know that the current log is a Hadoop log, according to historical experience, we perform preprocessing according to the Hadoop log format. For the preprocessed logs, based on the default strategy, by statistically analyzing the mining process data and results, we discover the basic characteristics of the current segment of logs and promptly adjust the mining strategy for the subsequent segment. For example, if there are frequently occurring parts in the log items corresponding to the same log template, then a frequent word mining strategy can be adopted for extraction.
[0064] Based on the above idea, the system structure is as Figure 2 shown. The system consists of a log preprocessing subsystem, a log template mining subsystem, and an evaluation and optimization subsystem. The log preprocessing subsystem is responsible for identifying the log type, stripping out relatively fixed components such as time from the original logs from the log items, and submitting the processed logs to the log template mining subsystem. The log template mining subsystem divides the logs into several consecutive windows according to the quantity, and implements mining using two basic methods: frequent item mining and similarity clustering. The evaluation and optimization subsystem evaluates the mining results within the window, adjusts the mining parameters according to the evaluation results, and outputs the log template.
[0065] The present invention completes log template mining by adopting the basic process of "trial mining → evaluation → re - mining". We use the window method to divide the log into adjacent but independent parts. After the raw log data is pre - processed by the log pre - processing subsystem, the input data to be mined will be constructed according to the window. The log template mining subsystem is responsible for the specific mining task, mining the log data of one window each time. The evaluation and optimization subsystem evaluates the mining results of the log template mining subsystem and adjusts and optimizes the mining strategy according to the evaluation results. The log template mining subsystem and the evaluation and optimization subsystem cooperate to complete the log template mining cycle after cycle.
[0066] The log pre - processing subsystem includes a log recognition module, a log element extraction module, and a log message processing module, which are responsible for filtering out unimportant and unnecessary components in the log items to form a data set to be processed. The log recognition module is the first module to receive log data, and it is responsible for identifying the source and type of the log to facilitate the subsequent modules to decide which algorithm to use. The log element extraction module and the log message processing module extract log elements according to the log type and re - process special content in the log message based on the formats and rules written in advance by humans.
[0067] The log template mining subsystem includes a log template filtering module, a frequent item mining module, and a similarity clustering module. The log template filtering module classifies some log messages into known log templates according to the log templates that have been mined, thereby reducing the subsequent mining workload. The frequent item mining module uses the method of frequent vocabulary mining to extract keywords from the log to generate log templates. The similarity clustering module clusters those log records without frequent words by using the method of length + vocabulary similarity.
[0068] The evaluation and optimization subsystem includes a mining result evaluation module, a mining parameter maintenance module, and a log template generation module. The mining result evaluation module is responsible for evaluating the mining results of the previous window, mainly examining the proportion of frequent items in the window and the existence of variable - length variables. The mining parameter maintenance module adjusts the mining strategy according to the evaluation results to adapt to the subsequent log mining and outputs the obtained log templates under certain conditions. The log template generation module periodically generates log templates and outputs the log templates after all log mining is completed.
[0069] 1 Log pre - processing subsystem
[0070] The log pre - processing subsystem is responsible for filtering out unimportant and unnecessary components in the log items to form a data set to be processed. The processing flow is as Figure 3As shown, it is mainly divided into three stages. Log identification is responsible for identifying the source and type of the original log, which facilitates subsequent preprocessing operations. Log element extraction extracts the elements of the log based on the source of the log to obtain the log message (Note: the present invention only focuses on log messages, but other elements can be used for other purposes). Log message preprocessing preprocesses special fields in the log message, such as IP addresses, and finally forms data to be processed by the log template mining subsystem.
[0071] 1.1 Log Identification
[0072] The purpose of log recognition is to identify the log type to facilitate subsequent log preprocessing operations. The diversity of logs also leads to the diversity of log preprocessing. Each log preprocessing rule is relatively fixed, but the processing rules between different types of logs often have obvious differences. If the log type cannot be identified, it will lead to the inability to determine the log preprocessing rules, which will further lead to unsatisfactory log preprocessing results and affect the subsequent log template mining.
[0073] Log recognition is done by identifying log signatures. Generally speaking, the log header format has a relatively fixed format, which is the basis for identifying logs. For example, the original header of a Hadoop log is as follows:
[0074] 2015-10-18 18:01:51 791 INFO[main]org.apache.hadoop.mapreduce.v2.app.client.MRClientService:
[0075] The header format can be summarized as: YYYY-DD-DD HH:MM:SS{enumeration value}[*]*hadoop*. The log item header format obtained after abstraction is the log feature code.
[0076] The method of identifying logs based on log signatures is as follows. Based on the characteristics of log signatures, the user decomposes the target log item header into multiple dimensions. For example, the signature code of a hadoop log can be decomposed into Figure 4 The five dimensions are shown in the figure. When identifying the log type, the log feature code of the target log is compared with the log feature codes of all types of logs in the log feature set, and the item that meets the minimum threshold and has the highest similarity is taken as the hit item to determine the log type. The log feature set is a collection of several log feature information. The log feature set is a tuple <log feature code, log type>, where the log type records the source information such as the software information and version number corresponding to the log. The log feature set is held by the system and records common log types.
[0077] The similarity determination method for each dimension is as follows. We divide it into several cases: format matching, value equality, same enumeration range, and containing specific identifiers. Format matching is a common type, such as the date YYYY-MM-DD and [*]. That is, as long as the formats match, they are considered similar, regardless of whether the content represented by the variable * is the same. Value equality is relatively rare and requires the format and value to be exactly the same. The same enumeration range mainly applies to enumerated variables. As long as the variable content does not exceed the enumeration range, it is considered similar. Containing specific identifiers means that the log entry contains specific words, which can be used to assist in identifying the log. For example, the word "hadoop" often appears in Hadoop logs.
[0078] The log constructs a vector based on the similarity of each dimension and calculates the cosine similarity with the log feature code vector to obtain the most similar log feature code and its corresponding type. The process is as follows: Let the vector of the log feature code be w = {1, 1, 1, 1, 1}. If the log feature code v to be analyzed is consistent with w in a certain dimension, then the vector corresponding to v in this dimension is 1; otherwise, it is 0. Note: Log feature codes are sensitive to position. YYYY-DD-DD HH:MM:SS and HH:MM:SS YYYY-DD-DD will be regarded as different. After comparison, the vector v can be obtained. Let's assume v = {v1, v2,..., v n}, and then calculate the cosine similarity between w and v according to Equation 1. If the cosine similarity cosine(w, v) is greater than the log similarity threshold (default 0.95), then the log is the log type corresponding to the log feature code.
[0079]
[0080] The following gives the log recognition algorithm.
[0081]
[0082]
[0083] 1.2 Log Entry Element Extraction
[0084] Log entry element extraction is to extract each element contained in the log entry and put them into their respective lists. The format of the log entry is mostly fixed, and each part uses the principle of fixed length or special character separation and is spliced to form a complete log. Therefore, after determining the log type, according to the log entry format, each element contained in the log entry can be directly extracted. These elements will be used or discarded according to the needs in the subsequent log template mining.
[0085] We describe the information of each element of the log entry format by length or special characters. The method based on length description is as follows: [Description: start position - end position]. For example, [Time: 1 - 19]: The 1st to 19th bytes represent the log time. The method based on special character description is as follows: [Description:'start character' - 'end character']. For example, [Time: '' - '']: The part from the space character to the next space character represents the log time (when there are multiple description units in one log entry using the same delimiter character, the delimiter characters are used in sequence. For example, [Date: '' - ''], [Time: '' - ''], indicating that the first space character to the second space character is the date, and the third space character to the fourth space character is the time). It is also possible to combine length and special characters to jointly describe the log entry header format. The method of mixed use is as follows: [Description: start position - 'end character'] indicates starting from the start position and ending before the end character. For example, [Time: 1 - '']: Starting from the 1st character to the next space character represents the log time.
[0086] Perform compliance checks on the description sets of all attribute information to form the description information of the log entry format. A typical format description is as follows: {[Date: * - *], [Time: * - *], [Level: * - *], [Message: * - *]}. The compliance check mainly examines the following in the format description of the log: (1) Whether there is overlapping description information; (2) Whether there are common necessary elements. For the former, it is necessary to traverse the information to check for repeated descriptions, such as having two descriptions about "Date"; for the latter, check whether there are descriptions of "Date" and "Message" because these are often essential components in the log.
[0087] Taking Hadoop logs as an example, illustrate the method of describing the Hadoop log header format. The existing Hadoop logs are as Figure 5 shown. Then the corresponding description information of the message element is as follows:
[0088] {[Date: 2015 - 10 - 18],
[0089] [Time: 18:01:51,791],
[0090] [Level: INFO],
[0091] [Program: main],
[0092] [Component: org.apache.hadoop.mapreduce.v2.app.client.MRClientService:],
[0093] [Message: Instantiated MRClientService at
[0094] MININT-FNANLI5.fareast.corp.microsoft.com / 10.86.169.121:62260]}。
[0095] Give the process of message item element processing:
[0096]
[0097] 1.3 Message content preprocessing
[0098] Message content preprocessing is the further processing of message content. Usually, there will be some information with special meanings and obvious characteristics in the message content. Although these information are helpful for restoring the events recorded in the log, for the event template mining in some specific scenarios, these information may be temporarily redundant and useless. In order to make the subsequent extraction of log templates and parameters more convenient, it is necessary to extract and replace these contents with a fixed format to make the log message content more compact and structured.
[0099] According to the characteristics of these information, preprocessing rules are formulated. Typical examples are the IP address characteristics and preprocessing rules: the IP address shows the characteristic of "0-255.0-255.0-255.0-255"; once such a string is found, the actual IP address is replaced with "[IP]". The port number characteristics and preprocessing rules: the port number is generally used in pair with the IP address. Once "IP address: number" is found, then the number part can be regarded as the port number; the actual port number value can be replaced with "[port number]". There are also such characteristics, "port number: 1024", and at this time, it can also be determined that 1024 is the port number and replaced with "[port number]". There may also be information such as process ID number and data block number in the message content, and these need to be manually written preprocessing rules according to the actual situation of the log.
[0100] Give the process of message content preprocessing:
[0101]
[0102] 2 Log template mining subsystem
[0103] The log template mining subsystem is mainly responsible for the specific template mining work. We sort the log messages to be analyzed in chronological order, and by default, 10,000 messages form 1 window. The log template mining subsystem processes one window each time, and the mining results will be evaluated and analyzed for use in the next window template mining.
[0104] 2.1 Log template filtering module
[0105] The log template filtering module performs preprocessing based on some log templates already available in the system, filters out the log messages belonging to the log templates, and no longer performs subsequent mining to improve the system efficiency. After a certain period of mining, the system will strictly authenticate some results to generate log templates. These log templates are equivalent to the intermediate results already obtained. Making full use of these intermediate results can filter out some log messages that can be regarded as the same log template, thereby reducing the workload of log mining.
[0106] The workflow based on log template filtering is as follows. Assume a log template S_c, and its corresponding longest common subsequence of log messages is LCS. Then, if a log message msg also has LCS and Len(msg) - Len(LCS) <= 2 (where Len(msg) represents the length of msg), then this log message is regarded as belonging to the log template S_c. For example: There is a log template S_c as Recalculating schedule,[*][*], and its corresponding longest common subsequence LCS is Recalculatingschedule. If the log message msg is Recalculating schedule,headroom=<memory:10240,vCores:-17>, which also has LCS and Len(msg) - Len(LCS) = 2, meeting the requirements. Then this log message is regarded as belonging to the log template S_c.
[0107] 2.2 Frequent item mining module
[0108] The frequent item mining module clusters similar log messages in the log messages based on frequent occurrence as the search basis. The frequent item mining module scans whether the log messages contain constants and divides the log messages into two categories: The log messages that do not contain constants will be directly handed over to subsequent mining, and the log messages that contain constants will continue to be mined. As Figure 6 shown, while scanning, the frequent item mining module classifies the log messages that contain constants according to the constants they contain; clusters the log messages in the same category, and only the clusters with the number of the same category reaching the minimum clustering threshold (denoted by t_clusterNum) are retained, and the remaining log messages are handed over to subsequent mining.
[0109] Continue to classify according to whether the log messages contain the same vocabulary. Generally speaking, the log messages in the same classification are likely to have more than 1 same vocabulary. By comparing the log messages pairwise, find the possible combinations of the same vocabulary, and use these combinations of vocabulary to guide subsequent classification, so as to minimize the computational cost and time overhead as much as possible.
[0110] Calculate the similarity of log messages within a group. We perform similarity analysis from two perspectives: the same words and word vectors. The starting point is that log messages within the same group generally do not have incorrect word order or the same words, but may have variable-length variables. Suppose the lengths of log message A and log message B are L1 and L2 respectively, and the number of their identical words is n. Then the similarity calculation method is Sim = 2n / (L1 + L2). For example: For log message (Invalid useradmin from [IP address]) and log message (Invalid user guest from [IP address]), the similarity between the two is Sim = 2 * 4 / (5 + 5) = 0.8.
[0111] The clustering method is as follows. Take any two log messages in the same group and calculate the similarity. If it exceeds the similarity threshold (denoted by t_clustValue, default 0.8), then form a cluster; take the longest word in the cluster as the identification message, and calculate the similarity between any log message entering the cluster and it. Only after meeting t_clustValue can it enter the cluster; when all log messages are processed, or the number of similarity calculations reaches 2n (n refers to the total number of log messages in the same group), stop clustering; within the cluster, if the number of log messages meets t_clusterNum, it is officially output as a template cluster, and the remaining log messages are left for subsequent mining (completed by the clustering mining module).
[0112] 2.3 Clustering Mining Module
[0113] To further mine log templates, as Figure 7 shown, the present invention adopts a method of first classifying by length, then classifying by anchor points, then classifying by word vector features, and finally clustering by similarity to realize the mining of the remaining log items. Classifying by length is to divide log items into several sets according to the length of the log items. Classifying by anchor points is to check whether there are anchor words inside the log items. If there are, classify them according to the anchor points. Of course, there may also be cases where there are no anchor words. In this case, cluster by similarity. Clustering by similarity means clustering according to the similarity of log items.
[0114] The method of classifying by length is as follows. Since the length of log items is faster than operations such as scanning anchor words, it can be completed in a shorter time, thus completing the classification as soon as possible and reducing excessive scanning operations. To further reduce matching operations, Figure 7In the middle tree, the address information of the length classification nodes is stored sequentially, so that after knowing the length of the log message, the target node can be directly located. To sum up, the process of classification by length is as follows: (1) Beforehand, the administrator estimates the maximum length of the log message and pre-sets the corresponding nodes in the classification tree; (2) During detection, for each log message (after preprocessing), scan its length and determine the corresponding length classification node.
[0115] The advantage of classifying using the anchor word is to improve accuracy and avoid the influence of variable-length variables on the mining results. The process of classification by anchor is as follows: (1) Check whether the first word (denoted as y for example) of the current log message (denoted as m) is an anchor word. If not, classify it into the non-anchor message class, and m becomes a descendant node of the non-anchor node and directly waits for the final clustering analysis; (2) Check whether there is an anchor (y) node in the tree. If not, create this node; (3) Make m a descendant node of the anchor (y) node and enter the word vector feature classification.
[0116] The word vector feature is for further classification to reduce the clustering operation after the classification by anchor word. Note: Log messages that do not contain anchor words will not be classified by word vector features and will directly go to the subsequent steps. The word vector feature adopts the average word length method, and other methods can also be used, but the present invention does not make requirements. The formula for calculating the average word length is c i is the i-th word in the log entry, len(c i ) represents the length of the word c i , and n is the number of words in the log message. That is, the average length of all words in the log entry is used as the word vector feature. The process of classification by word vector feature is as follows: Calculate the word vector features of all log entries under the same anchor node; Starting from the maximum word vector feature, every time the word vector feature is less than a threshold interval (default set to 1), it is regarded as a classification, enter the similar clustering layer of the classification tree, and complete the final clustering.
[0117] Finally, cluster by similarity in the similar clustering layer. To calculate the similarity, we vectorize the words in the log entry and use the cosine similarity to examine the similarity between the two. Considering the limited number of words in the log message and to reduce the calculation amount, we use the word length as the vectorization index. For example, the log messages (Invalid user admin from*) and (Invaliduser guest from*) can be vectorized as V(7,4,5,4,1) and W(7,4,5,4,1), and the method for calculating the similarity is as follows:
[0118] V = [v1, v2, …, v n
[0119] W = [w1, w2, …, wn ]
[0120]
[0121] Where W and V are log-term word length vectors, and Sc is the cosine similarity value. When the similarity Sc is greater than the threshold, we classify it into one category.
[0122] In order to minimize the impact of variable-length variables, we use segmented calculations based on anchor points and sum them up. The similarity of each segment is calculated separately with the anchor point as the separator, and the average value is taken as the final similarity. Assuming that two log messages x and y both contain anchor points, and the anchor points divide them into two parts, x1 and x2, and y1 and y2, then according to the cosine similarity method, the similarity of x1 and y1, and x2 and y2 is calculated respectively, and Sc1 and Sc2 are obtained, then the final similarity is Sc = (Sc1+Sc2) / 2. In the [Similarity Clustering Layer], if the similarity between log messages is higher than the similarity threshold (represented by t_clustValue, the default is 0.8), then they are considered to belong to the same cluster.
[0123] Finally, there may still be some log items that are not clustered, which requires further cross-cluster clustering. Theoretically, after completing the above operations, most of the log messages have been clustered, and there may still be some log messages in a free state, which need to be clustered and classified. Although these free log messages cannot be clustered in this group, these log messages may still be associated with log messages from other branches. To this end, we calculate the correlation of different log items again across clusters. Since the differences between log messages in different branches may be large, we no longer use length for vectorization, but compare words. Suppose the set of free log messages is x_Msg, the existing cluster set is C, and take the log message x_m i Try clustering with all the categories in C. The specific method is to take a category C i Any log message in c_m i , analyze x_m i and c_m i The specific process is as follows: (1) If x_m i and c_m i If the same anchor word is not included, then each word is matched one by one from left to right in the log message. If it matches, sim(pos)=1, otherwise sim(pos)=0. Pos indicates the position, that is, the number of words. If a log message is short, the subsequent position words are considered unmatched and sim(pos)=0. The final similarity is calculated as follows: Where len(max) is the maximum length of two log messages. (2) If x_m i and c_mi Contains the same anchor words. Starting from the anchor word and moving left and right, match words one by one (for example: a b c anchor word d, and a b c anchor word e, scan forward in order from the anchor word c→b→a). If a match is found, then sim(pos)=1; otherwise, sim(pos)=0. If a certain log message is short, then the subsequent words are considered non-matching, sim(pos)=0. The final similarity calculation is (3) If the similarity is higher than the similarity threshold (denoted by t_clustValue, default 0.8), then it is considered that they belong to the same cluster. (4) If x_m i does not belong to any classification, then x_m i is regarded as a new class and incorporated into C.
[0124]
[0125] 3 Evaluation and Optimization Subsystem
[0126] The evaluation and optimization subsystem is responsible for evaluating the mining results of the previous window, and focuses on adjusting two parameters that affect mining, namely constant words and anchor words. At the same time, within the log template review cycle, it periodically reviews the obtained classification results, and obtains some log templates in real time, so as to reduce subsequent mining work. When all log mining is completed, the evaluation and optimization subsystem outputs all classification results as log templates.
[0127] 3.1 Mining Result Evaluation Module
[0128] We evaluate the mining of the previous window to provide a basis for the processing of the current window. Generally speaking, the log items in adjacent windows have similar characteristics. Therefore, the characteristics of the current window can be inferred based on the characteristics of the recently processed window. It mainly includes: whether the log items in the current window show the characteristics of frequent item mining, and whether there are several variable-length variables. In this way, we can adaptively adjust the corresponding parameters when mining the log messages in this window.
[0129] To determine whether the log items in the current window show frequent characteristics, the main reference is the proportion of log items processed by the frequent item mining method in the previous window. We set the parameter frequent threshold to 50%. If the proportion of log items processed by frequent item mining in the previous window is lower than the frequent threshold, then we consider that there are difficulties in using frequent item mining for the log items in the current window and need to adjust it. Therefore, we start the frequent item mining process to expand the constant word set.
[0130] Determine whether there are several variable-length variables in the current window log. The main reference is the proportion of log items processed by the cross-cluster clustering method in the clustering mining module within the previous window. We set the parameter cross-cluster threshold to 10%. If the proportion of log items that finally used cross-cluster clustering in the previous window is higher than 10%, then we consider that there are relatively many variable-length variables in the current window log items or there are errors in classification by anchor points (there may be multiple anchor words). We start the anchor word sorting process and adjust the set of anchor words.
[0131] 3.2 Mining Parameter Maintenance Module
[0132] Both constant words and anchor words are words that appear more frequently within the window. However, from the perspective of occurrence frequency, the occurrence frequency of the constant set is higher than that of the anchor set. There are two thresholds related to frequent item mining: the constant threshold and the anchor threshold. Only words whose occurrence frequency exceeds the constant threshold can be called constants. In the present invention, t_content is used to represent the constant threshold, and the default setting is 0.03. Only words whose occurrence frequency exceeds the anchor threshold and is lower than the constant threshold can be called anchor words. In the present invention, t_anchor is used to represent the anchor threshold, and the default setting is 0.01.
[0133] When the system is initialized, the set of constant words and the set of anchor words are empty, and the process of generating constant words and anchor words is started. The methods for obtaining the constant set and the anchor set are as follows: First, traverse the log messages within the window, count the occurrence times of each word in the log messages, that is, the number of times a certain word appears within the window divided by the number of records within the window, and compare it with the constant threshold and the anchor threshold to obtain the constant set and the anchor set. The maximum capacity of the set of constant words and the set of anchor words is 100. When initially constructing the set of constant words and the set of anchor words, the sets may not be filled up. The above work will be completed in the next window until the number of constant words and anchor words meets the minimum capacity requirement (the number of words contained in the set of constant words and the set of anchor words is not less than 70).
[0134] The elimination process of constant words and anchor words. The number of hits num of each word is recorded in the set of constant words and the set of anchor words. According to num, the occurrence frequency of the word is converted. If it is higher than or lower than t_content or t_anchor, then the word will become / be eliminated as a constant word or an anchor word. The decreasing weight of the number of hits is d_weight (default 0.9), that is, if the number of hits in the i-th window is x, then when processing the next window, its number of hits decreases to 0.9*x. If it is hit once, then the new number of hits is 0.9*x + 1. After each window is processed, the system will automatically analyze num and execute the elimination process of constant words and anchor words. If the number of constant words and anchor words does not meet the minimum capacity requirement, when processing the next window, the process of generating constant words and anchor words will be started first.
[0135] Sorting process of constant words and anchor words. To reduce the computational complexity, after each round of window processing, according to the hit count num, sort the constant words and anchor words in descending order of the num value, so that the most frequently used words can be compared first in the next window processing. Since the order of constant words and anchor words is not greatly affected after each window processing and they are still in a relatively ordered state, the bubble sort method is used to sort the constant words and anchor words to reduce the computational complexity of sorting.
[0136] 3.3 Log Template Generation Module
[0137] We will obtain partial log templates based on the processing results of a certain number of windows. We set the log template review period to 10, that is, 10 windows. We will further analyze and confirm the log templates based on the current clustering results, which can effectively reduce the number of subsequent clusters and thus effectively improve the efficiency. There are two basic requirements for a log template: (1) This classification contains a large number of log messages; (2) The log messages in this classification are very similar, that is, there is no large deviation in the clustering center.
[0138] After each log template review period is completed, the process of mining log templates is as follows: (1) If the number of log messages contained in the current classification exceeds the log template quantity threshold (default is 3000), then continue with the subsequent process, otherwise consider this classification not a log template; (2) Statistically calculate the length of the common subsequence Len (LCS) of the log messages. If Len(x) - Len(LCS) <= 2, where x represents any log message, then consider this classification a log template.
[0139] If all log messages have been processed, then output all classifications as log templates. The method of generating log templates according to classifications is as follows: Extract the longest common subsequence of all log messages in the same classification as constants, and the other parts as variables, and output.
[0140] To better adapt to the changes in logs, the present invention proposes an online hybrid mining system for log templates for multi-source logs. The present invention combines two types of algorithms, frequent item mining-based and clustering mining-based. On the basis of identifying the log types, it adopts a working process of "trial mining → evaluation → re-mining", and can timely discover the characteristics of the logs through the method of segmented mining, and can timely adjust the subsequent mining strategies. For different types of logs, this system can adaptively adjust the mining strategies according to their characteristics, so as to minimize the degree of human participation in log preprocessing.
[0141] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An online hybrid mining system for log templates for multi-source logs, characterized in that The system consists of a log preprocessing subsystem, a log template mining subsystem, and an evaluation and optimization subsystem: The log preprocessing subsystem is used to identify the log type, strip the fixed components from the original log from the log items, and submit the processed log to the log template mining subsystem; The log template mining subsystem is used to divide the logs into several consecutive windows according to the quantity, and implement mining using two basic methods: frequent item mining and similarity clustering; The evaluation and optimization subsystem is used to evaluate the mining results within the window, adjust the mining parameters according to the evaluation results, and output the log template; The evaluation and optimization subsystem includes a mining parameter maintenance module and a log template generation module; the mining parameter maintenance module is used to adjust the mining strategy according to the evaluation results; the log template generation module is used to periodically generate log templates and output the log templates after all log mining is completed; The mining parameter maintenance module is specifically used for: At the initial stage of the system, the constant word set and the anchor word set are empty. The method for starting the constant word and anchor word generation process to obtain the constant set and the anchor set is as follows: First, traverse the log messages in the window, count the occurrence times of each word in the log messages, compare with the constant threshold and the anchor threshold, and obtain the constant set and the anchor set; After each window is processed, the system automatically analyzes the hit times of each word and executes the elimination process of constant words and anchor words. If the number of constant words and anchor words does not meet the minimum capacity requirement, when processing the next window, first start the constant word and anchor word generation process; After each round of window processing is completed, sort the constant words and anchor words according to the hit times of each word, sort them in descending order according to the hit times value, so that the most commonly used words can be compared first when processing the next window; after each window is processed, use the bubble sort method to sort the constant words and anchor words; The log template generation module is specifically used for: Set the log template review period. After each log template review period is completed, mine the log template in the following way: If the number of log messages included in the current classification exceeds the log template quantity threshold, continue the subsequent process, otherwise consider that this classification is not a log template; count the length of the common subsequence Len(LCS) of the log messages. If Len(x)-Len(LCS)<=2, where x represents any log message, then consider that this classification is a log template; If all log messages have been processed, then output all classifications as log templates in the following way: Extract the maximum common subsequence of all log messages in the same classification as the constant, and the other parts as variables, and output.
2. The online hybrid mining system for log templates facing multi-source logs according to claim 1, wherein The log preprocessing subsystem includes a log recognition module, a log element extraction module, and a log message processing module; The log recognition module is used to identify the source and type of the log; the log element extraction module is used to extract log elements according to the log type based on the pre-written format and rules; the log message processing module is used to reprocess the log messages.
3. The online hybrid mining system for log templates facing multi-source logs according to claim 1, characterized in that The log template mining subsystem includes a log template filtering module, a frequent item mining module, and a similarity clustering module; The log template filtering module is used to classify some log messages into known log templates according to the mined log templates; the frequent item mining module is used to extract keywords from the logs to generate log templates by using the method of frequent vocabulary mining; the similarity clustering module is used to cluster the log records without frequent vocabulary by using the method combining length and vocabulary similarity.
4. The online hybrid mining system for log templates facing multi-source logs according to claim 1, characterized in that, The evaluation and optimization subsystem further includes a mining result evaluation module; the mining result evaluation module is used to evaluate the mining result of the previous window, including examining the proportion of frequent items in the window and the existence of variable-length variables.
5. The log template online hybrid mining system for multi-source logs according to claim 2, wherein The log recognition module is specifically used for: Extract the header of the log entry, decompose the target log entry header into multiple dimensions, and construct the log feature code x of the log entry header; read all log feature codes y in the log feature set i , 1 < i < n, where n is the number of log feature information contained in the log feature set; detect the similarity between x and y in each dimension i ; according to the detection results, construct the vector v, by default the feature code vector in the log feature set is w, and calculate the cosine similarity between v and w; screen the log feature code y corresponding to the largest cosine similarity and meeting the lowest similarity threshold d ; if y d does not exist, return that the query is empty; if it exists, query the log type corresponding to y d in the log feature set and return it.
6. The online hybrid mining system for log templates facing multi-source logs according to claim 3, wherein The log template filtering module is specifically used for: Preprocess according to the log templates already available in the system, filter the log messages belonging to the log templates, and no longer perform subsequent mining: for the log template S_c, its corresponding log message's longest common subsequence is LCS. If the log message msg also has LCS and Len(msg)-Len(LCS)<=2, then msg belongs to the log template S_c, where Len(msg) represents the length of msg and Len(LCS) represents the length of LCS.
7. The online hybrid mining system for log templates facing multi-source logs according to claim 3, characterized in that The frequent item mining module is specifically used for: Scan whether the log messages contain constants. While scanning, group the log messages containing constants according to the contained constants, cluster the log messages in the same group, and retain the clusters whose number of the same type reaches the minimum clustering threshold: take any two log messages in the same group, calculate the similarity. If it exceeds the similarity threshold, form a cluster; take the word with the longest length in the cluster as the identification message, and calculate the similarity between any log message entering the cluster and it. Only after meeting the similarity threshold can it enter the cluster; when all log messages are processed, or the number of similarity calculations reaches 2n, where n refers to the total number of log messages in the same group, stop clustering; within the clusters, the log messages whose quantity meets the minimum clustering threshold are officially output as template clusters, and the remaining log messages are left for subsequent mining.
8. The online hybrid mining system for log templates for multi-source logs according to claim 3, wherein The similarity clustering module is specifically used for: Classify the message sequence set according to the length, and store the node information sequentially; Classify the node message sequences according to the anchor points and complete the following operations: Check whether the first word y of the current log message m is an anchor word; According to the above results, determine whether m is a child node of the non-anchor or anchor; Query whether there is an anchor node in the tree. If not, create the node; Classify the child nodes of the anchor node according to the word vector features and complete the following operations: Calculate the average word length of the logs in the node; Filter the log items whose average word length is in a word vector feature threshold range and regard them as a classification; Complete the following operations in the similar clustering layer of the classification tree: Construct the word length vector of the log items; Calculate the cosine similarity Sc pairwise. If both log items contain anchor words, then divide them into two parts, x1 and x2, and y1 and y2, on the left and right of the anchor respectively, and calculate x1 and y1, and x2 and y2 respectively. Finally, Sc is the average of the two values; Cluster the log entries whose calculated results exceed the similarity threshold into one cluster; Perform cross-cluster clustering on the unclustered log entries through word length similarity comparison and complete the following operations: Obtain the set x_Msg of free log messages and the cluster set C; Select any clustering classification C in C i and any log message c_m in it i to perform a similarity comparison with any log message x_m in x_Msg i ; Incorporate log messages with a similarity higher than the similarity threshold into the clustering. If it is lower than all cluster classifications, consider x_m i as a new class and incorporate it into C; Return the obtained cluster set C after processing.
Citation Information
Patent Citations
Automatic generation and online updating method and system for computer system log template
CN111435343A
System log analysis method based on N-gram and frequent pattern mining
CN112882997A