A Fast Log Parsing Method and System Oriented to Log Locality Features

Through a fast log analysis method for log locality characteristics, the longest common subsequence algorithm of multiple sequences and the log segmentation algorithm of the smallest random position word type, combined with the cached nearest matching algorithm, the problems of low log parsing efficiency and low accuracy in the existing technology are solved, and efficient and accurate log analysis is achieved.

CN115994157BActive Publication Date: 2025-06-20HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211602590.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2025-06-20
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

The existing log analysis methods are inefficient in parsing, and the parsing accuracy is not high enough. They fail to fully utilize the data characteristics of computer log data sets, and fail to fully mine the correlation between log messages and templates.

Method used

The fast log analysis method for log locality features is adopted, and the log data is segmented and template extracted through the multi-sequence longest common subsequence algorithm (MLCS) and the log segmentation algorithm based on the smallest word type of random locations. Combined with the cached nearest matching algorithm, the log matching process is optimized.

Benefits of technology

The efficiency and accuracy of log parsing are improved. By making full use of the local characteristics of log data, the template extraction and matching process is optimized, which significantly improves the analysis efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115994157B_ABST
    Figure CN115994157B_ABST
Patent Text Reader

Abstract

The present invention discloses a fast log parsing method oriented to log locality characteristics, which divides a log data set into a template extraction data set and a log matching data set. For the template extraction data set: the log messages in the template extraction data set are divided into different log groups according to the log length, and the multi-sequence longest common subsequence algorithm applicable to log data is executed on the log messages in the log group to extract the common words of the log group. It is judged whether the number of common words meets the threshold. When it is greater than or equal to the threshold, a log template is generated for the log group, and the template is added to the template library and the cache is updated. By dividing computer logs into log groups, executing the multi-sequence longest common subsequence algorithm applicable to log data on multiple log messages, extracting the common words of multiple log messages at one time to obtain the template, and using the cache-based nearest matching algorithm to match the template for each log message, the present invention improves the log parsing efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer log parsing, and more specifically, relates to a fast log parsing method and system for log locality characteristics. Background Art

[0002] With the development of technologies such as supercomputers, distributed systems, cloud services, and the Internet of Things, computer logs have shown an explosive growth trend, making traditional log parsing based on manual labor no longer applicable. In addition, the computer log formats generated by different owners and different systems are diverse, and these diverse computer logs often originate from the same architecture, such as a cloud service platform with numerous users, a supercomputer or distributed system with numerous nodes, and an Internet of Things system with numerous devices. This also makes script tool-based log parsing based on predefined rules no longer applicable.

[0003] Existing log parsing methods mainly include: an offline algorithm that uses frequent item mining technology to traverse and identify frequent items in computer logs, groups the computer logs according to the frequent items and their positions, and extracts the final log templates; a method that takes clustering as the central idea, establishes a log similarity formula to measure the similarity of different computer log statements, and performs clustering; and a method that judges the similarity of computer logs based on the Longest Common Subsequence (LCS) and classifies the computer logs.

[0004] However, the above several existing log parsing methods all have some non-negligible defects: First, due to their failure to fully utilize the data characteristics of the computer log data set itself, but only being limited to the pairwise comparison mode of computer logs, the parsing efficiency is low; second, the existing log parsing methods do not fully explore the correlation between log messages and templates, resulting in insufficient parsing accuracy. Summary of the Invention

[0005] In view of the above-mentioned defects or improvement requirements of the prior art, the present invention provides a fast log parsing method and system for log locality characteristics, aiming to solve the technical problems of low parsing efficiency and low parsing accuracy of the existing log parsing methods.

[0006] To achieve the above object, according to one aspect of the present invention, there is provided a fast log parsing method for log locality characteristics, including the following steps:

[0007] (1) Obtain a log data set composed of multiple log messages of a computer, and divide the obtained log data set according to a ratio to obtain a template extraction data set and a log matching data set;

[0008] (2) Divide the template extraction dataset obtained in step (1) into multiple log groups according to the log lengths of each log message in the dataset extracted according to the template;

[0009] (3) Set the counter i = 1;

[0010] (4) Determine whether i is greater than the number of log groups obtained in step (2). If so, go to step (14); otherwise, obtain the i-th log group, and use the multiple sequence longest common subsequence algorithm MLCS to process the i-th log group to obtain the common words corresponding to the i-th log group, and then go to step (5);

[0011] (5) Determine whether the number of common words corresponding to the i-th log group obtained in step (4) is greater than or equal to the threshold. If so, go to step (6); otherwise, go to step (8);

[0012] (6) Generate a template for the i-th log group, and update the template library and cache according to the generated template, and then go to step (7);

[0013] (7) Set i = i + 1, and return to step (4);

[0014] (8) Use the log splitting algorithm based on the minimum number of word types at random positions to process the current log group to obtain multiple sub-log groups;

[0015] (9) Set the counter j = 1;

[0016] (10) Determine whether j is greater than the number of sub-log groups obtained in step (8). If so, return to step (7); otherwise, obtain the j-th sub-log group, and use the multiple sequence longest common subsequence algorithm applicable to log data to process the j-th sub-log group to obtain the common words corresponding to the j-th sub-log group, and then go to step (11);

[0017] (11) Determine whether the number of common words corresponding to the j-th sub-log group obtained in step (10) is greater than or equal to the threshold. If so, go to step (12); otherwise, return to step (8);

[0018] (12) Generate a template for the j-th sub-log group and update the template library and cache, and then go to step (13);

[0019] (13) Set j = j + 1, and return to step (10);

[0020] (14) Use the cache-based nearest matching algorithm to match the template for each log message in the log matching dataset obtained in step (1), and update the template library and cache.

[0021] Preferably, the process of processing the log group in step (4) to obtain the common words corresponding to the log group includes the following sub-steps:

[0022] (4-1) Initialize the common word set CTokens (initially an empty string array) to represent the common word set;

[0023] (4-2) Set the counter p = 1;

[0024] (4-3) Determine whether p is greater than the log length L of each log message in the i-th log group. If so, the process ends; otherwise, go to step (4-4);

[0025] (4-4) Set the counter q = 2;

[0026] (4-5) Determine whether q is greater than the number N of log messages in the i-th log group. If so, it means that the p-th word in all log messages in the i-th log group is the same, and then go to step (4-7); otherwise, go to step (4-6);

[0027] (4-6) Determine whether the p-th word in the q-th log message in the i-th log group is equal to the p-th word in the first log message in the i-th log group. If so, go to step (4-8); otherwise, go to step (4-9)

[0028] (4-7) Add the p-th word in all log messages in the i-th log group to the common word set CTokens, and go to step (4-9);

[0029] (4-8) Set q = q + 1 and return to step (4-5);

[0030] (4-9) Set p = p + 1 and return to step (4-3).

[0031] Preferably, the calculation formula for the threshold Threshold in step (5) is:

[0032]

[0033] Preferably, step (6) includes the following sub-steps:

[0034] (6-1) Generate a template TPL1 for the i-th log group;

[0035] (6-2) Set the counter p = 1;

[0036] (6-3) Determine whether p is greater than the cache size. If so, go to step (6-7); otherwise, take out the p-th element cache p in the cache, and take out the cachep A template TPL2 is used to determine whether the template lengths of the template TPL1 obtained in step (6-1) and the template TPL2 are the same. If so, go to step (6-4); otherwise, go to step (6-6).

[0037] (6-4) Determine whether the template contents of the template TPL1 and the template TPL2 are the same. If so, it means that the template TPL1 and the template TPL2 match successfully, and go to step (6-5); otherwise, go to step (6-6).

[0038] (6-5) Add the template log IDs of the template TPL1 to the template log IDs of the template TPL2, and move the index number cache of the template TPL2 p to the first position of the cache, and the process ends.

[0039] (6-6) Set p = p + 1, and return to step (6-3).

[0040] (6-7) Set the counter q = 1.

[0041] (6-8) Determine whether q is greater than the number of templates in the template library. If so, go to step (6-13); otherwise, take out the qth template TPL3 in the template library and go to step (6-9).

[0042] (6-9) Determine whether the template lengths of the template TPL1 and the template TPL3 are the same. If so, go to step (6-10); otherwise, go to step (6-12).

[0043] (6-10) Determine whether the template contents of the template TPL1 and the template TPL3 are the same. If so, it means that the template TPL1 and the template TPL3 match successfully, and go to step (6-11); otherwise, go to step (6-12).

[0044] (6-11) Add the template log IDs of the template TPL1 to the template log IDs of the template TPL3, and add the index number of the template TPL3 to the first position of the cache, and the process ends.

[0045] (6-12) Set q = q + 1, and return to step (6-8).

[0046] (6-13) Add the template TPL1 to the template library, and add the index number of the template TPL1 in the template library to the first position of the cache, and the process ends.

[0047] Preferably, the process of generating a template for the ith log group in step (6-1) includes the following sub-steps:

[0048] (6-1-1) Create template TPL, initialize the template length of template TPL to the log length L of each log message in the i-th log group, initialize the template content of template TPL to an empty string array of size L, initialize the template constant positions of template TPL to an empty array, and initialize the template log IDs of template TPL to an empty array;

[0049] (6-1-2) Obtain the first log message log1 in the i-th log group;

[0050] (6-1-3) Set counter q = 1;

[0051] (6-1-4) Determine whether q is greater than L. If so, go to step (6-1-6). Otherwise, determine whether the q-th word of log1 exists in the common word set of the i-th log group. If so, set the q-th element of the template content of the template TPL created in step (6-1-1) to the q-th word of log1, and add q to the constant positions. Otherwise, set the q-th element of the template content of the template TPL to the wildcard "<*>". Then go to step (6-1-5);

[0052] (6-1-5) Set q = q + 1, and return to step (6-1-4);

[0053] (6-1-6) Add the log IDs of all log messages in the i-th log group to the log IDs of template TPL, and the process ends.

[0054] Preferably, step (8) includes the following sub-steps:

[0055] (8-1) Initialize an empty array Tss to save the word types at different positions;

[0056] (8-2) Generate a random number array RA with element values from 1 to L;

[0057] (8-3) Set counter p = 1;

[0058] (8-4) Determine whether p is greater than the log length L of each log message in the current log group. If so, go to step (8-16). Otherwise, go to step (8-5);

[0059] (8-5) Take out the p-th element RA in the array RA p ;

[0060] (8-6) Set an empty variable tokenType of dictionary type, whose key is used to store words, and the key value is used to store the index numbers of the log messages in the current log group that have the same word as the key for the RA p -th word;

[0061] (8 - 7) Set the counter q = 1;

[0062] (8 - 8) Determine whether q is greater than the number of log messages N in the current log group. If so, go to step (8 - 13); otherwise, go to step (8 - 9);

[0063] (8 - 9) Retrieve the RA p th word tokenRA p of the qth log message, and determine whether the word tokenRA p exists in the keys of the dictionary tokenType. If so, go to step (8 - 10); otherwise, go to step (8 - 11);

[0064] (8 - 10) Add q to the value array of the dictionary element in the dictionary tokenType whose key is the word tokenRA p , and go to step (8 - 12);

[0065] (8 - 11) Create a new element in the dictionary tokenType with the key being the word tokenRA p and the value being q, and go to step (8 - 12);

[0066] (8 - 12) Set q = q + 1, and return to step (8 - 8);

[0067] (8 - 13) Determine whether the number of key - value pairs of the dictionary variable tokenType is greater than 1 and less than the log length L. If so, go to step (8 - 14); otherwise, go to step (8 - 15);

[0068] (8 - 14) Add the dictionary tokenType to Tss, and determine whether the number of elements in the Tss array is equal to the preset value 3. If so, go to step (8 - 17); otherwise, go to step (8 - 15);

[0069] (8 - 15) Set p = p + 1, and return to step (8 - 4);

[0070] (8 - 16) Determine whether Tss is empty. If so, go to step (8 - 18); otherwise, go to step (8 - 17);

[0071] (8 - 17) Select the element minTokenType with the fewest key - value pairs in Tss, and split the log messages in the value array corresponding to each key in minTokenType into a sub - log group to form multiple sub - log groups, and the process ends.

[0072] (8 - 18) Each log message in the current log group forms a separate sub - log group, and the process ends.

[0073] Preferably, step (14) includes the following sub-steps:

[0074] (14-1) Set the counter k = 1;

[0075] (14-2) Determine whether k is greater than the number of log messages in the log matching dataset. If so, end the process; otherwise, retrieve the k-th log message log in the log matching dataset k , and initialize maxET_index = -1, indicating the index number of the template with the most common words with the log message log k , and initialize maxET = 0, indicating the number of common words between the log message log k and the maxET_index-th template in the template library, and enter step (14-3);

[0076] (14-3) Set the counter p = 1;

[0077] (14-4) Determine whether p is greater than the cache size. If so, enter step (14-13); otherwise, retrieve the p-th element cache in the cache p , and retrieve the cache p -th template TPL1 in the template library, and determine whether the log length of the k-th log message log obtained in step (14-2) is the same as the template length of the template TPL1. If so, enter step (14-5); otherwise, transfer to step (14-12); k

[0078] (14-5) Set ET = 0, which is used to represent the number of common words between the template content of the template TPL1 and the log message log k ;

[0079] (14-6) Set the counter r = 1;

[0080] (14-7) Determine whether r is greater than the array size of the template constant positions of the template TPL1. If so, enter step (14-10); otherwise, retrieve the r-th element constant of the template constant positions of the template TPL1 r and transfer to step (14-8);

[0081] (14-8) Determine whether the constant r -th word in the template content of the template TPL1 is the same as the constant k -th word in the log message log r . If so, set ET = ET + 1, and then transfer to step (14-9); otherwise, directly transfer to step (14-9);

[0082] ​(14 - 9) Set r = r + 1, and return to step (14 - 7);

[0083] (14 - 10) Determine whether ET is equal to the size of the array at the template constant positions of template TPL1. If so, it means that at all constant positions of template TPL1, the words in the template content of template TPL1 are the same as the words in the log message log k The words of, and the log message log k Can fully match template TPL1. Add the log ID of the log message log k To the log IDs of template TPL1, and move the index number cache of template TPL1 p To the first position of the cache, and the process ends. Otherwise, go to step (14 - 11);

[0084] (14 - 11) Determine whether ET is greater than or equal to the threshold and greater than maxET. If so, set maxET = ET, maxET_index = cache p , and then go to step (14 - 12). Otherwise, directly go to step (14 - 12);

[0085] (14 - 12) Set p = p + 1, and return to step (14 - 4);

[0086] (14 - 13) Set the counter q = 1;

[0087] (14 - 14) Determine whether q is greater than the number of templates in the template library. If so, go to step (14 - 22). Otherwise, take out the qth template TPL2 in the template library and determine whether the log length of the log message log k Is the same as the template length of template TPL2. If so, go to step (14 - 15). Otherwise, go to step (14 - 21);

[0088] (14 - 15) Set the counter r = 1;

[0089] (14 - 16) Determine whether r is greater than the size of the array at the template constant positions of template TPL2. If so, go to step (14 - 19). Otherwise, take out the rth element constant at the template constant positions of template TPL2 r And go to step (14 - 17);

[0090] (14 - 17) Determine the constant r The word at the of the template content of template TPL2 and the log message log k The constant rCheck if the words are the same. If so, set ET = ET + 1, and then go to step (14 - 18); otherwise, directly go to step (14 - 18).

[0091] (14 - 18) Set r = r + 1, and return to step (14 - 16).

[0092] (14 - 19) Determine if ET is equal to the size of the array at the template constant positions of template TPL2. If so, it means that at all constant positions of template TPL2, the words in the template content of template TPL2 are the same as the words in the log message log k The words of, and the log message log k Can fully match template TPL2. Add the log ID of the log message log k To the log IDs of template TPL2, move the index number q of template TPL2 to the first position of the cache, and the process ends; otherwise, go to step (14 - 20).

[0093] (14 - 20) Determine if ET is greater than or equal to the threshold and greater than maxET. If so, set maxET = ET, maxET_index = q, and then go to step (14 - 21); otherwise, directly go to step (14 - 21).

[0094] (14 - 21) Set p = p + 1, and return to step (14 - 14).

[0095] (14 - 22) Determine if maxET is greater than 0. If so, change the constants in the template content of template TPLm that are different from the log message log k To the wildcard "<*>" and delete them from the constant positions, add k to the template log IDs of template TPLm; otherwise, it means maxET has not been updated, and there is no template in the template library that matches the log message log k Then go to step (14 - 23).

[0096] (14 - 23) Add the log message log k As a separate template to the template library. The template length is the length of the log message log k The template content is each word of the log message log k The template constant positions are the positions of each word of the log message log k The template log IDs are the ID of the log message log k And add its index number in the template library to the first position of the cache, and the process ends.

[0097] According to another aspect of the present invention, a fast log parsing system for log locality characteristics is provided, including:

[0098] A first module, configured to obtain a log data set composed of a plurality of log messages of a computer, and divide the obtained log data set according to a ratio to obtain a template extraction data set and a log matching data set;

[0099] A second module, configured to divide the template extraction data set obtained by the first module into multiple log groups according to the log lengths of the log messages in the template extraction data set;

[0100] A third module, configured to set a counter i = 1;

[0101] A fourth module, configured to determine whether i is greater than the number of log groups obtained by the second module. If so, enter the fourteenth module; otherwise, obtain the i-th log group, and use the multi-sequence longest common subsequence algorithm MLCS to process the i-th log group to obtain the common words corresponding to the i-th log group, and then transfer to the fifth module;

[0102] A fifth module, configured to determine whether the number of common words corresponding to the i-th log group obtained by the fourth module is greater than or equal to a threshold. If so, enter the sixth module; otherwise, transfer to the eighth module;

[0103] A sixth module, configured to generate a template for the i-th log group, and update the template library and cache according to the generated template, and then enter the seventh module;

[0104] A seventh module, configured to set i = i + 1, and return to the fourth module;

[0105] An eighth module, configured to process the current log group using a log splitting algorithm based on the minimum number of word types at random positions to obtain multiple sub-log groups;

[0106] A ninth module, configured to set a counter j = 1;

[0107] A tenth module, configured to determine whether j is greater than the number of sub-log groups obtained by the eighth module. If so, return to the seventh module; otherwise, obtain the j-th sub-log group, and use the multi-sequence longest common subsequence algorithm applicable to log data to process the j-th sub-log group to obtain the common words corresponding to the j-th sub-log group, and then transfer to the eleventh module;

[0108] An eleventh module, configured to determine whether the number of common words corresponding to the j-th sub-log group obtained by the tenth module is greater than or equal to a threshold. If so, enter the twelfth module; otherwise, return to the eighth module;

[0109] The twelfth module is used to generate a template for the j-th sub-log group, update the template library and cache, and then enter the thirteenth module;

[0110] The thirteenth module is used to set j = j + 1 and return to the tenth module;

[0111] The fourteenth module is used to use the cache-based nearest matching algorithm to match templates for each log message in the dataset obtained by the first module, and update the template library and cache.

[0112] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0113] (1) Since the present invention adopts the multiple sequence longest common subsequence algorithm applicable to log data described in steps (4) and (10), it can make full use of the characteristic that "log messages belonging to the same template often appear continuously" in computer logs, extract common words from multiple log messages at one time, optimize the limitation of the traditional log parsing method of comparing log messages pairwise to judge similarity, and improve the parsing efficiency.

[0114] (2) Since the present invention adopts step (5), considering the characteristic that the constant ratio is lower in longer log messages, a calculation scheme of elastic threshold is used. For longer log messages, a threshold with a lower ratio based on the log length L is adopted, effectively improving the overall parsing accuracy.

[0115] (3) Since the present invention adopts step (6), when the extraction of common words in step (4) fails, the log messages can be classified using the minimum word type algorithm based on random positions to distinguish log messages belonging to different templates, split them into different log groups, and finally extract appropriate templates for each log message, improving the parsing accuracy.

[0116] (4) Since the present invention adopts step (14), secondary extraction of common words is performed when the log message matches the template. When a completely consistent template can be matched, the completely matching template is selected; otherwise, the template with the highest similarity (the most common words) meeting the threshold is selected. When both of the above do not meet the requirements, a new template is finally selected, improving the parsing accuracy.

[0117] (5) Since the present invention adopts the cache-based nearest matching algorithm in step (14), it makes full use of the characteristic of spatial locality of log messages belonging to the same template in log data, and preferentially matches the most recently matched template in the cache for each log message, improving the parsing efficiency. Description of the Drawings

[0118] Figure 1It is the flowchart of the fast log parsing method of the present invention facing the log locality feature;

[0119] Figure 2 It is a schematic diagram showing the log data locality, where Figure 2 (a) is an example of the locality of the log dataset BGL, Figure 2 (b) is an example of the locality of the log dataset HDFS, Figure 2 (c) is an example of the locality of the log dataset HPC, Figure 2 (d) is an example of the locality of the log dataset Zookeeper. Specific implementation manners

[0120] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0121] The present invention provides a fast log parsing method facing the log locality feature, which includes computer log splitting, a multiple sequence longest common subsequence algorithm applicable to computer log datasets, a log discrimination method based on the minimum number of word types at random positions, and a cache-based nearest matching algorithm in its implementation manner. Its main idea is to split the log dataset into a template extraction dataset and a log matching dataset according to a set ratio. For the template extraction dataset: split the log messages in the template extraction dataset into different log groups according to the log length, execute the multiple sequence longest common subsequence (MLCS) algorithm applicable to log data on the log messages in the log group to extract the common words of the log group, judge whether the number of common words meets the threshold. When it is greater than or equal to the threshold, generate a log template for the log group, add the template to the template library and update the cache. When it is less than the threshold, iteratively execute the log splitting method based on the minimum number of word types at random positions on the log group to split the log group into smaller sub-log groups, and re-execute the multiple sequence longest common subsequence algorithm applicable to log data on the sub-log groups; for the log matching dataset: match the template for each log message in the cache that stores the template index number. If the match hits, update the template library and the cache. Otherwise, perform template matching in the template library. If the match hits, update the template library and the cache. Otherwise, add the log message as a new template to the template library and update the template library and the cache.

[0122] The advantages of the present invention are as follows. By splitting computer logs into log groups and performing the multiple sequence longest common subsequence algorithm applicable to log data on multiple log messages to extract common words from multiple log messages at one time to obtain a template, both the efficiency and accuracy of log parsing are improved. If the extraction of common words fails, the log groups can be split again by an algorithm based on the minimum number of word types at random positions, effectively dividing the log messages into different sub-log groups and re-extracting common words for each sub-log group. In addition, through proximity matching based on caching, the speed of matching log messages with templates can be effectively increased.

[0123] As Figure 1 shown, the present invention provides a fast log parsing method oriented to log locality characteristics, including the following steps:

[0124] (1) Obtain a log data set composed of multiple log messages of a computer, and divide the obtained log data set according to a ratio to obtain a template extraction data set and a log matching data set;

[0125] Specifically, generally, the log data set is divided according to a ratio of 1:9. 1 / 10 of the log data is divided into the template extraction data set, which is used to extract templates, form a template library, and build a cache (Cache). 9 / 10 of the log data is divided into the log matching data set, and a suitable template is matched for each log message in the log matching data set in the cache and the template library based on the template library and cache obtained previously.

[0126] (2) Divide the template extraction data set obtained in step (1) into multiple log groups according to the log lengths of the log messages in the template extraction data set;

[0127] Specifically, the log length of a log message refers to the number of words that make up the log message. In this step, the log messages with the same log length in the template extraction data set are assigned to the same log group.

[0128] (3) Set a counter i = 1;

[0129] Specifically, the counter i is used to iterate through the log groups obtained in step (2).

[0130] (4) Determine whether i is greater than the number of log groups obtained in step (2). If so, go to step (14). Otherwise, obtain the i-th log group, process the i-th log group using the multiple sequence longest common subsequence algorithm (Multiple Longest Common Subsequence, abbreviated as MLCS) to obtain the common words corresponding to the i-th log group, and transfer to step (5);

[0131] Specifically, this step processes the log group to obtain the common words corresponding to the log group. This process includes the following sub-steps:

[0132] (4-1) Initialize the common word set CTokens (initially an empty string array) to represent the common word set;

[0133] (4-2) Set the counter p = 1;

[0134] (4-3) Determine whether p is greater than the log length L of each log message in the i-th log group. If so, the process ends; otherwise, go to step (4-4);

[0135] (4-4) Set the counter q = 2;

[0136] (4-5) Determine whether q is greater than the number N of log messages in the i-th log group. If so, it means that the p-th word in all log messages in the i-th log group is the same, and then go to step (4-7); otherwise, go to step (4-6);

[0137] (4-6) Determine whether the p-th word in the q-th log message in the i-th log group is equal to the p-th word in the first log message in the i-th log group. If so, go to step (4-8); otherwise, go to step (4-9)

[0138] (4-7) Add the p-th word in all log messages in the i-th log group to the common word set CTokens, and go to step (4-9);

[0139] (4-8) Set q = q + 1, and return to step (4-5);

[0140] (4-9) Set p = p + 1, and return to step (4-3).

[0141] The advantage of this step is that the improved multi-sequence longest common subsequence algorithm based on log data can be used to extract common words from multiple log messages at one time, improving the parsing efficiency.

[0142] (5) Determine whether the number of common words corresponding to the i-th log group obtained in step (4) is greater than or equal to the threshold. If so, go to step (6); otherwise, go to step (8);

[0143] Specifically, the calculation formula for the threshold Threshold in this step is:

[0144]

[0145] The advantage of this step is that a method for calculating the elastic threshold is proposed. Considering the characteristic that the proportion of constants in longer log messages is lower, the elastic threshold calculation scheme is used. For longer log messages, a threshold with a lower proportion based on the log length L is adopted, which improves the parsing accuracy.

[0146] (6) Generate a template for the i-th log group, update the template library (which is empty in the initial state) and the cache according to the generated template, and then proceed to step (7);

[0147] Specifically, taking the template node<*>detected a failed network<*>via interface alt0 as an example, the template described in this step consists of the following elements:

[0148] (a) Template length, that is, the number of words in the template. The template length of the example template is 10;

[0149] (b) Template content, an array of strings containing wildcards and constants. The template content of the example template is [node, <*>, detected, a, failed, network, <*>, via, interface, alt0], where <*> is called a wildcard and other words are called constants;

[0150] (c) Template constant positions, the index numbers of each constant in the template content. The template constant positions corresponding to the example template are [1, 3, 4, 5, 6, 8, 9, 10];

[0151] (d) Template log IDs, the set of log IDs of all log messages classified into this template.

[0152] Specifically, the template library described in this step consists of different templates, and the cache consists of the index numbers of the 10 most recently accessed templates in the template library. The cache capacity is a preset value of 10, the cache size is the number of template index numbers actually stored in the cache, the cache size is initially 0, and the maximum is 10.

[0153] Specifically, this step includes the following sub-steps:

[0154] (6-1) Generate a template TPL1 for the i-th log group;

[0155] Specifically, the process of generating a template for the i-th log group consists of the following steps:

[0156] (6-1-1) Create a template TPL, initialize the template length of the template TPL to the log length L of each log message in the i-th log group, initialize the template content of the template TPL to an empty string array of size L, initialize the template constant positions of the template TPL to an empty array, and initialize the template log IDs of the template TPL to an empty array;

[0157] (6-1-2) Obtain the first log message log1 in the i-th log group;

[0158] (6-1-3) Set the counter q = 1;

[0159] (6-1-4) Determine whether q is greater than L. If so, go to step (6-1-6). Otherwise, determine whether the q-th word of log1 exists in the common word set of the i-th log group. If so, set the q-th element of the template content of the template TPL created in step (6-1-1) to the q-th word of log1, and add q to the constant positions. Otherwise, set the q-th element of the template content of the template TPL to the wildcard "<*>". Then go to step (6-1-5);

[0160] (6-1-5) Set q = q + 1, and return to step (6-1-4);

[0161] (6-1-6) Add the log IDs of all log messages in the i-th log group to the log IDs of the template TPL, and the process ends.

[0162] (6-2) Set the counter p = 1;

[0163] (6-3) Determine whether p is greater than the cache size. If so, go to step (6-7). Otherwise, take out the p-th element cache p in the cache, and take out the cache p -th template TPL2 in the template library. Determine whether the template lengths of the template TPL1 obtained in step (6-1) and the template TPL2 are the same. If so, go to step (6-4). Otherwise, go to step (6-6);

[0164] (6-4) Determine whether the template contents of the template TPL1 and the template TPL2 are the same. If so, it means that the template TPL1 and the template TPL2 match successfully, and go to step (6-5). Otherwise, go to step (6-6);

[0165] (6-5) Add the template log IDs of the template TPL1 to the template log IDs of the template TPL2, and move the index number cache p of the template TPL2 to the first position in the cache, and the process ends;

[0166] (6-6) Set p = p + 1, and return to step (6-3);

[0167] (6-7) Set the counter q = 1;

[0168] (6 - 8) Determine whether q is greater than the number of templates in the template library. If so, proceed to step (6 - 13); otherwise, retrieve the q-th template TPL3 from the template library and transfer to step (6 - 9).

[0169] (6 - 9) Determine whether the template lengths of TPL1 and TPL3 are the same. If so, proceed to step (6 - 10); otherwise, transfer to step (6 - 12).

[0170] (6 - 10) Determine whether the template contents of TPL1 and TPL3 are the same. If so, it indicates that TPL1 and TPL3 match successfully, and proceed to step (6 - 11); otherwise, transfer to step (6 - 12).

[0171] (6 - 11) Add the template log IDs of TPL1 to the template log IDs of TPL3, and add the index number of TPL3 to the first position of the cache. The process ends.

[0172] (6 - 12) Set q = q + 1 and return to step (6 - 8).

[0173] (6 - 13) Add TPL1 to the template library, and add the index number of TPL1 in the template library to the first position of the cache. The process ends.

[0174] The advantage of this step is that when the extraction of common words fails in step (4), it can efficiently classify log messages to distinguish log messages belonging to different templates, split them into different sub - log groups, and finally extract appropriate templates for each log message, improving the parsing accuracy.

[0175] (7) Set i = i + 1 and return to step (4).

[0176] (8) Use a log splitting algorithm based on the minimum number of word types at random positions to process the current log group to obtain multiple sub - log groups.

[0177] It should be noted that when jumping from step (5) to this step (8), the current log group is the i - th log group; when returning from step (11) to this step (8), the current log group is the j - th sub - log group.

[0178] Specifically, this step includes the following sub - steps:

[0179] (8 - 1) Initialize an empty array Tss to save the number of word types at different positions.

[0180] (8 - 2) Generate a random number array RA (RandomArray) with element values ranging from 1 to L.

[0181] Set the counter p = 1;

[0182] Determine whether p is greater than the log length L of each log message in the current log group. If so, go to step (8-16); otherwise, go to step (8-5);

[0183] Take out the p-th element RA in the array RA p ;

[0184] Set an empty variable tokenType of dictionary type, whose key is used to store words, and the key value is used to store the index number of the log message in the current log group that has the same word as the key p for the RA

[0185] Set the counter q = 1;

[0186] Determine whether q is greater than the number of log messages N in the current log group. If so, go to step (8-13); otherwise, go to step (8-9);

[0187] Take out the RA p -th word tokenRA of the q-th log message p , and determine whether the word tokenRA p exists in the keys of the dictionary tokenType. If so, go to step (8-10); otherwise, go to step (8-11);

[0188] Add q to the key value array of the dictionary element whose key in the dictionary tokenType is the word tokenRA p , and go to step (8-12);

[0189] Create a new element in the dictionary tokenType with the key being the word tokenRA p and the key value being q, and go to step (8-12);

[0190] Set q = q + 1, and return to step (8-8);

[0191] Determine whether the number of key value pairs of the dictionary variable tokenType is greater than 1 and less than the log length L. If so, go to step (8-14); otherwise, go to step (8-15);

[0192] Add the dictionary tokenType to Tss, and determine whether the number of elements in the Tss array is equal to the preset value 3. If so, go to step (8-17); otherwise, go to step (8-15);

[0193] (8 - 15) Set p = p + 1, and return to step (8 - 4);

[0194] (8 - 16) Determine whether Tss is empty. If it is, enter step (8 - 18); otherwise, transfer to step (8 - 17);

[0195] (8 - 17) Select the element minTokenType with the fewest key - value pairs in Tss, and split the log messages in the key - value array corresponding to each key in minTokenType into a sub - log group, forming multiple sub - log groups, and the process ends.

[0196] (8 - 18) Each log message in the current log group forms a separate sub - log group, and the process ends.

[0197] The advantage of this step is that by splitting the log groups with the number of common words not meeting the threshold, the logs belonging to different templates are divided into different sub - log groups, and finally a suitable template is found for each log message, improving the parsing accuracy.

[0198] (9) Set the counter j = 1;

[0199] Specifically, the counter j is used to iterate over the sub - log groups obtained in step (8).

[0200] (10) Determine whether j is greater than the number of sub - log groups obtained in step (8). If it is, return to step (7); otherwise, obtain the j - th sub - log group, and use the multiple - sequence longest common subsequence algorithm applicable to log data to process the j - th sub - log group to obtain the common words corresponding to the j - th sub - log group, and then transfer to step (11);

[0201] (11) Determine whether the number of common words corresponding to the j - th sub - log group obtained in step (10) is greater than or equal to the threshold. If it is, enter step (12); otherwise, return to step (8);

[0202] The threshold in this step is exactly the same as the threshold in step (5), and will not be elaborated here.

[0203] (12) Generate a template for the j - th sub - log group and update the template library and cache, and then enter step (13);

[0204] (13) Set j = j + 1, and return to step (10);

[0205] (14) Use the cache - based nearest - neighbor matching algorithm to match the template for each log message in the log obtained in step (1) with the log messages in the dataset, and update the template library and cache.

[0206] Specifically, this step includes the following sub-steps:

[0207] (14-1) Set the counter k = 1;

[0208] (14-2) Determine whether k is greater than the number of log messages in the log matching dataset. If so, the process ends. Otherwise, retrieve the k-th log message log in the log matching dataset k , and initialize maxET_index = -1, indicating the index number of the template with the most common words with the log message log k , and initialize maxET = 0, indicating the number of common words between the log message log k and the maxET_index-th template in the template library, and proceed to step (14-3);

[0209] (14-3) Set the counter p = 1;

[0210] (14-4) Determine whether p is greater than the cache size. If so, proceed to step (14-13). Otherwise, retrieve the p-th element cache in the cache p , and retrieve the cache p -th template TPL1 in the template library, and determine whether the log length of the k-th log message log k obtained in step (14-2) is the same as the template length of the template TPL1. If so, proceed to step (14-5). Otherwise, transfer to step (14-12);

[0211] (14-5) Set ET = 0, which is used to represent the number of common words between the template content of the template TPL1 and the log message log k ;

[0212] (14-6) Set the counter r = 1;

[0213] (14-7) Determine whether r is greater than the array size of the template constant positions of the template TPL1. If so, proceed to step (14-10). Otherwise, retrieve the r-th element constant of the template constant positions of the template TPL1 r and transfer to step (14-8);

[0214] (14-8) Determine whether the constant r -th word in the template content of the template TPL1 is the same as the constant k -th word in the log message log r . If so, set ET = ET + 1, and then transfer to step (14-9). Otherwise, directly transfer to step (14-9);

[0215] (14-9) Set r = r + 1, and return to step (14-7);

[0216] (14-10) Determine whether ET is equal to the array size of the template constant positions of template TPL1. If so, it means that at all constant positions of template TPL1, the words in the template content of template TPL1 are the same as the words in the log message log k The log message log k Can exactly match template TPL1. Add the log ID of the log message log k To the log IDs of template TPL1, and move the index number cache p Of template TPL1 to the first position of the cache. The process ends. Otherwise, go to step (14-11);

[0217] (14-11) Determine whether ET is greater than or equal to the threshold and greater than maxET. If so, set maxET = ET, maxET_index = cache p , and then go to step (14-12). Otherwise, directly go to step (14-12);

[0218] The threshold in this step is exactly the same as the threshold in step (5), and will not be elaborated here.

[0219] (14-12) Set p = p + 1, and return to step (14-4);

[0220] (14-13) Set the counter q = 1;

[0221] (14-14) Determine whether q is greater than the number of templates in the template library. If so, go to step (14-22). Otherwise, take out the q-th template TPL2 in the template library and determine whether the log length of the log message log k Is the same as the template length of template TPL2. If so, go to step (14-15). Otherwise, go to step (14-21);

[0222] (14-15) Set the counter r = 1;

[0223] (14-16) Determine whether r is greater than the array size of the template constant positions of template TPL2. If so, go to step (14-19). Otherwise, take out the r-th element constant of the template constant positions of template TPL2 r Go to step (14-17);

[0224] (14-17) Determine the constant r Of the template content of template TPL2 and the log message log kthe constant of r Check if the words are the same. If so, set ET = ET + 1 and then go to step (14 - 18); otherwise, directly go to step (14 - 18).

[0225] (14 - 18) Set r = r + 1 and return to step (14 - 16).

[0226] (14 - 19) Check if ET is equal to the size of the array at the template constant positions of template TPL2. If so, it means that at all constant positions of template TPL2, the words in the template content of template TPL2 are the same as the words in the log message log k The words of k The log message log k Can fully match template TPL2. Add the log ID of the log message log

[0227] (14 - 20) Check if ET is greater than or equal to the threshold and greater than maxET. If so, set maxET = ET and maxET_index = q, and then go to step (14 - 21); otherwise, directly go to step (14 - 21).

[0228] (14 - 21) Set p = p + 1 and return to step (14 - 14).

[0229] (14 - 22) Check if maxET is greater than 0. If so, change the constants in the template content of template TPLm that are different from the log message log k To the wildcard "<*>" and delete them from the constant positions, add k to the template log IDs of template TPLm; otherwise, it means maxET has not been updated and there is no template in the template library that matches the log message log k Then go to step (14 - 23).

[0230] (14 - 23) Add the log message log k As a separate template to the template library. The template length is the length of the log message log k The template content is each word of the log message log k The template constant positions are the positions of each word of the log message log k The template log IDs are the ID of the log message log k And add its index number in the template library to the first position of the cache. The process ends.

[0231] The advantages of this step are as follows. First, taking advantage of the spatial locality of log data, when matching templates for each log message, the templates most recently accessed are first matched in the cache. Only when the entire cache misses are the templates in the entire template library matched, which improves the parsing speed. The spatial locality of log data is as shown in Figure 2 Figure 2. For each data set, 100 consecutive log messages are randomly selected to observe their template distributions. Figure 2 (a) shows that only 2 out of 377 templates appear in 100 randomly selected log messages in the BGL data set. Figure 2 (b) shows that only 6 out of 46 templates appear in 100 randomly selected log messages in the HDFS data set. Figure 2 (c) shows that only 2 out of 90 templates appear in 100 randomly selected log messages in the HPC data set. Figure 2 (d) shows that only 5 out of 87 templates appear in 100 randomly selected log messages in the Zookeeper data set. Therefore, using the cache to save the templates most recently used can effectively avoid traversing the entire template library. Second, during the matching of log messages to templates, common words are extracted a second time. When a template that exactly matches can be found, the exactly matching template is selected; otherwise, the template with the highest similarity (the largest number of common words) that meets the threshold is selected. When neither of the above two conditions is met, a new template is finally created, which improves the parsing accuracy.

[0232] All in all, the main idea of the present invention is to divide the log data set into a template extraction data set and a log matching data set according to a set ratio. For the template extraction data set: the log messages in the template extraction data set are divided into different log groups according to the log length, and the multi-sequence longest common subsequence algorithm applicable to log data is executed on the log messages in the log group to extract the common words of the log group. It is judged whether the number of common words meets the threshold. When it is greater than or equal to the threshold, a template is generated for the log group, the template is added to the template library and the cache is updated. When it is less than the threshold, the log splitting method based on the minimum number of word types at random positions is iteratively executed on the log group to split the log messages in the log group into smaller sub-log groups, and the multi-sequence longest common subsequence algorithm applicable to log data is re-executed on the sub-log groups. For the log matching data set: templates are matched for each log message in the cache that stores the template indexes. If the match hits, the template library and the cache are updated. Otherwise, template matching is performed in the template library. If the match hits, the template library and the cache are updated. Otherwise, the log message is added to the template library as a new template, and the template library and the cache are updated.

[0233] Experimental Results

[0234] The parsing accuracy and efficiency of the fast log parsing method (MLCS-Cache) for log locality features proposed in the present invention are verified through comparative experiments on real datasets. The present invention is compared with three of the most excellent recent co-workers on 4 real datasets, including Drain, IPLoM, and Spell. The metrics for the comparative experiments are: parsing efficiency (time / second), parsing accuracy (percentage), and F-measure (percentage). Among them, parsing efficiency refers to the time required for the parsing method to complete parsing, and the shorter the time, the higher the efficiency performance. Parsing accuracy refers to the percentage of correctly classified log messages among all log messages, and the higher the percentage, the better the accuracy performance. F-measure refers to the parsing accuracy, and the higher the percentage, the better. In addition, the present invention also verifies the effectiveness of the cache-based nearest matching method proposed in the present invention by comparing the parsing efficiency when using the cache and not using the cache.

[0235] Table 1 shows an overview of the datasets used in the comparative experiments.

[0236] Table 2 shows the parsing accuracies of four log parsing methods including the present invention on four datasets. The highest parsing accuracy is marked with bold numbers. It can be seen from the table that the method proposed in the present invention achieves the highest parsing accuracy on 3 datasets except Zookeeper, and the difference in parsing accuracy between the present invention and the method Drain with the highest parsing accuracy on Zookeeper is also almost negligible.

[0237] Table 3 shows the F-measures of four log parsing methods including the present invention on four datasets. It can be seen from the table that the method proposed in the present invention obtains the best scores on each dataset.

[0238] Table 4 shows the parsing efficiencies of four log parsing methods including the present invention on four datasets. The best parsing efficiency is marked with bold numbers. It can be seen from the table that the method proposed in the present invention achieves the best efficiency on two large datasets, BGL and HDFS. On Zookeeper, it achieves the best efficiency simultaneously with IPLoM. On HPC, although the efficiency of this method ranks second, the difference from IPLoM with the best efficiency is only 4.5%.

[0239] Table 5 shows the parsing efficiencies of the present invention when using the cache (MLCS-Cache) and not using the cache (MLCS) on four datasets. It can be seen from the table that the method of using the cache for nearest matching can effectively improve the parsing efficiency.

[0240] Table 1 Dataset Overview

[0241]

[0242] Table 2 Parsing Accuracy

[0243]

[0244]

[0245] Table 3 F-measure

[0246]

[0247] Table 4 Parsing Efficiency (Time / Second)

[0248]

[0249] Table 5 Efficiency Comparison between Without Using Cache and Using Cache (Time / Second)

[0250]

[0251] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A fast log parsing method for log locality characteristics, characterized in that, It includes the following steps: (1) Obtain a log dataset composed of multiple log messages of a computer, and divide the obtained log dataset according to a ratio to obtain a template extraction dataset and a log matching dataset; (2) Divide the template extraction dataset obtained in step (1) into multiple log groups according to the log lengths of the respective log messages in the template extraction dataset; (3) Set the counter i = 1; (4) Determine whether i is greater than the number of log groups obtained in step (2). If so, go to step (14). Otherwise, obtain the i-th log group, and use the multiple sequence longest common subsequence algorithm MLCS to process the i-th log group to obtain the common words corresponding to the i-th log group, and then go to step (5); (5) Determine whether the number of common words corresponding to the i-th log group obtained in step (4) is greater than or equal to the threshold. If so, go to step (6). Otherwise, go to step (8); (6) Generate a template for the i-th log group, and update the template library and cache according to the generated template, and then go to step (7); (7) Set i = i + 1, and return to step (4); (8) Use a log splitting algorithm based on the minimum number of word types at random positions to process the current log group to obtain multiple sub-log groups; (9) Set the counter j = 1; (10) Determine whether j is greater than the number of sub-log groups obtained in step (8). If so, return to step (7). Otherwise, obtain the j-th sub-log group, and use the multiple sequence longest common subsequence algorithm applicable to log data to process the j-th sub-log group to obtain the common words corresponding to the j-th sub-log group, and then go to step (11); (11) Determine whether the number of common words corresponding to the j-th sub-log group obtained in step (10) is greater than or equal to the threshold. If so, go to step (12). Otherwise, return to step (8); (12) Generate a template for the j-th sub-log group and update the template library and cache, and then go to step (13); (13) Set j = j + 1, and return to step (10); (14) Use a cache-based nearest matching algorithm to match templates for each log message in the log matching dataset obtained in step (1), and update the template library and cache.

2. The fast log parsing method for log locality characteristics according to claim 1, characterized in that, The process of processing the log group in step (4) to obtain the common words corresponding to the log group includes the following sub-steps: (4-1) Initialize the common word set CTokens, which is initially an empty string array and is used to represent the common word set; (4-2) Set the counter p = 1; (4-3) Determine whether p is greater than the log length L of each log message in the i-th log group. If so, the process ends. Otherwise, go to step (4-4); (4-4) Set the counter q = 2; (4-5) Determine whether q is greater than the number N of log messages in the i-th log group. If so, it means that the p-th word in all log messages in the i-th log group is the same, and then go to step (4-7). Otherwise, go to step (4-6); (4-6) Determine whether the p-th word in the q-th log message in the i-th log group is equal to the p-th word in the first log message in the i-th log group. If so, go to step (4-8); otherwise, go to step (4-9). (4-7) Add the p-th words in all log messages in the i-th log group to the common word set CTokens, and go to step (4-9). (4-8) Set q = q + 1, and return to step (4-5). (4-9) Set p = p + 1, and return to step (4-3).

3. The fast log parsing method for log locality characteristics according to claim 2, characterized in that, In step (5), the calculation formula for the threshold Threshold is as follows:

4. The fast log parsing method for log locality characteristics according to any one of claims 1 to 3, characterized in that, (6-1) Generate a template TPL1 for the i-th log group; (6-2) Set the counter p = 1; (6-3) Compare the template content of template TPL1 with the template content of template TPL2; (6-3) Determine whether p is greater than the cache size. If so, go to step (6-7); otherwise, retrieve the p-th element cache in the cache. p , and retrieve the cache-th p template TPL2 in the template library. Determine whether the template lengths of the template TPL1 obtained in step (6-1) and the template TPL2 are the same. If so, go to step (6-4); otherwise, go to step (6-6). (6-4) Determine whether the template content of template TPL1 is the same as that of template TPL2. If so, it means that template TPL1 and template TPL2 match successfully, and go to step (6-5); otherwise, go to step (6-6). (6-5) Add the template log IDs of template TPL1 to the template log IDs of template TPL2, and move the index number cache p of template TPL2 to the first position of the cache, and the process ends; (6-6) Set p = p + 1, and return to step (6-3); (6-7) Set the counter q = 1; (6-8) Determine whether q is greater than the number of templates in the template library. If so, go to step (6-13); otherwise, retrieve the q-th template TPL3 in the template library and go to step (6-9). (6-9) Determine whether the template lengths of template TPL1 and template TPL3 are the same. If so, go to step (6-10); otherwise, go to step (6-12). (6-10) Determine whether the template content of template TPL1 is the same as that of template TPL3. If so, it means that template TPL1 and template TPL3 match successfully, and go to step (6-11); otherwise, go to step (6-12). (6-11) Add the template log IDs of template TPL1 to the template log IDs of template TPL3, and add the index number of template TPL3 to the first position of the cache. The process ends. (6-12) Set q = q + 1, and return to step (6-8); (6-13) Add template TPL1 to the template library, and add the index number of template TPL1 in the template library to the first position of the cache. The process ends.

5. The fast log parsing method for log locality characteristics according to claim 4, wherein, (6-1) The process of generating a template for the i-th log group in step (6-1) includes the following sub-steps: (6-1-1) Create a template TPL, initialize the template length of template TPL to the log length L of each log message in the i-th log group, initialize the template content of template TPL to an empty string array of size L, initialize the template constant positions of template TPL to an empty array, and initialize the template log IDs of template TPL to an empty array; (6-1-2) Retrieve the first log message log1 in the i-th log group; (6-1-3) Set the counter q = 1; (6-1-4) Determine whether q is greater than L. If so, go to step (6-1-6); otherwise, determine whether the q-th word of log1 exists in the common word set of the i-th log group. If so, set the q-th element of the template content of the template TPL created in step (6-1-1) to the q-th word of log1, and add q to the constant positions; otherwise, set the q-th element of the template content of the template TPL to the wildcard "<*>”, and then go to step (6-1-5); (6-1-5) Set q = q + 1, and return to step (6-1-4); (6-1-6) Add the log IDs of all log messages in the i-th log group to the log IDs of the template TPL, and the process ends.

6. The fast log parsing method for log locality characteristics according to claim 1, wherein, (8) Step (8) includes the following sub-steps: (8-1) Initialize an empty array Tss to save the types of words at different positions; (8-2) Generate a random number array RA with element values ranging from 1 to L; (8-3) Set the counter p = 1; (8-4) Determine whether p is greater than the log length L of each log message in the current log group. If so, go to step (8-16); otherwise, go to step (8-5); Retrieve the p-th element RA in the array RA p ; (8-6) Set an empty variable tokenType of dictionary type, whose key is used to store words and whose key value is used to store the index number of the log message whose RA p th word in the current log group is the same as the key; (8-7) Set the counter q = 1; (8-8) Determine whether q is greater than the number of log messages N in the current log group. If so, go to step (8-13); otherwise, go to step (8-9); (8 - 9) Retrieve the RA-th word of the q-th log message p tokenRA p , and determine whether the word tokenRA p exists in the keys of the dictionary tokenType. If so, proceed to step (8 - 10); otherwise, go to step (8 - 11). (8 - 10) Add q to the value array of the dictionary element whose key in the dictionary tokenType is the word tokenRA p , and proceed to step (8 - 12); (8-11) Create a new element with the key as the word tokenRA and the key value as q in the dictionary tokenType, and proceed to step (8-12); p ​ (8-12) Set q = q + 1, and return to step (8-8); (8-13) Determine whether the number of key-value pairs of the dictionary variable tokenType is greater than 1 and less than the log length L. If so, go to step (8-14); otherwise, go to step (8-15); (8-14) Add the dictionary tokenType to Tss, and determine whether the number of elements in the Tss array is equal to the preset value 3. If so, go to step (8-17); otherwise, go to step (8-15); (8-15) Set p = p + 1, and return to step (8-4); (8-16) Determine whether Tss is empty. If so, go to step (8-18); otherwise, go to step (8-17); (8-17) Select the element minTokenType with the fewest key-value pairs in Tss, and split the log messages in the key-value array corresponding to each key in minTokenType into a sub-log group to form multiple sub-log groups, and the process ends; (8-18) Each log message in the current log group forms a separate sub-log group, and the process ends.

7. The fast log parsing method for log locality characteristics according to claim 1, wherein, (14) Step (14) includes the following sub-steps: (14-1) Set the counter k = 1; (14-2) Determine whether k is greater than the number of log messages in the log matching dataset. If so, end the process; otherwise, retrieve the k-th log message log from the log matching dataset k , initialize maxET_index = -1, indicating the index number of the template with the most common words with the log message log k , initialize maxET = 0, indicating the number of common words between the log message log k and the template at the maxET_index-th position in the template library, and proceed to step (14-3); (14-3) Set the counter p = 1; (14-4) Determine whether p is greater than the cache size. If so, go to step (14-13); otherwise, retrieve the p-th element cache from the cache. p , and retrieve the cache-th p template TPL1 from the template library, and determine whether the log length of the k-th log message log k obtained in step (14-2) is the same as the template length of template TPL1. If so, go to step (14-5); otherwise, go to step (14-12). (14-5) Set ET = 0, which is used to represent the number of common words between the template content of template TPL1 and the log message log k ; (14-6) Set the counter r = 1; (14-7) Determine whether r is greater than the size of the array of template constant positions of template TPL1. If so, go to step (14-10); otherwise, take out the r-th element constant of the template constant positions of template TPL1 r Go to step (14-8); (14-8) Determine whether the constant-th word of the template content of the template TPL1 is the same as the constant-th word of the log message log. If so, set ET = ET + 1, and then go to step (14-9); otherwise, directly go to step (14-9). r Otherwise, directly go to step (14-9). k Determine whether the constant-th word of the template content of the template TPL1 is the same as the constant-th word of the log message log. If so, set ET = ET + 1, and then go to step (14-9); otherwise, directly go to step (14-9). r Otherwise, directly go to step (14-9). (14-9) Set r = r + 1, and return to step (14-7); (14-10) Determine whether ET is equal to the array size of the template constant positions of template TPL1. If so, it means that for all constant positions of template TPL1, the words in the template content of template TPL1 are the same as the words in log message log k and the log message log k can fully match template TPL1. Add the log ID of log message log k to the log IDs of template TPL1, and move the index number cache of template TPL1 p to the first position of the cache. The process ends. Otherwise, go to step (14-11); (14-11) Determine whether ET is greater than or equal to the threshold and greater than maxET. If so, set maxET = ET and maxET_index = cache p , and then go to step (14-12). Otherwise, directly go to step (14-12); (14-12) Set p = p + 1, and return to step (14-4); (14-13) Set the counter q = 1; (14-14) Determine whether q is greater than the number of templates in the template library. If so, proceed to step (14-22). Otherwise, retrieve the q-th template TPL2 from the template library and determine whether the log length of the log message log k is the same as the template length of the template TPL2. If so, proceed to step (14-15). Otherwise, transfer to step (14-21); (14-15) Set the counter r = 1; (14-16) Determine whether r is greater than the size of the array of template constant positions of template TPL2. If so, proceed to step (14-19); otherwise, retrieve the r-th element constant of the template constant positions of template TPL2 r Proceed to step (14-17); (14-17) Determine whether the constant r th word of the template content of template TPL2 is the same as the constant k th word of the log message log r ; if so, set ET = ET + 1, then go to step (14-18), otherwise directly go to step (14-18); (14-18) Set r = r + 1, and return to step (14-16); (14-19) Determine whether ET is equal to the array size of the template constant positions of template TPL2. If so, it means that at all constant positions of template TPL2, the words in the template content of template TPL2 are the same as the words in log message log k ; the log message log k can fully match template TPL2. Add the log ID of the log message log k to the log IDs of template TPL2, move the index number q of template TPL2 to the first position of the cache, and the process ends. Otherwise, go to step (14-20); (14 - 20) Determine whether ET is greater than or equal to the threshold and greater than maxET. If so, set maxET = ET, maxET_index = q, and then transfer to step (14 - 21); otherwise, directly transfer to step (14 - 21). (14 - 21) Set p = p + 1 and return to step (14 - 14). (14-22) Determine whether maxET is greater than 0. If so, change the constants in the template content of template TPLm that are different from the log message log k to the wildcard "<*>" and delete them from their constant positions, and add k to the template log IDs of template TPLm. Otherwise, it means that maxET has not been updated and there is no template in the template library that matches the log message log k Then proceed to step (14-23); (14-23) Add the log message log k as a separate template to the template library, where the template length is the length of the log message log k and the template content is each word of the log message log k The template constant position is the position of each word of the log message log k The template log IDs are the IDs of the log message log k and add the index number of it in the template library to the first position of the cache, and the process ends.

8. A fast log parsing system for log locality characteristics, wherein, (14 - 22) Set p = p + 1 and return to step (14 - 14). (14 - 23) Set p = p + 1 and return to step (14 - 14). (14 - 24) Set p = p + 1 and return to step (14 - 14). (14 - 25) Set p = p + 1 and return to step (14 - 14). (14 - 26) Set p = p + 1 and return to step (14 - 14). (14 - 27) Set p = p + 1 and return to step (14 - 14). (14 - 28) Set p = p + 1 and return to step (14 - 14). (14 - 29) Set p = p + 1 and return to step (14 - 14). (14 - 30) Set p = p + 1 and return to step (14 - 14). (14 - 31) Set p = p + 1 and return to step (14 - 14). (14 - 32) Set p = p + 1 and return to step (14 - 14). (14 - 33) Set p = p + 1 and return to step (14 - 14). (14 - 34) Set p = p + 1 and return to step (14 - 14). (14 - 35) Set p = p + 1 and return to step (14 - 14). (14 - 36) Set p = p + 1 and return to step (14 - 14). (14 - 37) Set p = p + 1 and return to step (14 - 14). (14 - 38) Set p = p + 1 and return to step (14 - 14). (14 - 39) Set p = p + 1 and return to step (14 - 14). (14 - 40) Set p = p + 1 and return to step (14 - 14). (14 - 41) Set p = p + 1 and return to step (14 - 14). (14 - 42) Set p = p + 1 and return to step (14 - 14). (14 - 43) Set p = p + 1 and return to step (14 - 14). (14 - 44) Set p = p + 1 and return to step (14 - 14). (14 - 45) Set p = p + 1 and return to step (14 - 14). (14 - 46) Set p = p + 1 and return to step (14 - 14). (14 - 47) Set p = p + 1 and return to step (14 - 14). (14 - 48) Set p = p + 1 and return to step (14 - 14). (14 - 49) Set p = p + 1 and return to step (14 - 14). (14 - 50) Set p = p + 1 and return to step (14 - 14). (14 - 51) Set p = p + 1 and return to step (14 - 14). (14 - 52) Set p = p + 1 and return to step (14 - 14). (14 - 53) Set p = p + 1 and return to step (14 - 14). (14 - 54) Set p = p + 1 and return to step (14 - 14). (14 - 55) Set p = p + 1 and return to step (14 - 14). (14 - 56) Set p = p + 1 and return to step (14 - 14). (14 - 57) Set p = p + 1 and return to step (14 - 14). (14 - 58) Set p = p + 1 and return to step (14 - 14). (14 - 59) Set p = p + 1 and return to step (14 - 14). (14 - 60) Set p = p + 1 and return to step (14 - 14). (14 - 61) Set p = p + 1 and return to step (14 - 14). (14 - 62) Set p = p + 1 and return to step (14 - 14). (14 - 63) Set p = p + 1 and return to step (14 - 14). (14 - 64) Set p = p + 1 and return to step (14 - 14). (14 - 65) Set p = p + 1 and return to step (14 - 14). (14 - 66) Set p = p + 1 and return to step (14 - 14). (14 - 67) Set p = p + 1 and return to step (14 - 14). (14 - 68) Set p = p + 1 and return to step (14 - 14). (14 - 69) Set p = p + 1 and return to step (14 - 14). (14 - 70) Set p = p + 1 and return to step (14 - 14). (14 - 71) Set p = p + 1 and return to step (14 - 14). (14 - 72) Set p = p + 1 and return to step (14 - 14). (14 - 73) Set p = p + 1 and return to step (14 - 14). (14 - 74) Set p = p + 1 and return to step (14 - 14). (14 - 75) Set p = p + 1 and return to step (14 - 14). (14 - 76) Set p = p + 1 and return to step (14 - 14). (14 - 77) Set p = p + 1 and return to step (14 - 14). (14 - 78) Set p = p + 1 and return to step (14 - 14). (14 - 79) Set p = p + 1 and return to step (14 - 14). (14 - 80) Set p = p + 1 and return to step (14 - 14). (14 - 81) Set p = p + 1 and return to step (14 - 14). (14 - 82) Set p = p + 1 and return to step (14 - 14). (14 - 83) Set p = p + 1 and return to step (14 - 14). (14 - 84) Set p = p + 1 and return to step (14 - 14). (14 - 85) Set p = p + 1 and return to step (14 - 14). (14 - 86) Set p = p + 1 and return to step (14 - 14). (14 - 87) Set p = p + 1 and return to step (14 - 14). (14 - 88) Set p = p + 1 and return to step (14 - 14). (14 - 89) Set p = p + 1 and return to step (14 - 14). (14 - 90) Set p = p + 1 and return to step (14 - 14). (14 - 91) Set p = p + 1 and return to step (14 - 14). (14 - 92) Set p = p + 1 and return to step (14 - 14). (14 - 93) Set p = p + 1 and return to step (14 - 14). (14 - 94) Set p = p + 1 and return to step (14 - 14). (14 - 95) Set p = p + 1 and return to step (14 - 14). (14 - 96) Set p = p + 1 and return to step (14 - 14). (14 - 97) Set p = p + 1 and return to step (14 - 14). (14 - 98) Set p = p + 1 and return to step (14 - 14). (14 - 99) Set p = p + 1 and return to step (14 - 14). (14 - 100) Set p = p + 1 and return to step (14 - 14). (14 - 101) Set p = p + 1 and return to step (14 - 14). (14 - 102) Set p = p + 1 and return to step (14 - 14). (14 - 103) Set p = p + 1 and return to step (14 - 14). (14 - 104) Set p = p + 1 and return to step (14 - 14). (14 - 105) Set p = p + 1 and return to step (14 - 14). (14 - 106) Set p = p + 1 and return to step (14 - 14). (14 - 107) Set p = p + 1 and return to step (14 - 14). (14 - 108) Set p = p + 1 and return to step (14 - 14). (14 - 109) Set p = p + 1 and return to step (14 - 14). (14 - 110) Set p = p + 1 and return to step (14 - 14). (14 - 111) Set p = p + 1 and return to step (14 - 14). (14 - 112) Set p = p + 1 and return to step (14 - 14). (14 - 113) Set p = p + 1 and return to step (14 - 14). (14 - 114) Set p = p + 1 and return to step (14 - 14). (14 - 115) Set p = p + 1 and return to step (14 - 14). (14 - 116) Set p = p + 1 and return to step (14 - 14). (14 - 117) Set p = p + 1 and return to step (14 - 14). (14 - 118) Set p = p + 1 and return to step (14 - 14). (14 - 119) Set p = p + 1 and return to step (14 - 14). (14 - 120) Set p = p + 1 and return to step (14 - 14). (14 - 121) Set p = p + 1 and return to step (14 - 14). (14 - 122) Set p = p + 1 and return to step (14 - 14). (14 - 123) Set p = p + 1 and return to step (14 - 14). (14 - 124) Set p = p + 1 and return to step (14 - 14). (14 - 125) Set p = p + 1 and return to step (14 - 14). (14 - 126) Set p = p + 1 and return to step (14 - 14). (14 - 127) Set p = p + 1 and return to step (14 - 14). (14 - 128) Set p = p + 1 and return to step (14 -

Citation Information

Patent Citations

  • Online log analysis method and system and electronic terminal equipment thereof

    CN110888849A

  • Online analysis method and system of multi-source log, electronic equipment and storage medium

    CN115437877A