Log analysis method, device and equipment and computer readable storage medium
By building a hierarchical tree model and spanning tree technology, log analysis is automatically processed, and the problem of low automation in the existing technology is solved, and efficient and accurate log analysis is achieved.
Patent Information
- Application Number
- CN202510843904.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In the prior art, log parsing methods are low in automation, and require manual writing of parsing files and configuring regular expressions, resulting in inefficiency.
By obtaining the feature information of the log, a hierarchical tree model with the minimum spanning tree is constructed, clustered, and analytical template spanning tree is generated to realize automated log parsing.
It realizes the automation of the entire log analysis process, improves analysis efficiency, reduces manual intervention, and improves analysis accuracy and scope of application.
Smart Images

Figure CN120337865A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information processing technology, and in particular, to a log parsing method, apparatus, device, and computer-readable storage medium. Background Art
[0002] In the related art, the conventional process of parsing logs through ELK usually requires manually writing a parsing file according to the log content. This file consists of multiple modules, and each module is specifically used to parse and process logs of different formats or types. This method can almost handle all types of log parsing requirements, but its significant defect is that it requires manual writing of the parsing file and configuration of regular expressions for log parsing, etc. Some technical solutions parse logs by pre-setting a parsing rule library, and among them, the parsing rule library still requires manual intervention.
[0003] It can be seen that the log parsing method in the related art has the defect of low automation. Summary of the Invention
[0004] Embodiments of this application provide a log parsing method, apparatus, device, and computer-readable storage medium, which can solve the technical problem of low automation of the log parsing method.
[0005] In a first aspect, an embodiment of this application provides a log parsing method, including: Obtaining the feature information of each log in the log dataset to be parsed; Constructing a hierarchical tree model corresponding to the feature information based on the minimum spanning tree; Performing clustering processing on the logs in the log dataset to be parsed based on the hierarchical tree model, and obtaining the tree structure corresponding to each clustering type in the clustering result according to the clustering result; Performing a spanning tree process on the logs of the target clustering type according to the target clustering type and the target tree structure corresponding to the target clustering type, to obtain a parsing template spanning tree of the target clustering type, where the clustering types in the clustering result include the target clustering type; Parsing the logs of the target clustering type according to the parsing template spanning tree to obtain a log parsing result.
[0006] In a second aspect, an embodiment of this application further provides a log parsing apparatus, including: A first obtaining module, configured to obtain the feature information of each log in the log dataset to be parsed; A model construction module, configured to construct a hierarchical tree model corresponding to the feature information based on the minimum spanning tree; A clustering module, configured to perform clustering processing on the logs in the to-be-parsed log dataset based on the hierarchical tree model, and obtain the tree structure corresponding to each clustering type in the clustering result according to the clustering result; A spanning tree module, configured to perform spanning tree processing on the logs of the target clustering type according to the target clustering type and the target tree structure corresponding to the target clustering type, so as to obtain a parsing template spanning tree of the target clustering type, where the clustering types in the clustering result include the target clustering type; A log parsing module, configured to parse the logs of the target clustering type according to the parsing template spanning tree to obtain a log parsing result.
[0007] In a third aspect, an embodiment of the present application further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the steps in the log parsing method described above are implemented.
[0008] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps in the log parsing method described above are implemented.
[0009] In a fifth aspect, an embodiment of the present application further provides a computer program product, including computer instructions, where when the computer instructions are executed by a processor, the steps in the log parsing method described above are implemented.
[0010] In the embodiments of the present application, the feature information of each log in the log dataset to be parsed is obtained; a hierarchical tree model corresponding to the feature information is constructed based on the minimum spanning tree; the logs in the log dataset to be parsed are clustered based on the hierarchical tree model, and a tree structure corresponding to each clustering type in the clustering result is obtained according to the clustering result; according to the target clustering type and the target tree structure corresponding to the target clustering type, a spanning tree process is performed on the logs of the target clustering type to obtain a parsing template spanning tree of the target clustering type, where the clustering types in the clustering result include the target clustering type; the logs of the target clustering type are parsed according to the parsing template spanning tree to obtain a log parsing result. After obtaining the feature information of the logs, a hierarchical tree model corresponding to the feature information is constructed based on the minimum spanning tree, and the logs in the log dataset to be parsed are clustered based on this hierarchical tree model. Finally, based on the spanning tree technology, a parsing template spanning tree is automatically generated for the logs of each clustering type as the parsing rule applicable to the logs of this clustering type, so as to achieve efficient parsing of various types of logs. This method can realize the automation of the entire log parsing process and significantly improve the log parsing efficiency. Among them, the clustering process is optimized based on the idea of hierarchical clustering, so that the hierarchical tree model can automatically obtain the optimal clustering result only by giving the minimum samples included in the cluster, avoiding the process of manual parameter tuning. On the basis of improving the automation of the log parsing process, the accuracy and application scope of the hierarchical tree model can also be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the accompanying drawings required for describing the embodiments of the present application will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 is one of the flowcharts of the log parsing method provided by the embodiments of the present application; Figure 2 is a schematic diagram of the update process of the target tree structure provided by the embodiments of the present application; Figure 3 is another flowchart of the log parsing method provided by the embodiments of the present application; Figure 4 is a structural diagram of the log parsing device provided by the embodiments of the present application; Figure 5 is a structural diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0014] Next, in conjunction with the accompanying drawings, through specific embodiments and their application scenarios, the log parsing method, log parsing device, electronic device, computer-readable storage medium, and computer program product provided by the embodiments of the present application will be described in detail.
[0015] See Figure 1 , Figure 1 is the flowchart of the log parsing method provided by the embodiments of the present application. As Figure 1 shown, it includes the following steps: Step 101, obtain the feature information of each log in the log dataset to be parsed.
[0016] In some embodiments, the logs in the log dataset to be parsed may include various types of logs, such as system logs, application logs, security logs, etc.
[0017] In some embodiments, the logs in the log dataset to be parsed may include logs in diverse computing environments. Among them, the diverse computing environments include but are not limited to various hosts and server systems, as well as the operating systems, applications, and various software components running in these systems. A large amount of log information will be generated in this diverse computing environment, and the nature, application scope, and scale of these log data show diversity. The log files contain rich information resources, and for users, this information has high value. By analyzing different types of logs, users can extract diverse intelligence.
[0018] For the convenience of description, in the embodiments of the present application, the following email service logs are usually used as examples for illustration: Oct 6 00:00:13 localhost postfix / smtpd
[26072] : disconnect from n169-113.mail.139.com[120.232.169.113] ehlo=1 mail=1rcpt=9 / 10 data=1 quit=1 commands=13 / 14.
[0019] By parsing the above-mentioned email service logs, key information can be extracted, including the timestamp of the event occurrence, the client identifier that initiated the connection, the Internet Protocol (IP) address of the client, and the statistical data of relevant command executions, etc. However, unprocessed log data often has poor readability, which limits in-depth analysis and statistical research of its content. Therefore, in order to effectively utilize log information, it must be parsed.
[0020] In some embodiments, the log collection strategy in the log dataset can be mainly divided into two methods: local collection and remote collection according to different applications or components. Local collection involves directly extracting data from the log storage location using standard tools when the known log file storage path and corresponding permissions are available. Remote collection requires the target to be collected to configure a log tool, specify the local log source and the sending target (i.e., the central log server), and configure it on the central log server to complete the reception of the logs. The remote collection process includes installing and configuring a log collection tool on the log source device, establishing a secure transmission mechanism, and setting up reception rules on the central server to ensure the effective transmission and storage of log data.
[0021] In some embodiments, after the original log data is collected, these original log data can be preprocessed to obtain a log dataset to be parsed.
[0022] Among them, the preprocessing of the original log data can include at least one of the following: Filtering invalid data. Based on the fact that there may be a large amount of invalid data in the original log data, such as error records, duplicate records, blank lines, comments, duplicate log entries, etc., by cleaning the original log data, the efficiency of the logs in the log dataset to be parsed can be improved; Format unification. Unify the format of the logs. The sources of the original log data are diverse, and the formats of logs from different sources need to be unified, such as converting the logs into json or xml format uniformly; Outlier processing. The log data can be detected by an outlier detection tool to find the outliers in it, and the outliers can be changed or deleted; Null value processing. The missing null value data in the logs can be identified, and the null values can be filled or deleted, etc.
[0023] In some embodiments, for each log in the log dataset to be parsed, feature extraction algorithms, Artificial Intelligence (AI) models, machine learning models, etc. can be used to extract features from the logs to obtain feature information.
[0024] In some other embodiments, for each log in the log dataset to be parsed, the log can be divided into multiple segmented words, and the feature information of the log can be determined according to the word frequency and inverse document frequency index of the segmented words.
[0025] As an alternative embodiment, the obtaining of the feature information of each log in the log dataset to be parsed includes: Performing feature extraction processing on each log in the log dataset to be parsed respectively to obtain the feature information of each log; Among them, the feature extraction processing includes: Based on a regular word segmenter or a dictionary, performing first word segmentation processing on the log to obtain U segmented words of the log, where U is an integer greater than 1; Determining a first feature vector according to the feature values of the U segmented words, where the first feature vector is a U-dimensional vector; Performing dimensionality reduction processing on the first feature vector to obtain a second feature vector, where the feature information of each log includes the second feature vector, and the second feature vector is a two-dimensional feature vector.
[0026] In some embodiments, performing first word segmentation processing on the log based on a regular word segmenter or a dictionary can segment each word in the obtained log, and these segmented words can be segmented words with linguistic meanings, that is, these segmented words may not include character segments consisting entirely of symbols and numbers.
[0027] For example: for the following email service log: Oct 6 00:00:13 localhost postfix / smtpd
[26072] : disconnect from n169-113.mail.139.com[120.232.169.113] ehlo=1 mail=1rcpt=9 / 10 data=1 quit=1 commands=13 / 14.
[0028] After performing the first word segmentation processing, the U segmented words obtained are respectively: ['localhost', 'postfixsmtpd', 'disconnect', 'from', 'n169-113.mail.139.com', 'ehlo','mail', 'rcpt', 'data', 'quit', 'commands'], in this embodiment, U is equal to 11.
[0029] Then, calculate the feature value for each of the above segmented words one by one.
[0030] For example, the term frequency-inverse document frequency (TF-IDF) of each segmented word can be calculated using the following formula: tf-idf(t,d)=tf(t,d)×idf(t); Among them, tf(t,d) represents the term frequency of the segmented word t in the document d. For example, tf(t,d) = the number of occurrences of the segmented word t in the document d / the total number of segmented words in the document d. The document d can be the document composed of all the segmented words in the current log; idf(t) is the inverse document frequency of the segmented word t, which is used to measure the rarity of the segmented word t. For example, idf(t)=log(N / df(t)), where N represents the number of logs in the log dataset to be parsed, and df(t) represents the number of logs containing the segmented word t.
[0031] For example: Suppose each log in the log dataset to be parsed has the segmented word "localhost", then the TF-IDF eigenvalue of the segmented word "localhost" in the log is tf-idf = 1 / 11 * log(1) = 0.
[0032] It is worth noting that considering that TF-IDF is simple to implement, has low computational cost, and can meet the requirement that the log content contains less semantic information, using the TF-IDF value as the eigenvalue of the segmented word can improve the efficiency of extracting the first feature information.
[0033] In this embodiment, the first feature vector can be obtained by combining the eigenvalues of U segmented words. The dimension of the first feature vector is equal to U. Since most methods for extracting text feature vectors generally obtain text vectors with a very high dimension in the end, such as 10 dimensions, 20 dimensions or even higher, the dimension of the first feature vector is relatively high. And high-dimensional data is likely to lead to poor distinguishability of the data, resulting in overfitting problems in clustering analysis. By performing dimensionality reduction processing on the first feature vector, the dimension of the obtained second feature vector can be reduced to two dimensions. In this way, in subsequent processing based on the second feature vector, such as performing minimum spanning tree processing to construct a hierarchical tree model and clustering the logs in the log dataset to be parsed based on the hierarchical tree model, overfitting can be reduced and the efficiency of these processing processes can be improved.
[0034] In some embodiments, to perform dimensionality reduction processing on the first feature vector, the gradient descent method can be used to find the optimal solution of the second feature vector.
[0035] In some embodiments, the t-Distributed Stochastic Neighbor Embedding (t-SNE) algorithm can be used to perform dimensionality reduction on the first feature vectors. When this algorithm uses the gradient descent method to find the optimal solution of the second feature vectors, it uses a symmetric loss function, and uses the t-distribution to construct the data distribution in the low-dimensional space to solve the problem of data crowding.
[0036] For example: The following parameters can be input to the t-SNE algorithm: N first feature vectors of U dimensions {X1, X2, X3..., Xn}, the number of iterations T, the learning efficiency β, and the momentum a(t); then, N second feature vectors of two dimensions output by the t-SNE algorithm can be obtained.
[0037] Among them, the logical process of reducing the U-dimensional first feature vectors to two-dimensional second feature vectors based on the t-SNE algorithm is similar to the process of dimensionality reduction based on the t-SNE algorithm in the related art.
[0038] For example: First, calculate the conditional probability P in the high-dimensional space j|i , then use the normal distribution N(0, 10 -4 ), randomly initialize Y, iterate from round 1 to round T, and calculate the conditional probability q in the low-dimensional space ij and the gradient of the loss function C(y i ) with respect to y i , and update the value of Y accordingly, so that the similarity in the high-dimensional space (P j|i distribution) and the similarity in the low-dimensional space (q ij distribution) are as identical as possible. By minimizing the Kullback-Leibler (KL) divergence using the gradient descent method, it is ensured that the low-dimensional second feature vectors can retain the structure of the high-dimensional first feature vectors to the greatest extent. Step 102, construct a hierarchical tree model corresponding to the feature information based on the minimum spanning tree.
[0039] In some embodiments, taking the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm as an example, the clustering process of this DBSCAN algorithm does not depend on the number of clusters, can automatically discover the cluster structure in the data; and is not sensitive to outliers and noise. However, the two parameters of the minimum step size (Minpts) and the neighborhood radius (Eps) in its optimization process need to be continuously adjusted manually following the clustering process.
[0040] In this embodiment, the DBSCAN algorithm can be improved based on the idea of hierarchical clustering. A minimum spanning tree is used to construct a hierarchical tree model between the second feature vectors, enabling the model to automatically obtain the optimal clustering result only by specifying the minimum number of samples contained in a cluster, avoiding the complicated parameter tuning process, and improving the accuracy and applicable range of the hierarchical tree model.
[0041] Step 103: Cluster the logs in the log data set to be parsed based on the hierarchical tree model, and obtain the tree structure corresponding to each clustering type in the clustering result according to the clustering result.
[0042] In some embodiments, the second feature vectors of all the logs in the log data set to be parsed can be clustered based on the improved DBSCAN algorithm with the idea of hierarchical clustering in step 102. In this way, based on the low-dimensional characteristics of the second feature vectors, the efficiency of the clustering process can be improved, and the improved DBSCAN algorithm with the idea of hierarchical clustering enables the model to automatically obtain the optimal clustering result only by specifying the minimum number of samples contained in a cluster, avoiding the complicated parameter tuning process.
[0043] Of course, in some other embodiments, the feature information of the logs can also be other vectors or other forms of feature information other than vectors. In this case, an appropriate clustering method can be selected according to the specific type and size of the feature information, which is not specifically limited here.
[0044] Step 104: Perform a spanning tree process on the logs of the target clustering type according to the target clustering type and the target tree structure corresponding to the target clustering type to obtain the parsing template spanning tree of the target clustering type, where the clustering types in the clustering result include the target clustering type.
[0045] In some embodiments, after obtaining the clustering result of the logs in the log data set to be parsed, the logs can be divided into log types corresponding one by one to the clustering result, and each log type can correspond to its own tree structure, where the tree structure can represent the log parsing template of its corresponding log type.
[0046] In some embodiments, the tree structure can be related to the size, structural distribution characteristics, and the log type to which it belongs of the logs of its corresponding log type.
[0047] For example, logs of the same log type have the same number of segmented words, assumed to be L, where L is an integer greater than 1. At this time, the initialized tree structure can include L + 1 nodes. The first node is the root node, and the subsequent L nodes respectively correspond to the L segmented words in the log. Finally, it ends with the special symbol "$". That is to say, a path from the root node P to "$" in the tree structure corresponds to a sequence of segmented words in a log.
[0048] The parsing template generation tree can represent the parsing rules of this type of log. Based on the parsing template generation tree, this type of log can be parsed into readable structured information.
[0049] Step 105: Parse the log of the target clustering type according to the parsing template generation tree to obtain a log parsing result.
[0050] In this step, after determining the target clustering type of the log to be parsed, the parsing template generation tree corresponding to the target clustering type can be used for parsing to obtain the log parsing result of the log to be parsed, and this log parsing result is the readable structured information.
[0051] As an optional implementation manner, constructing the hierarchical tree model corresponding to the feature information based on the minimum spanning tree includes: Construct the hierarchical tree model corresponding to the feature information by using the minimum spanning tree, and optimize the minimum step size parameter and the neighborhood radius parameter in the density-based spatial clustering of applications with noise (DBSCAN); Clustering the logs in the log dataset to be parsed based on the hierarchical tree model, and obtaining the tree structure corresponding to each clustering type in the clustering result according to the clustering result, includes: Cluster the logs in the log dataset to be parsed based on the optimized DBSCAN to obtain a clustering result; Obtain the tree structure corresponding to the logs of each clustering type in the clustering result.
[0052] In some implementation manners, constructing the hierarchical tree model corresponding to the feature information by using the minimum spanning tree, optimizing the minimum step size parameter and the neighborhood radius parameter in the density-based spatial clustering of applications with noise (DBSCAN), and clustering the logs in the log dataset to be parsed based on the optimized DBSCAN to obtain a clustering result may include the following process: 1) Perform a spatial transformation on the feature information of the log according to the density and sparsity in the hierarchical tree model; 2) Based on the spatial transformation structure, construct the minimum spanning tree of the distance weighted graph.
[0053] 3) Construct the cluster hierarchy of the associated points in the minimum spanning tree.
[0054] 4) Construct the hierarchical structure of the size-compressed clusters of the minimum clusters.
[0055] 5) Extract stable clusters from the compressed tree as the clustering results.
[0056] In this embodiment, first, a hierarchical tree model corresponding to the feature information is constructed by using a minimum spanning tree, and the minimum step size parameter and the neighborhood radius parameter in the density-based clustering method DBSCAN with noise are optimized. In this way, the accuracy and the applicable range of the optimized DBSCAN can be improved, and it is avoided that during the clustering process, manual participation is required to continuously adjust the minimum step size parameter and the neighborhood radius parameter, which can improve the automation degree of the clustering process.
[0057] In some embodiments, the clustering results are determined based on the mutual reachability distance between logs.
[0058] For example: In the clustering algorithm, the distance between the feature information of different logs can be calculated according to the following formula: ; where d mr-k (A - B) refers to the mutual reachability distance between A and B; d AB refers to the Euclidean distance between A and B; d corek (A) is the distance from A to the core point of the clustering cluster; d corek (B) is the distance from B to the core point of the clustering cluster, where A and B respectively represent the feature information of two logs, such as the second feature vectors of the two logs respectively. d mr-k (A - B) takes the maximum value of d corek (A), d corek (B) and d AB .
[0059] In this way, by improving the distance calculation in the clustering process, during the clustering process, the distances from A and B to the core points of each clustering cluster and the Euclidean distance between A and B are fused, which can make the clustering results more accurate.
[0060] As an alternative embodiment, the generating a spanning tree process for the logs of the target clustering type according to the target clustering type and the target tree structure corresponding to the target clustering type to obtain the parsing template spanning tree of the target clustering type includes: Perform second word segmentation on the logs belonging to the target clustering type in the log dataset to be parsed, to obtain the word segmentation sequences of each log of the target clustering type, where the word segmentation sequence includes L word segments, L is an integer greater than 1, and the word segments in the word segmentation sequence include the first word segments common to the logs of the target clustering type and the second word segments not common to the logs of the target clustering type; Starting from the root node of the target tree structure corresponding to the target clustering type, sequentially insert the word segments in the target word segmentation sequence into the target tree structure to obtain the parsing template generation tree of the target clustering type, where the target word segmentation sequence includes the word segmentation sequences of each log of the target clustering type.
[0061] It is worth noting that the second word segmentation process in this embodiment needs to be distinguished from the first word segmentation process in the foregoing embodiment. Among them, the first word segmentation process is implemented based on a regular word segmenter or a dictionary, and the segmented words after word segmentation do not contain character segments of pure symbols or pure numbers. However, there is no such requirement for the second word segmentation process.
[0062] For example: Take the following email service log as an example: Oct 6 00:00:13 localhost postfix / smtpd
[26072] : disconnect from n169-113.mail.139.com[120.232.169.113] ehlo=1 mail=1rcpt=9 / 10 data=1 quit=1 commands=13 / 14。
[0063] Among them, [120.232.169.113], as a character segment composed of pure numbers and symbols, does not belong to the segmented words after the first word segmentation process, but belongs to the word segments after the second word segmentation process.
[0064] In other words, the log content can be roughly divided into two parts, namely the fixed part and the variable parameter part. Among them, the fixed part is usually a segmented word with linguistic meaning, and for logs of the same type, it has the same fixed part, while the variable parameter part may be a character segment composed of pure numbers and symbols. Even for logs of the same type, they may contain different variable parameters.
[0065] For example, the word segments belonging to the fixed part in the above email log include 'disconnect from', 'ehlo=', 'mail=', 'rcpt=', 'data=', 'quit=', 'commands='; the word segments belonging to the variable parameter part include 'n169-113.mail.139.com [120.232.169.113]', '1', '1', '9 / 10', '1', '1', '12 / 14'.
[0066] In some embodiments, starting from the root node of the target tree structure corresponding to the target clustering type, the word segments in the target word segment sequence are sequentially inserted into the target tree structure to obtain the parsing template generation tree of the target clustering type. The purpose is to form a log template with the word segments belonging to the fixed part in the log, that is, "disconnect from *ehlo=* mail=* rcpt=* data=* quit=* commands=*", where the wildcard "*" represents the word segments in the variable parameter part.
[0067] It should be noted that after obtaining the clustering result of the log, the second word segmentation process is performed again on all the logs of the same clustering type, so that the word segmentation results of each log have the same number of word segments, that is, for different logs of the same clustering type, the lengths of their word segment sequences are the same.
[0068] For example: the word segment sequence of the above email log is as follows: ['disconnect', 'from', 'n169-113.mail.139.com', '[120.232.169.113]', 'ehlo', '=', '1','mail', '=', '1', 'rcpt', '=', '9 / 10', 'data', '=', '1', 'quit', '=', '1', 'commands', '=', '13 / 14', ].
[0069] It should be noted that starting from the root node of the target tree structure corresponding to the target clustering type, sequentially inserting the word segments in the target word segment sequence into the target tree structure may be to match the word segments in the target word segment sequence one by one with the nodes starting from the root node of the target tree structure according to the arrangement order of the word segments in the target word segment sequence, and the search ranges of adjacent word segments in the target word segment sequence in the target tree structure are also adjacent nodes.
[0070] In some embodiments, taking the root node of the target tree structure corresponding to the target clustering type as a starting point, inserting the word segments in the target word segment sequence into the target tree structure in sequence to obtain the parsing template generation tree of the target clustering type includes: Taking the root node of the target tree structure corresponding to the target clustering type as a starting point, searching for a node in the target tree structure that matches the word segment in the target word segment sequence; Updating the target tree structure according to the search result; Determining the parsing template generation tree of the target clustering type according to the updated target tree structure.
[0071] In some embodiments, the search result may include that there is a node in the target tree structure that matches the word segment in the target word segment sequence, and there is no node in the target tree structure that matches the word segment in the target word segment sequence. At this time, updating the target tree structure according to the search result may be that when the search result is that there is a node in the target tree structure that matches the word segment in the target word segment sequence, the target tree structure is not updated, and continue to search for the next word segment in the target word segment sequence in the target tree structure until all the word segments in the target word segment sequence are searched; of course, when the search result is that there is no node in the target tree structure that matches the word segment in the target word segment sequence, the target tree structure needs to be updated. For example: a new branch can be added to the target tree structure according to the word segment for which no matching node is found.
[0072] In some embodiments, a node that matches the word segment in the target word segment sequence may be that the character stored in the node is the same as the word segment; a node that does not match the word segment in the target word segment sequence may be that the character stored in the node is different from the word segment.
[0073] In this embodiment, the target tree structure can be updated according to the matching result between the target tree structure and the word segments in the target word segment sequence, and finally obtain a parsing template generation tree that can match the fixed part of the logs of this type.
[0074] As an alternative embodiment, the updating the target tree structure according to the search result includes: When it is determined that the target tree structure includes a first node that matches a third word segment, searching for a node in the subtree corresponding to the first node that matches a fourth word segment in the target word segment sequence, where the target word segment sequence includes the third word segment and the fourth word segment, and the fourth word segment is the next word segment adjacent to the third word segment in the target word segment sequence; In the case where it is determined that the target tree structure does not include a node matching the fifth participle, a branch corresponding to the fifth participle is added to the subtree corresponding to the second node in the target tree structure, where the target participle sequence includes the fifth participle and the sixth participle, and the sixth participle is the previous participle adjacent to the fifth participle in the target participle sequence, and the second node matches the sixth participle.
[0075] Taking a log m c as an example, the participle sequence obtained after the second participle processing is T m , and taking the root node of the target tree structure corresponding to the target clustering type as the starting point, inserting the participles in the target participle sequence into the target tree structure in sequence may include the following process: Traverse the participle sequence T m , and insert the participle t m in it into the target tree structure in sequence. The special character '$' in the target tree structure is the mark of the last node of each branch, so the path passed from any '$' to the root node (root) represents a matching path of a log m c , where the depth of the target tree structure is L + 1, and the number of participles in the participle sequence T m is L; For the participle t m in the participle sequence T m , first search and traverse each path in the target tree structure. Start searching from the root node P to obtain the node to be searched that matches the participle t mi . If there is no matching node, insert the participle sequence T m into the tree structure as a new branch; if there is a matching node, select the subtree corresponding to the matching node in the target tree structure, and go to this subtree to continue searching for the matching node of the next participle in the participle sequence T m , until the search for L participles is completed, and the search for all the participles t m in a participle sequence T m ends.
[0076] For example: The process of constructing the target tree structure based on the above-mentioned participle sequence of the email log: ['disconnect', 'from', 'n169-113.mail.139.com', '[120.232.169.113]', 'ehlo', '=', '1','mail', '=', '1', 'rcpt', '=', '9 / 10', 'data', '=', '1', 'quit', '=', '1', 'commands', '=', '13 / 14', ] is as follows Figure 2As shown, where the initial target tree structure is as Figure 2 shown in the left tree structure in [reference], the initial target tree structure can be determined based on the number of word segments included in the word segmentation sequence of a log of the target clustering type. Thereafter, during the insertion of the first word segmentation sequence T m , if the second word segment t m2 in this sequence is found to be the same as the word segment stored in the 3rd node Q in the target tree structure, then update P = Q, and start searching for the word segmentation sequence T m for the next word segment t m3 . If t m3 fails to match the 4th node in the target tree structure, then record the node where the match fails, and insert the word segment t m3 as a new branch into the 4th node in the target tree structure, update the root node P to the 4th node, and continue to match the 3rd word segment t m in the word segmentation sequence T m3 with the 5th node in the target tree structure until the search for all word segments in the entire word segmentation sequence T m is completed, and the search process for one log ends. Then, continue to insert the next log into the target tree structure until the insertion of all logs of the target clustering type is completed, and the final parsing template generation tree is obtained.
[0077] In this embodiment, the word segmentation sequences of logs of the same clustering type can be inserted into the target tree structure in sequence to obtain a parsing template generation tree that can reflect the parsing rules of the logs of this clustering type. In this way, the parsing template generation tree can implement the parsing process of the logs of this clustering type.
[0078] For example: during the process of using the parsing template generation tree to parse logs, the characters in the log can be sequentially matched with the nodes in the parsing template generation tree, and the t m stored in the successfully matched nodes can be extracted, so as to obtain a structured log template and complete the log parsing process.
[0079] As Figure 2 shown, for the word segments of the fixed part, there is usually only one branch, and for the word segments of the variable parameter part, there are multiple branches.
[0080] In some embodiments, the target tree structure can also be verified, and when the target tree structure passes the verification, adding branches to the target tree structure is stopped. In this way, when the parsing degree of the target tree structure for logs meets the requirements, adding branches to the target tree structure based on this log is not performed, and the structural complexity of the target tree structure can be simplified while ensuring that the target tree structure can meet the parsing degree requirements of similar logs, thereby simplifying the complexity of the finally obtained parsing template generation tree and improving the parsing efficiency and training efficiency of the parsing template generation tree.
[0081] As an alternative implementation, the method further includes: Obtaining a first parameter and a second parameter, where the first parameter is the number of word segments for which no matching node is found in the target tree structure, and the second parameter is the number of wildcards included in the paths for finding each word segment in the target word segment sequence in the target tree structure; Determining an index parameter of the target tree structure according to the first parameter, the second parameter, and the length of the target word segment sequence; In the case where it is determined that the target tree structure does not include a node matching the fifth word segment, adding a branch corresponding to the fifth word segment to the subtree corresponding to the second node in the target tree structure includes: In the case where the index parameter meets a first preset condition, adding a branch corresponding to the third word segment to the subtree corresponding to the fourth node in the target tree structure; or, the method further includes: in the case where the index parameter does not meet the first preset condition, updating the node corresponding to the third word segment in the target tree structure to a wildcard.
[0082] In some embodiments, when sequentially searching for word segments in the word segment sequence in the target tree structure, a node with a successful match will feedback the word segment stored in the node, and a node with a failed match will return a wildcard '*'. In this way, an index parameter for determining whether the target tree structure can meet the parsing degree requirement of similar logs can be calculated based on the number of word segments included in a complete log, the number of wildcards '*' returned by the target tree structure, and the number of wildcards included in the matching paths for searching the word segment sequence in the target tree structure.
[0083] Among them, if the index parameter meets the first preset condition, it means that the target tree structure cannot meet the parsing degree requirement of similar logs; if the index parameter does not meet the first preset condition, it means that the target tree structure can meet the parsing degree requirement of similar logs.
[0084] For example: The index parameter can be calculated through the following formula: X = ∑unmatch(t mi ) / L - θ; where L represents the length of the target word segment sequence; ∑unmatch(t mi ) represents the first parameter; θ represents the second parameter; and the first condition is X ≥ λ.
[0085] Thus, when X≥λ, a branch corresponding to the third word segment is added to the subtree corresponding to the fourth node in the target tree structure; when X<λ, a branch corresponding to the third word segment is not added to the subtree corresponding to the fourth node in the target tree structure.
[0086] In this embodiment, when the parsing degree of the log by the target tree structure meets the requirements, branches can be not added to the target tree structure based on the log, which can simplify the structural complexity of the target tree structure while ensuring that the target tree structure can meet the parsing degree requirements of similar logs, and further simplify the complexity of the finally obtained parsing template generation tree, improving the parsing efficiency and training efficiency of the parsing template generation tree.
[0087] In the embodiments of the present application, the feature information of each log in the log dataset to be parsed is obtained; a hierarchical tree model corresponding to the feature information is constructed based on the minimum spanning tree; the logs in the log dataset to be parsed are clustered based on the hierarchical tree model, and each tree structure corresponding to each clustering type in the clustering result is obtained according to the clustering result; the logs of the target clustering type are subjected to a spanning tree process according to the target clustering type and the target tree structure corresponding to the target clustering type to obtain a parsing template generation tree for the target clustering type, where the clustering types in the clustering result include the target clustering type; the logs of the target clustering type are parsed according to the parsing template generation tree to obtain a log parsing result. After obtaining the feature information of the log, a hierarchical tree model corresponding to the feature information is constructed based on the minimum spanning tree, and the logs in the log dataset to be parsed are clustered based on this hierarchical tree model. Finally, based on the spanning tree technology, a parsing template generation tree is automatically generated for the logs of each clustering type as the parsing rule applicable to the logs of this clustering type to achieve the efficient parsing of various types of logs. This method can realize the automation of the entire log parsing process and significantly improve the log parsing efficiency. Among them, the clustering process is optimized based on the idea of hierarchical clustering, so that the hierarchical tree model can automatically obtain the optimal clustering result only by giving the minimum samples included in the cluster, avoiding the process of manual parameter tuning. On the basis of improving the automation of the log parsing process, the accuracy and applicability range of the hierarchical tree model can also be improved.
[0088] Refer to Figure 3 , the embodiments of the present application also provide a log parsing method, as Figure 3 shown, this log parsing method includes the following steps: Step 301, obtain a large amount of log data and preprocess the log data.
[0089] In some embodiments, the preprocessing in this step may include at least one of the following: Clean the log data to remove error records, duplicate records, blank lines, comments, and duplicate log entries; Unify the formats of log data from different sources; Modify or delete outliers; Fill or delete null values.
[0090] In some embodiments, for outliers in the log data, they can be marked, and the logs with outliers can be used as a separate log type to perform clustering and construct a parsing template generation tree, so as to implement the parsing of outliers using the parsing template generation tree. In this way, compared with the method of modifying or deleting outliers in the foregoing manner, the parsing of outliers can be achieved, and the omission of the parsing of outliers can be avoided.
[0091] In some embodiments, the log data set to be parsed can be obtained regularly, such as once every 1 day or 1 week.
[0092] In some embodiments, the computing power load of the electronic device executing the log parsing method of the embodiments of the present application can be detected, and the log parsing process can be executed during the period when the computing power load of the electronic device is low. For example, log parsing is performed at 2:00 in the morning every day.
[0093] Step 302: Use natural language processing (Natural Language Processing, NLP) technology to extract log features and convert the logs into feature vectors.
[0094] In some embodiments, the log can be subjected to a first word segmentation process based on a regular word segmenter or a dictionary to obtain U segmented words of the log, and a first feature vector can be determined according to the feature values of the U segmented words. Finally, the first feature vector is subjected to dimensionality reduction processing to obtain a second feature vector.
[0095] Among them, the dimensionality reduction processing of the first feature vector can be to use the t-SNE algorithm to reduce the dimension of the first feature vector, avoiding the overfitting problem caused by using the high-dimensional first feature vector for clustering processing.
[0096] Step 303: Use the DBSCAN algorithm to cluster the obtained feature vectors.
[0097] In this step, the DBSCAN clustering algorithm can be improved based on the idea of a hierarchical tree model, so that the improved DBSCAN algorithm can automatically obtain the optimal clustering result only by giving the minimum number of samples contained in the cluster, avoiding the cumbersome parameter tuning process.
[0098] Step 304: Generate corresponding parsing template generation trees for logs of different clustering types.
[0099] In this step, according to the clustering results, the spanning tree method can be used to generate the tree structure corresponding to the logs of each clustering type, and update the tree structure based on the logs of the same type, so as to obtain the parsing template generation tree of each clustering type.
[0100] Step 305: Generate structured data.
[0101] In this step, the parsing template generation tree of each clustering type is used to parse the logs of this clustering type, so as to obtain the structured data of each log in the log dataset to be parsed, and complete the parsing of each log in the log dataset to be parsed.
[0102] The log parsing method of this embodiment has at least the following beneficial effects: Compared with the way of manually configuring the log parsing file in the related technology, and compared with the log parsing method that realizes partial automation in the related technology, this application realizes the automation of the entire log parsing process, can improve the parsing efficiency of log parsing, and save labor costs; Compared with the log parsing process in the related technology, this solution improves the calculation method of feature vectors and the clustering method, can improve the efficiency and applicability of the clustering process, and reduce the complexity of parameter adjustment of the clustering algorithm; Compared with the way of pre-setting parsing rules to parse logs in the related technology, this application classifies the logs in advance and then parses them, which improves the efficiency in the parsing process and reduces the consumption of computing resources.
[0103] This application embodiment also provides a log parsing device. See Figure 4 , Figure 4 is the structural diagram of the log parsing device 400 provided by this application embodiment. Since the principle of the log parsing device 400 to solve problems is similar to that of the log parsing method in this application embodiment, the implementation of this log parsing device can refer to the implementation of the method, and the repeated parts will not be described again.
[0104] As Figure 4 shown, the log parsing device 400 includes: The first acquisition module 401 is used to acquire the feature information of each log in the log dataset to be parsed; The model construction module 402 is used to construct a hierarchical tree model corresponding to the feature information based on the minimum spanning tree; The clustering module 403 is used to perform clustering processing on the logs in the log dataset to be parsed based on the hierarchical tree model, and obtain the tree structure corresponding to each clustering type in the clustering results according to the clustering results; A spanning tree module 404, configured to perform spanning tree processing on the logs of the target clustering type according to the target clustering type and the target tree structure corresponding to the target clustering type, so as to obtain a parsing template spanning tree of the target clustering type, where the clustering types in the clustering result include the target clustering type; A log parsing module 405, configured to parse the logs of the target clustering type according to the parsing template spanning tree to obtain a log parsing result.
[0105] Optionally, the first acquisition module 401 is specifically configured to: Perform feature extraction processing on each log in the log dataset to be parsed respectively to obtain the feature information of each log; Wherein, the feature extraction processing includes: Perform first word segmentation processing on the log based on a regular word segmenter or a dictionary to obtain U segmented words of the log, where U is an integer greater than 1; Determine a first feature vector according to the feature values of the U segmented words, where the first feature vector is a U-dimensional vector; Perform dimensionality reduction processing on the first feature vector to obtain a second feature vector, where the feature information of each log includes the second feature vector, and the second feature vector is a two-dimensional feature vector.
[0106] Optionally, the model construction module 402 is specifically configured to: Construct a hierarchical tree model corresponding to the feature information by using a minimum spanning tree, and optimize the minimum step size parameter and the neighborhood radius parameter in the density-based spatial clustering of applications with noise (DBSCAN); The clustering module 403 includes: A clustering processing unit, configured to perform clustering processing on the logs in the log dataset to be parsed based on the optimized DBSCAN to obtain a clustering result; An acquisition unit, configured to acquire the tree structure corresponding to each log of each clustering type in the clustering result.
[0107] Optionally, the clustering result is determined based on the mutual reachability distance between the logs.
[0108] Optionally, the spanning tree module 404 includes: A word segmentation unit, configured to perform second word segmentation processing on the logs belonging to the target clustering type in the log dataset to be parsed to obtain a word segmentation sequence of each log of the target clustering type, where the word segmentation sequence includes L word segments, L is an integer greater than 1, and the word segments in the word segmentation sequence include a first word segment common to the logs of the target clustering type and a second word segment not common to the logs of the target clustering type; An insertion unit is configured to start from the root node of the target tree structure corresponding to the target clustering type, and sequentially insert the word segments in the target word segment sequence into the target tree structure to obtain a parsing template generation tree of the target clustering type, where the target word segment sequence includes the word segment sequences of each log of the target clustering type.
[0109] Optionally, the insertion unit includes: A search subunit is configured to start from the root node of the target tree structure corresponding to the target clustering type, and search for a node in the target tree structure that matches the word segment in the target word segment sequence; An update subunit is configured to update the target tree structure according to the search result; A determination subunit is configured to determine a parsing template generation tree of the target clustering type according to the updated target tree structure.
[0110] Optionally, the update subunit is specifically configured to: In the case where it is determined that the target tree structure includes a first node that matches the third word segment, search for a node in the subtree corresponding to the first node that matches the fourth word segment in the target word segment sequence, where the target word segment sequence includes the third word segment and the fourth word segment, and the fourth word segment is the next word segment adjacent to the third word segment in the target word segment sequence; In the case where it is determined that the target tree structure does not include a node that matches the fifth word segment, add a branch corresponding to the fifth word segment in the subtree corresponding to the second node in the target tree structure, where the target word segment sequence includes the fifth word segment and the sixth word segment, and the sixth word segment is the previous word segment adjacent to the fifth word segment in the target word segment sequence, and the second node matches the sixth word segment.
[0111] Optionally, the log parsing device 400 further includes: A second acquisition module is configured to acquire a first parameter and a second parameter, where the first parameter is the number of word segments for which no matching node is found in the target tree structure, and the second parameter is the number of wildcards included in the path for searching each word segment in the target word segment sequence in the target tree structure; A determination module is configured to determine an index parameter of the target tree structure according to the first parameter, the second parameter, and the length of the target word segment sequence; The update subunit is further configured to: When the index parameter satisfies the first preset condition, a branch corresponding to the third word segmentation is added to the subtree corresponding to the fourth node in the target tree structure; alternatively, the log parsing device 400 further includes an update module, configured to update the node corresponding to the third word segmentation in the target tree structure to a wild card symbol when the index parameter does not satisfy the first preset condition.
[0112] The log parsing device 400 provided by the embodiments of the present application can execute the above method embodiments, and its implementation principle and technical effects are similar. Details are not described herein again in this embodiment.
[0113] Embodiments of the present application further provide an electronic device. Since the principle of the electronic device to solve problems is similar to the log parsing method in the embodiments of the present application, the implementation of the electronic device can refer to the implementation of the method, and repeated parts are not described again. As Figure 5 shown, the electronic device of the embodiments of the present application includes: a processor 500, configured to read a program in a memory 520 and execute the following processes: Obtain the feature information of each log in the log dataset to be parsed; Construct a hierarchical tree model corresponding to the feature information based on a minimum spanning tree; Perform clustering processing on the logs in the log dataset to be parsed based on the hierarchical tree model, and obtain a tree structure corresponding to each clustering type in the clustering result according to the clustering result; Perform a spanning tree process on the logs of the target clustering type according to the target clustering type and the target tree structure corresponding to the target clustering type, to obtain a parsing template spanning tree of the target clustering type, where the clustering types in the clustering result include the target clustering type; Parse the logs of the target clustering type according to the parsing template spanning tree to obtain a log parsing result.
[0114] Among them, in Figure 5 , the bus architecture may include any number of interconnected buses and bridges, specifically, various circuits represented by one or more processors represented by the processor 500 and a memory represented by the memory 520 are linked together. The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits together, which are well known in the art, and thus are not further described herein. The bus interface provides an interface. The processor 500 is responsible for managing the bus architecture and general processing, and the memory 520 can store data used by the processor 500 when executing operations.
[0115] Optionally, the processor 500 is further configured to read a program in the memory 520 and execute the following steps: Perform feature extraction processing on each log in the log data set to be parsed to obtain the feature information of each log; Among them, the feature extraction processing includes: Based on a regular word segmenter or dictionary, perform first word segmentation processing on the log to obtain U segmented words of the log, where U is an integer greater than 1; Determine a first feature vector according to the feature values of the U segmented words, where the first feature vector is a U-dimensional vector; Perform dimensionality reduction processing on the first feature vector to obtain a second feature vector. The feature information of each log includes the second feature vector, and the second feature vector is a two-dimensional feature vector.
[0116] Optionally, the processor 500 is further configured to read a program in the memory 520 and execute the following steps: Construct a hierarchical tree model corresponding to the feature information by using a minimum spanning tree, and optimize the minimum step size parameter and the neighborhood radius parameter in the density-based spatial clustering of applications with noise (DBSCAN); Perform clustering processing on the logs in the log data set to be parsed based on the hierarchical tree model, and obtain the tree structure corresponding to each clustering type in the clustering result, including: Perform clustering processing on the logs in the log data set to be parsed based on the optimized DBSCAN to obtain a clustering result; Obtain the tree structure corresponding to the logs of each clustering type in the clustering result.
[0117] Optionally, the clustering result is determined based on the mutual reachability distance between logs.
[0118] Optionally, the processor 500 is further configured to read a program in the memory 520 and execute the following steps: Perform second word segmentation processing on the logs belonging to the target clustering type in the log data set to be parsed to obtain the word segmentation sequence of each log of the target clustering type. Among them, the word segmentation sequence includes L word segments, where L is an integer greater than 1, and the word segments in the word segmentation sequence include the first word segments common to the logs of the target clustering type and the second word segments not common to the logs of the target clustering type; Starting from the root node of the target tree structure corresponding to the target clustering type, insert the word segments in the target word segmentation sequence into the target tree structure in sequence to obtain the parsing template generation tree of the target clustering type, where the target word segmentation sequence includes the word segmentation sequences of each log of the target clustering type.
[0119] Optionally, the processor 500 is further configured to read a program in the memory 520 and execute the following steps: Starting from the root node of the target tree structure corresponding to the target clustering type, search for nodes in the target tree structure that match the words in the target word segmentation sequence; Update the target tree structure according to the search result; Determine the parsing template generation tree of the target clustering type according to the updated target tree structure.
[0120] Optionally, the processor 500 is further configured to read the program in the memory 520 and execute the following steps: In the case where it is determined that the target tree structure includes a first node that matches the third word, search for a node in the subtree corresponding to the first node that matches the fourth word in the target word segmentation sequence, where the target word segmentation sequence includes the third word and the fourth word, and the fourth word is the next word adjacent to the third word in the target word segmentation sequence; In the case where it is determined that the target tree structure does not include a node that matches the fifth word, add a branch corresponding to the fifth word in the subtree corresponding to the second node in the target tree structure, where the target word segmentation sequence includes the fifth word and the sixth word, and the sixth word is the previous word adjacent to the fifth word in the target word segmentation sequence, and the second node matches the sixth word.
[0121] Optionally, the processor 500 is further configured to read the program in the memory 520 and execute the following steps: Obtain a first parameter and a second parameter, where the first parameter is the number of words for which no matching node is found in the target tree structure, and the second parameter is the number of wildcards included in the path for searching each word in the target word segmentation sequence in the target tree structure; Determine the metric parameter of the target tree structure according to the first parameter, the second parameter, and the length of the target word segmentation sequence; In the case where the metric parameter satisfies a first preset condition, add a branch corresponding to the third word in the subtree corresponding to the fourth node in the target tree structure; or, in the case where the metric parameter does not satisfy the first preset condition, update the node corresponding to the third word in the target tree structure to a wildcard.
[0122] The terminal provided in the embodiments of the present application can execute the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0123] In addition, the computer-readable storage medium of the embodiments of the present application is used to store a computer program, and the computer program can be executed by a processor to implement the following steps: Obtain the feature information of each log in the log dataset to be parsed; Construct a hierarchical tree model corresponding to the feature information based on the minimum spanning tree; Perform clustering processing on the logs in the log dataset to be parsed based on the hierarchical tree model, and obtain the tree structure corresponding to each clustering type in the clustering result according to the clustering result; Perform a spanning tree process on the logs of the target clustering type according to the target clustering type and the target tree structure corresponding to the target clustering type, and obtain the parsing template spanning tree of the target clustering type, where the clustering types in the clustering result include the target clustering type; Parse the logs of the target clustering type according to the parsing template spanning tree to obtain a log parsing result.
[0124] Optionally, the computer program can also be executed by a processor to implement the following steps: Perform feature extraction processing on each log in the log dataset to be parsed to obtain the feature information of each log; Among them, the feature extraction processing includes: Perform a first word segmentation process on the log based on a regular word segmenter or dictionary to obtain U segmented words of the log, where U is an integer greater than 1; Determine a first feature vector according to the feature values of the U segmented words, and the first feature vector is a U-dimensional vector; Perform dimensionality reduction processing on the first feature vector to obtain a second feature vector, and the feature information of each log includes the second feature vector, and the second feature vector is a two-dimensional feature vector.
[0125] Optionally, the computer program can also be executed by a processor to implement the following steps: Use the minimum spanning tree to construct a hierarchical tree model corresponding to the feature information, and optimize the minimum step size parameter and the neighborhood radius parameter in the density-based spatial clustering of applications with noise (DBSCAN); The performing clustering processing on the logs in the log dataset to be parsed based on the hierarchical tree model, and obtaining the tree structure corresponding to each clustering type in the clustering result according to the clustering result, includes: Perform clustering processing on the logs in the log dataset to be parsed based on the optimized DBSCAN to obtain a clustering result; Obtain the tree structure corresponding to the logs of each clustering type in the clustering result.
[0126] Optionally, the clustering result is determined based on the mutual reachability distance between logs.
[0127] Optionally, the computer program can also be executed by a processor to implement the following steps: Perform second word segmentation on the logs belonging to the target clustering type in the log data set to be parsed, to obtain a word segmentation sequence for each log of the target clustering type, where the word segmentation sequence includes L word segments, L is an integer greater than 1, and the word segments in the word segmentation sequence include first word segments common to the logs of the target clustering type and second word segments not common to the logs of the target clustering type; Starting from the root node of the target tree structure corresponding to the target clustering type, sequentially insert the word segments in the target word segmentation sequence into the target tree structure, to obtain a parsing template generation tree for the target clustering type, where the target word segmentation sequence includes the word segmentation sequences of each log of the target clustering type.
[0128] Optionally, the computer program can also be executed by a processor to implement the following steps: Starting from the root node of the target tree structure corresponding to the target clustering type, search for nodes in the target tree structure that match the word segments in the target word segmentation sequence; Update the target tree structure according to the search result; Determine a parsing template generation tree for the target clustering type according to the updated target tree structure.
[0129] Optionally, the computer program can also be executed by a processor to implement the following steps: In the case where it is determined that the target tree structure includes a first node that matches a third word segment, search for a node in the subtree corresponding to the first node that matches a fourth word segment in the target word segmentation sequence, where the target word segmentation sequence includes the third word segment and the fourth word segment, and the fourth word segment is the next word segment adjacent to the third word segment in the target word segmentation sequence; In the case where it is determined that the target tree structure does not include a node that matches a fifth word segment, add a branch corresponding to the fifth word segment in the subtree corresponding to a second node in the target tree structure, where the target word segmentation sequence includes the fifth word segment and a sixth word segment, and the sixth word segment is the previous word segment adjacent to the fifth word segment in the target word segmentation sequence, and the second node matches the sixth word segment.
[0130] Optionally, the computer program can also be executed by a processor to implement the following steps: Obtain a first parameter and a second parameter, where the first parameter is the number of word segments for which no matching node is found in the target tree structure, and the second parameter is the number of wildcards included in the path for searching each word segment in the target word segmentation sequence in the target tree structure; Determine the index parameter of the target tree structure according to the first parameter, the second parameter, and the length of the target word segmentation sequence; When the index parameter meets the first preset condition, add a branch corresponding to the third word segmentation to the subtree corresponding to the fourth node in the target tree structure; or, when the index parameter does not meet the first preset condition, update the node corresponding to the third word segmentation in the target tree structure to a wild card symbol.
[0131] The embodiments of the present application also provide a computer program product, including computer instructions, which when executed by a processor, implement the steps of the foregoing embodiments of the log parsing method, and can achieve the same beneficial effects as the foregoing embodiments of the log parsing method. To avoid repetition, they will not be elaborated here.
[0132] In several embodiments provided by the present application, it should be understood that the disclosed method and device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.
[0133] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can be physically included separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0134] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit stored in a storage medium includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute some steps of the transceiver method described in the various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0135] The above are the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle described in the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A log parsing method, characterized in that, Including: Obtain the feature information of each log in the log dataset to be parsed; Construct a hierarchical tree model corresponding to the feature information based on the minimum spanning tree; Perform clustering processing on the logs in the log dataset to be parsed based on the hierarchical tree model, and obtain the tree structure corresponding to each clustering type in the clustering result according to the clustering result; Perform a spanning tree process on the logs of the target clustering type according to the target clustering type and the target tree structure corresponding to the target clustering type, and obtain the parsing template spanning tree of the target clustering type, where the clustering types in the clustering result include the target clustering type; Parse the logs of the target clustering type according to the parsing template spanning tree to obtain a log parsing result.
2. The method according to claim 1, wherein The obtaining the feature information of each log in the log dataset to be parsed includes: Perform feature extraction processing on each log in the log dataset to be parsed respectively to obtain the feature information of each log; Among them, the feature extraction processing includes: Perform a first word segmentation process on the log based on a regular word segmenter or dictionary to obtain U segmented words of the log, where U is an integer greater than 1; Determine a first feature vector according to the feature values of the U segmented words, and the first feature vector is a U-dimensional vector; Perform dimensionality reduction processing on the first feature vector to obtain a second feature vector, and the feature information of each log includes the second feature vector, and the second feature vector is a two-dimensional feature vector.
3. The method according to claim 2, wherein The constructing the hierarchical tree model corresponding to the feature information based on the minimum spanning tree includes: Use the minimum spanning tree to construct a hierarchical tree model corresponding to the feature information, and optimize the minimum step size parameter and the neighborhood radius parameter in the density-based spatial clustering of applications with noise (DBSCAN); The performing clustering processing on the logs in the log dataset to be parsed based on the hierarchical tree model and obtaining the tree structure corresponding to each clustering type in the clustering result according to the clustering result includes: Perform clustering processing on the logs in the log dataset to be parsed based on the optimized DBSCAN to obtain a clustering result; Obtain the tree structure corresponding to the logs of each clustering type in the clustering result.
4. The method according to claim 3, wherein Among them, The clustering result is determined based on the mutual reachability distance between logs.
5. The method according to any one of claims 1 to 4, characterized in that The performing a spanning tree process on the logs of the target clustering type according to the target clustering type and the target tree structure corresponding to the target clustering type to obtain the parsing template spanning tree of the target clustering type includes: Perform a second word segmentation process on the logs belonging to the target clustering type in the log dataset to be parsed to obtain a word segmentation sequence of each log of the target clustering type, where the word segmentation sequence includes L word segments, L is an integer greater than 1, and the word segments in the word segmentation sequence include a first word segment common to the logs of the target clustering type and a second word segment not common to the logs of the target clustering type; Starting from the root node of the target tree structure corresponding to the target clustering type, insert the word segments in the target word segment sequence into the target tree structure in sequence to obtain the parsing template generation tree of the target clustering type, where the target word segment sequence includes the word segment sequences of each log of the target clustering type.
6. The method according to claim 5, wherein The step of starting from the root node of the target tree structure corresponding to the target clustering type and inserting the word segments in the target word segment sequence into the target tree structure in sequence to obtain the parsing template generation tree of the target clustering type includes: Starting from the root node of the target tree structure corresponding to the target clustering type, search for a node in the target tree structure that matches the word segment in the target word segment sequence; Update the target tree structure according to the search result; Determine the parsing template generation tree of the target clustering type according to the updated target tree structure.
7. The method according to claim 6, wherein The step of updating the target tree structure according to the search result includes: When it is determined that the target tree structure includes a first node that matches the third word segment, search for a node in the subtree corresponding to the first node that matches the fourth word segment in the target word segment sequence, where the target word segment sequence includes the third word segment and the fourth word segment, and the fourth word segment is the next word segment adjacent to the third word segment in the target word segment sequence; When it is determined that the target tree structure does not include a node that matches the fifth word segment, add a branch corresponding to the fifth word segment in the subtree corresponding to the second node in the target tree structure, where the target word segment sequence includes the fifth word segment and the sixth word segment, and the sixth word segment is the previous word segment adjacent to the fifth word segment in the target word segment sequence, and the second node matches the sixth word segment.
8. The method according to claim 7, characterized in that, The method further includes: Obtain a first parameter and a second parameter, where the first parameter is the number of word segments for which no matching node is found in the target tree structure, and the second parameter is the number of wildcards included in the path for searching each word segment in the target word segment sequence in the target tree structure; Determine the index parameter of the target tree structure according to the first parameter, the second parameter, and the length of the target word segment sequence; The step of adding a branch corresponding to the fifth word segment in the subtree corresponding to the second node in the target tree structure when it is determined that the target tree structure does not include a node that matches the fifth word segment includes: When the index parameter meets the first preset condition, add a branch corresponding to the third word segment in the subtree corresponding to the fourth node in the target tree structure; or, the method further includes: when the index parameter does not meet the first preset condition, update the node corresponding to the third word segment in the target tree structure to a wildcard.
9. A log parsing device, characterized in that, including: A first acquisition module for acquiring the feature information of each log in the log dataset to be parsed; A model construction module for constructing a hierarchical tree model corresponding to the feature information based on the minimum spanning tree; A clustering module, configured to perform clustering processing on the logs in the to-be-parsed log dataset based on the hierarchical tree model, and obtain a tree structure corresponding to each clustering type in the clustering result according to the clustering result; A tree generation module, configured to perform tree generation processing on the logs of the target clustering type according to the target clustering type and the target tree structure corresponding to the target clustering type, so as to obtain a parsing template generation tree of the target clustering type, where the clustering types in the clustering result include the target clustering type; A log parsing module, configured to parse the logs of the target clustering type according to the parsing template generation tree to obtain a log parsing result.
10. An electronic device, characterized in that, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the steps in the log parsing method according to any one of claims 1 to 8 are implemented.
11. A computer-readable storage medium for storing a computer program, characterized in that, When the computer program is executed by the processor, the steps in the log parsing method according to any one of claims 1 to 8 are implemented.
12. A computer program product, characterized in that, Including computer instructions, wherein when the computer instructions are executed by the processor, the steps in the log parsing method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Log template extraction method based on online hierarchical clustering
CN109981625A
Online log analysis method and system and electronic terminal equipment thereof
CN110888849A
Log semantic vectorization and hierarchical clustering-based log analysis method
CN117520033A
Real-time power grid alarm situation awareness system based on clustering algorithm
CN117892163A
Clustering iteration method and device, electronic equipment and computer readable storage medium
CN118245825A