NLP-based log clustering method and device
Through the NLP-based log clustering method and combining multiple algorithms to process log data, the massive log classification problem is solved, rapid fault location and abnormal detection are achieved, and storage costs are reduced.
Patent Information
- Application Number
- PCT/CN2024/133378
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-26
- Filing Date
- 2024-11-21
- Publication Date
- 2025-07-31
AI Technical Summary
The existing technology is difficult to efficiently classify and process massive heterogeneous logs, resulting in the submerged key error logs, making it difficult to quickly grasp the overall picture of the log and locate problems.
The NLP-based log clustering method is used, combined with Punctuation, DBSCAN, LCS and jFlex algorithms, and the log clustering results are generated through word segmentation, pattern expression matching, clustering and keyword extraction.
Aggregation of similar logs is realized, and the rules and common problems in the log are discovered, storage costs are reduced, and fault location and abnormal detection are facilitated.
Smart Images

Figure CN2024133378_31072025_PF_FP_ABST
Abstract
Description
A log clustering method and device based on NLP Technical Field
[0001] The embodiments of the present application relate to, but are not limited to, the field of natural language processing technology, and in particular to a log clustering method and device based on NLP. Background Art
[0002] Natural Language Processing (NLP) is a key technology in computer science and artificial intelligence. It primarily uses computers to understand, analyze, process, and generate human language, enabling interaction and application with human language. As enterprise software systems become increasingly large and complex, for example, a service may generate a large number of alarms and multiple error logs in a short period of time. Key error logs in a particular category may contain fewer entries and are easily overwhelmed by other error logs. Classifying and processing the massive amount of heterogeneous logs generated by the system to quickly understand the full picture for subsequent problem location and anomaly detection has become a pressing technical challenge. Summary of the Invention
[0003] To address the above-mentioned problems in the prior art, the present invention proposes a log clustering method and device based on NLP. The technical solution adopted in this application is as follows:
[0004] A log clustering method based on NLP, the method comprising:
[0005] Step 1: Get the original log data and obtain the historical custom pattern metadata from the database. Determine whether the custom pattern metadata expression is successfully matched in the original log data. If so, jump to step 2; if not, jump to step 3.
[0006] Step 2: Assign a unique identifier to the successfully matched custom pattern metadata, and use the jFlex algorithm to extract variables from the corresponding original log data. Based on the extracted variables, the corresponding original log data is decomposed into multiple word segmentation tokens and labeled respectively;
[0007] Step 3: Process the original log data using the punctuation algorithm to generate a log pattern expression based on the type, number, and order of punctuation marks in the original log data. Match the generated log pattern expression with the historical pattern expression. If no match is found, return to step 2. If a match is found, group the pattern expressions according to the LCS algorithm to obtain the log clustering result.
[0008] Furthermore, in the above step 2, the multiple segmentation tokens are respectively labeled to indicate the type of the segmentation token, which includes time type, number type and email type.
[0009] Furthermore, in step 3, if there is no match, the unmatched pattern expression and the corresponding word segmentation token are persisted in the database and become part of the historical pattern expression.
[0010] Furthermore, in step 3, the pattern expressions are grouped and processed according to the LCS algorithm. The specific processing includes: comparing the word segmentation tokens of multiple logs with the matched pattern expressions. If the similarity of the tokens between the logs is greater than or equal to 80%, and the similarity of the longest common string in the pattern expression is greater than or equal to 80%, then they are aggregated into pattern groups, and the pattern groups and the correspondence between the pattern expressions and the pattern groups are synchronized to the database. A single pattern group corresponds to one or more pattern expressions; if the pattern expression corresponding to the newly accessed log belongs to the pattern group, the grouping variable is extracted, and the data in the database where the pattern group expression is located is updated at the same time; if it does not belong, the pattern variable is extracted and the historical pattern expression is updated.
[0011] A log clustering method based on NLP, the method comprising:
[0012] Step 1: Get the original log data and generate log patterns based on the type, quantity, and order of punctuation marks. Save the metadata of these patterns to the database for persistent storage.
[0013] Step 2: Tokenize the raw log data using punctuation and jFlex. A specification file with a set of regular expressions and corresponding operations is fed into the lexical analyzer. The input is matched against the regular expressions in the specification file. When the regular expressions match, the matched text is labeled with the text type, and the raw log data is broken down into multiple tokens.
[0014] Step 3: Extract variables from the multiple tokens that have been decomposed, extract different character strings that appear at the same position in the pattern expression as variables, iterate and calculate the longest common string through LCS, merge the pattern expressions and form pattern groups;
[0015] Step 4: Based on the pattern expression, the DBSCAN algorithm is used to perform secondary clustering and intelligent pattern grouping expressions;
[0016] Step 5: Use the TF-IDF algorithm to extract keywords. Based on the original log text that matches the pattern, use the natural language processing algorithm to name the grouped patterns.
[0017] Furthermore, in the above step 2, the text type is determined through the marking process, and the text type includes a time type, a number type, and an email type.
[0018] Furthermore, in step 3, the longest common string is iteratively calculated using LCS, patterns are merged and grouped, specifically including:
[0019] Step 301: Match existing groups in free mode and update group expressions based on the matching results.
[0020] Step 302: Determine the sampling pattern and calculate the distance matrix; generate new groups according to the clustering pattern and generate grouping expressions; use the new groups to match the remaining patterns, update the grouping expressions, and cluster the new groups;
[0021] Step 303: Calculate the distance matrix based on the clustered new groups, generate new groups based on the clustered groups, generate grouping expressions, match existing groups, update the grouping expressions, and return the updated groups.
[0022] Furthermore, in step 5 above, the algorithm formula of the term frequency-inverse document frequency algorithm TF-IDF is as follows:
[0023] tf ij Indicates the frequency of word i in text j, idf i represents the inverse text frequency of word i, n ij Indicates the number of times word i appears in text j; n kj represents the total number of words in text j, D i Indicates the number of documents containing the word i.
[0024] A log clustering device based on NLP includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above method is implemented.
[0025] A log clustering device based on NLP includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned another method when executing the computer program.
[0026] Through the embodiments of the present application, the following technical effects can be achieved: the present application combines multiple algorithms such as Punctuation, DBSCAN, LCS, and JFLEX to perform data processing on log text to obtain log clustering results, clustering logs with high similarity together, which is conducive to discovering patterns and common problems in the logs, and facilitating problem troubleshooting and fault location from massive logs; massive logs only require a small number of log patterns to represent, extracting the common parts to retain independent information and reduce storage costs.
[0027] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be achieved and obtained through the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application, a brief introduction is given below to the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0029] Figure 1 is a schematic diagram of the algorithm flow of the log clustering method;
[0030] Figure 2 is a schematic diagram of the algorithm composition used in the log clustering method;
[0031] Figure 3 is a business logic diagram of the log clustering method;
[0032] FIG4 is a schematic diagram of the generated pattern identification;
[0033] Figure 5 is a schematic diagram of word segmentation effect;
[0034] Figure 6 is a schematic diagram of the variable extraction process;
[0035] FIG7 is a schematic diagram of the iterative calculation process;
[0036] Figure 8 is a schematic diagram of pattern distance calculation;
[0037] Figure 9 is a schematic diagram of the TF-IDF algorithm process. DETAILED DESCRIPTION
[0038] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0039] It should be understood that in the description of the embodiments of this application, "multiple" (or multiple) means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, and "above," "below," and "within" are understood to include the number itself. The use of "first," "second," and the like in the description is solely for the purpose of distinguishing technical features and is not to be construed as indicating or implying relative importance, or implicitly indicating the number of the indicated technical features, or implicitly indicating the order of the indicated technical features.
[0040] First, let’s explain the technical terms:
[0041] DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is a representative density-based clustering algorithm. Unlike partitioning and hierarchical clustering methods, it defines a cluster as the largest set of density-connected points. It can cluster areas with sufficiently high density and can discover clusters of arbitrary shapes in noisy spatial databases.
[0042] NLP is the abbreviation of Neuro-Linguistic Programming. N (Neuro) refers to the nervous system, including the brain and thinking process; L (Linguistic) refers to language, or more precisely, the process from the input of sensory signals to the formation of meaning; P (Programming) refers to a set of specific instructions to be executed to produce a certain consequence.
[0043] TF-IDF (term frequency-inverse document frequency) is a commonly used weighting technique used in information retrieval and data mining. TF stands for term frequency, and IDF stands for inverse document frequency.
[0044] This application proposes an implementation architecture of a log clustering method based on NLP. The overall algorithm flow of the log clustering method is shown in Figure 1. This method combines the Punctuation, DBSCAN, LCS, and jFlex algorithms, as shown in Figure 2. The received log data is grouped according to the symbol sequence, and various log patterns are generated based on Punctuation's LCS; under various log patterns, the TF-IDF algorithm is used to perform keyword extraction processing on the grouped log data to obtain log labels, and jFlex is used to perform word segmentation processing on the grouped log data to obtain log variable distribution; clusters of log labels are generated through automatic clustering DBSCAN and manual adjustment of pattern distance, and based on Punctuation's MLCS processing.
[0045] In one embodiment, the process of the log clustering method is shown in FIG3 , and the method includes the following steps:
[0046] Step 1: Get the original log data and obtain the historical custom pattern metadata from the database. Determine whether the custom pattern metadata expression is successfully matched in the original log data. If so, jump to step 2; if not, jump to step 3.
[0047] Step 2: Assign a unique identifier to the successfully matched custom pattern metadata, and use the jFlex algorithm to extract variables from the corresponding original log data. Based on the extracted variables, the corresponding original log data is decomposed into multiple word segmentation tokens and labeled respectively;
[0048] In step 2 above, the multiple segmentation tokens are labeled to indicate the type of the segmentation token, which includes time type, number type, and email type;
[0049] Step 3: Process the original log data using the punctuation algorithm, and generate a log pattern expression (pattern CODE) based on the type, number, and order of punctuation marks in the original log data; match the generated log pattern expression in the historical pattern expression. If no match is found, return to step 2 and execute; if a match is found, group the pattern expression according to the LCS algorithm to obtain the log clustering result.
[0050] In step 3, if there is no match, the unmatched pattern expression and the corresponding word segmentation token are persisted to the database and become part of the historical pattern expression;
[0051] In step 3, the pattern expressions are grouped and processed according to the LCS algorithm. The specific processing includes: comparing the word segmentation tokens of multiple logs with the matched pattern expressions. If the similarity of the tokens between the logs is greater than or equal to 80%, and the similarity of the longest common string in the pattern expression is greater than or equal to 80%, it is considered that they can be aggregated into pattern groups. The pattern groups and the correspondence between pattern expressions and pattern groups (a single pattern group corresponds to one or more pattern expressions) are synchronized to the database. If the pattern expression corresponding to the newly accessed log belongs to the pattern group, the grouping variable is extracted and the data in the database where the pattern group expression is located is updated at the same time; if it does not belong to the pattern group, the pattern variable is extracted and the historical pattern expression is updated.
[0052] In another embodiment, the log clustering method includes the following steps:
[0053] Step 1: Get the original log data and generate log patterns based on the type, quantity, and order of punctuation marks. Save the metadata of these patterns to the database for persistent storage.
[0054] In the above step 1, the pattern identifier is generated by using the punctuation technology at the bottom layer, and the generated pattern identifier is shown in FIG4 .
[0055] Step 2: Tokenize the raw log data using punctuation and jFlex. A specification file with a set of regular expressions and corresponding operations is fed into the lexical analyzer. The input is matched against the regular expressions in the specification file. When the regular expressions match, the matched text is labeled with the text type, and the raw log data is broken down into multiple tokens.
[0056] In the above step 2, the text type is determined by the marking process, and the text type includes a time type, a number type, and an email type;
[0057] In the above steps, word segmentation uses the word library to perform word segmentation, displaying the variables more intuitively and making it easier to understand the meaning of the variables, as shown in Figure 5. The lexical analyzer generator takes a specification with a set of regular expressions and corresponding operations as input. It generates a program (lexical analyzer) that reads the input, matches the input with the regular expressions in the specification file, and runs the corresponding operation when the regular expression matches. The lexical analyzer is usually the first front-end step in the compiler, matching keywords, comments, operators, etc., and generating an input token stream for the parser.
[0058] Step 3: Extract variables from the multiple tokens that have been decomposed, extract different character strings that appear at the same position in the pattern expression as variables, iterate and calculate the longest common string through LCS, merge the patterns and form groups;
[0059] In the above steps, variable extraction extracts different strings that appear at the same position in the pattern into variables, initially represented by *. The effect is shown in Figure 5. After tokenization of the original log using punct and jFlex, the original text is broken down into multiple tokens. The longest common string is calculated using LCS, and the patterns are merged to form groups. The effect is shown in Figure 6. To speed up the merging of patterns, iterative calculation is used. The specific calculation process is shown in Figure 7.
[0060] In step 3, the longest common string is iteratively calculated using LCS, patterns are merged and grouped, specifically including:
[0061] Step 301: Match existing groups in free mode and update group expressions based on the matching results.
[0062] Step 302: Determine the sampling pattern and calculate the distance matrix; generate new groups according to the clustering pattern and generate grouping expressions; use the new groups to match the remaining patterns, update the grouping expressions, and cluster the new groups;
[0063] Step 303: Calculate the distance matrix based on the clustered new groups, generate new groups based on the clustered groups, generate grouping expressions, match existing groups, update the grouping expressions, and return the updated groups.
[0064] Step 4: Based on the patterns, use the DBSCAN algorithm to perform secondary clustering, intelligent grouping, and reduce the number of patterns;
[0065] In the above steps, the density-based spatial clustering method DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is used. It is a density-based spatial clustering algorithm. This algorithm divides areas with sufficient density into clusters and discovers clusters of any shape in a noisy spatial database. It defines a cluster as the largest set of density-connected points, as shown in Figure 8. The above example of using the DBSCAN algorithm for secondary clustering and intelligent grouping is as follows:
[0066] The pattern corresponding to log A is: This is lili
[0067] The pattern corresponding to log B is: This is nana, hello balabalabala
[0068] The pattern corresponding to log C is: This is beibei,hi balabalabala
[0069] Calculate according to the DBSCAN formula:
[0070] Distance between AB: 1-(1+0.5) / 2=0.25
[0071] AC distance: 1-(1+0.5) / 2=0.25
[0072] Distance between BC: 1-(0.9+0.9) / 2=0.1
[0073] From the above calculation results, it can be determined that the distance between BC is closer and it is easier to merge into the same pattern group.
[0074] Step 5: Use the TF-IDF algorithm to extract keywords. Based on the original log text that matches the pattern, use the natural language processing algorithm to name the grouped patterns.
[0075] By naming, these grouping results have a default name, which makes it easier for users to distinguish and identify them.
[0076] In the above steps, the algorithm formula of the term frequency-inverse document frequency algorithm TF-IDF is as follows:
[0077] tf ij It represents the frequency of word i in text j, that is, term frequency, which reflects the relationship between the number of times word i appears in text j and the length of text j; idfi represents the inverse document frequency of word i, which is used to measure the importance of word i in the entire text collection. The more texts word i appears in, the smaller its value is; n ij Indicates the number of times word i appears in text j; n kj represents the total number of words in the j text, all words are represented by k; D represents the total number of texts; D j Indicates the number of documents containing the word i.
[0078] TFIDF is actually TF*IDF, which stands for Term Frequency and Inverse Document Frequency. If a word or phrase appears frequently in one article (TF) and rarely in other articles, it is considered to have good class distinction and is suitable for classification. TF represents the frequency of a term in document d. The key idea behind IDF is that the fewer documents containing term t (that is, the smaller n), the higher the IDF, indicating that term t has good class distinction. If the number of documents in a certain category C containing term t is m, and the total number of documents containing t in other categories is k, then the total number of documents containing t is n = m + k. When m is large, n is also large, and the IDF value obtained according to the IDF formula will be small, indicating that term t has poor class distinction. However, if a term appears frequently in documents of a certain category, it indicates that it well represents the characteristics of the documents in that category. Such terms should be given higher weights and selected as feature words for that category to distinguish it from documents of other categories. In a given document, term frequency (TF) refers to how often a given word appears in the document. This number is normalized to prevent it from being biased towards long documents.
[0079] Inverse document frequency (IDF) is a measure of the general importance of a term. The IDF of a specific term can be calculated by dividing the total number of documents by the number of documents containing the term and taking the base-10 logarithm of the resulting quotient. A high term frequency within a specific document combined with a low term frequency across the entire document collection will yield a high-weighted TF-IDF. Therefore, TF-IDF tends to filter out common terms and retain important ones. This statistically based calculation method is often used to assess the importance of a term to a document within a document collection. This function clearly meets the requirements of keyword extraction: the more important a term is to a document, the more likely it is to be a keyword for that document. The TF-IDF algorithm is often used for keyword extraction. An example of the algorithm process is shown in Figure 9.
[0080] In summary, this application uses the log clustering method to classify logs of the same pattern into one category. When generating an alarm, log clustering is used to summarize and abstract the error logs, which enables developers to quickly grasp the overall picture of the logs, and key error information is not easily ignored. Log clustering can group logs with high similarity together and extract common log PATTERNs, which is a direct benefit. It is conducive to discovering patterns and common problems in logs, and it is convenient to troubleshoot and locate faults from massive logs; massive logs only need a small number of log patterns to represent them, and the common parts are extracted to retain independent information and reduce storage costs. Log clustering is very helpful for subsequent functions such as log anomaly detection. Anomaly detection of error logs needs to be based on log clustering. When performing anomaly detection, this application uses deep learning-based log anomaly detection to distinguish between normal system operation error logs and system failure logs through log clustering analysis, the temporal changes in the frequency of occurrence of various types of logs, whether new types of log errors appear after the change is released, etc., and conducts root cause analysis of the cause of the anomaly. For example, SYSLOG logs are used to assist in locating machine failures, and DMSG logs are used to inspect potential abnormal sub-machine failures. Log clustering can not only provide an overview of log data, but also facilitate problem location and anomaly detection. For example, a service generates a large number of alarms in a short period of time, and generates multiple types of error logs at the same time. A certain type of key error log may have fewer entries, which can easily be overwhelmed by other error logs. If log clustering can be used to summarize and abstract error logs at the same time as alarms are generated, developers can quickly grasp the full picture of the logs, and key error information is not easily overlooked.
[0081] The present application also provides a log clustering device. In an exemplary embodiment, the log clustering device includes one or more processors and a memory, with one processor and a memory being used as an example. The processor and the memory may be connected via a bus or other means.
[0082] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs, such as the data processing method in the above-mentioned embodiments of the present application. The processor implements the data processing method in the above-mentioned embodiments of the present application by running the non-transitory software programs and programs stored in the memory.
[0083] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data required for executing the data processing method in the above-mentioned embodiment of the present application, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may include a memory remotely arranged relative to the processor, and these remote memories may be connected to the data processing device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0084] Those skilled in the art will appreciate that all or some of the steps and systems in the method disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). As known to those skilled in the art, the term computer storage media is included in any method or technology for storing information (such as a computer-readable program, a data structure, a program module, or other data) and is volatile and non-volatile, removable, and non-removable. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVD), or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage, or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically embodies computer-readable programs, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0085] The above is a specific description of several implementations of the present application, but the present application is not limited to the above-mentioned implementation methods. Technical personnel familiar with the art can also make various equivalent modifications or substitutions under the sharing conditions that do not violate the essence of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.
Claims
1. A log clustering method based on NLP, characterized in that, The method includes: Step 1: Obtain the original log data, and obtain the historical custom pattern metadata from the database. Determine whether the expression of the custom pattern metadata is successfully matched in the original log data. If so, jump to and execute Step 2; if not, jump to and execute Step 3; Step 2: Assign a unique identifier to the successfully matched custom pattern metadata, and use the jFlex algorithm to extract variables from the corresponding original log data. Decompose the corresponding original log data into multiple tokenized words according to the extracted variables and perform tagging on each of them; Step 3: Perform punctuation algorithm processing on the original log data. Generate a pattern expression for the log according to the punctuation symbol type, quantity, and order in the original log data. Match the generated pattern expression for the log in the historical pattern expressions. If not matched, return to and execute Step 2; if matched, perform grouping processing on the pattern expression according to the LCS algorithm to obtain the log clustering result.
2. The method according to claim 1, characterized in that, In the above Step 2, by performing tagging on each of the decomposed multiple tokenized words to indicate the type of the tokenized word, and the type includes time type, number type, and email type.
3. The method according to claim 1, characterized in that, In Step 3, if not matched, persist the unmatched pattern expression and the corresponding tokenized words into the database and make them part of the historical pattern expressions.
4. The method according to claim 1, characterized in that In Step 3, perform grouping processing on the pattern expression according to the LCS algorithm. The specific processing includes: Compare the tokenized words of multiple logs with the matched pattern expressions. If the similarity of the tokens between logs is greater than or equal to 80%, and the similarity of the longest common substring in the pattern expressions is greater than or equal to 80%, then aggregate them into a pattern group, and synchronize the pattern group and the corresponding relationship between the pattern expression and the pattern group to the database. A single pattern group corresponds to one or more pattern expressions; If the pattern expression corresponding to the newly accessed log belongs to a pattern group, extract the grouped variables and simultaneously update the data in the database where the pattern group expression is located; If not, extract the pattern variables and update the historical pattern expressions.
5. A log clustering method based on NLP, characterized in that, The method includes: Step 1: Obtain the original log data, and generate patterns for the logs based on the punctuation symbol type, quantity, and order. Save the metadata of these patterns to the database for persistent storage; Step 2: Perform word segmentation processing on the original log data based on punctuation and jFlex. Input the specification with a set of regular expressions and corresponding operations into the lexical analyzer. Match the input with the regular expressions in the specification file, and perform tagging processing on the matched text according to the text type when the regular expression is matched, and decompose the original log data into multiple tokenized words; Step 3: Extract variables from the decomposed multiple tokenized words. Extract different strings that appear at the same position in the pattern expression as variables. Perform iterative calculation on the longest common substring through LCS, and merge the pattern expressions to form a pattern group; Step 4: Based on the pattern expression, perform secondary clustering using the DBSCAN algorithm to obtain the intelligent pattern grouping expression; Step 5: Extract keywords using the TF-IDF algorithm. According to the log original text matched in the pattern, use the natural language processing algorithm to identify the name for the grouped patterns.
6. The method according to claim 5, characterized in that In the above Step 2, the text type is determined through the marking process, and the text type includes time type, number type, and email type.
7. The method according to claim 5, wherein In Step 3, iterative calculation of the longest common substring is performed through LCS to merge patterns and form groups, specifically including: Step 301: Match existing groups in the free mode and update the grouping expression according to the matching result; Step 302: Determine the sampling pattern and calculate the distance matrix; generate new groups according to the clustering pattern and generate grouping expressions; use the new groups to match the remaining patterns, update the grouping expressions, and cluster the new groups; Step 303: Calculate the distance matrix according to the newly clustered groups, cluster the grouped generations to generate new groups, generate grouping expressions, match the existing groups, update the grouping expressions, and return the updated groups.
8. The method according to claim 5, wherein In the above step 5, the algorithm formula of the term frequency-inverse document frequency algorithm TF-IDF is as follows: tf ij represents the frequency of the word i appearing in the text j, idf i represents the inverse document frequency of the word i, n ij represents the number of times the word i appears in the text j; n kj represents the total number of words in the text j, D i represents the number of documents containing the word i.
9. A log clustering device based on NLP, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the method described in any one of claims 1 to 4 is implemented.
10. A log clustering device based on NLP, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the method described in any one of claims 5 to 8 is implemented.
Citation Information
Patent Citations
Chameleon real-time log clustering method based on LCS (Local Clustering System)
CN111400500A
Log classification method, log classification device, equipment and medium
CN116578700A
Log classification method and device, computer equipment, storage medium and program product
CN116932753A
NLP-based log clustering method and device
CN117892152A
Communication log with extracted keywords from speech-to-text processing
US8606576B1