Log classification method and device, computer device, storage medium and program product
By using word vectors and similarity algorithms to process log words in log classification, the problem of unreliable classification of unknown types of logs in traditional technologies is solved, and more accurate log type identification and system fault analysis are achieved.
Patent Information
- Application Number
- CN202310854488.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-12
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2043-07-12
AI Technical Summary
In traditional technologies, regular matching cannot detect error logs of unknown types, resulting in unreliable log classification methods.
By obtaining the target log to be classified, extracting log words and filling them into the preset initial log template, and using word vectors and similarity algorithms to determine the log type, including part of speech, word frequency, synonyms and antonyms processing, the text vector similarity is calculated.
It improves the reliability and accuracy of log classification, can automatically analyze potential errors in software systems, implement log root cause analysis, and ensure system security.
Smart Images

Figure CN116932753B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data technology, and in particular to a log classification method, apparatus, computer equipment, storage medium, and program product. Background Art
[0002] Error logs are text files used by software systems to record runtime error information. Real-time classification of error logs can provide operations and maintenance personnel with clearer fault characteristics, helping managers and operations and maintenance personnel debug software system faults and analyze anomalies.
[0003] In traditional technology, error logs are classified through regular matching. However, regular matching cannot detect error logs of unknown types, resulting in the inability to classify such error logs. As a result, the log classification method of traditional technology is unreliable. Summary of the Invention
[0004] Embodiments of the present application provide a log classification method, apparatus, computer device, storage medium, and program product, which can improve the reliability of log classification.
[0005] In a first aspect, the present application provides a log classification method, which includes: obtaining a target log to be classified, and extracting log words from the target log, filling the extracted log words into a preset initial log template to obtain a target log template; obtaining a first text vector of the target log template based on the word vector of each log word in the target log template; determining a target text vector whose similarity meets preset conditions from multiple second text vectors based on the first text vector, wherein the multiple second text vectors correspond to different log types; determining the log type of the target log based on the log type corresponding to the target text vector.
[0006] In one embodiment, a first text vector of a target log template is obtained based on a word vector of each log word in the target log template, including: determining a weight of each log word based on the part of speech and word frequency of each log word in the target log template; and performing a weighted sum operation on the weight and word vector of each log word in the target log template to obtain the first text vector.
[0007] In one embodiment, the weight of each log word is determined based on the part of speech and word frequency of each log word in the target log template, including: for each log word in the target log template, determining the first word frequency of the log word appearing in the log template of the target type, and determining the second word frequency of the log word appearing in the target log template; for each log word in the target log template, determining the part-of-speech weight based on the part of speech of the log word; and determining the weight of each log word in the target log template based on the first word frequency, second word frequency and part-of-speech weight corresponding to each log word in the target log template.
[0008] In one of the embodiments, before obtaining the first text vector of the target log template according to the word vectors of the log words in the target log template, the method further comprises: obtaining synonyms and antonyms of the log words in the target log template; and determining the word vectors of the log words in the target log template according to the synonyms and antonyms of the log words in the target log template.
[0009] In one of the embodiments, the determining the word vectors of the log words in the target log template according to the synonyms and antonyms of the log words in the target log template comprises: dividing the log words in the target log template into first log words and second log words, the first log words being the words existing in the historical log word database, and the second log words being the words not existing in the historical log word database; performing vector conversion processing based on the vector algorithm corresponding to the first log words and the synonyms and antonyms of the first log words to obtain the word vectors of the first log words; and performing vector conversion processing based on the vector algorithm corresponding to the second log words and the synonyms and antonyms of the second log words to obtain the word vectors of the second log words.
[0010] In one of the embodiments, the method further comprises: taking the first text vector and the second text vector as variables of a cosine similarity function to obtain a cosine similarity by the cosine similarity function; taking the first text vector and the second text vector as variables of a text similarity function to obtain a text similarity by the text similarity function; and obtaining the similarity according to the cosine similarity and the text similarity.
[0011] In one of the embodiments, the method further comprises: pre-processing the target log, the pre-processing comprising at least one of removing time stamp, conjunction separation, word segmentation, command entity recognition, and case conversion.
[0012] In a second aspect, the present application provides a log classification device, the device comprising: an extraction module configured to obtain a target log to be classified, and extract log words from the target log, and fill the extracted log words into a preset initial log template to obtain a target log template; a first determination module configured to obtain a first text vector of the target log template according to word vectors of the log words in the target log template; a second determination module configured to determine a target text vector satisfying a preset condition in similarity from a plurality of second text vectors according to the first text vector, wherein the plurality of second text vectors correspond to different log types; and a third determination module configured to determine a log type of the target log according to a log type corresponding to the target text vector.
[0013] In one embodiment, the first determination module is specifically used to determine the weight of each log word according to the part of speech and word frequency of each log word in the target log template; and perform a weighted sum operation on the weight and word vector of each log word in the target log template to obtain a first text vector.
[0014] In one embodiment, the first determination module is specifically used to determine, for each log word in the target log template, a first word frequency of the log word appearing in the target type of log template, and determine a second word frequency of the log word appearing in the target log template; for each log word in the target log template, determine a part-of-speech weight according to the part-of-speech of the log word; and determine the weight of each log word in the target log template according to the first word frequency, second word frequency and part-of-speech weight corresponding to each log word in the target log template.
[0015] In one embodiment, the first determination module is further configured to obtain synonyms and antonyms of each log word in the target log template; and determine a word vector of each log word in the target log template based on the synonyms and antonyms of each log word in the target log template.
[0016] In one embodiment, the first determination module is specifically used to divide the log words in the target log template into first log words and second log words, the first log words are words existing in the historical log word database, and the second log words are words that do not exist in the historical log word database; based on the vector algorithm corresponding to the first log word and the synonyms and antonyms of the first log word, vector conversion processing is performed to obtain the word vector of the first log word; based on the vector algorithm corresponding to the second log word and the synonyms and antonyms of the second log word, vector conversion processing is performed to obtain the word vector of the second log word.
[0017] In one embodiment, the second determination module is further used to use the first text vector and the second text vector as variables of a cosine similarity function to obtain cosine similarity through the cosine similarity function; use the first text vector and the second text vector as variables of a text similarity function to obtain text similarity through the text similarity function; and obtain similarity based on the cosine similarity and the text similarity.
[0018] In one embodiment, the device further includes a preprocessing module for preprocessing the target log, wherein the preprocessing includes at least one of removing timestamps, separating conjunctions, segmenting words, recognizing command entities, and converting uppercase and lowercase letters.
[0019] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the methods described in the first aspect when executing the computer program.
[0020] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the methods described in the first aspect.
[0021] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, characterized in that when the computer program is executed by a processor, the steps of the method described in any one of the above-mentioned first aspects are implemented.
[0022] The above-mentioned log classification method, apparatus, computer equipment, storage medium and program product obtain a target log to be classified, extract log words from the target log, fill the extracted log words into a preset initial log template to obtain a target log template, and then obtain a first text vector of the target log template based on the word vector of each log word in the target log template; then, based on the first text vector, determine a target text vector whose similarity meets preset conditions from multiple second text vectors, wherein the multiple second text vectors correspond to different log types, and then determine the log type of the target log based on the log type corresponding to the target text vector. In this way, the log type of the target log is determined by the similarity between the first text vector and the second text vector corresponding to the target log, which can improve the reliability of log classification compared with traditional technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 A flow chart of a log classification method according to an embodiment;
[0024] Figure 2 Schematic diagram of a flow chart of a method for determining word vectors for each log word in a target log template in one embodiment;
[0025] Figure 3 Schematic diagram of a flow chart of a method for determining the weight of each log word in one embodiment;
[0026] Figure 4 1 is a flow chart of another log classification method according to an embodiment;
[0027] Figure 5 1 is a flow chart of yet another log classification method according to an embodiment;
[0028] Figure 6 is a structural block diagram of a log classification device in one embodiment;
[0029] Figure 7 The figure is a diagram of the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0031] Logs play a crucial role in IT (Information Technology) operations and maintenance. Logs record detailed information about software system runtimes and contain a wealth of system information. Software system developers and operations personnel can monitor software systems based on logs and analyze abnormal behavior and errors. Real-time error log classification provides operations personnel with clearer fault characteristics, helping them debug software system failures and analyze anomalies.
[0032] Current error log classification algorithms have the following problems: 1. Regular matching is often used to discover log variables, but regular matching requires prior knowledge of the error log type and the definition of a strict regular expression. 2. When a failure occurs, the number of error logs increases exponentially, requiring the logging system to classify the logs and detect unknown errors. However, regular matching cannot detect error logs of unknown types, making it impossible to classify these types of error logs. Consequently, traditional log classification methods are unreliable.
[0033] Based on the above-mentioned traditional technology, an embodiment of the present application provides a log classification method, which obtains a target log to be classified and extracts log words from the target log, fills the extracted log words into a preset initial log template to obtain a target log template, and then obtains a first text vector of the target log template based on the word vector of each log word in the target log template; then, based on the first text vector, a target text vector whose similarity meets preset conditions is determined from multiple second text vectors, wherein the multiple second text vectors correspond to different log types, and then the log type of the target log is determined according to the log type corresponding to the target text vector, so as to improve the reliability of log classification.
[0034] It should be noted that the beneficial effects or technical problems solved by the embodiments of the present application are not limited to this one, but may also include other implicit or related problems. For details, please refer to the description of the following embodiments.
[0035] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0036] In one embodiment, Figure 1 As shown, a log classification method is provided, which is described by taking the application of the method to a server as an example, and includes the following steps:
[0037] Step 101 : obtaining a target log to be classified, extracting log words from the target log, and filling the extracted log words into a preset initial log template to obtain a target log template.
[0038] The target log includes an error log. The method for selecting the preset initial log template includes: identifying a recording mode of the target log, and selecting a preset initial log template from a plurality of initial log templates according to the recording mode.
[0039] Optionally, the target log is intelligently identified based on the log template extraction method to extract log words from the target log, and the extracted log words are used to replace the log variables in the preset initial log template to fill the extracted log words into the preset initial log template, thereby obtaining the target log template.
[0040] In this step, we extract log terms from the target log and populate them into the pre-set initial log template, avoiding issues with incomplete and inaccurate regular expression recognition. Furthermore, using the target log template instead of the target log for subsequent processing avoids issues with word frequency anomalies caused by a large number of repeated log terms.
[0041] Step 102 : Obtain a first text vector of the target log template based on the word vector of each log word in the target log template.
[0042] The word vector of a log word refers to the semantic vector of the log word.
[0043] Optionally, the log words are transformed into vectors using a vector algorithm to obtain word vectors for the log words. The vector algorithm can be one of the following: BERT (Bidirectional Encoder Representation from Transformers) algorithms, GPT (Generative Pretrained Transformer) algorithms, Log2Vec, doc2vec, etc. The word vectors of each log word constitute the first text vector of the target log template.
[0044] In another optional embodiment, the log words are divided into first log words and second log words. Vector conversion processing is performed based on the vector algorithm corresponding to the first log words to obtain word vectors for the first log words. Vector conversion processing is performed based on the vector algorithm corresponding to the second log words to obtain word vectors for the second log words. The first log words are words that exist in the historical log word database, and the second log words are words that do not exist in the historical log word database. The word vectors of each first log word and each second log word constitute the first text vector of the target log template.
[0045] Step 103 : determining a target text vector whose similarity meets a preset condition from a plurality of second text vectors according to the first text vector, wherein the plurality of second text vectors correspond to different log types.
[0046] Step 104 : Determine the log type of the target log according to the log type corresponding to the target text vector.
[0047] The second text vector is a text vector that annotates a log template. The log type of the annotated log template is pre-annotated, so the second text vector corresponds to the log type. There are many preset conditions that can be set as needed, and they are not limited here. The following examples illustrate three optional methods.
[0048] Optionally, the second text vector having the greatest similarity to the first text vector is used as the target text vector.
[0049] In another optional embodiment, the similarities between all second text vectors and the first text vector are ranked, and the annotated log templates corresponding to the second text vectors ranked at the top of a preset ranking are output for selection by the operation and maintenance personnel, and the second text vector corresponding to the annotated log template selected by the operation and maintenance personnel is used as the target text vector.
[0050] In this embodiment, the annotated log template corresponding to the second text vector ranked at the top of the preset ranking is output, and the second text vector corresponding to the annotated log template selected by the operation and maintenance personnel is used as the target text vector. This can achieve more accurate determination of the target text vector when the gap between multiple similarities is small, thereby more accurately determining the log type of the target log.
[0051] In yet another optional embodiment, a second text vector having a similarity with the first text vector greater than a preset threshold is used as the target text vector.
[0052] In summary, by obtaining a target log to be classified, and extracting log words from the target log, the extracted log words are filled into a preset initial log template to obtain a target log template, and then a first text vector of the target log template is obtained according to the word vectors of the log words in the target log template; then, a target text vector satisfying a preset condition in similarity is determined from a plurality of second text vectors according to the first text vector, wherein the plurality of second text vectors correspond to different log types, and then the log type of the target log is determined according to the log type corresponding to the target text vector. In this way, the log type of the target log is determined by the similarity between the first text vector corresponding to the target log and the second text vector, which can improve the reliability of log classification compared with traditional techniques.
[0053] In one embodiment, before obtaining the first text vector of the target log template according to the word vectors of the log words in the target log template, the method further comprises: obtaining synonyms and antonyms of the log words in the target log template; and determining the word vectors of the log words in the target log template according to the synonyms and antonyms of the log words in the target log template.
[0054] The synonyms and antonyms of the log words can be obtained from a historical log word database, wherein the historical log word database includes a plurality of corresponding relationships between a plurality of log words and a plurality of synonyms, and a plurality of corresponding relationships between a plurality of log words and a plurality of antonyms. In addition, the antonyms of the log words can also be obtained by directly deleting specific words, for example, for the log word "not found", its antonym can be obtained by directly deleting "not" to obtain its antonym "found".
[0055] Optionally, as shown in Figure 2 A method for determining the word vectors of the log words in the target log template is provided, which comprises:
[0056] In step 201, the log words in the target log template are divided into first log words and second log words, the first log words are words existing in a historical log word database, and the second log words are words not existing in the historical log word database.
[0057] The historical log word database includes a plurality of words.
[0058] Optionally, for a log word in the target log template, if the log word can be found in the historical log word database, the log word is classified as a first log word, that is, the log word is an existing word; if the log word cannot be found in the historical log word database, the log word is classified as a second log word, that is, the log word is a new word.
[0059] Step 202 : Perform vector conversion processing based on the vector algorithm corresponding to the first log word and the synonyms and antonyms of the first log word to obtain a word vector for the first log word.
[0060] The vector algorithm corresponding to the first log word includes one of the BERT-type algorithm and the GPT-type algorithm. The BERT-type algorithm, also known as the BERT-type model, utilizes MLM for pre-training and employs a deep, bidirectional Transformer component to construct the entire model, ultimately generating a deep, bidirectional language representation that can integrate left and right context information. Thus, the word vector for the first log word obtained using the BERT-type algorithm is more accurate. The GPT-type algorithm, also known as the GPT-type model, whose two-stage pre-training and fine-tuning process has been proven to be highly effective in achieving state-of-the-art results in various natural language processing tasks, can adapt to various tasks through transfer learning capabilities and requires relatively little additional training data. Therefore, the GPT algorithm suitable for vector conversion of the first log word is easily accessible, thereby reducing the difficulty of implementing this method.
[0061] Optionally, the first log word, synonyms, and antonyms of the first log word are used as inputs of a vector algorithm corresponding to the first log word, thereby obtaining a word vector for the first log word.
[0062] Step 203 : Perform vector conversion processing based on the vector algorithm corresponding to the second log word and the synonyms and antonyms of the second log word to obtain a word vector for the second log word.
[0063] The vector algorithm corresponding to the second log word includes one of Log2Vec and doc2vec. Log2Vec is a modular approach to cyberspace threat detection based on heterogeneous graph embedding. It uses a heuristic method to represent various relationships between log records as a heterogeneous graph, and uses an improved graph embedding to represent log records in the heterogeneous graph as low-dimensional vectors. Log2Vec can be used to obtain the word vector of the second log word. Doc2vec is an unsupervised algorithm that can obtain vectors for sentences, paragraphs, and text. Doc2vec can be used to obtain the word vector of the second log word.
[0064] Optionally, the second log word, its synonyms, and its antonyms are used as inputs to a vector algorithm corresponding to the second log word, thereby obtaining a word vector for the second log word.
[0065] In this embodiment, word vectors of log words are obtained based on synonyms and antonyms, which can expand the amount of information of the first text vector. Therefore, in the process of determining a target text vector whose similarity meets preset conditions from multiple second text vectors based on the first text vector, the second text vector similar to the first text vector can be accurately determined, thereby increasing the robustness of log classification.
[0066] In one embodiment, a first text vector of a target log template is obtained based on a word vector of each log word in the target log template, including: determining a weight of each log word based on the part of speech and word frequency of each log word in the target log template; and performing a weighted sum operation on the weight and word vector of each log word in the target log template to obtain the first text vector.
[0067] Among them, part of speech refers to nouns, verbs, adjectives, etc.; the frequency of log words refers to the frequency of occurrence of the log words.
[0068] Optional, such as Figure 3 As shown, a method for determining the weight of each log word is provided, which determines the weight of each log word according to the part of speech and word frequency of each log word in the target log template, including:
[0069] Step 301 : For each log word in the target log template, determine a first word frequency of the log word in the target type of log template, and determine a second word frequency of the log word in the target log template.
[0070] Among them, the target type can be selected for preset or randomly selected. If the target type is preset, the preset initial log template will be used as the log template of the target type; if the target type is randomly selected, since multiple types of log templates are stored in the historical log word database, a type of log template is randomly selected from the historical log word database as the log template of the target type.
[0071] Optionally, the calculation method of the first word frequency includes: counting the number of times a log word appears in a log template of the target type to obtain a first number of occurrences; counting the sum of the number of times all words appear in the log template of the target type to obtain the sum of the first number of occurrences; and taking the ratio of the first number of occurrences to the sum of the first number of occurrences as the first word frequency.
[0072] The first word frequency is expressed using the mathematical formula as follows:
[0073]
[0074] wherein n i,j represents the number of times that the log word i appears in the target log template, n k,j represents the number of times that the word k appears in the target log template, and N represents the number of words in the target log template.
[0075] Optionally, the method for calculating the second word frequency comprises: counting the number of times that the log word appears in the target log template to obtain a second number of times; counting the sum of the number of times that all log words appear in the target log template to obtain a second sum of the number of times; and taking the ratio of the second number of times to the second sum of the number of times as the first word frequency.
[0076] The second word frequency is expressed by a mathematical formula as follows:
[0077]
[0078] wherein n i represents the number of times that the log word i appears in the target log template, n z represents the number of times that the log word z appears in the target log template, and M represents the number of log words in the target log template.
[0079] In step 302, for each log word in the target log template, the part-of-speech weight of the log word is determined according to the part of speech of the log word.
[0080] Optionally, the part-of-speech weight of the log word can be determined according to the part of speech of the log word and a preset initial part-of-speech weight. For example, the initial part-of-speech weight of a noun is 0.6, the initial part-of-speech weight of a verb is 0.3, and the initial part-of-speech weight of an adjective is 0.1, i.e., noun > verb > adjective. Therefore, if the part of speech of the log word is a noun, the part-of-speech weight of the log word is 0.6.
[0081] In another optional embodiment, the parts of speech of the log words in the target log template are distinguished, and the number of each part of speech is counted. Then, for each part of speech, the ratio of the number of the part of speech to the sum of the number of all parts of speech is taken as the weight of the part of speech, so that the weight of each part of speech can be obtained, and then the part-of-speech weight of the log word can be determined according to the part of speech of the log word, i.e., if the log word is a noun, the part-of-speech weight of the log word is the weight of the noun.
[0082] For example, the weight of a noun is expressed by a mathematical formula as follows:
[0083]
[0084] wherein m 名词 , m 动词 , and m 形容词They represent the number of nouns, verbs, and adjectives in the target log template respectively.
[0085] Step 303 : Determine the weight of each log word in the target log template according to the first word frequency, the second word frequency, and the part-of-speech weight corresponding to each log word in the target log template.
[0086] Optionally, for each log word, the inverse of the second word frequency of the log word is used as the variable of the log function to obtain the logarithmic difference of the log word through the log function, and the product of the first word frequency of the log word and the logarithmic difference of the log word and the part-of-speech weight of the log word is used as the weight of the log word.
[0087] The weight of log words is expressed using the following mathematical formula:
[0088]
[0089] Among them, w i,词性 Represents the part-of-speech weight of log word i.
[0090] According to the above steps 301 to 303, the weight of each log word in the target log template can be obtained. According to the above steps 201 to 203, the word vector of each log word in the target log template can be obtained. Then, the weight of each log word in the target log template and the word vector are multiplied, that is, a weighted sum operation is performed to obtain the first text vector.
[0091] In this embodiment, the target log template is processed using the weights of log words, taking into account part of speech and word frequency information, which effectively reflects the importance of log words and the distribution of feature words, and is more suitable for log scenarios.
[0092] In one embodiment, the method further includes: using the first text vector and the second text vector as variables of a cosine similarity function to obtain cosine similarity through the cosine similarity function; using the first text vector and the second text vector as variables of a text similarity function to obtain text similarity through the text similarity function; and obtaining similarity based on the cosine similarity and the text similarity.
[0093] Optionally, A represents the first text vector and B represents the second text vector, and the calculation formula of cosine similarity cosθ is as follows:
[0094]
[0095] The text similarity J(A, B) is calculated using the Jaccard similarity coefficient formula, which is as follows:
[0096]
[0097] The similarity is obtained according to the cosine similarity and the text similarity, and includes:
[0098] The calculation formula of the similarity X is as follows:
[0099] X = 0.618 * cos 0 + (1-0.618) * J(A, B)
[0100] 0.618 is a golden section value, which can be set according to needs, and is not limited here.
[0101] In this embodiment, compared with obtaining the similarity by using only the cosine similarity or the text similarity, the similarity is obtained by using the cosine similarity and the text similarity, so that the accuracy of the similarity calculation can be improved.
[0102] In one of the embodiments, the method further includes: preprocessing the target log, and the preprocessing includes at least one of removing a timestamp, conjunction separation, word segmentation, command entity recognition, and case conversion.
[0103] The removing of the timestamp refers to removing variables belonging to time in the target log.
[0104] The conjunction separation refers to separating words without spaces, punctuation marks, or the like between the words, such as mysqlconnector.
[0105] The word segmentation refers to segmenting the target log by spaces, punctuation marks, or the like, such as segmenting the target log by using NLP (Natural Language Processing) technology.
[0106] The command entity recognition refers to converting a website, an address, a parameter, or the like into an entity name.
[0107] The case conversion refers to uniformly converting letters in the target log into upper case or lower case.
[0108] In another optional embodiment, as shown in Figure 4 Another log classification method is provided:
[0109] Step 401, real-time acquisition of a target log to be classified.
[0110] Step 402, preprocessing of the target log.
[0111] Step 403, extraction of log words from the preprocessed target log, filling of the extracted log words into a preset initial log template, and obtaining of a target log template.
[0112] Step 404, calculating the weight of each log word in the target log template, including determining the first word frequency of the log word appearing in the log template of the target type, wherein the log template is a labeled log template pre-labeled with a log type.
[0113] Step 405, log template vectorization, that is, calculating the word vector of each log word in the target log template, and performing weighted summation operation according to the weight and word vector of each log word in the target log template to obtain a first text vector.
[0114] Step 406, calculating the similarity of the first text vector and each second text vector, including calculating the cosine similarity and the text similarity, wherein the second text vector is a text vector of a labeled log template, and the log type of the labeled log template is pre-labeled.
[0115] Step 407, determining the log type of the second text vector with the largest similarity to the first text vector as the log type of the target log.
[0116] In another optional embodiment, as shown in Figure 5 , another log classification method is provided, which includes two parts, one part is calculating the weight of each log word in the target log template, and the other part is calculating the first text vector of the target log template.
[0117] For calculating the weight of each log word in the target log template, it includes:
[0118] For each log word, the part-of-speech weight of the log word is calculated; the weight of the log word is calculated according to the part-of-speech weight of the log word;
[0119] And for each log word, the synonyms and antonyms of the log word are obtained.
[0120] For calculating the first text vector of the target log template, it includes:
[0121] Dividing the log words in the target log template into first log words and second log words;
[0122] For each first log word, performing vector conversion processing based on the vector algorithm corresponding to the first log word and the synonyms and antonyms of the first log word to obtain the word vector of the first log word;
[0123] For each second log word, performing vector conversion processing based on the vector algorithm corresponding to the second log word and the synonyms and antonyms of the second log word to obtain the word vector of the second log word;
[0124] Log template vectorization is to perform a weighted sum operation based on the weight and word vector of each log word in the target log template to obtain the first text vector.
[0125] Furthermore, this application enables automated log analysis during log classification, enabling early detection of potential errors in software system operation and proactively addressing any inappropriate behavior. Furthermore, this application can also perform root cause analysis of logs, helping administrators and maintenance personnel debug system failures and analyze anomalies, ensuring software system security.
[0126] It should be understood that although Figures 1-5 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figures 1-5 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The order of execution of these steps or stages is not necessarily one by one, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.
[0127] In one embodiment, Figure 6 As shown, a log classification device is provided. The log classification device 600 includes: an extraction module 601, a first determination module 602, a second determination module 603 and a third determination module 604, wherein:
[0128] The extraction module 601 is used to obtain a target log to be classified, extract log words from the target log, and fill the extracted log words into a preset initial log template to obtain a target log template.
[0129] The first determining module 602 is configured to obtain a first text vector of the target log template according to the word vector of each log word in the target log template.
[0130] The second determining module 603 is configured to determine a target text vector whose similarity satisfies a preset condition from a plurality of second text vectors according to the first text vector, wherein the plurality of second text vectors correspond to different log types.
[0131] The third determining module 604 is configured to determine the log type of the target log according to the log type corresponding to the target text vector.
[0132] In one embodiment, the first determination module 602 is specifically configured to determine the weight of each log word according to the part of speech and word frequency of each log word in the target log template; and perform a weighted sum operation on the weight and word vector of each log word in the target log template to obtain a first text vector.
[0133] In one embodiment, the first determination module 602 is specifically used to determine, for each log word in the target log template, a first word frequency of the log word appearing in the target type of log template, and determine a second word frequency of the log word appearing in the target log template; for each log word in the target log template, determine a part-of-speech weight according to the part-of-speech of the log word; and determine the weight of each log word in the target log template according to the first word frequency, second word frequency and part-of-speech weight corresponding to each log word in the target log template.
[0134] In one embodiment, the first determination module 602 is further configured to obtain synonyms and antonyms of each log word in the target log template; and determine a word vector of each log word in the target log template based on the synonyms and antonyms of each log word in the target log template.
[0135] In one embodiment, the first determination module 602 is specifically used to divide the log words in the target log template into first log words and second log words, the first log words are words existing in the historical log word database, and the second log words are words that do not exist in the historical log word database; based on the vector algorithm corresponding to the first log word and the synonyms and antonyms of the first log word, vector conversion processing is performed to obtain the word vector of the first log word; based on the vector algorithm corresponding to the second log word and the synonyms and antonyms of the second log word, vector conversion processing is performed to obtain the word vector of the second log word.
[0136] In one embodiment, the second determination module 603 is further configured to use the first text vector and the second text vector as variables of a cosine similarity function to obtain cosine similarity through the cosine similarity function; use the first text vector and the second text vector as variables of a text similarity function to obtain text similarity through the text similarity function; and obtain similarity based on the cosine similarity and the text similarity.
[0137] In one embodiment, the device further includes a preprocessing module for preprocessing the target log, wherein the preprocessing includes at least one of removing timestamps, separating conjunctions, segmenting words, recognizing command entities, and converting uppercase and lowercase letters.
[0138] For the specific definition of the log classification device, please refer to the definition of the log classification method above, which will not be repeated here. The various modules in the above-mentioned log classification device can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0139] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store multiple second text vectors and different log types corresponding to the multiple second text vectors. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a compilation method is implemented.
[0140] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0141] In one embodiment, a computer device is provided, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any of the above method embodiments when executing the computer program.
[0142] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.
[0143] An embodiment of the present application further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the steps of any of the above method embodiments.
[0144] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0145] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.
[0146] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be pointed out that for those skilled in the art, without departing from the concept of the present application, some modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A log classification method, characterized in that: The method comprises: Obtaining a target log to be classified, extracting log words from the target log, and filling the extracted log words into a preset initial log template to obtain a target log template; the preset initial log template is selected from multiple initial log templates based on the recording mode of the target log by identifying the recording mode; Obtaining a first text vector of the target log template according to the word vector of each log word in the target log template; Determining, based on the first text vector, a target text vector whose similarity satisfies a preset condition from a plurality of second text vectors, wherein the plurality of second text vectors correspond to different log types; Determining the log type of the target log according to the log type corresponding to the target text vector; Obtaining a first text vector of the target log template based on the word vector of each log word in the target log template includes: determining a weight of each log word according to the part of speech and word frequency of each log word in the target log template; performing a weighted sum operation on the weight and word vector of each log word in the target log template to obtain the first text vector; The determining the weight of each log word according to the part of speech and word frequency of each log word in the target log template includes: for each log word in the target log template, determining a first word frequency of the log word appearing in a log template of the target type, and determining a second word frequency of the log word appearing in the target log template; for each log word in the target log template, determining a part-of-speech weight according to the part of speech of the log word; and determining the weight of each log word in the target log template according to the first word frequency, the second word frequency, and the part-of-speech weight corresponding to each log word in the target log template.
2. The method according to claim 1, characterized in that The determining the weight of each log word in the target log template according to the first word frequency, the second word frequency, and the part-of-speech weight corresponding to each log word in the target log template includes: For each of the log words, the inverse of the second word frequency of the log word is used as the variable of the log function to obtain the logarithmic difference of the log word through the log function, and the product of the first word frequency of the log word and the logarithmic difference of the log word and the part-of-speech weight of the log word is used as the weight of the log word.
3. The method according to claim 1, characterized in that Before obtaining the first text vector of the target log template based on the word vector of each log word in the target log template, the method further includes: Obtaining synonyms and antonyms for each log word in the target log template; Determine a word vector for each log word in the target log template based on synonyms and antonyms of each log word in the target log template.
4. The method according to claim 3, characterized in that The step of determining a word vector for each log word in the target log template based on the synonyms and antonyms of each log word in the target log template includes: Dividing the log words in the target log template into first log words and second log words, wherein the first log words are words existing in a historical log word database, and the second log words are words not existing in the historical log word database; Performing vector conversion processing based on a vector algorithm corresponding to the first log word and synonyms and antonyms of the first log word to obtain a word vector for the first log word; Based on the vector algorithm corresponding to the second log word and the synonyms and antonyms of the second log word, vector conversion processing is performed to obtain a word vector for the second log word.
5. The method according to claim 1, wherein The method further comprises: Using the first text vector and the second text vector as variables of a cosine similarity function to obtain cosine similarity through the cosine similarity function; Using the first text vector and the second text vector as variables of a text similarity function to obtain text similarity through the text similarity function; The similarity is obtained according to the cosine similarity and the text similarity.
6. The method according to claim 1, characterized in that The method further comprises: The target log is preprocessed, where the preprocessing includes at least one of removing timestamps, separating conjunctions, segmenting words, recognizing command entities, and converting uppercase and lowercase letters.
7. A log classification device, characterized in that: The device comprises: an extraction module, configured to obtain a target log to be classified, extract log terms from the target log, and fill the extracted log terms into a preset initial log template to obtain the target log template; the preset initial log template is selected from a plurality of initial log templates based on the recording pattern of the target log by identifying the recording pattern; A first determining module, configured to obtain a first text vector of the target log template based on a word vector of each log word in the target log template; a second determining module, configured to determine, based on the first text vector, a target text vector whose similarity satisfies a preset condition from a plurality of second text vectors, wherein the plurality of second text vectors correspond to different log types; A third determining module, configured to determine the log type of the target log according to the log type corresponding to the target text vector; The first determination module is specifically configured to determine the weight of each log word according to the part of speech and word frequency of each log word in the target log template; perform a weighted sum operation on the weight and word vector of each log word in the target log template to obtain the first text vector; The first determination module is specifically configured to determine, for each log word in the target log template, a first word frequency of the log word appearing in the target type log template, and a second word frequency of the log word appearing in the target log template; determine, for each log word in the target log template, a part-of-speech weight according to the part-of-speech of the log word; and determine a weight of each log word in the target log template according to the first word frequency, the second word frequency, and the part-of-speech weight corresponding to each log word in the target log template.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Abnormality tracing method combining system log and origin graph
CN112765603A
Application log analysis method and device, equipment and storage medium
CN114610881A
Log classification method and device, electronic equipment and storage medium
CN115982366A