A log classification method and device
By combining machine learning and knowledge engineering methods, using word frequency and frequency modulation models to adjust the number of occurrences of characteristic words in log classification, the problem of low log classification accuracy caused by uneven training data is solved, and higher classification accuracy and less labor costs are achieved.
Patent Information
- Application Number
- CN201911060648.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-01
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2039-11-01
AI Technical Summary
The existing log classification model has low classification accuracy due to uneven training data, especially the small number of high-level error log samples, which affects the classification effect of the model.
Combining machine learning algorithms and knowledge engineering, the number of occurrences of characteristic words under different log classifications is adjusted through word frequency model and frequency modulation model, and the frequency modulation matrix is used to amplify the characteristic word frequency of log categories with small samples to simulate the effect of increasing sample number, thereby reducing the problem of inaccurate model training.
It improves the accuracy of the log classification model, reduces the impact of sample imbalance, saves labor costs, and corrects classification errors through regression analysis.
Smart Images

Figure CN110929028B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of financial technology (Fintech), and in particular, to a method and device for log classification. Background Art
[0002] With the development of computer technology, more and more technologies are applied in the financial field. The traditional financial industry is gradually transforming into financial technology, and machine learning technology is no exception. However, due to the security and real-time requirements of the financial and payment industries, higher requirements are also put forward for machine learning technology.
[0003] Currently, the common idea for log classification is based on the text classification algorithm of machine learning. The text classification algorithm is based on statistical theory, and uses the algorithm to enable the machine to have an automatic learning ability similar to that of humans, that is, to perform statistical analysis on the known training data to obtain rules, and then use the rules to perform predictive analysis on the unknown data. Since machine learning technology has good practical performance in the field of text classification, it has become the mainstream in the field of log analysis and classification.
[0004] When training a classification model, the problem of unbalanced training data is usually encountered. Taking error logs as an example, the higher and more serious the level of the error, the smaller the general probability of occurrence, and thus the smaller the number of samples of this type. Using an unbalanced sample set for model training often cannot obtain good results, and the accuracy of model classification is relatively low. Summary of the Invention
[0005] Embodiments of the present invention provide a method and device for log classification, which combines machine learning algorithms with knowledge engineering to overcome the problem of unbalanced training data in the sample set, thereby improving the accuracy of model classification.
[0006] A method for log classification provided by an embodiment of the present invention includes:
[0007] Determine the number of occurrences of each feature word in the log to be classified;
[0008] Determine the log classification to which the log to be classified belongs according to the number of occurrences of each feature word in the log to be classified and the classification model; the classification model is determined according to the conditional probability of each feature word in each log classification in the sample log;
[0009] Wherein, the conditional probability of each feature word in each log classification is determined according to the word frequency model and the frequency modulation model; the word frequency model includes the number of occurrences of each feature word in each log classification, and the frequency modulation model includes the adjustment parameter of each feature word in each log classification, and the adjustment parameter is used to adjust the number of times of the corresponding feature word in the corresponding log classification.
[0010] Optionally, the conditional probability of each feature word under each log classification is determined according to the word frequency model and the frequency modulation model, including:
[0011] Perform the following operations for each feature word under each log classification:
[0012] Determine the sum of the occurrences of each feature word under the log classification;
[0013] According to the number of times the feature word appears in the word frequency model, the adjustment parameter of the feature word in the frequency modulation model, and the sum of the occurrences of each feature word under the log classification, determine the conditional probability of the feature word under the log classification.
[0014] Optionally, the word frequency model is a word frequency matrix of m rows × n columns, and the frequency modulation model is a frequency modulation matrix of m rows × n columns; the log classification corresponding to the i-th row in the word frequency matrix is the same as the log classification corresponding to the i-th row in the frequency modulation matrix, and the feature word corresponding to the j-th column in the word frequency matrix is the same as the feature word corresponding to the j-th column in the frequency modulation matrix; 0 < i ≤ m, 0 < j ≤ n;
[0015] The determining the conditional probability of the feature word under the log classification according to the number of times the feature word appears in the word frequency model, the adjustment parameter of the feature word in the frequency modulation model, and the sum of the occurrences of each feature word under the log classification includes:
[0016] Determine the conditional probability of the feature word under the log classification according to formula (1);
[0017] The formula (1) is:
[0018]
[0019] where, x j is the feature word in the j-th column; T i is the log classification in the i-th row; P(x j |T i ) is the conditional probability of x i under T j ; A(i, j) is the number of times the feature word corresponding to the j-th column appears in the log classification corresponding to the i-th row; B(i, j) is the adjustment parameter of the feature word corresponding to the j-th column in the log classification corresponding to the i-th row; count(T i ) is the sum of the occurrences of each feature word under T i ; α is the smoothing coefficient; n is the number of columns of the frequency modulation matrix or the word frequency matrix.
[0020] Optionally, the classification model is determined according to the conditional probability of each feature word in each log classification in the sample log, including:
[0021] For each feature word, determine the sum of the conditional probabilities of the feature word under each log classification; determine the feature weight of the feature word under each log classification as the ratio of the conditional probability of the feature word under each log classification to the sum of the conditional probabilities of the feature word under each log classification.
[0022] Form a feature weight matrix with the feature weights of each feature word in each log classification, and use the feature weight matrix as the classification model.
[0023] In the above technical solution, a frequency modulation matrix is used to adjust the word frequency of feature words in some log categories with fewer sample logs, so as to amplify the word frequency of the feature word under this log classification, simulating the effect of increasing the number of sample logs in this log category, thereby reducing the problem of inaccurate model training caused by the imbalance of sample logs corresponding to each log category.
[0024] Correspondingly, an embodiment of the present invention further provides a log classification device, including:
[0025] A determination unit, a classification unit, and a training unit;
[0026] The determination unit is configured to determine the number of occurrences of each feature word in the log to be classified.
[0027] The classification unit is configured to determine the log classification to which the log to be classified belongs according to the number of occurrences of each feature word in the log to be classified and the classification model; the classification model is determined by the training unit according to the conditional probability of each feature word in each log classification of the sample logs.
[0028] Among them, the conditional probability of each feature word in each log classification is determined by the training unit according to the word frequency model and the frequency modulation model; the word frequency model includes the number of occurrences of each feature word in each log classification, and the frequency modulation model includes the adjustment parameter of each feature word in each log classification, and the adjustment parameter is used by the training unit to adjust the number of times of the corresponding feature word in the corresponding log classification.
[0029] Optionally, the training unit is specifically configured to:
[0030] Perform the following operations for each feature word in each log classification:
[0031] Determine the sum of the number of occurrences of each feature word in the log classification.
[0032] According to the number of times of the feature word in the word frequency model, the adjustment parameter of the feature word in the frequency modulation model, and the sum of the number of occurrences of each feature word in the log classification, determine the conditional probability of the feature word in the log classification.
[0033] Optionally, the word frequency model is a word frequency matrix of m rows × n columns, and the frequency modulation model is a frequency modulation matrix of m rows × n columns; the log classification corresponding to the i-th row in the word frequency matrix is the same as the log classification corresponding to the i-th row in the frequency modulation matrix, and the feature word corresponding to the j-th column in the word frequency matrix is the same as the feature word corresponding to the j-th column in the frequency modulation matrix; 0 < i ≤ m, 0 < j ≤ n;
[0034] The training unit is specifically configured to:
[0035] Determine the conditional probability of the feature word under the log classification according to formula (1);
[0036] The formula (1) is:
[0037]
[0038] where x j is the feature word in the j-th column; T i is the log classification in the i-th row; P(x j |T i ) is the conditional probability of x i under T j ; A(i, j) is the number of times the feature word corresponding to the j-th column appears in the log classification corresponding to the i-th row; B(i, j) is the adjustment parameter of the feature word corresponding to the j-th column in the log classification corresponding to the i-th row; count(T i ) is the sum of the number of times each feature word appears under T i ; α is the smoothing coefficient; n is the number of columns of the frequency modulation matrix or the frequency modulation matrix.
[0039] Optionally, the training unit is specifically configured to:
[0040] For each feature word, determine the sum of the conditional probabilities of the feature word under each log classification; determine the ratio of the conditional probability of the feature word under each log classification to the sum of the conditional probabilities of the feature word under each log classification as the feature weight of the feature word under each log classification;
[0041] Form a feature weight matrix with the feature weights of each feature word in each log classification, and use the feature weight matrix as the classification model.
[0042] Correspondingly, an embodiment of the present invention further provides a computing device, including:
[0043] A memory for storing program instructions;
[0044] A processor for calling the program instructions stored in the memory and executing the above log classification method according to the obtained program.
[0045] Correspondingly, an embodiment of the present invention further provides a computer-readable non-volatile storage medium, including computer-readable instructions. When a computer reads and executes the computer-readable instructions, the computer is caused to execute the above-mentioned log classification method. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 A schematic diagram of a system architecture provided by an embodiment of the present invention;
[0048] Figure 2 A schematic flowchart of a log classification method provided by an embodiment of the present invention;
[0049] Figure 3 A schematic flowchart of a method for determining conditional probability provided by an embodiment of the present invention;
[0050] Figure 4 A schematic flowchart of a method for determining feature weights provided by an embodiment of the present invention;
[0051] Figure 5 A schematic flowchart of another log classification method provided by an embodiment of the present invention;
[0052] Figure 6 A schematic diagram of the structure of a log classification device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0054] In order to better explain the embodiments of the present invention, the Naive Bayes classification algorithm involved in the embodiments of the present invention is explained as follows:
[0055] There are many common classification algorithms at present, such as Bayesian, neural network, decision tree, KNN (K-Nearest Neighbor), SVM (Support Vector Machine), etc. Among them, Bayesian classification is a general term for a class of classification algorithms. These algorithms are all based on Bayes' theorem, so they are collectively called Bayesian classification. And naive Bayesian classification is the simplest and most common classification method in Bayesian classification. Bayes' theorem is named after the British mathematician Bayes and is used to solve the relationship between two conditional probabilities. Simply put, it is how to obtain the probability of P(B|A) when P(A|B) is known. Naive Bayes assumes that the feature P(A) is independent under a specific result P(B). The Bayesian algorithm calculates the probability of P(B|A) occurring through the three probabilities of P(A|B), P(A), and P(B) that are known. Its calculation method can be attributed to Bayes' formula, and Bayes' formula is shown in formula (2).
[0056]
[0057] In the above Bayes' formula, each probability has a specific name:
[0058] P(B) is the probability that event B occurs in the sample space, also called the prior probability of event B.
[0059] P(A) is the probability that event A occurs in the sample space, also called the prior probability of event A.
[0060] P(A|B) is the conditional probability of A given that B has occurred, called the likelihood function.
[0061] P(B|A) is the conditional probability of B given that A has occurred, called the posterior probability.
[0062] P(A|B) / P(A) is the adjustment factor, also called the standard likelihood.
[0063] The basic method of naive Bayes: Based on statistical data, according to the conditional probability formula, calculate the probability that the sample of the current feature belongs to a certain classification, and select the classification with the largest probability. For the given item to be classified, solve the probability that each category appears under the condition that this item appears. Whichever is the largest, it is considered that this item to be classified belongs to that category.
[0064] The calculation process is as follows:
[0065] (1) x = {a1, a2, ……, am} is the item to be classified, and each a is a feature attribute of x;
[0066] (2) There is a category set C = {y1, y2, ……, yn};
[0067] (3) Calculate P(y1∣x), P(y2∣x), ……, P(yn∣x) respectively;
[0068] (4) P(yk∣x) = max{P(y1∣x), P(y2∣x), ……, P(yn∣x)};
[0069] Figure 1 An exemplary system architecture applicable to the log classification method provided by the embodiments of the present invention is shown. The system architecture may include a data source module, a front - end module, a back - end module, a classification algorithm module, and a database. The functions of each module are as follows:
[0070] Data source module: Provides the error log text used for training the model in the embodiments of the present invention, which can also be referred to as the source error log.
[0071] Front - end module: Responsible for providing a Web interface, mainly used to display log classification information and provide operation entrances such as data management for users.
[0072] Back - end module: Mainly used for log processing, responsible for pulling the original log text from the data source, and cleaning it (filtering out worthless text content in ways such as regular expression matching), removing duplicates (merging samples with too high similarity), and finally storing the generated sample set (training set) in the database. In addition, the back - end module is also responsible for providing data operation interfaces, automatically calling the classification algorithm module for model training, and storing the model parameters in the database.
[0073] Classification algorithm module: Responsible for the training of the classifier model and the classification function of sample logs.
[0074] Database: Used to store various types of data such as processed standard sample logs (error sample log sets), frequency modulation matrix information, configuration data, and classification information.
[0075] Based on the above description, Figure 2 An exemplary flow of a log classification method provided by the embodiments of the present invention is shown. This flow can be executed by a log classification device, which can be located in the classification algorithm module and can be the classification algorithm module.
[0076] As Figure 2 shown, this flow specifically includes:
[0077] Step 201, determine the number of occurrences of each feature word in the log to be classified;
[0078] Step 202, determine the log classification to which the log to be classified belongs according to the number of occurrences of each feature word in the log to be classified and the classification model.
[0079] In an embodiment of the present invention, a feature word refers to a word or phrase determined according to multiple sample logs in a sample set. Since sample logs are essentially in text format and cannot directly participate in calculations, it is necessary to vectorize the sample logs first. In one implementation, the bag-of-words model can be used to vectorize the sample logs. Taking words as the basic processing unit, all the words in the sample set are first summarized to obtain a vocabulary of size N, and each sample log in the sample set is mapped into an N-dimensional vector. The value on each dimension represents the number of feature words present in the sample log (which can also be said to be the word frequency of the feature words present in the sample log). This N-dimensional vector reflects the word frequency information in the sample log.
[0080] For example, assume that a sample set is summarized to generate a vocabulary of size 10: ("async", "at", "connection", "db", "error", "jdbc", "mysql", "redis", "timeout", "user");
[0081] Now, vectorize a sample log "mysql jdbc connection timeout error" according to the above rules to generate a vector of length 10: (0 0 1 0 1 1 1 0 1 0);
[0082] In the above example, when using the bag-of-words model for vectorization, since the word frequency is counted for each individual word, there will be a problem of loss of word order information: for example, the phrase "dead lock" will be split into two independent features "dead" and "lock" for statistics, and the semantics of the phrase itself is lost. To solve this problem, in an embodiment of the present invention, when vectorizing the text, the text can be split in the way of combining n words. The adjacent words of length n are combined into a new feature and added to the vocabulary. Here, n can be set according to experience. For example, when n is set to 2, two consecutive words in the sample log can be used as a word combination to obtain new feature words.
[0083] Still taking the above sample log as an example, in the case of n = 2, the following feature words will be generated:
[0084] ("mysql", "jdbc", "connection", "timeout", "error", "mysql jdbc", "jdbc connection", "connection timeout", "timeout error");
[0085] Using the way of combining n words for text splitting can effectively retain semantic feature words.
[0086] After determining the number of occurrences of each feature word in the log to be classified, the log to be classified can also be vectorized, such as generating a vector of length 10: (0 1 1 0 1 1 1 1 1 0), and then, based on the vector generated from the log to be classified and the classification model, in combination with the Bayesian classification algorithm, determine the log classification to which the log to be classified belongs.
[0087] In the embodiments of the present invention, the classification model is determined according to the conditional probability of each feature word in the sample log under each log classification, where the conditional probability of each feature word under each log classification is determined according to the word frequency model and the frequency modulation model.
[0088] Specifically, the word frequency model includes the number of occurrences of each feature word in each log classification. The word frequency model can be presented in the form of a word frequency matrix, or in the form of a word frequency array or other forms. The word frequency model can be determined according to the feature words in each sample log in the sample set.
[0089] Taking the determination of the word frequency matrix according to the feature words in each sample log in the sample set as an example, the description is as follows:
[0090] There are sample logs in the sample set as shown in Table 1, that is, there are three log classifications in the sample set, namely httperror, db error, and redis error; http error includes sample log 1, sample log 2, and sample log 3, db error includes sample log 4, sample log 5, sample log 6, and sample log 7, and redis error includes sample log 8 and sample log 9. And each sample log corresponds to its own vector, such as the vector corresponding to sample log 1 is (2 0 3 0 4 0 0 0 3).
[0091] Table 1
[0092]
[0093] Statistical analysis is performed on each sample log in Table 1 to determine the total number of occurrences of each feature word in each log classification. The generated word frequency matrix after statistics can be as shown in Table 2. For example, the number of occurrences of async in http error is 5 times, the number of occurrences of async in db error is 0 times, and the number of occurrences of async in redis error is 1 time. It can be observed that if the number of occurrences of a feature word in a certain log classification is very high, its correlation with this classification is generally also very high.
[0094] Table 2 Word Frequency Matrix
[0095] async at connection db error jdbc mysql redis timeout http error 5 2 10 0 12 0 0 0 8 db error 0 5 8 10 15 22 22 0 12 redis error 1 12 4 0 8 0 0 20 5
[0096] After determining the word frequency model, the frequency modulation model can be determined according to the word frequency model. The frequency modulation model includes the adjustment parameters of each feature word under each log classification, and the adjustment parameters are used to adjust the number of times of the corresponding feature word under the corresponding log classification. The frequency modulation model can be represented in the form of a frequency modulation matrix, or in the form of a frequency modulation array or other forms.
[0097] The frequency modulation matrix is an adjustment of the word frequency matrix, and its number of rows and columns is the same as that of the word frequency matrix. The frequency modulation matrix is used to improve the naive Bayes classification algorithm. As shown in Table 3, the frequency modulation matrix includes the adjustment parameters of each feature word under each log classification, and the adjustment parameters are used to adjust the number of times (word frequency) of the feature word under the corresponding log classification according to artificial rules. For example, features such as jdbc and mysql will appear in the log information of db error in most cases. Generally, if such a type of feature word appears, it can be determined that this log information belongs to the db error classification. Therefore, we can increase the word frequency of the feature by configuring a preset value, thereby increasing the weight of the feature under the db error classification, and making the log information containing these features more likely to be classified into db error. On the contrary, we can also reduce the word frequency of the feature by configuring a preset value. For example, by configuring an adjustment parameter less than 1, the number of times the feature word appears under a certain classification can be reduced, thereby reducing the weight of the feature word under this classification.
[0098] The frequency modulation matrix is a matrix representation of artificial rules, and the initial parameter of each item is 1, that is, no adjustment is made by default. We can adjust the adjustment parameters of each item in the frequency modulation matrix to precisely control the weight of each feature word under a specific classification, combine the existing knowledge rules with the naive Bayes classification algorithm, and thereby improve the classification accuracy of the model.
[0099] Table 3 Frequency Modulation Matrix
[0100] async at connection db error jdbc mysql redis timeout http error 1 1 1 1 1 1 1 1 1 db error 0.2 1 1 20 1 20 20 1 1 redis error 1 1 1 1 1 1 1 1 1
[0101] After determining the word frequency matrix and the frequency modulation matrix, the conditional probability of each feature word under each log classification can be determined according to the word frequency matrix and the frequency modulation matrix. For the convenience of description, any one of the feature words under any one of the log classifications can be used as an example for illustration, such as Figure 3 shown in the flowchart:
[0102] Step 301, determine the sum of the number of times each feature word appears under the log classification;
[0103] As shown in formula (3):
[0104]
[0105] where, Ti is for log classification; count(T i ) is the sum of the occurrences of each feature word under T i ; A(i, j) is the number of occurrences of the keyword x i under T j , that is, the word frequency.
[0106] Taking Table 2 as an example, if T i is http, then the sum of the occurrences of each feature word under T i is: count(http) = 5 + 2 + 10 + 0 + 12 + 0 + 0 + 0 + 8 = 37. Similarly, when T i is db, the sum of the occurrences of each feature word count(db) is 94; when T i is redis, the sum of the occurrences of each feature word count(redis) is 50.
[0107] Step 302: Determine the conditional probability of the feature word under the log classification according to the number of times of the feature word in the word frequency model, the adjustment parameter of the feature word in the frequency modulation model, and the sum of the occurrences of each feature word under the log classification.
[0108] In one implementation, the word frequency model is a word frequency matrix of m rows × n columns, and the frequency modulation model is a frequency modulation matrix of m rows × n columns. The log classification corresponding to the i-th row in the word frequency matrix is the same as the log classification corresponding to the i-th row in the frequency modulation matrix, and the feature word corresponding to the j-th column in the word frequency matrix is the same as the feature word corresponding to the j-th column in the frequency modulation matrix; 0 < i ≤ m, 0 < j ≤ n. When determining the conditional probability of the feature word under the log classification according to the number of times of the feature word in the word frequency model, the adjustment parameter of the feature word in the frequency modulation model, and the sum of the occurrences of each feature word under the log classification, it can be determined according to formula (1).
[0109] Among them, formula (1) is:
[0110]
[0111] Among them, x j is the feature word of the j-th column;
[0112] T i is the log classification of the i-th row;
[0113] P(x j |T i ) is the conditional probability of x i under T j ;
[0114] A(i, j) is the number of occurrences of the feature word corresponding to the j-th column in the log classification corresponding to the i-th row;
[0115] B(i, j) is the adjustment parameter of the feature word corresponding to the j-th column in the log classification corresponding to the i-th row;
[0116] count(T i ) is the sum of the occurrences of each feature word under T i ;
[0117] α is the smoothing coefficient, which adds a small word frequency value to all feature words to reduce the negative impact on classification calculation caused by the conditional probability being 0 when the word frequency is 0.
[0118] n is the number of columns of the frequency modulation matrix or the frequency modulation matrix.
[0119] Taking the word frequency matrix in Table 2 and the frequency modulation matrix in Table 3 as an example, according to formula (1), the conditional probability of each feature word in each log classification can be determined as shown in Table 4. Here, it can be assumed that α = 1.
[0120] Table 4 Conditional Probability Matrix
[0121] async at connection db error jdbc mysql redis timeout http error 0.13 0.07 0.24 0.02 0.28 0.02 0.02 0.02 0.20 db error 0.01 0.06 0.09 1.95 0.16 4.28 4.28 0.01 0.13 redis error 0.03 0.22 0.08 0.02 0.15 0.02 0.02 0.36 0.10
[0122] In one implementation, the conditional probability matrix composed of the conditional probabilities of each feature word in each log classification can be used as the classification model. At this time, the classification model can be as shown in Table 4. In another implementation, considering that the values of conditional probabilities are generally very small, often in the order of 10e-3, the conditional probabilities are normalized to obtain a new matrix, which is used to better reflect the influence degree of each feature word under different classifications. We call it the weight of the feature word. The higher the weight of a feature word under a certain classification, the higher the probability that the sample log carrying this feature word will be classified into this category. After determining the conditional probability matrix, the feature weight matrix can be extracted, specifically as shown in Figure 4 the flowchart shown.
[0123] Step 401, for each feature word, determine the sum of the conditional probabilities of the feature word in each log classification; determine the ratio of the conditional probability of the feature word in each log classification to the sum of the conditional probabilities of the feature word in each log classification as the feature weight of the feature word in each log classification;
[0124] The feature weights of each feature word in each log classification can be determined according to formula (4), where formula (4) can be:
[0125]
[0126] Among them, W(i, j) is the feature weight of x j under T i ; m is the number of rows of the word frequency matrix or the frequency modulation matrix.
[0127] Step 402: Compose a feature weight matrix with the feature weights of each feature word in each log category, and use the feature weight matrix as the classification model.
[0128] Combined with the conditional probability matrix in Table 4, the feature weight matrix is determined as shown in Table 5.
[0129] Table 5 Feature Weight Matrix
[0130] async at connection db error jdbc mysql redis timeout http error 0.75 0.19 0.58 0.01 0.48 0.01 0.01 0.06 0.46 db error 0.06 0.17 0.21 0.98 0.26 0.99 0.99 0.03 0.30 redis error 0.19 0.64 0.21 0.01 0.26 0.00 0.00 0.92 0.24
[0131] After the model training is completed, the classification prediction of the logs to be classified can be started. Consistent with the model training process, the logs to be classified also need to be vectorized before classification. However, when vectorizing the logs to be classified, the vocabulary generated during the vectorization of the sample set must be used. After the vectorization is completed, the calculation process of the classification probability is no different from the Naive Bayes classification process. The Bayes formula is directly used to calculate the probability of the logs to be classified under each log category, and the log category with the maximum probability is taken as the final classification result, which will not be elaborated here.
[0132] To better explain the embodiments of the present invention, another log classification process is provided below, as Figure 5 shown, specifically as follows:
[0133] The left - hand part of the process is the model training process. Obtain the training set, which includes each sample log, vectorize the text of each sample log, determine the word frequency of each feature word under each log category, and calculate the conditional probability of each feature word under each log category, and then generate the classification model.
[0134] The right - hand part of the process is the model usage process. Obtain the logs to be classified, vectorize the text of the logs to be classified, combine the classification model and use the Bayes formula to calculate the probability of the logs to be classified under each log category, and then determine the log category corresponding to the maximum probability as the log category to which the logs to be classified belong.
[0135] The frequency - modulation matrix adopted in the embodiments of the present invention has the following beneficial effects:
[0136] (1) Reduce the influence brought by sample imbalance through the frequency - modulation matrix.
[0137] Sample imbalance is a common problem in the field of machine learning. Taking classification as an example, ideally, the number of samples of different classes in the sample set should be evenly distributed, that is, ensuring that each class has enough samples for model training. However, under real-world conditions, the imbalance of sample distribution is widespread. In the field of log classification, the frequencies of logs at different levels and of different types are often different. For example, http connect time out is a common network request exception with a high occurrence probability and may occur every day; while the OOM (out of memory) of the JVM (Java Virtual Machine) is a very rare but extremely serious error. In the sample set, it is obvious that the number of abnormal http samples is much larger than that of JVM abnormal samples, which causes the problem of sample imbalance and further affects the classification accuracy of JVM abnormal samples.
[0138] In this method, we can set a very high adjustment parameter for those classes with too few samples through the frequency modulation matrix, so as to amplify the word frequency of feature words under this log classification, simulating the effect of adding samples of this class to the sample set, and then reducing the impact brought by sample imbalance. Taking JVM exceptions as an example, we can adjust the adjustment parameters corresponding to the most significant feature words in JVM exception samples, such as features like "out of memory".
[0139] (2) Perform rapid sample annotation through the frequency modulation matrix to save labor costs.
[0140] Sample annotation is a very headache-inducing thing. To train a high-quality model, the size of the sample set is a very crucial decisive factor. In the past, sample annotation had to be carried out manually one by one. For hundreds or thousands of samples, it would consume a significant amount of labor costs.
[0141] In the embodiment of the present invention, we can use the determined frequency modulation matrix to initialize the annotation of the sample set that needs to be classified later, effectively reducing the workload of manual annotation. In the error sample log, most of the samples have features that can significantly distinguish the categories, such as "mysql", "redis", "gns", "http", "timeout", "out ofmemory", etc. Basically, as long as these keywords appear, it can be determined that the sample log belongs to a certain category. We call this type of feature initialization feature. After collecting enough initialization features, we use the frequency modulation matrix to set a very large adjustment parameter (for example, more than 1000) for this type of feature, and then classify the sample set to be classified, and use the result as the classification label; most samples can correctly fall into the corresponding category, and a small number of samples that do not contain the initialization feature fall into the default unknown category, and then manually mark them.
[0142] (3) Perform regression analysis on misclassified samples and adjust features based on the frequency modulation matrix.
[0143] Model classification may be wrong. Under the naive Bayes classification algorithm based on the word frequency model, a problem will arise: we find that a sample is classified into the wrong category, then we manually correct this sample, put it into the sample set, retrain the model, and classify the sample again - the result model still gives the previous wrong classification. This is because the word frequency model counts the word frequency of all samples under the same category, and adjusting a single sample is just a drop in the bucket and cannot achieve the purpose of correcting the model.
[0144] In an embodiment of the present invention, we can use the feature weight matrix obtained by model training to perform regression analysis on the sample. For example, there is a processed sample log "Bank report sys TransDAO certNo query timeout costTime", which should belong to the category of "external partner business abnormality", but the model classifies it into the default "unknown" category. We query through the feature weight matrix that the five features with the highest weights in the "unknown" category of this sample are as follows:
[0145] ("Bank Report", 0.8652419428703651)
[0146] ("Query timed out", 0.5142907974010534)
[0147] ("sys",0.5142907974010534)
[0148] ("timeout", 0.15651730037704084)
[0149] ("costtime", 0.1881949392920741)
[0150] From this, we found that the feature of "bank report" has the highest weight under the "unknown" classification. In fact, "bank" obviously belongs to the category of features of external partners. Therefore, we need to adjust the weights of this feature in these two categories. We can combine the frequency modulation matrix, lower the adjustment parameter of this feature word "bank report" under the "unknown" classification, and raise the adjustment parameter under the "classification of abnormal business of external partners". After the adjustment, retrain the model and conduct the classification test again. As a result, this sample has been successfully classified under the "classification of abnormal business of external partners".
[0151] Based on the same inventive concept, Figure 6 Exemplarily, the structure of a log classification device provided by an embodiment of the present invention is shown. This device can execute the process of the log classification method.
[0152] This device includes:
[0153] A determination unit 601, a classification unit 602, and a training unit 603;
[0154] The determination unit 601 is used to determine the number of occurrences of each feature word in the log to be classified;
[0155] The classification unit 602 is used to determine the log classification to which the log to be classified belongs according to the number of occurrences of each feature word in the log to be classified and the classification model; the classification model is determined by the training unit 603 according to the conditional probability of each feature word in each log classification in the sample log;
[0156] Among them, the conditional probability of each feature word in each log classification is determined by the training unit 603 according to the word frequency model and the frequency modulation model; the word frequency model includes the number of occurrences of each feature word in each log classification, and the frequency modulation model includes the adjustment parameter of each feature word in each log classification. The adjustment parameter is used by the training unit 603 to adjust the number of times of the corresponding feature word in the corresponding log classification.
[0157] Optionally, the training unit 603 is specifically used for:
[0158] Perform the following operations for each feature word in each log classification:
[0159] Determine the sum of the number of occurrences of each feature word in the log classification;
[0160] Determine the conditional probability of the feature word under the log classification according to the number of times of the feature word in the word frequency model, the adjustment parameter of the feature word in the frequency modulation model, and the sum of the number of times each feature word appears under the log classification.
[0161] Optionally, the word frequency model is a word frequency matrix of m rows × n columns, and the frequency modulation model is a frequency modulation matrix of m rows × n columns; the log classification corresponding to the i-th row in the word frequency matrix is the same as the log classification corresponding to the i-th row in the frequency modulation matrix, and the feature word corresponding to the j-th column in the word frequency matrix is the same as the feature word corresponding to the j-th column in the frequency modulation matrix; 0 < i ≤ m, 0 < j ≤ n;
[0162] The training unit 603 is specifically configured to:
[0163] Determine the conditional probability of the feature word under the log classification according to formula (1);
[0164] The formula (1) is:
[0165]
[0166] where x j is the feature word in the j-th column; T i is the log classification in the i-th row; P(x j |T i ) is the conditional probability of x i under T j ; A(i, j) is the number of times the feature word corresponding to the j-th column appears in the log classification corresponding to the i-th row; B(i, j) is the adjustment parameter of the feature word corresponding to the j-th column in the log classification corresponding to the i-th row; count(T i ) is the sum of the number of times each feature word appears under T i ; α is a smoothing coefficient; n is the number of columns of the frequency modulation matrix or the word frequency matrix;
[0167] Optionally, the training unit 603 is specifically configured to:
[0168] For each feature word, determine the sum of the conditional probabilities of the feature word under each log classification; determine the ratio of the conditional probability of the feature word under each log classification to the sum of the conditional probabilities of the feature word under each log classification as the feature weight of the feature word under each log classification;
[0169] Form a feature weight matrix with the feature weights of each feature word in each log classification, and use the feature weight matrix as the classification model.
[0170] Based on the same inventive concept, an embodiment of the present invention further provides a computing device, including:
[0171] A memory for storing program instructions;
[0172] A processor for calling the program instructions stored in the memory and executing the above-mentioned log classification method according to the obtained program.
[0173] Based on the same inventive concept, an embodiment of the present invention also provides a computer-readable non-volatile storage medium, including computer-readable instructions, which, when read and executed by a computer, cause the computer to execute the above-mentioned log classification method.
[0174] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0175] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that realizes the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0177] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.
[0178] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications.
Claims
1. A log classification method, characterized in that, Including: Determine the number of occurrences of each feature word in the log to be classified; Determine the log classification to which the log to be classified belongs according to the number of occurrences of each feature word in the log to be classified and the classification model; The classification model is determined according to the conditional probability of each feature word in the sample log under each log classification; Wherein, for each feature word under each log classification, the following operations are performed: Determine the total number of occurrences of each feature word under the log classification; Determine the conditional probability of the feature word under the log classification according to the number of occurrences of the feature word in the word frequency model, the adjustment parameter of the feature word in the frequency modulation model, and the total number of occurrences of each feature word under the log classification; The word frequency model includes the number of occurrences of each feature word under each log classification, which is a word frequency matrix of m rows × n columns, and the frequency modulation model includes the adjustment parameter of each feature word under each log classification, which is a frequency modulation matrix of m rows × n columns; the log classification corresponding to the i-th row in the word frequency matrix is the same as the log classification corresponding to the i-th row in the frequency modulation matrix, and the feature word corresponding to the j-th column in the word frequency matrix is the same as the feature word corresponding to the j-th column in the frequency modulation matrix; 0 < i ≤ m, 0 < j ≤ n; the adjustment parameter is used to adjust the number of occurrences of the corresponding feature word under the corresponding log classification; The determining the conditional probability of the feature word under the log classification according to the number of occurrences of the feature word in the word frequency model, the adjustment parameter of the feature word in the frequency modulation model, and the total number of occurrences of each feature word under the log classification includes: Determine the conditional probability of the feature word under the log classification according to formula (1); The formula (1) is: where x j is the feature word in the j-th column; T i is the log classification in the i-th row; P(x j |T i ) is the conditional probability of x i under T j ; A(i, j) is the number of times the feature word corresponding to the j-th column appears in the log classification corresponding to the i-th row; B(i, j) is the adjustment parameter of the feature word corresponding to the j-th column in the log classification corresponding to the i-th row; count(T i ) is the sum of the number of times each feature word appears under T i ; α is the smoothing coefficient; n is the number of columns of the smoothing matrix or the smoothing matrix.
2. The method according to claim 1, characterized in that, The classification model is determined according to the conditional probability of each feature word in the sample log under each log classification, including: For each feature word, determine the sum of the conditional probabilities of the feature word under each log classification; determine the ratio of the conditional probability of the feature word under each log classification to the sum of the conditional probabilities of the feature word under each log classification as the feature weight of the feature word under each log classification; Form a feature weight matrix with the feature weights of each feature word in each log classification, and use the feature weight matrix as the classification model.
3. A log classification device, characterized in that, Including: A determination unit, a classification unit and a training unit; The determination unit is used to determine the number of occurrences of each feature word in the log to be classified; The classification unit is used to determine the log classification to which the log to be classified belongs according to the number of occurrences of each feature word in the log to be classified and the classification model; The classification model is determined by the training unit according to the conditional probability of each feature word in the sample log under each log classification; Among them, the conditional probability of each feature word under each log classification is determined by the training unit according to the word frequency model and the frequency modulation model; the word frequency model includes the number of occurrences of each feature word under each log classification, which is a word frequency matrix of m rows × n columns, and the frequency modulation model includes the adjustment parameters of each feature word under each log classification, which is a frequency modulation matrix of m rows × n columns; the log classification corresponding to the i-th row in the word frequency matrix is the same as the log classification corresponding to the i-th row in the frequency modulation matrix, and the feature word corresponding to the j-th column in the word frequency matrix is the same as the feature word corresponding to the j-th column in the frequency modulation matrix; 0 < i ≤ m, 0 < j ≤ n; the adjustment parameter is used for the training unit to adjust the number of times of the corresponding feature word under the corresponding log classification; Specifically, the training unit is configured to perform the following operations for each feature word under each log classification: determine the sum of the number of occurrences of each feature word under the log classification; determine the conditional probability of the feature word under the log classification according to the number of times of the feature word in the word frequency model, the adjustment parameter of the feature word in the frequency modulation model, and the sum of the number of occurrences of each feature word under the log classification; Specifically, the training unit is configured to determine the conditional probability of the feature word under the log classification according to formula (1); The formula (1) is: where x j is the feature word of the j-th column; T i is the log classification of the i-th row; P(x j |T i ) is the conditional probability of x i under T j ; A(i, j) is the number of times the feature word corresponding to the j-th column appears in the log classification corresponding to the i-th row; B(i, j) is the adjustment parameter of the feature word corresponding to the j-th column in the log classification corresponding to the i-th row; count(T i ) is the sum of the number of times each feature word appears under T i ; α is the smoothing coefficient; n is the number of columns of the smoothing matrix or the smoothing matrix.
4. The device according to claim 3, characterized in that, Specifically, the training unit is configured to: For each feature word, determine the sum of the conditional probabilities of the feature word under each log classification; determine the ratio of the conditional probability of the feature word under each log classification to the sum of the conditional probabilities of the feature word under each log classification as the feature weight of the feature word under each log classification; Form a feature weight matrix with the feature weights of each feature word in each log classification, and use the feature weight matrix as the classification model.
5. A computing device, characterized in that, Comprising: A memory for storing program instructions; A processor for calling the program instructions stored in the memory and executing the method according to claim 1 or 2 according to the obtained program.
6. A computer-readable non-volatile storage medium, characterized in that, Including computer-readable instructions, when a computer reads and executes the computer-readable instructions, the computer is caused to execute the method according to claim 1 or 2.
Citation Information
Patent Citations
Log based computer system fault diagnosis method and device
CN103761173A
Big data based work order type identification method and system and computing device
CN108897754A