Log processing method and device, electronic equipment and storage medium

By constructing graph data and using graph convolution model, the difficulty of quickly positioning problems in massive log data is solved, efficient classification and analysis of abnormal log data is achieved, and log processing efficiency is improved.

CN120216468APending Publication Date: 2025-06-27中国邮政储蓄银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510255382.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When processing massive log data, it is difficult for the existing technology to quickly locate the real cause of the problem, resulting in wasting time and energy for development and operation and maintenance personnel.

Method used

By constructing graph data, including subgraphs of words and words, subgraphs of tags and tags, subgraphs of words and tags, subgraphs of words and tags, and subgraphs of words and tags, the graph convolution model is used to classify exception information with associations in the exception log data.

Benefits of technology

It realizes efficient classification of texts with related relationships, improves log analysis efficiency, and reduces the working time and energy of development and operation and maintenance personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216468A_ABST
    Figure CN120216468A_ABST
Patent Text Reader

Abstract

The invention discloses a log processing method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining abnormal log data; according to the abnormal log data, graph data is constructed, and the graph data at least comprises one of words and a first sub-graph of the words, tags and a second sub-graph of the tags, and words and a third sub-graph of the tags; and taking the word and the first sub-graph of the word, the tag and the second sub-graph of the tag, and the word and the third sub-graph of the tag as an adjacent matrix of the three sub-graphs as input of a graph convolution model, and participating in training to obtain a classifier so as to classify abnormal information with an association relationship in the abnormal log data. According to the method, on one hand, classification and marking of the texts are achieved, on the other hand, a new graph construction method is provided, and text and graph construction is decoupled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of text classification, and particularly to a log processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of network technology, text information has seen an explosive growth. How to process massive text resources and efficiently classify and process these text resources is one of the key research areas in the field of natural language processing. At the same time, multi-label text classification technology emerged to solve the problem that a single text is labeled with multiple different labels in text classification, and it has a very wide range of application fields.

[0003] Taking the cluster logs of a data platform as an example, the construction of the data platform involves website development and maintenance, distributed cluster transformation and maintenance, and the maintenance and optimization of big data computing and storage services. The maintenance work of these contents requires developers and operators to analyze the logs generated during the operation of the program, cluster, and website, so as to accurately locate the problems.

[0004] However, in these application scenarios, the output logs often have problems such as a large quantity, redundancy, and associated errors. In most cases, developers and operators rely on the logs to search for keywords and timestamps and then analyze them line by line, and they need to search through multiple different logs to analyze the cause of the error. It is very difficult to quickly locate the real cause of the problem, which wastes the time and energy of developers and operators and is inefficient. Summary of the Invention

[0005] Embodiments of this application provide a log processing method, apparatus, electronic device, and storage medium to achieve text classification and tagging.

[0006] Embodiments of this application adopt the following technical solutions:

[0007] In a first aspect, embodiments of this application provide a log processing method, where the processing method includes:

[0008] Obtain abnormal log data;

[0009] Construct graph data according to the abnormal log data, where the graph data includes at least one of the following: a first subgraph of words and words, a second subgraph of labels and labels, and a third subgraph of words and labels;

[0010] Take the adjacency matrices of the three subgraphs obtained from the first subgraph of words and words, the second subgraph of labels and labels, and the third subgraph of words and labels as the input of a graph convolutional model, and participate in training to obtain a classifier to classify the abnormal information with an associated relationship in the abnormal log data.

[0011] In some embodiments, obtaining the adjacency matrices of the three subgraphs of the word and the first subgraph of the word, the label and the second subgraph of the label, and the word and the third subgraph of the label as the input of the graph convolutional model, and participating in training to obtain a classifier for classifying the abnormal information with an association relationship in the abnormal log data, includes:

[0012] Performing convolution according to the word and the first subgraph of the word to obtain an embedding matrix of the word;

[0013] Performing convolution according to the label and the second subgraph of the label to obtain an embedding matrix of the label;

[0014] Taking the embedding matrix of the word and the embedding matrix of the label as the input of the third subgraph of the word and the label, and then obtaining the embedding matrix of the word and the label through a graph convolutional network.

[0015] In some embodiments, the embedding matrix of the word and the label includes:

[0016] According to the symmetric normalized adjacency matrix and the label feature matrix of the third subgraph of the word and the label, obtaining the output result of the single-layer graph convolutional model, that is, the finally trained embedding matrix of the word and the label.

[0017] In some embodiments, the method further includes:

[0018] Splitting the finally trained embedding matrix of the word and the label into two matrices respectively representing the word embedding matrix and the label embedding matrix;

[0019] Multiplying the word embedding matrix by the 0, 1 vector matrix of the text to obtain a text embedding matrix;

[0020] According to the matrix multiplication of the text embedding matrix and the label embedding matrix, obtaining the similarity between the text and the label, and the probability prediction obtained by the text passing through a linear network;

[0021] Weighting the similarity between the text and the label and the probability prediction obtained by the text passing through a linear network, and mapping the one-dimensional features in each text into the final classification result.

[0022] In some embodiments, constructing graph data according to the abnormal log data, where the graph data includes at least one of the following: the first subgraph of the word and the word, the second subgraph of the label and the label, the third subgraph of the word and the label, includes:

[0023] The first subgraph of the word and the word includes a graph composed of only word nodes and edges connecting between words;

[0024] Assume any two different words i and j in the graph. If i and j appear in the same exception log, then there is an edge between word i and word j. If i and j do not appear together in an error, then there is no edge between the words, and the adjacency matrix A of the first subgraph of words and words is obtained. ww Definition

[0025]

[0026] Among them, PMI represents point mutual information, which is used to measure the correlation between words.

[0027] When i and j are any two different words, the PMI metric is used to measure the correlation between word i and j. When i = j, 1 is used to represent the weight of the node itself. In other cases, the value of the adjacency matrix element is represented by 0. The definition of PMI point mutual information is as follows:

[0028]

[0029] In some embodiments, constructing graph data according to the exception log data, the graph data includes at least one of the following: the first subgraph of words and words, the second subgraph of labels and labels, the third subgraph of words and labels, including:

[0030] The second subgraph of labels and labels includes a graph composed of only label nodes and edges connecting labels to labels.

[0031] Assume any two different labels i and j in the graph. If i and j label the same set of error reports, then there is an edge between label i and label j. If i and j do not belong to the same associated error report, then there is no edge between the nodes, and the adjacency matrix A of the second subgraph of labels and labels is obtained. ll Definition

[0032]

[0033] PMI is used to measure the similarity between two labels. If two labels have appeared together, there must be an edge. When i and j are any two different labels, the correlation is represented by PMI. When i = j, 1 is used to represent the weight of the node itself. In other cases, the value of the adjacency matrix element is represented by 0.

[0034] In some embodiments, constructing graph data according to the exception log data, the graph data includes at least one of the following: the first subgraph of words and words, the second subgraph of labels and labels, the third subgraph of words and labels, including:

[0035] The third subgraph of words and labels includes a subgraph with both word nodes and label nodes.

[0036] Assume that for any two nodes i and j, if i is a word, j is a label node, and the word i and the label j appear in the same log, then there is an edge between the word i and the label j; if i and j are two nodes of the same attribute, both being word nodes or both being label nodes, then there is no edge between the nodes, and the adjacency matrix A of the third subgraph of words and labels is obtained. wl Definition

[0037]

[0038] Among them, TF-IDF is the term frequency-inverse document frequency.

[0039] In a second aspect, an embodiment of the present application further provides a log processing device, where the processing device includes:

[0040] An acquisition module, configured to acquire abnormal log data;

[0041] A graph construction module, configured to construct graph data according to the abnormal log data, where the graph data includes at least one of the following: a first subgraph of words and words, a second subgraph of labels and labels, and a third subgraph of words and labels;

[0042] A classification module, configured to use the first subgraph of words and words, the second subgraph of labels and labels, and the third subgraph of words and labels to obtain the adjacency matrices of the three subgraphs as the input of a graph convolutional model, and participate in training to obtain a classifier to classify the abnormal information with an association relationship in the abnormal log data.

[0043] In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor; and a memory arranged to store computer-executable instructions, where the executable instructions, when executed, cause the processor to execute the above method.

[0044] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device is caused to execute the above method.

[0045] The above at least one technical solution adopted in the embodiments of the present application can achieve the following beneficial effects: Obtain abnormal log data, and construct graph data according to the abnormal log data. Since the graph data includes at least one of the following: words and the first subgraph of words, labels and the second subgraph of labels, words and the third subgraph of labels, it can be used for subsequent model training. Take the first subgraph of words and words, the second subgraph of labels and labels, and the third subgraph of words and labels to obtain the adjacency matrices of the three subgraphs as the input of the graph convolution model, and participate in training to obtain a classifier to classify the abnormal information with an associated relationship in the abnormal log data. Through the above method, the classification of text with an associated relationship can be realized. Description of the Drawings

[0046] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:

[0047] Figure 1 It is a schematic flow chart of the log processing method in the embodiments of the present application;

[0048] Figure 2 It is a schematic principle diagram of the log processing method in the embodiments of the present application;

[0049] Figure 3 It is a schematic structural diagram of the log processing device in the embodiments of the present application;

[0050] Figure 4 It is a schematic structural diagram of an electronic device in the embodiments of the present application. Detailed Embodiments

[0051] To make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the protection scope of the present application.

[0052] For existing algorithms for text multi-label classification, some algorithms simply use statistical features to select empirical features or the model algorithms used do not consider the relevance and similarity between labels and between texts. Or some use graph convolutional neural networks, so they have their own defects. For example, the newly obtained text to be classified needs to be used as part of the graph to construct and retrain, which is time-consuming and laborious.

[0053] In view of the above problems, the embodiments of the present application provide a log processing method, which utilizes the relevance between tags and texts and decouples the construction of texts and graphs, so as to extract the association relationships of errors contained in different error log files, and better process the situation where large-scale applications or clusters contain many log files.

[0054] The following will describe in detail the technical solutions provided by the embodiments of the present application with reference to the accompanying drawings.

[0055] The embodiments of the present application provide a log processing method, as Figure 1 shown in the schematic flowchart of the log processing method in the embodiments of the present application. The method at least includes the following steps S110 to S130:

[0056] Step S110, obtain abnormal log data.

[0057] In the text preprocessing stage, the word segmentation result in the error log data can be obtained by obtaining the log data and performing relevant preprocessing.

[0058] Taking the database log file as an example,

[0059] java.io.FileNotFoundException: / Hadoop / hdfs / journal / test / last-proised-epoch.tmp(No such file or directory)

[0060] After processing, java io FileNotFoundException Hadoop hdfs journaltest last-proised-epoch tmp No such file or directory is obtained.

[0061] In this way, relying on the file path and file name, the component information and the log information from different log sources can be added to each error log set text as a training sample.

[0062] It should be noted that in the embodiments of the present application, all the log texts associated with a single error are regarded as an error log set, and an error log set is a piece of training data. The purpose of such processing is that all associated different log file names are included in the set, and the connections between the associated error logs and error logs can also be included in the training text, so that they can be learned as a kind of relevance.

[0063] Step S120: Construct graph data based on the abnormal log data. The graph data includes at least one of the following: the first sub-graph of words and words, the second sub-graph of labels and labels, and the third sub-graph of words and labels.

[0064] Construct graph data as the input of the model, and disassemble the original heterogeneous graph of words-labels into two homogeneous sub-graphs and one heterogeneous sub-graph, namely, the word-word sub-graph, the label-label sub-graph, and the word-label sub-graph.

[0065] It can be understood that the first sub-graph of words and words is a graph composed only of word nodes and edges connecting words to words, representing the relationship between words. The second sub-graph of labels and labels is a graph composed only of label nodes and edges connecting labels to labels, representing the relationship between labels. The third sub-graph of words and labels is a sub-graph with both word nodes and label nodes, representing the relationship between words and labels.

[0066] Step S130: Take the first sub-graph of words and words, the second sub-graph of labels and labels, and the third sub-graph of words and labels to obtain the adjacency matrices of the three sub-graphs as the input of the graph convolution model, and participate in the training to obtain a classifier to classify the abnormal information with associated relationships in the abnormal log data.

[0067] Input the constructed graph data into the heterogeneous lightweight graph convolution model to obtain the trained word vectors and classifier. Input the two-layer lightweight graph convolution model to train and obtain the embedding matrix, and obtain the embedding matrix of the text to be classified. According to the embedding matrix of the text to be classified, obtain the multi-label annotation result of the predicted text, that is, classify the abnormal information with associated relationships in the abnormal log data.

[0068] Through the above method, obtain abnormal log data; construct graph data according to the abnormal log data; take the first sub-graph of words and words, the second sub-graph of labels and labels, and the third sub-graph of words and labels to obtain the adjacency matrices of the three sub-graphs as the input of the graph convolution model, and participate in the training to obtain a classifier to classify the abnormal information with associated relationships in the abnormal log data. Different from other neural networks, using graph neural networks can consider the similarities between labels and between texts in the case of multiple labels and multiple text files through graph construction, providing guidance for multi-label text classification.

[0069] Through the above method, automatically perform feature screening through graph neural networks. Not only use statistical information as part of the features, but also optimize the feature extraction ability. Utilize the characteristic that graph neural networks can better extract the associated relationships between adjacent nodes, and better classify and associate similar log texts together.

[0070] Different from the GCN (Graph Convolutional Network) used in the related art, a neural network model for processing graph-structured data, which requires training the new text as part of the graph every time the text is classified to correctly classify the new text. Through the method proposed in this application, the text and the graph are decoupled. Only the words and labels are used to construct the graph, and the embedding vectors of the words and labels can be obtained through training. The text vector can be obtained only by multiplying the text [0, 1] vector matrix by the word embedding matrix. In this way, when there is a new text that needs to be multi-labeled, there is no need to reconstruct the graph with the new text and put it into training again.

[0071] In an embodiment of the present application, obtaining the adjacency matrices of the three subgraphs of the word and the first subgraph of the word, the label and the second subgraph of the label, and the word and the third subgraph of the label as the input of the graph convolution model, and participating in training to obtain a classifier to classify the abnormal information with an association relationship in the abnormal log data, including: performing convolution according to the first subgraph of the word and the word to obtain the embedding matrix of the word; performing convolution according to the second subgraph of the label and the label to obtain the embedding matrix of the label; using the embedding matrix of the word and the embedding matrix of the label as the input of the third subgraph of the word and the label, and then obtaining the embedding matrix of the word and the label through the graph convolution network.

[0072] It can be understood that the graph convolution model can be a graph convolution model well-known to those skilled in the art and in the scenario, and is not specifically limited in the embodiments of the present application. The graph construction method does not require constructing a new graph structure and retraining every time there is a new text, and takes into account the problem of the correlation between labels.

[0073] The process of obtaining the adjacency matrices of the three subgraphs of the word and the first subgraph of the word, the label and the second subgraph of the label, and the word and the third subgraph of the label as the input of the graph convolution model and participating in training to obtain a classifier is a process of performing convolution separately.

[0074] As Figure 2 shown, first, perform graph convolution on the constructed word-word graph to obtain the embedding matrix of the word, and the calculation formula is as shown in Equation (8):

[0075]

[0076] Where represents the symmetric normalized adjacency matrix of the word-word graph; Z ww ∈R m*(m+n) represents the output result of the single-layer graph convolution model, that is, the word embedding matrix; the input word feature matrix is set as X ww =[I, 0] ∈ Rm*(m+n) , where \(I\in R\) m*m . \(m\) represents the number of words, and \(n\) represents the number of texts.

[0077] As Figure 2 shown, then, perform graph convolution on the constructed label-label graph to obtain the embedding matrix of labels. The calculation formula is as shown in Equation (9):

[0078]

[0079] where represents the symmetric normalized adjacency matrix of the label-label subgraph; \(Z\) LL \(\in R\) n*(m+n) represents the output result of the single-layer graph convolution model, that is, the label embedding matrix; the input label feature matrix is set as \(X\) LL \(=[0, I]\in R\) n*(m+n) , where \(I\in R\) n*n .

[0080] As Figure 2 shown, it can be seen from the above that through the two models of Equations (8) and (9), the word embedding matrix \(Z\) WW \(\in R\) m*(m+n) of the first-order information and the label embedding matrix \(Z\) LL \(\in R\) n*(m+n) of the first-order information are obtained respectively. Combine the two in the first dimension to obtain the input feature matrix \(X\) WL \(=[Z\) WW ; \(Z\) LL \(\in R\) (m+n)*(m+n) .

[0081] Please continue to refer to Figure 2 . From these two parts of the graph convolution model, combined with their input feature matrices, it can be known that the information contained in the trained \(X\) WL is completely trained based on the characteristics of the graph structure, because their initializations are one-hot vectors that do not contain any information and are only used to distinguish each other, and mainly contain the structural information of the first-hop neighbor nodes. It can be understood that a one-hot vector means using \(N\) bits of 0 or 1 to encode \(N\) states, each state has its independent representation form, and only one bit is 1, and the other bits are 0.

[0082] As Figure 2 shown, the word-label subgraph needs to use the outputs obtained from the first two submodels as inputs, and then pass through a lightweight graph convolution network to finally obtain the final embedding representations of words and labels. The formula of the second-layer lightweight graph convolution network is as shown in Equation (10):

[0083]

[0084] Among them represents the symmetric normalized adjacency matrix of the word-label subgraph; Z WL ∈R (m+n)*(m+n) represents the output result of the single-layer graph convolutional model, that is, the final embedding matrix of words and labels obtained through training, denoted by Z word and Z label respectively; the input label feature matrix is X WL ∈R (m+n)*(m+n) .

[0085] Please continue to refer to Figure 2 , where there is also only one layer of graph convolutional layer, that is to say, the embedding representation of the output label nodes in the graph will only collect information from one-hop neighbors. However, although there is only one layer here, actually combining the information obtained from the previous layer, the information collected by the nodes in the graph is the information of the second hop.

[0086] In the word nodes in the adjacency matrix of Equation (10), the embedding information of directly connected labels will be collected, and the embedding information carried by these labels is the node information of adjacent labels collected in the previous layer Equation (9). Similarly, in the second layer, the label nodes will collect the information of directly connected word nodes, and at the same time, the word nodes also collect the information of adjacent word nodes in the model of Equation (8).

[0087] Therefore, through Equation (10), the final output label and word embedding matrices collect the information of indirectly connected graph nodes, which also conforms to the original intention of the graph neural network, that is, collecting the information of indirect nodes and direct neighbor nodes and converging it to the central node.

[0088] In an embodiment of the present application, the embedding matrix of the word and the label includes: obtaining the output result of the single-layer graph convolutional model, that is, the final embedding matrix of the word and the label obtained through training, according to the symmetric normalized adjacency matrix of the third subgraph of the word and the label and the label feature matrix.

[0089] Such as the calculation formula

[0090]

[0091] Among them represents the symmetric normalized adjacency matrix of the word-label subgraph;

[0092] Among them, Z WL ∈R (m+n)*(m+n) represents the output result of the single-layer graph convolutional model, that is, the final embedding matrix of the word and the label obtained through training, denoted by Z word and Z label respectively;

[0093] Among them, the input label feature matrix is X WL ∈R (m+n)*(m+n) .

[0094] Automatically perform feature screening through the graph neural network. Not only use statistical information as part of the features to optimize the feature extraction ability, but also utilize the characteristic that the graph neural network can better extract the correlation relationships between adjacent nodes, better classify similar log texts together, and obtain similar text embedding representations by training similar texts.

[0095] In an embodiment of the present application, the method further includes: splitting the final embedding matrix of words and labels into two matrices respectively representing the word embedding matrix and the label embedding matrix; multiplying the word embedding matrix by the 0, 1 vector matrix of the text to obtain the text embedding matrix; performing matrix multiplication according to the text embedding matrix and the label embedding matrix to obtain the similarity between the text and the label, and the probability prediction obtained by the text passing through a linear network; weighting the similarity between the text and the label and the probability prediction obtained by the text passing through a linear network, and mapping the one-dimensional features in each text into the final classification result.

[0096] Split Z WL into two matrices respectively representing word embedding Z W and label embedding Z L , by multiplying the word embedding by the 0, 1 vector matrix of the text, where the text here refers to a set of error logs at a time, the text matrix ∈R t*(m+n) , where t represents the number of texts, multiplying by the word embedding matrix to obtain the text embedding matrix, denoted by Z text . Equation (11) represents that the text embedding matrix and the label embedding matrix perform matrix multiplication to obtain the similarity between the text and the label, and the probability prediction obtained by the text passing through a linear network. The two are weighted and passed through the Softmax(·) function to map the one-dimensional features of each text into the final classification.

[0097]

[0098] Among them represents the predicted label matrix for the text to be classified, and c represents the number of the final labels; represents the parameter matrix of the linear classifier.

[0099] In one embodiment of the present application, constructing graph data according to the abnormal log data, the graph data includes at least one of the following: a first sub-graph of words and words, a second sub-graph of labels and labels, a third sub-graph of words and labels, including: the first sub-graph of words and words includes a graph composed of only word nodes and edges connecting between words; assuming any two different words i, j in the graph, if i and j appear in the same abnormal log, then there is an edge between word i and word j, if i and j do not appear together in an error, then there is no edge between words, and the adjacency matrix A of the first sub-graph of words and words is obtained ww Definition

[0100]

[0101] Among them, PMI represents pointwise mutual information, which is used to measure the correlation between words; when i and j are any two different words, this index of PMI is used to measure the correlation between word i and j; when i = j, 1 is used to represent the weight of the node itself; in other cases, the value of the adjacency matrix element is represented by 0, and the definition of the PMI pointwise mutual information is as follows:

[0102]

[0103] The word-word sub-graph is a graph composed of only word nodes and edges connecting between words. Assuming any two different words i, j in the graph, if i and j appear in the same error, then there is an edge between word i and word j, if i and j do not appear together in an error, then there is no edge between words. Among them, the adjacency matrix A of the graph ww The definition is as formula (1):

[0104]

[0105] Among them, PMI (PMI, Pointwise Mutual Information) represents pointwise mutual information, and this index is used to measure the correlation between words;

[0106] When i is the set of relevant error logs of a certain error and j is a word, the term frequency-inverse document frequency index is used to measure the correlation between the text and the word; when i = j, 1 is used to represent the weight of the node itself; in other cases, the value of the adjacency matrix element is represented by 0. The definition of the pointwise mutual information is as follows:

[0107]

[0108] In one embodiment of the present application, based on the abnormal log data, graph data is constructed, and the graph data includes at least one of the following: a first sub-graph of words and words, a second sub-graph of labels and labels, and a third sub-graph of words and labels, including: the second sub-graph of labels and labels includes a graph composed of only label nodes and edges connecting between labels; assuming any two different labels i and j in the graph, if i and j are marked with the same set of error reports, then there is an edge between label i and label j, and if i and j do not belong to the same associated error report, then there is no edge between the nodes, and the adjacency matrix A of the second sub-graph of labels and labels is obtained ll Definition

[0109]

[0110] PMI is used to measure the similarity between two labels. If two labels have co-occurred, there must be an edge. When i and j are any two different labels, the correlation is represented by PMI; when i = j, 1 is used to represent the weight of the node itself; in other cases, the value of the adjacency matrix element is represented by 0

[0111] The label-label sub-graph is a graph composed of only label nodes and edges connecting between labels. Assuming any two different labels i and j in the graph, if i and j are marked with the same set of error reports, then there is an edge between label i and label j, and if i and j do not belong to the same associated error report, then there is no edge between the nodes. Finally, the adjacency matrix A of the label-label sub-graph ll The definition is as shown in formula (3):

[0112]

[0113] Among them, PMI is used to measure the similarity between two labels. If two labels have co-occurred, there must be an edge. When i and j are any two different labels, the correlation is represented by PMI; when i = j, 1 is used to represent the weight of the node itself; in other cases, the value of the adjacency matrix element is represented by 0

[0114] In one embodiment of the present application, based on the abnormal log data, graph data is constructed, and the graph data includes at least one of the following: a first sub-graph of words and words, a second sub-graph of labels and labels, and a third sub-graph of words and labels, including: the third sub-graph of words and labels includes a sub-graph with both word nodes and label nodes; assuming any two nodes i and j are words, if I is a word and j is a label node and word i and label j appear in the same log, then there is an edge between word i and label j; if i and j are two nodes with the same attribute, that is, both are word nodes or both are label nodes, then there is no edge between the nodes, and the adjacency matrix A of the third sub-graph of words and labels is obtainedwl Definition

[0115]

[0116] Among them, TF-IDF is the term frequency-inverse document frequency.

[0117] The word-label subgraph is a subgraph that has both word nodes and label nodes. Most importantly, this is a bipartite graph, that is, in the entire graph, there are no connections between word nodes and no connections between label nodes. The nodes in the graph are all connected by the edges between words and labels.

[0118] Assume any two nodes i and j. If i is a word, j is a label node, and the word i and the label j appear in the same log, that is, if one of them is a word node and the other is a label node, and the word i appears in the text of the label j, then there is an edge between the word i and the label j. If i and j are two nodes with the same attribute, that is, both are word nodes or both are label nodes, then there is no edge between the nodes. Finally, the adjacency matrix A of the word-label subgraph wl The definition is as shown in formula (4):

[0119]

[0120] In formula (4), the definition of the term frequency–inverse document frequency (TF-IDF) is as follows:

[0121] TF-IDF = TF * IDF (5)

[0122]

[0123] Among them, the term frequency TF represents the frequency of a word appearing in a certain text, and is used to measure the importance of the word in the article.

[0124] Among them, n i,j represents the number of times the word i appears in a certain set j of error logs, and ∑ k n k,j then represents the total number of words contained in a single set j of error logs.

[0125] The inverse document frequency IDF measures the degree of general importance of the word in all sets of error logs. |D| represents the total number of texts contained in the entire corpus, and |j:t i ∈d j | represents the total number of texts j in which the word i appears. Here, the text is a set of error logs.

[0126] The embodiment of the present application also provides a log processing device 300, as follows Figure 3 As shown, a schematic structural diagram of the log processing device in the embodiment of the present application is provided. The log processing device 300 at least includes: an acquisition module 310, a graph construction module 320, and a classification module 330, where:

[0127] In an embodiment of the present application, the acquisition module 310 is specifically configured to: acquire abnormal log data.

[0128] In the text preprocessing stage, by acquiring log data and performing relevant preprocessing, the word segmentation result in the error log data can be obtained.

[0129] Taking the database log file as an example,

[0130] java.io.FileNotFoundException: / Hadoop / hdfs / journal / test / last-proised-epoch.tmp(No such file or directory)

[0131] After processing, java io FileNotFoundException Hadoop hdfs journaltest last-proised-epoch tmp No such file or directory is obtained.

[0132] In this way, relying on the file path and file name, the component information and the log information from different log sources can be added to each error log set text as a training sample.

[0133] It should be noted that in the embodiment of the present application, all log texts associated with a single error are used as an error log set, and an error log set is a piece of training data. The purpose of such processing is that all associated different log file names are included in the set, and the connection between the associated error logs can also be included in the training text, so that it can be learned as a kind of relevance.

[0134] In an embodiment of the present application, the graph construction module 320 is specifically configured to: construct graph data according to the abnormal log data, and the graph data at least includes one of the following: a first subgraph of words and words, a second subgraph of labels and labels, and a third subgraph of words and labels.

[0135] Constructing graph data as the model input, disassembling the original heterogeneous graph of word-label into two defined homogeneous subgraphs and one heterogeneous sub Figure 3Three heterogeneous graphs, namely, word-word subgraph, label-label subgraph, and word-label subgraph.

[0136] It can be understood that the first subgraph of words is a graph composed only of word nodes and edges connecting words to words, representing the relationship between words; the second subgraph of labels is a graph composed only of label nodes and edges connecting labels to labels, representing the relationship between labels; the third subgraph of words and labels is a subgraph with both word nodes and label nodes, representing the relationship between words and labels.

[0137] In one embodiment of the present application, the classification module 330 is specifically configured to: obtain the adjacency matrices of the three subgraphs, namely, the first subgraph of words, the second subgraph of labels, and the third subgraph of words and labels, as the input of the graph convolutional model, and participate in training to obtain a classifier for classifying the abnormal information with an association relationship in the abnormal log data.

[0138] Input the constructed graph data into the heterogeneous lightweight graph convolutional model to obtain the trained word vectors and classifier. Input the two-layer lightweight graph convolutional model to train and obtain the embedding matrix, and obtain the embedding matrix of the text to be classified. According to the embedding matrix of the text to be classified, obtain the multi-label annotation result of the predicted text, that is, classify the abnormal information with an association relationship in the abnormal log data.

[0139] It can be understood that the above log processing device can implement each step of the log processing method provided in the foregoing embodiment. The relevant explanations of the log processing method are applicable to the log processing device and will not be elaborated here.

[0140] Figure 4 This is a schematic structural diagram of an electronic device according to an embodiment of the present application. Please refer to Figure 4 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.

[0141] The processor, network interface, and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 4 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0142] Memory, used to store programs. Specifically, the program can include program code, and the program code includes computer operation instructions. The memory can include a memory and a non-volatile memory, and provide instructions and data to the processor.

[0143] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a log processing device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:

[0144] Obtain abnormal log data;

[0145] According to the abnormal log data, construct graph data, which at least includes one of the following: a first sub-graph of words and words, a second sub-graph of labels and labels, and a third sub-graph of words and labels;

[0146] Take the first sub-graph of words and words, the second sub-graph of labels and labels, and the third sub-graph of words and labels, obtain the adjacency matrices of the three sub-graphs as the input of the graph convolution model, and participate in training to obtain a classifier to classify the abnormal information with an association relationship in the abnormal log data.

[0147] The above as in this application Figure 1The method executed by the log processing device disclosed in the illustrated embodiment can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or by instructions in software form. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0148] The electronic device can also execute Figure 1 the method executed by the log processing device in Figure 1 the illustrated embodiment and implement the functions of the log processing device in

[0149] Embodiments of the present application also propose a computer-readable storage medium that stores one or more programs. The one or more programs include instructions that, when executed by an electronic device including multiple application programs, can enable the electronic device to execute Figure 1 the method executed by the log processing device in the illustrated embodiment and specifically used to execute:

[0150] Obtain abnormal log data;

[0151] According to the abnormal log data, construct graph data, where the graph data includes at least one of the following: a first sub-graph of words and words, a second sub-graph of labels and labels, and a third sub-graph of words and labels;

[0152] The adjacency matrices of the three subgraphs are obtained from the word and the first subgraph of the word, the label and the second subgraph of the label, and the third subgraph of the word and the label, and used as the input of the graph convolution model to participate in training to obtain a classifier for classifying the abnormal information with an association relationship in the abnormal log data.

[0153] Those skilled in the art will understand that the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0154] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or a combination of blocks.

[0155] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or a combination of blocks.

[0156] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or a combination of blocks.

[0157] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0158] The memory may include non - permanent memory in the form of computer - readable media, such as random access memory (RAM) and / or non - volatile memory, such as read - only memory (ROM) or flash RAM. The memory is an example of computer - readable media.

[0159] Computer - readable media includes both permanent and non - permanent, removable and non - removable media and can store information by any method or technology. The information can be computer - readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase - change memory (PRAM), static random - access memory (SRAM), dynamic random - access memory (DRAM), other types of random - access memory (RAM), read - only memory (ROM), electrically erasable programmable read - only memory (EEPROM), flash memory or other memory technologies, compact disc read - only memory (CD - ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non - transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer - readable media does not include transitory media such as modulated data signals and carrier waves.

[0160] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non - exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0161] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer - usable storage media (including but not limited to disk storage, CD - ROM, optical storage, etc.) that contain computer - usable program code.

[0162] The above - mentioned are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A log processing method, wherein: The processing method comprises: Get exception log data; Constructing graph data according to the abnormal log data, wherein the graph data includes at least one of the following: a first subgraph of words and words, a second subgraph of labels and labels, and a third subgraph of words and labels; The adjacency matrix of the three subgraphs obtained by taking the words and the first subgraph of the words, the labels and the second subgraph of the labels, and the third subgraph of the words and the labels as the input of the graph convolution model is trained to obtain a classifier to classify the abnormal information with correlation in the abnormal log data.

2. The method of claim 1, wherein: The adjacency matrix of the three subgraphs obtained by using the words and the first subgraph of the words, the labels and the second subgraph of the labels, and the third subgraph of the words and the labels is used as the input of the graph convolution model, and the classifier is obtained through training to classify the abnormal information with correlation relationship in the abnormal log data, including: Perform convolution on the word and the first subgraph of the word to obtain an embedding matrix of the word; Perform convolution on the label and the second subgraph of the label to obtain an embedding matrix of the label; The embedding matrix of the word and the embedding matrix of the label are used as input of the third subgraph of the word and the label, and then the embedding matrix of the word and the label is obtained through a graph convolutional network.

3. The method of claim 2, wherein the embedding matrix of the words and tags comprises: According to the symmetric normalized adjacency matrix and label feature matrix of the third subgraph of the words and labels, the output result of the single-layer graph convolution model, that is, the final embedding matrix of the words and labels obtained through training, is obtained.

4. The method according to claim 3, further comprising: Split the final word and label embedding matrix into two matrices, representing the word embedding matrix and the label embedding matrix respectively; The text embedding matrix is ​​obtained by multiplying the word embedding matrix with the 0 and 1 vector matrix of the text; Performing matrix multiplication on the text embedding matrix and the label embedding matrix to obtain the similarity between the text and the label, and the probability prediction of the text after passing through a layer of linear network; The similarity between the text and the label and the probability prediction obtained by passing the text through a layer of linear network are weighted, and the one-dimensional features in each text are mapped into the final classification result.

5. The method of claim 1, wherein: The step of constructing graph data according to the abnormal log data, wherein the graph data includes at least one of the following: a first subgraph of words and words, a second subgraph of labels and labels, and a third subgraph of words and labels, including: The words and the first subgraph of the words include a graph consisting of only word nodes and edge connections between words; Assuming any two different words i and j in the graph, if i and j appear in the same exception log, there is an edge between word i and word j. If i and j do not appear in the same error, there is no edge between the words. The adjacency matrix A of the word and the first subgraph of the word is obtained. ww Definition of Among them, PMI stands for point mutual information, which is used to measure the correlation between words; When i and j are any two different words, the PMI indicator is used to measure the correlation between words i and j; when i = j, 1 is used to represent the weight of the node itself; in other cases, the value of the adjacency matrix element is represented by 0, where the PMI point mutual information is defined as follows:

6. The method of claim 1, wherein: The step of constructing graph data according to the abnormal log data, wherein the graph data includes at least one of the following: a first subgraph of words and words, a second subgraph of labels and labels, and a third subgraph of words and labels, including: The second subgraph of the label and the label includes a graph consisting of only label nodes and edge connections between labels; Assume that there are any two different labels i and j in the graph. If i and j are labeled with the same set of errors, there is an edge between labels i and j. If i and j do not belong to the same associated error, there is no edge between the nodes. The adjacency matrix A of the second subgraph of the label and the label is obtained. ll Definition of PMI is used to measure the similarity between two labels. If two labels appear together, there must be an edge. When i and j are any two different labels, the correlation is represented by PMI. When i=j, 1 is used to represent the weight of the node itself. In other cases, the value of the adjacency matrix element is represented by 0.

7. The method of claim 1, wherein: The step of constructing graph data according to the abnormal log data, wherein the graph data includes at least one of the following: a first subgraph of words and words, a second subgraph of labels and labels, and a third subgraph of words and labels, including: The third subgraph of the words and labels includes a subgraph having both word nodes and label nodes; Assume that for any two nodes i and j, if i is a word node and j is a label node and word i and label j appear in the same log, there is an edge between word i and label j; if i and j are two nodes with the same attributes, that is, both are word nodes or both are label nodes, there is no edge between the nodes, and the adjacency matrix A of the third subgraph of words and labels is obtained wl Definition of Among them, TF-IDF is term frequency-inverse document frequency.

8. A log processing device, wherein: The processing device comprises: Acquisition module, used to obtain abnormal log data; A graph building module, used to build graph data according to the abnormal log data, wherein the graph data includes at least one of the following: a first subgraph of words and words, a second subgraph of labels and labels, and a third subgraph of words and labels; The classification module is used to use the adjacency matrix of the three subgraphs, namely, the words and the first subgraph of the words, the labels and the second subgraph of the labels, and the third subgraph of the words and the labels, as the input of the graph convolution model, and participate in the training to obtain a classifier to classify the abnormal information with correlation in the abnormal log data.

9. An electronic device, comprising: processor; as well as A memory arranged to store computer executable instructions, which when executed cause the processor to perform the method of any one of claims 1 to 7.

10. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, causes the electronic device to execute any one of the methods of claims 1 to 7.