A variable recommendation method for logging
The graph structure information is extracted through the graph neural network and integrated log semantic information, and the log variables are recommended, which solves the problem of unreasonable logging in the existing technology, and improves the logging quality and development efficiency.
Patent Information
- Application Number
- CN202210453072.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-04-27
- Filing Date
- 2022-04-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-04-27
AI Technical Summary
The prior art cannot effectively obtain log variables through logging statements, resulting in unreasonable logging, affecting performance and development efficiency.
Graph neural network is used to extract graph structure information, and integrate graph structure information and log semantic information extracted by pre-trained models to recommend log variables.
Directly help developers write high-quality logging statements, solve the problem of unreasonable log variables, and improve performance and development efficiency.
Smart Images

Figure CN114780065B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of log mining, and more particularly, to a method for recommending variables for log records. Background Art
[0002] Refer to Figure 1 In the automatic log analysis flowchart shown, when the software runs (i.e., runs the source code), logs are generated by the log generation subsystem, and the logs are compressed in the log compression subsystem to output compressed logs; in the log parsing subsystem, the logs or the compressed logs can be parsed into a set of log templates; finally, log features are mined in the log mining subsystem. Generally speaking, a log consists of three parts - log level, static (contextualized) text, and dynamic content (the runtime state of the software recorded by log variables). Some empirical studies have shown that in the open-source systems studied, more than 25% of the log statement changes are related to the recorded variables. On the one hand, recording too many log variables in the logging statement may lead to performance overhead and prevent developers from magnifying (or focusing) on the real problems. On the other hand, the lack of recording important log variables will increase the burden on developers to perform many corrective software maintenance tasks. For example, developers may not be able to well understand the root cause of the problems affecting the deployed system because there is no unified tool for recording log variables. This highlights the need for a tool that can recommend log variables to developers to help them write high-quality logging statements. Refer to Figure 2 As shown, there is at least a logging statement in the running source code, and there may be log variables in the logging statement. Obtain the runtime state of the software from the log variables characterized in the log.
[0003] A graph is a data structure that can model a set of objects (nodes) and their relationships (edges). In recent years, due to the powerful expressive ability of graphs, the research on using machine learning to analyze graphs has received increasing attention. That is, graphs can be used to represent a large number of systems across different fields, including social sciences (social networks), natural sciences (physical systems and protein-protein interaction networks), knowledge graphs, and many other research fields. As a unique non-Euclidean (Euclidean) data structure for machine learning, graph analysis focuses on node classification, link prediction, and clustering. Graph neural networks (GNNs) are deep learning-based methods that operate in the graph domain and aggregate information in the graph structure based on CNNs and graph embeddings. In standard neural networks, the dependent information is only regarded as the features of nodes. However, GNNs can propagate through the graph structure without using it as part of the features. Due to its convincing performance and high interpretability, GNNs have recently become a widely used graph analysis method.
[0004] Recently, a large amount of work has shown that pre-trained models (PTMs) on large corpora can learn general language representations, which are very beneficial for downstream natural language processing (NLP) tasks and can avoid training new models from scratch. The first-generation pre-trained models were designed to learn good word embeddings, such as Skip-Gram and GloVe. Since these models themselves are no longer needed for downstream tasks, their computational efficiency is usually weak. Although these pre-trained embeddings can capture the semantics of words, they have no context and cannot capture high-level concepts in context, such as ambiguity resolution, syntactic structure, semantic roles, and anaphora. The second-generation pre-trained models focus on learning contextualized word embeddings, such as CoVe, ELMo, OpenAI GPT, and BERT. These learned models still need to represent words in context through downstream tasks. Among them, BERT even swept 11 NLP tasks at its birth. It would be meaningful to use pre-trained models such as BERT or improved pre-trained models for variable recommendation of logging statements. Summary of the Invention
[0005] Aiming at the problem that the current log generation subsystem cannot obtain log variables from logging statements, the present invention proposes a variable recommendation method for logging. First, use graph neural networks to extract graph structure information, and then fuse the graph structure information and the log semantic information extracted by using pre-trained models to recommend log variables. This method can directly help developers write high-quality logging statements to solve the technical problem of unreasonable log variables recorded in the source code.
[0006] A variable recommendation method for logging according to the present invention includes the following steps:
[0007] Step 1: Obtain the tags corresponding to the flags in each code segment from the source code;
[0008] Step 2: Construct a heterogeneous graph of code segments;
[0009] Step 3: Calculate the tag values of the flags;
[0010] Step 4: Encode the graph structure information of the code segments;
[0011] Step 5: Fusion of graph structure information based on BERT;
[0012] The advantages of the variable recommendation method for log records in the present invention are as follows:
[0013] ① First, select the source code of high-quality open-source Java projects, use the Java Development Tool (JDT) to construct multiple log variable recommendation extraction rules, and accurately and efficiently extract and process the tags of various flags required for variable recommendation from them. Then, convert the extracted and processed data into a dataset that can be used by subsequent models.
[0014] ② The present invention proposes a new method for recommending which log variables should be logged, using both the semantic information and the structural information of the code segments for variable recommendation. Specifically, the model of the present invention can efficiently encode semantic information and use a neural network to encode graph structure information. Given a code segment without log record statements, the present invention first uses a graph neural network to extract graph structure information, and then fuses the graph structure information and the log semantic information extracted using a pre-trained model to recommend log record variables. Brief Description of the Drawings
[0015] Figure 1 is a schematic diagram of the overall framework of automatic log analysis.
[0016] Figure 2 is a schematic diagram of the source code and log record statements.
[0017] Figure 3 is a flowchart of the variable recommendation for log records according to the present invention.
[0018] Figure 4 is a schematic diagram of an example of log record statements and corresponding code segments in the embodiment.
[0019] Figure 5 is a flowchart of the application of Step 1 in the embodiment. Detailed Embodiment
[0020] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0021] See Figure 3As shown, a variable recommendation method for log records of the present invention is stored in the log generation subsystem, and it includes the following steps:
[0022] Step 1, obtain the tags corresponding to the flags in each code segment from the source code;
[0023] In the present invention, for the source code of an open-source Java project, the Java Development Tools (JDT) are used to construct multiple log variable recommendation extraction rules.
[0024] Step 11, obtain the log record statements and the corresponding code segments from the source code;
[0025] In the present invention, the log record statements are obtained from the source code and denoted as ls; multiple log record statements ls form a log record statement set, denoted as LS = {ls 1 , ls 2 , …, ls n , …, ls N}. Each log record statement ls corresponds to a code segment cs, so there is a code segment set, denoted as CS = {cs 1 , cs 2 , …, cs n , …, cs N}.
[0026] ls 1 represents the first log record statement.
[0027] ls 2 represents the second log record statement.
[0028] ls n represents the nth log record statement.
[0029] ls N represents the last log record statement.
[0030] cs 1 represents the code segment corresponding to ls 1 , that is, the first code segment.
[0031] cs 2 represents the code segment corresponding to ls 2 , that is, the second code segment.
[0032] cs n represents the code segment corresponding to ls n , that is, the nth code segment.
[0033] cs N represents the code segment corresponding to ls N , that is, the last code segment.
[0034] The subscript n represents the identification number of the log record statement. For the sake of convenience of explanation, ls n is also referred to as any log record statement. Corresponding to the said ls n the said cs n is also referred to as any code segment.
[0035] For example, see Figure 2 Lines 3 and 5 in the source code shown are log record statements respectively.
[0036] Step 12, formulate the log variable extraction rules;
[0037] In the present invention, the log variable extraction rule set is denoted as RULE = {rule 1 , rule 2 , …, rule γ}.
[0038] rule 1 represents the first log variable extraction rule.
[0039] rule 2 represents the second log variable extraction rule.
[0040] rule γ represents the last log variable extraction rule.
[0041] For the sake of convenience of explanation, the said rule γ is also referred to as any log variable extraction rule, and the subscript γ represents the identification number of the log variable extraction rule.
[0042] The formulation of the log variable extraction rules refers to Section 4 of IEEE Transactions on Software Engineering, September 2019.
[0043] Step 13, extract the log variables in the log record statement;
[0044] Any log record statement ls n extracts log variables according to RULE = {rule 1 , rule 2 , …, rule γ}, and the obtained log variable set is denoted as and the subscript δ is the total number of log variables in ls n .
[0045] represents the first log variable of ls n .
[0046] Represents the second log variable of ls n .
[0047] Represents the last log variable of ls n .
[0048] Similarly, it can be obtained that the log record statement ls 1 According to RULE = {rule 1 , rule 2 , …, rule γ} is used to extract the set of log variables, which is denoted as and the subscript g is the total number of log variables in ls 1 .
[0049] Represents the first log variable of ls 1 .
[0050] Represents the second log variable of ls 1 .
[0051] Represents the last log variable of ls 1 .
[0052] Similarly, it can be obtained that the log record statement ls 2 According to RULE = {rule 1 , rule 2 , …, rule γ} is used to extract the set of log variables, which is denoted as and the subscript h is the total number of log variables in ls 2 .
[0053] Represents the first log variable of ls 2 .
[0054] Represents the second log variable of ls 2 .
[0055] Represents the last log variable of ls 2 .
[0056] Similarly, it can be obtained that the log record statement ls N According to RULE = {rule 1 , rule 2 , …, rule γ} The set of log variables extracted is denoted as and The subscript u is the total number of log variables in ls N .
[0057] Denotes the first log variable in ls N .
[0058] Denotes the second log variable in ls N .
[0059] Denotes the last log variable in ls N .
[0060] Step 14, determine the label corresponding to each flag in the code segment;
[0061] In the present invention, multiple flags token form a code segment cs.
[0062] In the present invention, the set of flags belonging to cs n is denoted as For any flag n in the code segment cs , the label is denoted as When the exists in the , then the corresponding is a log variable. When the does not exist in the , then the corresponding is not a log variable. The label assignment is denoted as
[0063] Denotes the first flag in cs n .
[0064] Denotes the second flag in cs n .
[0065] Denotes the y-th flag in cs n .
[0066] Denotes the last flag in cs n .
[0067] Similarly, it can be obtained that the set of flags belonging to cs N is denoted as In the code segment cs NAny one of the flags The label of is denoted as When the exists in the then The corresponding is a log variable (i.e., the value is 1). When the does not exist in the then The corresponding is not a log variable (i.e., the value is 0).
[0068] represents the first flag in cs N
[0069] represents the second flag in cs N
[0070] represents the b-th flag in cs N
[0071] represents the last flag in cs N
[0072] Similarly, it can be obtained that: the set of flags belonging to cs 1 is denoted as Then there is: for any one flag 1 in the code segment cs The label of is denoted as When the exists in the then The corresponding is a log variable (i.e., the value is 1). When the does not exist in the then The corresponding is not a log variable (i.e., the value is 0).
[0073] Similarly, it can be obtained that: the set of flags belonging to cs 2 is denoted as Then there is: for any one flag 2 in the code segment cs The label of is denoted as When the exists in the then The corresponding is a log variable (i.e., the value is 1). When the does not exist in the then corresponding is not a log variable (i.e., the value is 0).
[0074] In the present invention, for cs 1 the set of flags is denoted as for cs 2 the set of flags is denoted as for cs n the set of flags is denoted as for cs N the set of flags is denoted as the set of source code segment flags is denoted as
[0075] Step two, construct a heterogeneous graph of code segments;
[0076] Step 21, randomly select a set of code segments;
[0077] Randomly select 70% of the code segments from CS = {cs 1 , cs 2 , …, cs n , …, cs N} and denote them as the training code segment set TCS = {tcs 1 , tcs 2 , …, tcs α , …, tcs β}.
[0078] tcs 1 represents the first selected code segment.
[0079] tcs 2 represents the second selected code segment.
[0080] tcs α represents the α-th selected code segment.
[0081] tcs β represents the β-th selected code segment.
[0082] From randomly select 70% of the set of flags and denote it as the training set of flags
[0083] represents the training set of flags selected from TKK corresponding to tcs 1 .
[0084] represents the training set of flags selected from TKK corresponding to tcs 2 .
[0085] represents the training set of flags selected from TKK corresponding to tcsα Training flag set.
[0086] Indicates the corresponding tcs selected from TKK β Training flag set.
[0087] In the present invention, TCS and TTKK are selected in one-to-one correspondence.
[0088] Step 22, duplicate removal of the training flag set;
[0089] Since there are the same flag tokens in different training code segments tcs, in order to facilitate the selection of edges, a flag set without duplicates is added. By removing the duplicate flag tokens to construct a flag set without duplicates. The flags after removing the duplicate flags are denoted as the flag training sequence UT, and UT = [token 1 , token 2 , …, token i , …, token j , …, token K ;
[0090] token 1 Represents the first training flag.
[0091] token 2 Represents the second training flag.
[0092] token i Represents the i-th training flag.
[0093] token j Represents the j-th training flag.
[0094] token K Represents the last training flag.
[0095] The subscript K represents the total number of training flags in the flag training sequence UT.
[0096] Step 23, selection of nodes of the code segment heterogeneous graph;
[0097] In the present invention, the code segment heterogeneous graph is denoted as G(V, E). In the G(V, E), V represents the set of nodes in the graph, and E represents the set of edges in the graph.
[0098] In the present invention, TCS and UT are used as the nodes of G(V, E).
[0099] Step 24, selection of edges of the code segment heterogeneous graph;
[0100] In the present invention, the point mutual information method (PMI method) and the term frequency - inverse document frequency method (TF-IDF method) are used to calculate the weights of the edges of G(V, E).
[0101] The code segment heterogeneous graph G(V, E) is composed of nodes and edges. The graph constructed is called a heterogeneous graph because there are two types of nodes in the constructed graph, one is TCS and the other is UT. There are also two types of edges in the graph, namely the edges between token i and token j , and the edges between tcs α and token i . The weights of the edges between token i and token j are calculated according to the PMI method, while the weights of the edges between tcs α and token i are calculated according to the TF-IDF method.
[0102] In the present invention, the edge weight value calculated by the PMI method is PMI(i, j), that is:
[0103]
[0104]
[0105]
[0106]
[0107] f(i) represents the frequency of the sliding window containing token i .
[0108] f(j) represents the frequency of the sliding window containing token j .
[0109] f(i, j) represents the frequency of the sliding window containing both token i and token j .
[0110] S represents the total number of sliding windows.
[0111] s(i) represents the total number of sliding windows containing token i .
[0112] s(j) represents the total number of sliding windows containing token i .
[0113] s(i, j) represents the total number of sliding windows containing both token i and tokenj The total number of sliding windows.
[0114] In the present invention, the edge weight value calculated by the TF-IDF method is TF-IDF(i,α), that is:
[0115] TF-IDF(i,α) = TF(i,α) × IDF(i) (5)
[0116]
[0117]
[0118] TF(i,α) represents the word frequency of the token i in tcs α .
[0119] λ represents the number of tokens i in tcs α ..
[0120] η represents the number of flags α in tcs
[0121] IDF(i) represents the inverse logarithmic frequency of the code segments containing the token i .
[0122] K represents the total number of training flags.
[0123] represents the number of code segments containing the token i in TCS.
[0124] Step 3, calculate the label value of the flag;
[0125] In the present invention, the proportion M_label(i) of the number of times the label assignment is 1 is used to calculate the label value of any flag token i .
[0126] The calculation formula is:
[0127]
[0128] M_label(i) represents the proportion of the number of times the label of the token i is 1.
[0129] |true i | represents the number of tokens i whose label is 1.
[0130] |i| represents the number of tokens i in TCS = {tcs1 , tcs 2 , …, tcs α , …, tcs β} in the number of occurrences.
[0131] In the present invention, 0 ≤ M_label(i) ≤ 1.
[0132] Step 4, encode the graph structure information of the code segment;
[0133] The traditional graph convolutional model (GCN model) refers to Equation 2 in "SEMI - SUPERVISED CLASSIFICATION WITH GRAPH CONVOLUTIONAL NETWORKS", April 2017, ICLR conference.
[0134] In the present invention, a ReLU activation function and a sigmoid classifier are added to the traditional GCN model, which is called the improved GCN model.
[0135] In the present invention, the heterogeneous graph G(V, E) of the code segment is input into the improved GCN model for processing, and the output token i The predicted label value, denoted as M_label 预测 (i).
[0136] In the present invention, the improved GCN model is trained using 2 - layer graph convolutional layers. The activation function of the first graph convolutional layer is ReLU, the output result of the second graph convolutional layer is fed to the sigmoid function, and the loss function is the mean square error loss function MSE. Then, for each output token i The predicted label value, denoted as M_label 预测 (i).
[0137]
[0138]
[0139] σ(·) is the activation function.
[0140] is the sum of the adjacency matrix and the identity matrix.
[0141] is the degree matrix.
[0142] X is the vertex representation matrix in G(V, E).
[0143] W 1 is the weight matrix of the first graph convolutional layer.
[0144] W 2Is the weight matrix of the second graph convolutional layer.
[0145] In the present invention, after the code segment heterogeneous graph G(V,E) is trained by the improved GCN model, the embedding representations of each node in G(V,E) are obtained, denoted as The Is also called the graph structure information encoding of the code segment.
[0146] Step five, graph structure information fusion based on BERT;
[0147] The traditional BERT model refers to pages 46-50 of "Intelligent Summarization and Deep Learning", author Gao Yang, Beijing Institute of Technology Press, 2019.07.
[0148] Step 51, construct the SE-BERT model;
[0149] In the present invention, Is added to the embedding representation layer of the BERT model to realize the fusion of graph structure information and semantic information, denoted as the fusion model (i.e., the SE-BERT model).
[0150] In the present invention, due to the word fragment tokenizer used in the BERT model, so when Adding with the word embedding representation, segment embedding representation and position embedding representation, a token i Of Is only added to the first word fragment, and the corresponding remaining word fragments are added with zero vectors.
[0151] In the present invention, the segment embedding representation of any tcs α Is set to 0.
[0152] Step 52, use the SE-BERT model for label prediction;
[0153] In the present invention, the cross-entropy loss function is used to adjust the model parameters of the SE-BERT model to obtain the predicted label of each α In the code segment tcs Of
[0154] The cross-entropy loss function of the present invention
[0155] In the present invention, through steps one to five, a log variable recommendation model (i.e., the REVAL model) and a log variable recommendation database are obtained. Applying the REVAL model of the present invention to the source code can determine which flag tokens in the source code are log variables.
[0156] Example 1
[0157] SeeFigure 4 , Figure 5 As shown in Figure 4 the source code shown below, verify the log variable recommendation model (i.e., the REVAL model) of the present invention. Figure 5 is the detailed process of Figure 4 processing the source code shown below in Step 1.
[0158] Eight rules for extracting log variables from log record statements:
[0159]
[0160] Evaluate the log variable recommendation model of the present invention using three metrics: Hits@1, MRR, and MAP.
[0161] (1), Hits@1 means whether the first variable in the recommended ordered variable list is actually logged. Given a sequence of code variables, if the label of the first variable is '1' (i.e., actually logged), then it is considered a correct recommendation and assigned a value of 1; if the label of the first variable is '0' (i.e., not logged), then it is considered not a recommendation and assigned a value of 0. Therefore, Hits@1 is the average of the correct scores for all code segments. The higher Hits@1 is, the better the recommendation method.
[0162] (2), MRR (Mean Reciprocal Rank) represents the reciprocal of the rank of the flag of the first correctly predicted code segment. MRR is a commonly used metric for evaluating information retrieval methods. Given a sequence of code variables and a predicted ordered list, search the ordered list from the beginning until the position of the first variable that is actually logged. The reciprocal of the number of searches to find the actually logged variable is the reciprocal rank. The mean reciprocal rank is the average of the reciprocal ranks of all code variables. A higher MRR value means that the first correctly predicted variable is ranked higher in the ordered prediction list.
[0163] (3), MAP is a single - metric quality measure that has been proven to have good discriminability and stability when evaluating ranking techniques. MAP takes into account all correctly predicted code variables and can be regarded as a measure of average performance. Given a series of code variables and their ordered prediction lists, MAP is the average of the average precisions of all code segments. Since there may be multiple code variables in a log record statement, this metric (MAP) is essential. The higher MAP is, the more it takes into account all the recommended and logged variables.
[0164] Evaluation and comparison results
[0165] Based on the three metrics of Hits@1, MRR, and MAP, the data is evaluated using models such as RG (random guess), IR-comp, IR-flat, IR-mix, RNN_Attn, BERT, and REVAL. The evaluation results are shown in Tables 1 to 3.
[0166] (1) It can be found from the evaluation results that models like RG, IR-comp, IR-flat, and IR-mix, which neither encode semantic information nor encode graph structure information, cannot reach the performance of models that only encode semantic information (BERT, pretraining_BERT, CodeBERT) in each metric, nor can they reach the performance of the REVAL model of the present invention that encodes both semantic information and graph structure information. This shows that semantic information is quite important for the problems targeted in the present invention.
[0167] (2) When comparing the average performance of REVAL and BERT in each metric, it can be found that REVAL is better than BERT in Hits@1, MRR, and MAP. Better Hits@1 and MRR mean that the first log variable recommended by REVAL (if there are multiple recommended variables) is more accurate than that of BERT. Better MAP means that the overall performance of REVAL is better than that of other models; in other words, regardless of whether the log statement has only one log variable or multiple log variables, the recommendation quality of REVAL will be better. This also shows that the graph structure information additionally encoded by REVAL compared to BERT is beneficial, enabling REVAL to utilize more information in the code segment and being more conducive to variable recommendation.
[0168] Table 1 Comparison of RG, IR-comp, IR-flat, IR-mix, RNN_Attn, BERT, and REVAL in Hits@1
[0169] Projects RG IR-comp IR-flat IR-mix RNN_Attn BERT REVAL ActiveMQ 0.179 0.2227 0.1703 0.1659 0.6026 0.6026 0.6463 Camel 0.2024 0.3306 0.2893 0.2893 0.6942 0.686 0.7066 Cassandra 0.4722 0.0833 0.1389 0.1389 0.5556 0.6667 0.75 CloudStack 0.2152 0.1525 0.1166 0.1211 0.509 0.5202 0.5628 DirectoryServer 0.2973 0.4054 0.2297 0.2432 0.6216 0.7162 0.7162 Hadoop 0.2278 0.15 0.1167 0.1315 0.6407 0.6759 0.6778 HBase 0.1982 0.1586 0.1189 0.1189 0.5286 0.6167 0.6255 Hive 0.1667 0.175 0.15 0.1542 0.6125 0.6083 0.6292 Zookeeper 0.2979 0.3829 0.2128 0.3404 0.5957 0.5532 0.6383 average 0.2507 0.229 0.1715 0.1893 0.5956 0.6273 0.6614
[0170] Table 2 Comparison of RG, IR-comp, IR-flat, IR-mix, RNN_Attn, BERT, and REVAL in MRR
[0171] Projects RG IR-comp IR-flat IR-mix RNN_Attn BERT REVAL ActiveMQ 0.3132 0.2238 0.1779 0.1736 0.7366 0.7617 0.778 Camel 0.3464 0.3337 0.2903 0.2903 0.808 0.8115 0.8246 Cassandra 0.5798 0.0833 0.1389 0.1389 0.7077 0.7841 0.8221 CloudStack 0.328 0.1599 0.1177 0.1222 0.6483 0.6686 0.7033 DirectoryServer 0.4797 0.4054 0.2297 0.2432 0.7542 0.8175 0.8186 Hadoop 0.3788 0.1553 0.1252 0.1415 0.7689 0.7982 0.7993 HBase 0.3201 0.1601 0.1215 0.1215 0.6904 0.7565 0.7553 Hive 0.2847 0.1823 0.1597 0.1618 0.7415 0.7419 0.7537 Zookeeper 0.4479 0.383 0.2128 0.3404 0.7164 0.7225 0.7543 average 0.3865 0.2319 0.1749 0.1926 0.7302 0.7625 0.7788
[0172] Table 3 Comparison of RG, IR-comp, IR-flat, IR-mix, RNN_Attn, BERT, and REVAL in MAP
[0173] Projects RG IR-comp IR-flat IR-mix RNN_Attn BERT REVAL ActiveMQ 0.2634 0.222 0.1789 0.1746 0.7155 0.7463 0.7626 Camel 0.314 0.3333 0.29 0.2903 0.7848 0.805 0.8106 Cassandra 0.5575 0.0794 0.135 0.135 0.671 0.7616 0.8036 CloudStack 0.2939 0.1598 0.1174 0.1217 0.6255 0.652 0.6838 DirectoryServer 0.4359 0.4054 0.2297 0.2432 0.7441 0.8072 0.8113 Hadoop 0.343 0.1548 0.1238 0.1399 0.7394 0.7674 0.7677 HBase 0.297 0.1598 0.1208 0.1208 0.6611 0.7321 0.7222 Hive 0.2669 0.1724 0.1553 0.1578 0.7004 0.716 0.7171 Zookeeper 0.3885 0.383 0.2128 0.3404 0.7164 0.7172 0.7472 average 0.3511 0.23 0.1737 0.1915 0.7064 0.745 0.7585
[0174] In this embodiment, the log variable recommendation is as follows Figure 4 As shown, it can be found from the log record statement in the figure that the log variable is "i", and the recommended variable of the REVAL model of the present invention is also i, while the variables recommended by other models are "me" and "context", which are not accurate.
[0175] Summary of log recording methods
[0176]
Claims
1. A variable recommendation method for log records, characterized in that it includes the following steps: Step 1, obtain the tags corresponding to the flags in each code segment from the source code; Step 11, obtain the log record statements and the corresponding code segments from the source code; Step 12, formulate log variable extraction rules; Step 13, extract the log variables in the log record statements; Step 14, determine the tags corresponding to each flag in the code segment; Step 2, construct a heterogeneous graph of code segments; Step 21, randomly select a set of code segments; Step 22, remove duplicates from the training flag set; Step 23, select the nodes of the heterogeneous graph of code segments; The heterogeneous graph of code segments is denoted as G(V,E); in the G(V,E), V represents the set of nodes in the graph, and E represents the set of edges in the graph; Take TCS and UT as the nodes of G(V,E); TCS represents the training code segment set; UT represents the flag training sequence; Step 24, select the edges of the heterogeneous graph of code segments; There are also two types of edges in the code segment heterogeneous graph G(V, E), namely the edges between token i and token j , and the edges between tcs α and token i ; The weights of the edges between token i and token j are calculated according to the PMI method, while the weights of the edges between tcs α and token i are calculated according to the TF-IDF method; token i represents the i-th training flag; token j represents the j-th training flag; tcs α Indicates the α-th selected code segment; Step 3, calculate the tag value of the flag; Use the proportion M_label(i) of the number of times the label assignment is 1 to calculate the label value of any flag token i ; Step 4, encode the graph structure information of the code segment; Input the heterogeneous graph G(V, E) of code segments into the improved GCN model for processing, and output tokens i The predicted label value, denoted as M_label 预测 (i); The improved GCN model is trained using 2 layers of graph convolutional layers. The activation function of the first graph convolutional layer is ReLU, and the output result of the second graph convolutional layer is given to the sigmoid function. The loss function is the mean square error loss function MSE; After the improved GCN model processes the code segment heterogeneous graph G(V, E), the embedding representations of each node in G(V, E) are obtained, denoted as The is also called the graph structure information encoding of the code segment; Step 5, fuse the graph structure information based on BERT; Step 51, construct the SE-BERT model; Add to the embedding representation layer of the BERT model to achieve the fusion of graph structure information and semantic information, denoted as the SE-BERT model; Step 52, use the SE-BERT model for label prediction; Adjust the model parameters of the SE-BERT model using the cross-entropy loss function to obtain the selected code segment tcs α in the flag predicted label of i.e., the log variable; Cross-entropy loss function K represents the total number of training flags; In code segment cs n the flag is labeled as 2. A variable recommendation method for log records according to claim 1, characterized in that: In step 11, obtain a set of logging statements from the source code, denoted as LS = {ls 1 , ls 2 , …, ls n , …, ls N}; each logging statement ls corresponds to a code segment cs, so there is a set of code segments, denoted as CS = {cs 1 , cs 2 , …, cs n , …, cs N}; ls 1 Indicates the first logging statement; ls 2 Indicates the second logging statement; ls n represents the nth logging statement; ls N Indicates the last logging statement; cs 1 Indicates ls 1 The corresponding code segment, i.e., the first code segment; cs 2 Indicates ls 2 The corresponding code segment, i.e., the second code segment; cs n denotes ls n the corresponding code segment, i.e., the nth code segment; cs N Indicates ls N The corresponding code segment, that is, the last code segment.
3. A variable recommendation method for log records according to claim 2, characterized in that: The log variable extraction rule set in step 12 is denoted as RULE = {rule 1 , rule 2 , …, rule γ}; rule 1 Indicates the first log variable extraction rule; rule 2 Indicates the second log variable extraction rule; rule γ Indicates the last log variable extraction rule.
4. A variable recommendation method for log records according to claim 3, characterized in that: The logging statement ls in step 13 described above 1 According to RULE = {rule 1 , rule 2 , …, rule γ}, the set of log variables extracted is denoted as and the subscript g is the total number of log variables in ls 1 ; Log record statement ls 2 Extracted according to RULE = {rule 1 , rule 2 , …, rule γ}, the set of log variables is denoted as And The subscript h is the total number of log variables in ls 2 ; Log record statement ls n According to RULE = {rule 1 , rule 2 , …, rule γ}, log variables are extracted, and the obtained set of log variables is denoted as and the subscript δ is the total number of log variables in ls n ; Log record statement ls N Extracted according to RULE = {rule 1 , rule 2 , …, rule γ}, the set of log variables is denoted as And The subscript u is the total number of log variables in ls N ; Indicates ls 1 The first log variable of; Indicates ls 1 The second log variable of; Indicates the last log variable of ls 1 ; Indicates ls 2 The first log variable of; Indicates ls 2 The second log variable of; Indicates ls 2 The last log variable; Indicates ls n The first log variable; Indicates ls n The second log variable; Indicates the last log variable of ls n ; Indicates ls N The first log variable of; Indicates ls N The second log variable of; Indicates ls N The last log variable of 5. A variable recommendation method for log records according to claim 4, characterized in that: In the step 14, multiple flag tokens form a code segment cs; Belonging to cs 1 The set of flags is denoted as Then there is: in the code segment cs 1 Any one flag The label of, is denoted as When the Exists in the Then The corresponding Is a log variable; when the Does not exist in the Then The corresponding Is not a log variable; Belonging to cs 2 The set of flags is denoted as Then there is: in the code segment cs 2 Any one of the flags The label of, is denoted as When the Exists in the Then The corresponding Is a log variable; when the Does not exist in the Then The corresponding Is not a log variable; Belonging to cs n The set of flags is denoted as In the code segment cs n Any one of the flags The label is denoted as When the Exists in the Then The corresponding Is a log variable; when the Does not exist in the Then The corresponding Is not a log variable; Belonging to cs N The set of flags is denoted as In the code segment cs N For any one flag The label is denoted as When the Exists in the Then The corresponding Is a log variable; when the Does not exist in the Then The corresponding Is not a log variable; Indicates the first flag in cs n ; Indicates cs n The second flag in; Indicates the y-th flag in cs n ; Indicates cs n The last flag in; Indicates the first flag in cs N ; Indicates cs N The second flag in; Indicates cs N the b-th flag in; Indicates cs N the last flag in; The flag sets of each code segment in the statistical code segment set CS = {cs 1 , cs 2 , …, cs n , …, cs N} are denoted as the source code segment flag set 6. A variable recommendation method for log records according to claim 5, characterized in that: In step 21 described above, 70% of the code segments are randomly selected from CS = {cs 1 , cs 2 , …, cs n , …, cs N} and denoted as the training code segment set TCS = {tcs 1 , tcs 2 , …, tcs α , …, tcs β}; From Randomly select 70% of the set of markers and denote it as the training set of markers TCS and TTKK are selected in a one-to-one correspondence; tcs 1 Indicates the first selected code segment; tcs 2 Indicates the second selected code segment; tcs α Indicates the α-th selected code segment; tcs β Indicates the β-th selected code segment; Indicates the training flag set corresponding to tcs selected from TKK 1 ; Indicates the training flag set corresponding to tcs selected from TKK 2 ; Indicates the corresponding training flag set selected from TKK for tcs α ; Indicates the training flag set corresponding to tcs selected from TKK β .
7. A variable recommendation method for log records according to claim 6, characterized in that: In step 22, remove duplicate flag tokens in 1 to construct a set of non-duplicate flags; the flags after removing duplicate flags are denoted as the flag training sequence UT, and UT = [token 2 , token i , …, token j , …, token K ; token 1 Indicates the first training flag; token 2 Indicates the second training flag; token i represents the i-th training flag; token j represents the j-th training flag; token K Indicates the last training flag; The subscript K represents the total number of training flags in the flag training sequence UT.
8. A variable recommendation method for log records according to claim 7, characterized in that: In the step 24, the edge weight value calculated by the PMI method is PMI(i,j), that is: f(i) represents the frequency of the sliding window that contains the token i ; f(j) represents the frequency of the sliding window that contains the token j ; f(i,j) represents the frequency of the sliding window that contains both the token i and the token j ; S represents the total number of sliding windows; s(i) represents the total number of sliding windows containing the token i ; s(j) represents the total number of sliding windows containing the token i ; s(i, j) represents the total number of sliding windows that contain both the token i and the token j .
9. A variable recommendation method for log records according to claim 8, characterized in that: In the step 24, the edge weight value calculated by the TF-IDF method is TF-IDF(i,α), that is: TF-IDF(i,α)=TF(i,α)×IDF(i) (5) TF(i,α) represents the token i in tcs α the term frequency; λ represents a token i in tcs α the number of; η represents the number of flags in tcs α ; IDF(i) represents the inverse log frequency of the code segment containing the token i ; K represents the total number of training flags; Indicates the number of code segments containing tokens in the TCS i .
10. A variable recommendation method for log records according to claim 9, characterized in that: In step 3, the proportion M_label(i) of the number of times the label assignment is 1 is used to calculate the label value of any flag token i ; The calculation formula is: M_label(i) represents the token i The label is the proportion of the number of times the label is 1; and 0 ≤ M_label(i) ≤ 1; |true i |represents the number of tokens i with a label of 1; |i| represents the token i in TCS = {tcs 1 , tcs 2 , …, tcs α , …, tcs β}} the number of occurrences.
11. A variable recommendation method for log records according to claim 10, characterized in that: Output each token in Step 4 i The predicted label value, denoted as M_label 预测 (i); σ(·) is an activation function; is the sum of the adjacency matrix and the identity matrix; is the degree matrix; X is a vertex representation matrix in G(V, E); W 1 is the weight matrix of the first graph convolutional layer; W 2 is the weight matrix of the second convolutional layer of the graph.
12. A variable recommendation method for log records according to claim 11, characterized in that: In step 51, due to the word-piece tokenizer used by the BERT model, when adding to the word embedding representation, segment embedding representation, and position embedding representation, a token i of is only added to the first word-piece, and zero vectors are added to the corresponding remaining word-pieces; Set the segment embedding representation of any tcs α to 0.
Citation Information
Patent Citations
Log level prediction method and device and storage medium
CN110806962A
Log mode extraction method and system for log training of cloud native system
CN111190873A