A hybrid masking method for self-supervised training of language models for log anomaly detection

By using a hybrid mask method in the BERT model to analyze high-frequency phrases and subsequences in the log text, the problem that the BERT model in the prior art fails to make full use of word correlation in the log text is solved, and more accurate log anomaly detection is achieved.

CN119248618BActive Publication Date: 2025-05-23INFORMATION & COMMNUNICATION BRANCH STATE GRID JIANGXI ELECTRIC POWER CO
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411761577.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-05-23
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

The existing BERT model fails to fully utilize the correlation between words in log text in log anomaly detection, resulting in the failure of understanding of log words to utilize context information.

Method used

The mixed mask method is used to self-supervise training of log text. By analyzing high-frequency phrases and subsequences, and combining word masks, the performance of BERT model for log text analysis is improved.

Benefits of technology

Through the hybrid masking method, the BERT model can better understand the context information of the log text, enhance the understanding of log structure and semantics, and improve the ability to capture long-distance dependencies of words, thereby more accurately identifying potential abnormal patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119248618B_ABST
    Figure CN119248618B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of log anomaly detection, and discloses a mixed masking method in self-supervised training of a log anomaly detection language model, which parses a log database to form a log template library, extracts a constant word subsequence library and a variable word subsequence library; disassembles long words in constants and variables, updates the constant word subsequence library and the variable word subsequence library, and constructs a log word library; analyzes high-frequency phrases in word subsequences in a way of phrase frequency; performs mixed masking of log text based on words, high-frequency phrases and subsequences to obtain masked word sequences; constructs the input of a Transformer encoder based on the masked word sequence, adopts a group query attention mechanism in the Transformer encoder, and performs self-supervised training of a BERT model by predicting masked words. The present invention can better train the BERT model's ability to understand phrases and subsequences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of log anomaly detection and relates to a mixed mask method in self-supervised training of a log anomaly detection language model. Background Art

[0002] The log anomaly detection method based on the BERT model is currently a better method for log anomaly detection. The BERT model is an improved model of the language analysis Transformer model. Referring to the processing of natural language by the Transformer model, each word in the log text is encoded separately, and the encoder is used to analyze the features between words. In the self-supervised training process of the BERT model, the randomly selected words are masked by using a mask, and the BERT model is trained to predict the masked words. If the predicted words are consistent with expectations, the log is considered normal, otherwise the log is considered abnormal. Mask prediction can train the BERT model's ability to understand log text.

[0003] Since there is a correlation between words in log texts, and the existing BERT model fails to fully utilize the correlation between words when learning log texts, the model's understanding of log words fails to utilize contextual information. Summary of the invention

[0004] Aiming at the characteristics of log text, the present invention provides a hybrid masking method in self-supervised training of a log anomaly detection language model, analyzes high-frequency phrases and subsequences in log text, and adopts a hybrid strategy when masking the log text based on the high-frequency phrases and subsequences, so that the model can extract the features of both single words and high-frequency phrases and subsequences, thereby improving the performance of the BERT model in log text analysis.

[0005] The present invention is implemented by the following technical solution: A hybrid masking method in self-supervised training of a log anomaly detection language model, comprising the following steps:

[0006] Step 1: Use a log parser to perform statistical analysis on each log text in the log database, analyze the log template contained in the log text, and form a log template library;

[0007] Step 2: Extract the word combinations corresponding to all constants in the log template library to form a constant word subsequence library; extract the word combinations corresponding to all variables to form a variable word subsequence library;

[0008] Step 3: Disassemble the long words in constants and variables, reconstruct word combinations, update the constant word subsequence library and variable word subsequence library, and build a log word library;

[0009] Step 4: Segment the word subsequences in the constant word subsequence library and the variable word subsequence library, and analyze the high-frequency phrases contained in the word subsequences; select the subsequences and count the frequencies of G consecutive words. , when greater than the threshold When G consecutive words are combined into a phrase As a high frequency phrase;

[0010] Step 5: Mix the log text based on words, high-frequency phrases and subsequences to obtain the masked word sequence;

[0011] Step 6: Construct the input of the Transformer encoder based on the masked word sequence and perform self-supervised training of the BERT model.

[0012] Specifically, the log template library , , is the number of log templates, For the Log templates, , The number of variables / constants. A constant in a log template contains a fixed number of word combinations, and the length of the words contained in the variable is not fixed.

[0013] Specifically, in step 2, for each log text in the log database, it is matched with each log template, constants are matched according to the longest match principle, and words between constants are regarded as variables, so as to identify constants and variables contained in the log text.

[0014] Specifically, in step 3, the constant word subsequence library and variable word subsequence library All words in the log are taken out to form a log word library , is the first word in the log word library, The first words, Represents the number of all words in the log text base;

[0015] ;

[0016] In the formula It means taking different words from the subsequence and removing the same words.

[0017] Specifically, in step 4, a corresponding relationship between each subsequence and a high-frequency phrase in the constant word subsequence library and the variable word subsequence library is established to obtain the high-frequency phrases contained in each subsequence.

[0018] Specifically, in step 5, a log text data is taken from the training data set to match with the log template, the constant word subsequence and variable word subsequence contained in the log text are analyzed, and the long words are disassembled to obtain the subsequence set , this log text data contains P subsequences, Representative subsequences, , or , is a constant word subsequence library, Variable word subsequence library At the same time, the log text data is also parsed by word to obtain the word sequence , the log text data sequence contains words, Indicates The word in position, , ;

[0019] The process of mixed masking of log text based on words, high-frequency phrases and subsequences is as follows:

[0020] Step a: Randomly select a location ( );

[0021] Step b: If the position The word and its adjacent words are in a high-frequency phrase. Words The high-frequency phrases to which it belongs are used as the words to be masked;

[0022] Step c: If the position The word and its adjacent words are not in the high-frequency phrase, and the position is changed to Words As the word to be covered, the other 50% probability will be located at Words the subsequence it belongs to As the word to be covered;

[0023] Step d: Calculate the number of words to be masked MN, if , then return to step a and continue to randomly select a position. If , then the selection of the masked word is completed;

[0024] Step e: For the selected words to be masked, there is an 80% probability of directly replacing the original words with the [MASK] mark; there is a 10% probability of replacing the original words with random words; there is a 10% probability of keeping the original words unchanged;

[0025] Step f: After mixed masking, a word sequence of log text data There are three output , , , where MN words are replaced by the [MASK] tag , the word sequence after MN words are replaced by random words , the word sequence after MN words keep the original words unchanged ;in is the first word replaced by the [MASK] tag. is the MNth word replaced by the [MASK] tag, is the first word replaced by a random word, is the MNth word replaced by a random word.

[0026] Further preferably, in step 6, the output after the mixed mask is , , , construct the word embedding matrix formed by each word ,in is the word embedding matrix formed by all words in the log database, is the dimension of the word embedding of a word, is the number of all words in the log database; construct the position encoding matrix Z of each word ,in represents the position encoding matrix formed by each word in the current log text, T represents the number of words in the current log text, is the dimension of the position encoding of a word, a word in the log text Expressed as word embedding and position encoding , [MASK] mark word The word embedding is , Positional encoding , the sum of word embedding and position encoding forms the input representation , , , the input represents the set As input to the Transformer encoder in the BERT model.

[0027] Further optimization, when the BERT model is trained in self-supervision, the Transformer encoder is used as the input representation set The feature extraction module After the features are obtained, the feature map ,according to Predict the content of the masked word.

[0028] Further preferably, the self-attention mechanism in the Transformer encoder is replaced by a group query attention mechanism.

[0029] Further preferably, in the process of group query attention mechanism, the number of query heads is set to 8, the number of groups is set to 4, and each group of query heads shares the same keyword mapping and value mapping obtained by linear transformation of the input.

[0030] The present invention randomly masks the fragments of consecutive words in the word sequence of the input log text and trains the BERT model to predict the masked words, which not only enhances the BERT model's understanding of the log structure and semantics, but also improves the BERT model's ability to capture long-distance dependencies between words. By predicting the masked words, the BERT model can better understand the contextual information of the log text, thereby more accurately identifying potential abnormal patterns. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 Schematic diagram of the hybrid masking process.

[0032] Figure 2 Schematic diagram of the input processing of the Transformer encoder. DETAILED DESCRIPTION

[0033] A large amount of log data is exported from a system, and the log data is disassembled into several log text data to form a log database. It is necessary to determine whether a certain log text data is an abnormal log text. Log anomaly detection based on the language model BERT is currently a better method. In the process of self-supervised model training, the masking method can effectively improve the performance of the model. The present invention proposes a hybrid masking method in the self-supervised training of the log anomaly detection language model for the text encoding process and model training process in the BERT model for log anomaly detection, including the following steps 1-5.

[0034] Step 1: Use a log parser to perform statistical analysis on each log text in the log database, analyze the log templates contained in the log text, and form a log template library.

[0035] When a log text describes an event, a suitable log template will be selected according to the specific event, and the event information will be filled in according to the log template. The log database exported from the system mostly appears in the form of text. When the log template is unknown, it is necessary to obtain the log template from the log database so that the fields contained in the subsequent analysis of a log text data can be analyzed according to the log template. For the log database formed by the exported log text data, the present invention uses a log parser to perform statistical analysis on all log text data in the log database, analyze the log templates contained in the log database, and form a log template library. , , is the number of log templates, For the Log templates, , The number of variables / constants. A constant in a log template contains a fixed number of word combinations, and the number of words in a variable is not fixed, which is represented by * in the log template. The constant is a description of the nature of the variable. In a log text, the constant is represented as the log attribute name, and the variable is represented as the log attribute value.

[0036] Step 2: Constant word subsequence and variable word subsequence analysis.

[0037] Extract the word combinations corresponding to all constants in the log template library to form a constant word subsequence library ,in, is the first constant word subsequence, For the A constant subsequence of words.

[0038] For each log text in the log database, match it with each log template, match constants according to the longest match principle, and treat the words between constants as variables, so as to identify the constants and variables contained in the log text. Extract the word combinations corresponding to all variables to form a variable word subsequence library , is the first variable word subsequence, For the variable word subsequences.

[0039] Step 3: Disassemble the long words in constants and variables, reconstruct word combinations, update the constant word subsequence library and variable word subsequence library, and build a log word library;

[0040] In order to include rich information in one word in the log text, some words are connected with the "_" connection symbol to form a long word. When analyzing the information contained in the log text, the long word needs to be disassembled and reconstructed into a word combination. and variable word subsequence library Long words in the text file are decomposed into word combinations according to the connection symbols, and the constant word subsequence library is updated. , update the variable word subsequence library .

[0041] Constant word subsequence library and variable word subsequence library All words in the log are taken out to form a log word library , is the first word in the log word library, The first words, Represents the number of all words in the log text base.

[0042] ;

[0043] In the formula It means taking different words from the subsequence and removing the same words.

[0044] Step 4: Perform word segmentation on the subsequences in the constant word subsequence library and the variable word subsequence library;

[0045] Use the word frequency method to segment some repeated phrases in the subsequence. Select all subsequences with more than or equal to 4 words in the constant word subsequence library and the variable word subsequence library, and count the frequency of G consecutive words. , when greater than the threshold When G consecutive words are combined into a phrase As high-frequency phrases, where G ≧ 3, G < the maximum number of words in the subsequence.

[0046] The corresponding relationship between each subsequence and the high-frequency phrase in the constant word subsequence library and the variable word subsequence library is established to obtain the high-frequency phrases contained in each subsequence.

[0047] Step 5: Mix the log text based on words, high-frequency phrases, and subsequences to obtain the masked log text.

[0048] Take a log text data from the training data set and match it with the log template, analyze the constant word subsequences and variable word subsequences contained in the log text, and disassemble the long words to obtain the subsequence set , this log text data contains P subsequences, Representative subsequences, , or At the same time, the log text data is also parsed by word to obtain the word sequence , the log text data contains in order words, Indicates The word in position, ,word will belong to a subsequence of the subsequence set, that is .

[0049] Reference Figure 1 , the process of mixed masking of log text based on words, high-frequency phrases and subsequences is as follows:

[0050] Step a: Randomly select a location ( );

[0051] Step b: If the position The word and its adjacent words are in a high-frequency phrase. Words The high-frequency phrases to which it belongs are used as the words to be masked;

[0052] Step c: If the position The word and its adjacent words are not in the high-frequency phrase, and the position is changed to Words As the word to be covered, the other 50% probability will be located at Words the subsequence it belongs to As the word to be covered;

[0053] Step d: Calculate the number of words to be masked MN, if , then return to step a and continue to randomly select a position. If , then the selection of the masked word is completed;

[0054] Step e: For the selected words to be masked, there is an 80% probability of directly replacing the original words with the [MASK] mark; there is a 10% probability of replacing the original words with random words; there is a 10% probability of keeping the original words unchanged;

[0055] Step f: After mixed masking, a word sequence of log text data There are three output , , , where MN words are replaced by the [MASK] tag , the word sequence after MN words are replaced by random words , the word sequence after MN words keep the original words unchanged ;in is the first word replaced by the [MASK] tag. is the MNth word replaced by the [MASK] tag, is the first word replaced by a random word, is the MNth word replaced by a random word.

[0056] Step 6: Construct the input of the Transformer encoder based on the masked word sequence and perform self-supervised training of the BERT model.

[0057] (1) Input processing

[0058] Referring to the way the BERT model processes text, the words in the current input log text are embedded and encoded.

[0059] Output after mixing mask , , , construct the word embedding matrix formed by each word ,in is the word embedding matrix formed by all words in the log database, is the dimension of the word embedding of a word, is the number of all words in the log database; construct the position encoding matrix Z of each word ,in represents the position encoding matrix formed by each word in the current log text, T represents the number of words in the current log text, is the dimension of the position encoding of a word, which is the same as the dimension of the word embedding of a word. Can be expressed as word embedding and position encoding , [MASK] mark word The word embedding is , Positional encoding , the sum of word embedding and position encoding forms the input representation , , , ,like Figure 2 As shown. The input represents the set As input to the Transformer encoder in the BERT model.

[0060] (2) Feature extraction and word prediction

[0061] Transformer encoder as input representation set The feature extraction module After the features are obtained, the feature map ,according to Predict the content of the masked word.

[0062] The self-attention mechanism in the Transformer encoder is replaced with a group query attention mechanism to capture the correlation between subsequences and subsequences.

[0063] In the process of group query attention mechanism, the number of query heads is set to 8 and the number of groups is set to 4. Each group of query heads shares the same keyword mapping and value mapping obtained by linear transformation of the input.

[0064] (3) Iterative Optimization

[0065] Compare the difference between the predicted masked words and the original words, and calculate the cross entropy loss. Continue to input the log text in the training data set and continuously iterate and optimize the BERT model.

[0066] This embodiment provides a computer-readable storage medium having computer instructions stored thereon. When the computer instructions are executed by a processor, the mixed mask method in the self-supervised training of the log anomaly detection language model is implemented.

[0067] The invention described above only expresses the implementation methods of the embodiments of the present invention, and cannot be understood as limiting the scope of the invention patent, nor does it impose any form of limitation on the structure of the embodiments of the present invention. It should be pointed out that for ordinary technicians in this field, several changes and improvements can be made without departing from the concept of the embodiments of the present invention, which all belong to the protection scope of the embodiments of the present invention.

Claims

1. A hybrid masking method for self-supervised training of a language model for log anomaly detection, characterized in that: The following steps are involved: Step 1: Use a log parser to perform statistical analysis on each log text in the log database, analyze the log template contained in the log text, and form a log template library; Step 2: Extract the word combinations corresponding to all constants in the log template library to form a constant word subsequence library; extract the word combinations corresponding to all variables to form a variable word subsequence library; Step 3: Disassemble the long words in constants and variables, reconstruct word combinations, update the constant word subsequence library and variable word subsequence library, and build a log word library; Step 4: Perform word segmentation on the word subsequences in the constant word subsequence library and the variable word subsequence library, and analyze the high-frequency phrases contained in the word subsequences; Select a subsequence by counting the frequency of G consecutive words , when greater than the threshold When G consecutive words are combined into a phrase As a high frequency phrase; Step 5: Mix the log text based on words, high-frequency phrases and subsequences to obtain the masked word sequence; Step 6: Construct the input of the Transformer encoder based on the masked word sequence and perform self-supervised training of the BERT model; In step 5, a log text data is taken from the training data set to match with the log template, the constant word subsequence and variable word subsequence contained in the log text are analyzed, and the long words are disassembled to obtain the subsequence set , this log text data contains P subsequences, Representative subsequences, , or , is a constant word subsequence library, Variable word subsequence library ; At the same time, the log text data is also parsed by word to obtain the word sequence , the log text data sequence contains words, Indicates The word in position, , ; The process of mixed masking of log text based on words, high-frequency phrases and subsequences is as follows: Step a: Randomly select a location , ; Step b: If the position The word and its adjacent words are in a high-frequency phrase. Words The high-frequency phrases to which it belongs are used as the words to be masked; Step c: If the position The word and its adjacent words are not in the high-frequency phrase, and the position is changed to Words As the word to be covered, the other 50% probability will be located at Words the subsequence it belongs to As the word to be covered; Step d: Calculate the number of words to be masked MN, if , then return to step a and continue to randomly select a position. If , then the selection of the masked word is completed; Step e: For the selected words to be masked, there is an 80% probability of directly replacing the original words with the [MASK] mark; there is a 10% probability of replacing the original words with random words; there is a 10% probability of keeping the original words unchanged; Step f: After mixed masking, a word sequence of log text data There are three output , , , where MN words are replaced by the [MASK] tag , the word sequence after MN words are replaced by random words MN words keep the original words unchanged after the word sequence ;in is the first word replaced by the [MASK] tag, is the MNth word replaced by the [MASK] tag, is the first word replaced by a random word, is the MNth word replaced by a random word.

2. According to claim 1, a hybrid masking method in self-supervised training of a log anomaly detection language model is characterized in that: Log Template Library , , is the number of log templates, For the Log templates, , The number of variables / constants. A constant in a log template contains a fixed number of word combinations, and the length of the words contained in the variable is not fixed.

3. According to claim 1, a hybrid masking method in self-supervised training of a log anomaly detection language model is characterized in that: In step 2, for each log text in the log database, it is matched with each log template, and the constants are matched according to the longest match principle. The fields between the constants are regarded as variables, so as to identify the constants and variables contained in the log text.

4. According to claim 1, a hybrid masking method in self-supervised training of a log anomaly detection language model is characterized in that: In step 3, the constant word subsequence library and variable word subsequence library All words in the log are taken out to form a log word library , is the first word in the log word library, The first word in the log word library words, Represents the number of all words in the log text base; ; In the formula It means taking different words from the subsequence and removing the same words.

5. According to claim 1, a hybrid masking method in self-supervised training of a log anomaly detection language model is characterized in that: In step 4, a correspondence between each subsequence and a high-frequency phrase in the constant word subsequence library and the variable word subsequence library is established to obtain the high-frequency phrases contained in each subsequence.

6. According to claim 1, a hybrid masking method in self-supervised training of a log anomaly detection language model is characterized in that: In step 6, the output after mixing mask , , , construct the word embedding matrix formed by each word ,in is the word embedding matrix formed by all words in the log database, is the dimension of the word embedding of a word, is the number of all words in the log database; construct the position encoding matrix Z of each word ,in represents the position encoding matrix formed by each word in the current log text, T represents the number of words in the current log text, is the dimension of the position encoding of a word, a word in the log text Expressed as word embedding and position encoding , [MASK] mark word The word embedding is , Positional encoding , the sum of word embedding and position encoding forms the input representation , , , the input represents the set As input to the Transformer encoder in the BERT model.

7. The hybrid masking method in self-supervised training of a log anomaly detection language model according to claim 6, characterized in that: When the BERT model is trained in a self-supervised manner, the Transformer encoder is used as an input representation set The feature extraction module After the features are obtained, the feature map ,according to Predict the content of the masked word.

8. The hybrid masking method in self-supervised training of a log anomaly detection language model according to claim 6, characterized in that: Replace the self-attention mechanism in the Transformer encoder with a group attention mechanism.

9. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by the processor, the hybrid mask method in the self-supervised training of the log anomaly detection language model described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Method and apparatus of anomaly detection of system logs based on self-supervised learning

    US20240078320A1