Log anomaly prediction method based on prompt enhancement
Through template average sampling and RoBERTa model parsing logs, combined with the LSTM network of time-aware context attention mechanism, the efficiency and accuracy of log anomaly detection in large-scale log data is solved, and end-to-end log anomaly prediction and early exception prompts are realized.
Patent Information
- Application Number
- CN202510552054.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-29
AI Technical Summary
Existing log anomaly detection methods are difficult to efficiently parse and predict exceptions when facing large-scale log data, especially due to the exponential growth of log data scale and the difficulty of extracting log template parameters, resulting in insufficient generalization capabilities of model and high cost of manual labeling.
The log exception prediction method based on prompt enhancement is adopted, and the log is parsed through template average sampling and RoBERTa model, combined with the LSTM network of the time-aware context attention mechanism, the semantics, sequence, time and template parameter characteristics of the log sequence are extracted to achieve end-to-end log exception prediction.
It improves the efficiency and accuracy of log exception detection, reduces manual labeling costs, simplifies the model structure, enhances the prediction ability of log exceptions, and realizes early abnormal prompts.
Smart Images

Figure CN120386686A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of log anomaly detection, specifically a log anomaly prediction method based on prompt enhancement. Background Art
[0002] Log data, as an important resource for system maintenance, has long been widely used by operation and maintenance personnel and researchers. In the era when the scale of logs was small, operation and maintenance personnel mainly relied on manual operations and some basic algorithms to process logs, which could basically meet the needs of system maintenance. However, with the rapid development of science and technology, modern computer systems have undergone huge changes, and the scale of log data has also increased exponentially. For example, Alibaba's cloud computing system can generate about 30 - 50 GB of trace logs per hour. Facing such a large amount of log data, researchers have started to introduce machine learning algorithms, including various algorithms such as deep learning, to improve the efficiency of log analysis and anomaly detection.
[0003] Log anomaly detection methods based on deep learning are mainly divided into three categories: unsupervised, supervised, and semi - supervised anomaly detection methods. These methods identify abnormal situations by using the patterns and features of log data, significantly improving the detection efficiency.
[0004] The log anomaly prediction method based on prompt enhancement also belongs to the above three categories. Different from other log anomaly detection methods that only consider semantic information or sequence information, the log anomaly prediction method based on prompt enhancement completes the entire process of log processing from end to end. Only with the original log data, subsequent parsing and anomaly prediction can be carried out, and when predicting, it comprehensively considers the semantic, sequential, temporal, and template information of the logs, grasps the specificity of the logs, and deeper considers the relationships between logs. Summary of the Invention
[0005] Aiming at the increasingly complex log anomaly detection problem, the present invention aims to propose a log anomaly prediction method based on prompt enhancement. This method uses a deep learning model for log parsing and anomaly prediction, utilizes the powerful ability of the deep learning model to capture the temporal dependence between sequences, extracts the complex connections between log sequences, predicts upcoming anomalies, and discovers log anomalies as early as possible and gives a prompt.
[0006] To solve the above - mentioned technical problems, the present invention provides the following technical solutions:
[0007] A log anomaly prediction method based on prompt enhancement, the method includes:
[0008] Few - shot log parsing based on template average sampling;
[0009] Log representation technology based on the fusion of log multi - features and template parameters;
[0010] Log anomaly prediction model with context attention mechanism integrating time perception.
[0011] Preferably, few-shot log parsing based on template average sampling includes:
[0012] S101. Log preprocessing:
[0013] Logs are usually semi-structured texts used to record the state of the system and important events. Logs usually contain a constant part and a parameter part. The purpose of log parsing is to extract timestamp, level, component, log template, and parameter information from the logs. Among them, timestamp, level, and component information can be easily obtained through simple regular expressions, but the extraction of log templates and parameters is more difficult than the former.
[0014] In the training stage, a small amount of labeled log data is obtained as the training dataset, and the Drain method is used to extract the training log templates; the effectiveness, robustness, and efficiency of this method are better than other existing log parsing methods.
[0015] For each log template, a diverse and evenly distributed sample set is obtained using the template average sampling algorithm, and all the extracted samples are integrated to form the training set D train ; specifically, the algorithm identifies different templates existing in the log data. For each identified template, the algorithm performs equal-proportion random sampling, that is, a fixed number n of samples are extracted from the set of log messages corresponding to each template. This method ensures that each template has enough samples for training, thereby improving the generalization performance and robustness of the model.
[0016] During the sampling process, the algorithm ensures the diversity and even distribution of samples through random selection. In this way, the model bias problem caused by too few samples of some templates can be avoided. Finally, all the extracted samples are integrated to form the training set D train , which is used for subsequent model training and optimization. This method not only improves the accuracy of the model but also reduces the burden caused by the high cost of manual annotation.
[0017] S102. Template parameter embedding generation:
[0018] Pre-trained language models have demonstrated excellent performance in many NLP tasks. These models are usually pre-trained on large-scale unlabeled datasets and then fine-tuned on specific downstream tasks. Recent research results show that these pre-trained models can effectively understand the semantic meaning of log messages, thus providing strong support for various log analysis tasks. Given that RoBERTa is one of the most widely used models, we choose it as the pre-trained model of the present invention. In addition, multiple studies have also confirmed the effectiveness of RoBERTa in the field of log analysis.
[0019] Get the pre-trained language model RoBERTa and set the training set D train The samples in are input into the pre-trained language model RoBERTa for pre-training.
[0020] S103, Template parameter embedding learning: The log message is input into the trained model. The model first tokenizes the input into a set of tokens and then predicts their corresponding target tokens. If a token is predicted to be "TEMPLATE PARAM", it will be integrated into the parameter list. Otherwise, it will remain in the log template.
[0021] Preferably, in S102, the training set D train The samples in are input into the pre-trained language model RoBERTa for pre-training, including:
[0022] Given an input sequence X, use a separator to separate the sequence X={x1,x2,...,x n}, and dynamically expand the RoBERTa vocabulary; predict the template virtual label token "TEMPLATE PARAM" at the log parameter position i of each template through the pre-trained language model RoBERTa; where x i Represents a parameter in the sequence X;
[0023] From the training set Let’s start by using the pre-trained language model RoBERTa to predict the probability distribution of each log parameter position i and obtain the probability distribution of each template tag token “TEMPLATE PARAM”:
[0024] Specifically, each sample (X, Y) is input into RoBERTa to obtain the parameter position x of each log in the log message X. i The probability distribution of x i The first n prediction tokens of are used as the initial parameter indicator set V ini ;
[0025] From the initial label word set V ini Starting from this, we select the top-k high-frequency words for each template as the feature words of the template parameters; calculate each token t∈V ini In the dataset D, each template T i Frequency And select the top-k most frequent words by ranking:
[0026] After obtaining V, the template embedding vector is assigned to the template virtual label token "TEMPLATEPARAM" by calculating the average vector of all tokens in V, and it is added to the language model RoBERTa.
[0027] Preferably, the template parameter embedding learning in S103 includes:
[0028] Given the input log message X = {x1, x2,..., x n}, the target sequence Y = {y1, y2,..., y n} is constructed by replacing the parameter with the template virtual label token "TEMPLATE PARAM" at position j and keeping the original word at the keyword position;
[0029] The conditional probability P(Y|X) of the target sequence Y is maximized by using the trained language model RoBERTa; based on this, the loss function is defined as:
[0030] where K is the number of labeled training samples, x j represents the parameter information at position j in the log message X, and y j represents the parameter information at position j in the target sequence Y.
[0031] During the fine-tuning process, the present invention reuses the entire pre-trained model RoBERTa. The entity-oriented objective is similar to the language model objective based on masked token prediction, which helps to reduce the gap between pre-training and fine-tuning, so that the model retains the knowledge learned by the pre-trained language model.
[0032] During the testing process, the present invention directly inputs the log message into the trained model. The model first tokenizes the input into a set of tokens, and then predicts their corresponding target tokens. If a token is predicted as "TEMPLATEPARAM", it will be integrated into the parameter list; otherwise, it will remain in the log template. The overall framework diagram is as Figure 2 .
[0033] Preferably, the log representation technology based on log multi-feature and template parameter fusion includes:
[0034] S201. Design a prompt-word-embedding-based scheme to tightly combine the semantic, sequential, temporal, and template parameter features of the log to construct template information:
[0035] T = {[semantic], <sem>,[id], <id>,[time], <tm>,
[0036] [param],<TEMPLATE PARAM>}
[0037] Among them, the four prompt words "semantic", "id", "time" and "param" are set as trainable prompt embeddings, so that they can express the core features of log sequence information, namely semantics ( <sem>),order( <id>) Relative time ( <tim>) and template information (<TEMPLATE PARAM> );
[0038] S202. Based on the data characteristics of structured logs, template, ID, and timestamp information are extracted and processed as key features. The two core text features, template and ID, are embedded using the RoBERTa model in the log parsing process. Timestamps are processed as relative times relative to the first log in the window. This method explicitly represents the time differences between different logs, helping the model understand the fluidity of time and the order of events. Template parameter features have their embedded information saved in the log parsing process, and the corresponding template parameter embedding is obtained based on the log ID.
[0039] S203. Sequentially concatenate the various log feature information into a template, and apply position embedding to each part of the content: First, set the template format according to the log feature structure, and concatenate the ID, template content, time, and template parameter features in a fixed order; second, to help the model better capture the relative position and hierarchical structure of each feature, apply position embedding to each part of the template after concatenation;
[0040] Specifically, after concatenation, the ID information, template content, timestamp, and template parameter features are processed through position embedding, allowing the model to clearly identify the location of each feature. Position embedding is not only used to locate each feature, but also combined with the semantics, sequence, time, and template parameter information of the log, giving the model a deeper understanding capability. The overall framework diagram is shown in the figure below. Figure 3 .
[0041] Preferably, the log anomaly prediction model integrating the time-aware contextual attention mechanism includes:
[0042] S301. Use the sliding window method to segment and construct the log data:
[0043] The log sequence is divided into two continuous windows with absolute time ranges, called "observation window" and "prediction window" respectively;
[0044] The data in the observable window serves as the model's input, helping it learn the semantics, sequence, timing, and template parameter characteristics of the logs. The prediction window is a period of time after the observable window, encompassing future log events. The data within this window serves as the model's prediction target, measuring the model's ability to predict future behavior.
[0045] The sliding window slides along the timeline in steps of 1 second, ensuring that the model can learn features at different time points. For example, if the window size is set to 2 seconds and the step size is 1 second, each time the model slides, it obtains a new set of training samples, that is, a new observation window and corresponding prediction window. This approach not only increases the number of training samples and improves the model's generalization ability, but also better captures the contextual information before and after the anomaly.
[0046] S302. Build a log anomaly prediction model that integrates a time-aware contextual attention mechanism. This model is based on a single-layer LSTM structure and introduces a contextual enhancement module to accurately model key behaviors and potential anomalies in log sequences.
[0047] The model inputs the embedded log sequence into the LSTM network to extract its temporal dependency features;
[0048] After extracting the temporal features of LSTM, the model introduces the contextual attention mechanism and the time-aware attention mechanism to further improve the ability to perceive abnormal patterns.
[0049] Preferably, the model in S302 inputs the embedded log sequence into the LSTM network to extract its temporal dependency features, including:
[0050] Assume that the input log sequence is: X = {x1, x2, ..., x T };
[0051] After LSTM processing, the hidden state sequence is obtained: H = {h1,h2,...,h T } = LSTM(X);
[0052] Among them, x t represents the embedding representation of the t-th log, h t is the hidden state at this time step, which contains the timing information of the current and previous logs.
[0053] Preferably, the context attention mechanism is introduced in S302, including:
[0054] Based on the LSTM output of each log, calculate its "importance score" in the entire sequence:
[0055] Set the hidden state of the LSTM output at time t to h t , then the hidden state of each time step is linearly transformed: α t =W α h t +b a ;
[0056] Among them, W α is the weight matrix, b a is the bias term, α t represents the context importance score at this time step;
[0057] Perform Softmax normalization on several importance scores to obtain the normalized weight α t :
[0058]
[0059] where T is the sequence length, and α t represents the attention weight corresponding to time step t; Through this mechanism, the model can identify which logs are more representative in the global context, thus focusing more attention on key events that may be related to anomalies.
[0060] Preferably, a time-aware attention mechanism is introduced in S302, including:
[0061] When calculating the attention weight, this mechanism explicitly considers the time interval information between adjacent logs;
[0062] Set Δt t-1,t represents the time interval between log t and t - 1, then adjust the attention weight by introducing the time interval:
[0063] where λ represents the adjustment coefficient of the time interval, and Δt t-1,t represents the time difference term, reflecting the time interval between adjacent logs;
[0064] After adding the time-aware factor, perform Softmax normalization again to obtain the new time-aware weight
[0065] The time-aware attention mechanism enables the model to automatically identify time anomaly patterns such as sudden changes in event density and changes in time delay, so as to respond more strongly when abnormal behavior occurs;
[0066] Finally, the model will fuse the attention weights of context information and time information:
[0067] Perform weighted summation with the LSTM output to obtain the weighted representation of each log:
[0068] Finally, sum the weighted representations of all time steps and then send them to the fully connected layer:
[0069] Obtain the prediction result y of the prediction window.
[0070] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0071] The present invention uses an algorithm of template average sampling to obtain a training set by average sampling according to the number of templates in the log, and uses the RoBERTa model to obtain the top-k probability distribution of parameter positions in each template. The average vector of them is used as the parameter "TEMPLATE PARAM" of each template to be input into RoBERTa to obtain the ability of log parsing. The present invention first obtains the structured pattern of a small amount of log data through Drain, and then uses the algorithm of template average sampling according to the number of templates M, extracts N logs for each template, and then takes N*M logs as the training set to be input into the model for training. The present invention combines keywords by using semantic, sequential, temporal and template parameter information, and adds position embedding to construct a template vector for each log. The observable window and the predicted log window are divided by a sliding window, and the observable window template sequence vector is input into the LSTM model, and combined with the context attention mechanism to enhance the weight of the log sequence strongly related to the prediction. Finally, the model is trained for log anomaly prediction. The present invention first uses the RoBERTa model to construct log template parameters, obtains the embedding representations of semantics, sequence, time and template parameters of the log through keyword splicing, and combines position embedding to enrich the log embedding representation. Through a single-layer LSTM combined with the context attention mechanism, the embedding representation of the log sequence is directly processed, which not only simplifies the model structure and improves the training efficiency, but also automatically learns the importance weights of different positions in the sequence through the attention mechanism, realizing end-to-end anomaly prediction. Description of the Drawings
[0072] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0073] Figure 1 is a flowchart of the method for log anomaly prediction based on prompt enhancement of the present invention;
[0074] Figure 2 is a framework diagram of a few-shot log parsing based on template average sampling of the present invention;
[0075] Figure 3 is a framework diagram of prompt enhancement technology based on log multi-feature and template parameter fusion of the present invention;
[0076] Figure 4 is a log anomaly prediction model that fuses time-aware context attention of the present invention. Detailed Embodiments
[0077] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0078] Please refer to Figures 1-4 , the present invention provides a technical solution:
[0079] Embodiment 1: The present invention proposes a log anomaly prediction method based on prompt enhancement, aiming to process log data in large-scale systems with deep learning technology, mine potential knowledge and rules therein, so as to meet the requirements of modern operation and maintenance systems for process automation and intelligent decision-making. This method first parses the original logs, extracts log templates, and selects a small number of representative samples from each template through the average sampling algorithm to construct a few-shot training set. For the content at the parameter positions in each sampling template, a word prediction method is used to generate a candidate word set, and the average vector of its top-k predicted words is taken as the vector representation of the parameter position in the template, and then it is added to the vocabulary of the large model for predicting the parameter positions of the logs, thereby realizing the generation of structured logs. On this basis, the semantic features, time information, event order, and template parameter representations of the logs are fused through keyword splicing, and a sliding window strategy is used to divide them into observable windows and prediction windows. The model learns the patterns in the observable windows based on the context attention mechanism, so as to accurately predict the status of the logs in the subsequent prediction windows, and realize the early warning and detection of system anomalies. The specific implementation steps of the log anomaly prediction method based on prompt enhancement proposed by the present invention are as follows:
[0080] Training process of the log anomaly prediction system:
[0081] Step 1: Parse the original unstructured log data into a structured format and extract log template information. The template average sampling algorithm is used to randomly extract n representative logs and their corresponding templates from each log template to construct a training data set.
[0082] Step 2: Input the constructed training set into RoBERTa to predict the parameter positions in the log messages, and obtain the probability distribution of each position in each log being a parameter. Select the top-k predicted words with the highest probability as the initial parameter indication set. Based on this initial word set, select the keywords with higher occurrence frequencies for each template as its representative parameter feature words, calculate the average vector of these feature words and use it as the embedding representation of the virtual label token "TEMPLATE PARAM", and further add it to the model vocabulary.
[0083] Step 3: Under the guidance of the prompt words, fuse multi-dimensional features such as the semantic information, chronological order, timestamps, and template parameters of the logs to construct an enhanced template representation. For each type of the above information, prompt word embeddings are introduced to strengthen the model's perception ability of this information type.
[0084] Step 4: Adopt a sliding window mechanism to divide the log sequence into two consecutive windows with a fixed absolute time span along the time axis, namely the "observable window" and the "prediction window".
[0085] Step 5: Construct an LSTM network model integrating time-aware attention. This model takes the log embedding sequence in the "observable window" as the input. First, a single-layer LSTM extracts the temporal features. Subsequently, an attention mechanism is introduced to generate time-aware attention weights by combining the time interval information, which is used to highlight the key positions related to anomalies. Finally, the abnormal prediction results for the "prediction window" are output through a fully connected layer.
[0086] Usage process of the log anomaly prediction system:
[0087] Step 1: Parse the incoming unstructured logs, extract the log templates, predict the parameter positions, obtain the template parameter vectors, add them to the model vocabulary, and train the model to predict the templates.
[0088] Step 2: Fuse multi-dimensional features such as the semantic information, chronological order, timestamps, and template parameters of the logs, and use positional embeddings as the input. Use a sliding window to divide the "observable window" and the "prediction window".
[0089] Step 3: Input the "observable window" into the log anomaly prediction model, and the model will output the prediction results for the "prediction window" according to the input log sequence.
[0090] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.< / tim> < / id> < / sem> < / tm> < / id> < / sem>
Claims
1. A log anomaly prediction method based on prompt enhancement, characterized in that: The method includes: Few-shot log parsing based on template average sampling; Log representation technology based on the fusion of log multivariate features and template parameters; Log anomaly prediction model integrating time-aware context attention mechanism.
2. The method for predicting log anomalies based on prompt enhancement according to claim 1, wherein The few-shot log parsing based on template average sampling includes: S101. Obtain a small amount of labeled log data as a training data set, and use the Drain method to extract training log templates; for each log template, use the template average sampling algorithm to obtain a diverse and evenly distributed sample set, and integrate all the extracted samples to form a training set D train ; S'102. Obtain the pre-trained language model RoBERTa, and input the samples in the training set D train into the pre-trained language model RoBERTa for pre-training; S103. Template parameter embedding learning: Input the log message into the trained model. The model first tokenizes the input into a set of tokens and then predicts their corresponding target tokens. If a token is predicted as "TEMPLATE PARAM", it will be integrated into the parameter list; otherwise, it will remain in the log template.
3. The method for predicting log anomalies based on prompt enhancement according to claim 2, wherein, In step S102, the samples in the training set D train are input into the pre-trained language model RoBERTa for pre-training, including: Given an input sequence X, use a delimiter to separate the sequence X = {x1, x2,..., x n}, and dynamically expand the RoBERTa vocabulary; predict the template virtual label token "TEMPLATE PARAM" at the log parameter position i of each template through the pre-trained language model RoBERTa; where x i represents a certain parameter in the sequence X; From the training set Starting from it, use the pre-trained language model RoBERTa to predict the probability distribution of each log parameter position i, and obtain the probability distribution of each template label token "TEMPLATE PARAM": Specifically, each sample (X, Y) is input into RoBERTa to obtain the parameter position x of each log in the log message X i 's probability distribution, and x i 's top n predicted tokens are selected as the initial parameter indication set V ini ; Starting from the initial label word set V ini For each template, select the Top-k high-frequency words as the feature words for the parameters of that template; calculate for each token t ∈ V ini in the dataset D for each template T i frequency and select the most Top-k frequent words by ranking: After obtaining V, assign a template embedding vector to the template virtual label token "TEMPLATEPARAM" by calculating the average vector of all tokens in V and add it to the language model RoBERTa.
4. The method for predicting log anomalies based on prompt enhancement according to claim 2, wherein, The template parameter embedding learning in S103 includes: Given the input log message X = {x1, x2,..., x n}, the target sequence Y = {y1, y2,..., y n} is constructed by replacing the parameter with the template dummy label token "TEMPLATE PARAM" at position j and keeping the original words at the keyword positions; Maximize the conditional probability P(Y∣X) of the target sequence Y using the trained language model RoBERTa; based on this, the loss function is defined as: where K is the number of labeled training samples, x j represents the parameter information at position j in the log message X, and y j represents the parameter information at position j in the target sequence Y.
5. The method for predicting log anomalies based on prompt enhancement according to claim 1, wherein, The log representation technology based on the fusion of log multivariate features and template parameters includes: S201. Design a prompt-word embedding-based scheme to tightly combine the semantic, sequential, temporal, and template parameter features of the log and construct template information: T = {[semantic], <sem>,[id], <id>,[time], <tim> ,< / tim> < / id> < / sem> [param],<TEMPLATE PARAM>} Among them, the four cue words "semantic", "id", "time", and "param" are set as trainable cue embeddings, enabling them to respectively express the core features of the log sequence information, namely semantics( <sem>)、Sequence( <id>) Relative time( <tim>) and template information (<TEMPLATE PARAM>);< / tim> < / id> < / sem> S202. Based on the data characteristics of structured logs, extract and process the template, ID, and timestamp information as key features. Among them, for the two core text features of the template and ID, use the RoBERTa model in the log parsing part for embedding processing. The timestamp is processed as the relative time with respect to the first log within the window. In this way, the time difference between different logs is explicitly represented, helping the model understand the time fluidity and the sequence of events; S203. Orderly splice the various feature information of the log in the form of a template and apply positional embedding to each part of the content: First, set the template format according to the feature structure of the log, and splice the ID, template content, time, and template parameter features in a fixed order. Second, to help the model better capture the relative positions and hierarchical structures of each feature, apply positional embedding to each part of the content after template splicing.
6. The method for predicting log anomalies based on prompt enhancement according to claim 1, wherein, The log anomaly prediction model integrating time-aware context attention mechanism includes: S301. Use the sliding window method to segment and construct the log data: The log sequence is segmented into two consecutive windows with absolute time ranges, called the "observable window" and the "prediction window" respectively; Among them, the data in the observable window is used as the input of the model to help the model learn the semantic, sequential, temporal, and template parameter features of the log. The prediction window is a time period after the observable window, containing the log events that will occur in the future. The data within this window is used as the prediction target of the model to measure the model's prediction ability for future behaviors; At the same time, the sliding window slides along the time axis with a step size of 1 second to ensure that the features at different time points can be learned by the model; S302. Build a log anomaly prediction model with a context attention mechanism integrated with time awareness: The model inputs the log sequence after embedding representation into the LSTM network to extract its temporal dependence features; After extracting the temporal features of the LSTM, the model introduces a context attention mechanism and a time awareness attention mechanism respectively to further enhance the perception ability of abnormal patterns.
7. The method for predicting log anomalies based on prompt enhancement according to claim 6, wherein In S302, the model inputs the log sequence after embedding representation into the LSTM network to extract its temporal dependence features, including: Set the input log sequence as: X = {x1, x2,..., x T}; Then, after LSTM processing, a hidden state sequence is obtained: H = {h1, h2,..., h T} = LSTM(X); Among them, x t represents the embedding representation of the t-th log, and h t is the hidden state at this time step, which contains the temporal information of the current and previous logs.
8. The method for predicting log anomalies based on hint enhancement according to claim 6, wherein, In S302, introducing a context attention mechanism includes: According to the LSTM output of each log, calculate its "importance score" in the whole sequence: Set the hidden state at the t-th moment of the LSTM output as h t , then perform a linear transformation on the hidden state for each time step: a t = W a h t + b a ; Among them, W a is the weight matrix, b a is the bias term, and a t represents the context importance score at this time step; Softmax normalization is performed on a number of importance scores to obtain the normalized weight α t : where T is the sequence length, and α t represents the attention weight corresponding to time step t.
9. The method for predicting log anomalies based on prompt enhancement according to claim 6, wherein In S302, introducing a time awareness attention mechanism includes: Set Δt t-1,t represents the time interval between log t and t - 1, and the attention weights are adjusted by introducing the time interval: where λ represents the adjustment coefficient of the time interval, and Δt t-1,t represents the time difference term; After adding the time perception factor, perform Softmax normalization again to obtain the new time perception weights The time awareness attention mechanism enables the model to automatically identify time anomaly patterns such as sudden changes in event density and changes in time delay, so as to respond more strongly when abnormal behaviors occur; Finally, the model will fuse the attention weights that incorporate context information and temporal information: Perform weighted summation with the LSTM output to obtain the weighted representation of each log: Finally, sum the weighted representations of all time steps and then feed them into the fully connected layer: Obtain the prediction result y of the prediction window.
Citation Information
Cited By
Application log intelligent inspection method and system based on large model
CN120763004A