Serial port log analysis method and device and storage medium
By converting the serial port logs in the supercomputing system into log vectors, and using multiple machine learning models to calculate confidence scores, combined with sliding window technology, the problems of high false alarm rate and large resource utilization of serial log analysis in the existing technology are solved, and high-precision and low false alarm analysis effect is achieved.
Patent Information
- Application Number
- CN202510638296.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-19
AI Technical Summary
In supercomputing systems, it is difficult for the prior art to analyze serial logs with high accuracy and low false alarms, especially when processing complex long sequence data, it is difficult to capture cross-row abnormal patterns, resulting in limited detection accuracy.
The abnormal state of the log text is determined by converting the log to a log vector and utilizing a pre-trained classic machine learning model and end-to-end network model. This method combines sliding window technology to reduce resource usage.
High-precision analysis of serial port logs is realized, the false alarm rate is reduced, the wide applicability of analysis is improved, and resource utilization is reduced.
Smart Images

Figure CN120162235A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of system operation and maintenance, and particularly to a serial port log analysis method, device, and storage medium. Background Art
[0002] In a large-scale supercomputer system, the serial port cache startup logs of computing nodes are an important source of information for debugging and operation and maintenance. These logs record various types of information during the startup process of computing nodes in a form that does not intrude on the operating system resources of the computing nodes themselves. The serial port logs are stored in an unstructured long text format, and the complete logs of each node startup usually exceed 2000 lines, containing rich semantic information, which is the key to quickly locating system failures and abnormal behaviors.
[0003] Currently, log anomaly detection methods are mainly divided into: traditional machine learning methods and deep learning-based methods. However, machine learning methods perform well on structured or semi-structured log data, but due to their dependence on log parsers, they are still limited by the insufficient processing ability of the parsers for unstructured logs. Although deep learning methods have high detection accuracy, due to the need for large-scale data training and the dependence on high-performance hardware resources, it is difficult for them to be adapted and deployed on the board-level monitoring and management unit of a supercomputer system with limited computing resources. Therefore, for log data related to hardware such as serial port logs in a supercomputer system with unfixed word order, there is a problem of high false alarm rate, and it is difficult to capture cross-line abnormal patterns when processing complex long sequence data, resulting in limited detection accuracy.
[0004] In view of this, the present invention is specifically proposed. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention provides a serial port log analysis method, device, and storage medium, which realizes high-precision and low-false-alarm analysis of serial port logs, reduces the resource occupancy when analyzing serial port logs, and improves the wide applicability of serial port log analysis.
[0006] An embodiment of the present invention provides a serial port log analysis method, which includes:
[0007] Determine a log vector corresponding to each line of log text in the to-be-analyzed log according to the to-be-analyzed log;
[0008] Determine a first confidence score corresponding to each log vector according to each log vector and a pre-trained classical machine learning model, and determine a startup prediction vector from each log vector according to each first confidence score;
[0009] Divide each log vector starting from the startup prediction vector into each prediction window according to a preset prediction window parameter;
[0010] For each prediction window, at least two second confidence scores corresponding to the log vectors adjacent to and after the prediction window are determined according to at least two pre-trained end-to-end network models, the abnormal states of the log texts in each row within the prediction window, and the log vectors.
[0011] According to the first confidence scores corresponding to the respective log vectors and at least two second confidence scores, the abnormal states of the log texts in each row of the log to be analyzed are determined.
[0012] An embodiment of the present invention provides an electronic device, which includes:
[0013] a processor and a memory;
[0014] The processor is configured to execute the steps of the serial port log analysis method according to any one of the embodiments by calling a program or instruction stored in the memory.
[0015] An embodiment of the present invention provides a computer-readable storage medium, which stores a program or instruction, and the program or instruction causes a computer to execute the steps of the serial port log analysis method according to any one of the embodiments.
[0016] The embodiments of the present invention have the following technical effects:
[0017] By determining the log vector corresponding to each log text row in the log to be analyzed according to the log to be analyzed, determining the first confidence score corresponding to each log vector according to each log vector and a pre-trained classical machine learning model, analyzing in the first way, and determining a start prediction vector from each log vector according to each first confidence score to judge when to introduce the second way for analysis, dividing the log vectors starting from the start prediction vector into each prediction window according to the preset prediction window parameters, for each prediction window, determining at least two second confidence scores corresponding to the log vectors adjacent to and after the prediction window according to at least two pre-trained end-to-end network models, the abnormal states of the log texts in each row within the prediction window, and the log vectors, for analyzing in the second way, using a sliding window method to reduce resource occupancy, and determining the abnormal states of the log texts in each row of the log to be analyzed according to the first confidence scores corresponding to the respective log vectors and at least two second confidence scores, achieving the effects of improving log analysis accuracy, reducing resource occupancy, and improving wide applicability by comprehensively using multiple log analysis methods at an appropriate time. Description of the Drawings
[0018] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 is a flowchart of a serial port log analysis method provided by an embodiment of the present invention;
[0020] Figure 2 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific Embodiments
[0021] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present invention.
[0022] The serial port log analysis method provided by the embodiments of the present invention is mainly applicable to the situation of analyzing and processing unstructured serial port logs without a fixed format in a supercomputer system with high precision and low resource consumption. The serial port log analysis method provided by the embodiments of the present invention can be executed by an electronic device.
[0023] Embodiment 1
[0024] Figure 1 is a flowchart of a serial port log analysis method provided by an embodiment of the present invention. Refer to Figure 1 , the serial port log analysis method specifically includes:
[0025] S110. Determine the log vector corresponding to each line of log text in the log to be analyzed according to the log to be analyzed.
[0026] Among them, the log to be analyzed is all or part of the serial port cache startup log of the supercomputer system. The log text is the unstructured long text data of each line in the log to be analyzed. The log vector is the result of vectorizing the log text.
[0027] Specifically, for each line of log text in the log to be analyzed, it is vectorized using Word2vec (Word to Vector, a related model for generating word vectors). First, an initial semantic vector is generated for each line of log text. An inverse document probability weight is assigned to each word in the log to be analyzed. An embedding matrix is constructed by combining the initial semantic vector corresponding to the log text and each inverse document probability weight. Each column is the feature vector obtained by weighting each word contained in a line of log text, that is, the log vector corresponding to that line of log text.
[0028] Based on the above example, before determining the log vector corresponding to each line of log text in the log to be analyzed according to the log to be analyzed, the serial port logs during the startup process of the supercomputer operating system can also be grouped according to different startup process types to obtain each log to be analyzed, which is convenient for subsequent grouped analysis. Specifically, it can be:
[0029] Obtain the serial port logs during the startup process of the supercomputer operating system, and perform data cleaning on the serial port logs during the startup process of the supercomputer operating system to obtain the logs to be grouped;
[0030] Input the logs to be grouped into a pre-trained language model to determine the text features corresponding to the logs to be grouped;
[0031] According to the text features corresponding to the logs to be grouped and a pre-trained classifier, the logs to be grouped are divided into the logs to be analyzed corresponding to each category of serial port logs in each startup process.
[0032] Among them, the serial port logs during the startup process of the supercomputer operating system are the original logs that need to be analyzed. The logs to be grouped are the logs obtained after data cleaning of the serial port logs during the startup process of the supercomputer operating system. The pre-trained language model is an artificial intelligence model used to understand natural language, usually obtaining knowledge through unsupervised learning on a large-scale text corpus. The pre-trained language model can be, for example, BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language model based on the Transformer architecture), etc. The text features are the features obtained by extracting text features from the logs to be grouped using the pre-trained language model. The classifier can be a linear classifier, etc., and is used to group the logs to be grouped according to the categories of serial port logs in each startup stage of the supercomputer system. The category of serial port logs during the startup process is the category of serial port output logs for the startup process of the supercomputer operating system, such as kernel loading, hardware initialization, hardware self-check, etc.
[0033] Specifically, obtain the serial port log during the startup process of the supercomputer operating system, clean the data of the serial port log during the startup process of the supercomputer operating system, and use the cleaned part as the log to be grouped. Input the log to be grouped into the pre-trained language model, and extract the text features of the log for the understanding of general semantics, so as to obtain the text features in the log to be grouped. Furthermore, input the text features corresponding to the log to be grouped into the pre-trained classifier, and group the log to be grouped according to the categories of the serial port log during the startup process, so as to divide the log to be grouped into the logs to be analyzed corresponding to each category of the serial port log during the startup process.
[0034] Among them, in the above process, the calculation complexity can be reduced and the inference process can be accelerated by freezing some parameters in the pre-trained language model, such as some BERT weights, etc.
[0035] Exemplarily, perform a cleaning operation on the serial port log during the startup process of the supercomputer operating system. For example: use regular expressions to remove redundant information in the brackets and their contents in the log, and skip the lines with specific error codes at the same time, and retain the log text lines with key features. For incomplete or short-line text, ensure the continuity and context consistency of the log content by appending it to the previous line. For duplicate log content, filter out important information through the difference comparison technology based on similarity calculation to reduce the processing burden of redundant data.
[0036] S120. According to each log vector and the pre-trained classical machine learning model, determine the first confidence score corresponding to each log vector, and determine the startup prediction vector from each log vector according to each first confidence score.
[0037] Among them, the classical machine learning model can be models such as isolation forest and decision tree that are not neural networks. The first confidence score is the evaluation value of the analysis result of the classical machine learning model corresponding to each line of log vector. The startup prediction vector is the vector determined by the second confidence score when the first confidence scores of the log vectors of consecutive preset lines meet the normal requirements.
[0038] Specifically, input each log vector into the pre-trained classical machine learning model respectively, and the first confidence score corresponding to each log vector can be obtained. Analyze the first confidence scores of consecutive preset quantities in sequence according to the time sequence. If the first confidence scores of consecutive preset quantities all meet the normal requirements (for example, higher than a certain preset score), then stop continuing to judge whether they all meet the normal requirements, and determine that the calculation of the subsequent second confidence score can be started. Therefore, the log vector corresponding to the first first confidence score among these consecutive preset quantities of first confidence scores is determined as the startup prediction vector.
[0039] Based on the above examples, the first confidence score corresponding to each log vector can be determined according to each log vector and a pre-trained classical machine learning model, and the startup prediction vector can be determined from each log vector according to each first confidence score:
[0040] For each log vector, input the log vector into the pre-trained classical machine learning model to determine the deviation gap corresponding to the log vector, and normalize the deviation gap to obtain the first confidence score corresponding to the log vector;
[0041] Take the starting vector in each log vector as the candidate vector. Starting from the candidate vector, according to the window length in the preset prediction window parameter, determine the target window corresponding to the candidate vector, and judge whether the first confidence scores corresponding to each log vector in the target window are all less than the first threshold;
[0042] If so, take the candidate vector as the startup prediction vector;
[0043] If not, update the candidate vector according to the moving step length in the preset prediction window parameter, and return to execute the step of determining the target window corresponding to the candidate vector starting from the candidate vector according to the window length in the preset prediction window parameter until the first confidence scores corresponding to each log vector in the target window are all less than the first threshold or the number of log vectors in the target window is less than the window length.
[0044] Among them, the deviation gap can be the gap from the normal log. The preset prediction window parameter includes the window length and the moving step length. The window length is used to describe the amount of data that can be included in the window, and the moving step length is the distance when moving the window each time. The starting vector is the first vector in each log vector. The target window is the window for judging each first confidence score within the window. The candidate vector is the first vector in the target window. The first threshold is a preset threshold for determining whether the log vector is considered normal when using the classical machine learning model.
[0045] Specifically, for each log vector, the log vector is input into a pre-trained classical machine learning model, and the model can output the deviation gap corresponding to the log vector. Normalizing the deviation gap means normalizing it to between 0 and 1, where 1 indicates an anomaly, and taking the value after normalization as the first confidence score corresponding to the log vector. Using the starting vector in each log vector as a candidate vector, starting from the candidate vector, according to the window length in the preset prediction window parameter, obtain the target window corresponding to the candidate vector, and determine whether the first confidence scores corresponding to the log vectors in the target window are all less than the first threshold. If so, it means that the processing of the second confidence score can be started. Therefore, the candidate vector can be used as the starting prediction vector. If not, it means that the processing of the second confidence score cannot be started yet, and the target window needs to be moved. That is, according to the moving step length in the preset prediction window parameter, move the window, take the first log vector in the moved target window as the new candidate vector, and return to execute the step of determining the target window corresponding to the candidate vector starting from the candidate vector according to the window length in the preset prediction window parameter until the first confidence scores corresponding to the log vectors in the target window are all less than the first threshold or the number of log vectors in the target window is less than the window length.
[0046] S130. Divide each log vector starting from the starting prediction vector into each prediction window according to the preset prediction window parameter.
[0047] Among them, the prediction window is each window obtained by moving starting from the starting prediction vector according to the preset prediction window parameter.
[0048] Specifically, starting from the starting prediction vector, according to the preset prediction window parameter, perform window sliding, and multiple prediction windows can be obtained, or the log vectors within each prediction window can be determined.
[0049] S140. For each prediction window, determine at least two second confidence scores corresponding to the log vector adjacent to and after the prediction window according to at least two pre-trained end-to-end network models, the abnormal state of each line of log text in the prediction window, and the log vector.
[0050] Among them, the end-to-end network model is a deep learning network model. For example, it can be a Transformer prediction model (a deep neural network model based on the self-attention mechanism), a VAE (Variational Autoencoder), etc. The abnormal state includes normal and abnormal. The second confidence score is the result evaluation value for analyzing the log vector calculated through the prediction results corresponding to each end-to-end network model.
[0051] Specifically, for each prediction window, similar processing is performed using each pre-trained end-to-end network model. Therefore, one of them can be taken as an example for illustration. Update the log vectors of each row within the prediction window according to the corresponding anomaly status. For each row of log vectors, if its anomaly status is abnormal, use the prediction vector of this row of log vectors by the previous end-to-end network model; if its anomaly status is normal, directly use this log vector. Input the updated vectors in the prediction window into the pre-trained end-to-end network model to obtain the output result of the end-to-end network model, that is, the prediction vector corresponding to the log vector adjacent to and after the prediction window. Calculate the distance metric value between the two and perform normalization processing to obtain the second confidence score output by the end-to-end network model for the log vector adjacent to and after the prediction window.
[0052] Based on the above example, the at least two second confidence scores corresponding to the log vector adjacent to and after the prediction window can be determined according to at least two pre-trained end-to-end network models, the anomaly status of each row of log texts within the prediction window, and the log vectors in the following manner:
[0053] For each pre-trained end-to-end network model, determine the vectors to be input corresponding to each row of log texts in the prediction window according to the anomaly status of each row of log texts within the prediction window;
[0054] Input each vector to be input into the end-to-end network model to determine the prediction vector corresponding to the prediction window;
[0055] Determine the second confidence score of the log vector adjacent to and after the prediction window according to the prediction vector and the log vector adjacent to and after the prediction window.
[0056] Among them, the anomaly status of each log vector in the prediction window that starts the prediction vector is normal. The vector to be input is the vector used for subsequent prediction by inputting into the end-to-end network model, which may be the corresponding log vector or the prediction vector corresponding to the corresponding log vector. The prediction vector is the result output by the same end-to-end network model for the log vector.
[0057] Specifically, for each pre-trained end-to-end network model, in combination with the abnormal status of each line of log text within the prediction window, it is determined that the input vector to be corresponding to each line of log text in the prediction window is either a log vector or a corresponding prediction vector. That is, if the abnormal status is abnormal, the input vector to be corresponding to this line of log text is the prediction vector output by the end-to-end network model corresponding to this line of log text; if the abnormal status is normal, the input vector to be corresponding to this line of log text is the log vector corresponding to this line of log text. Input the input vectors to be in the prediction window into the end-to-end network model for time series prediction, and the prediction vector corresponding to the prediction window can be obtained. Furthermore, calculate the distance metric value between the prediction vector and the log vector adjacent to the prediction window and located after the prediction window, and normalize the obtained distance metric to the interval 0-1. This normalized value is the second confidence score of the log vector adjacent to the prediction window and located after the prediction window.
[0058] Based on the above example, the input vectors to be corresponding to each line of log text in the prediction window can be determined in the following way according to the abnormal status of each line of log text in the prediction window:
[0059] For each line of log text in the prediction window, when the abnormal status of the log text is normal, use the log vector corresponding to the log text as the input vector to be corresponding to the log text;
[0060] For each line of log text in the prediction window, when the abnormal status of the log text is abnormal, use the prediction vector corresponding to the log text as the input vector to be corresponding to the log text.
[0061] Specifically, for each line of log text in the prediction window, if the abnormal status of this line of log text is normal, the log vector corresponding to this line of log text can be directly used as the corresponding input vector to be; if the abnormal status of this line of log text is abnormal, when judging the abnormal status of this line of log text, use the prediction vector output by the end-to-end network model as the corresponding input vector to be.
[0062] Based on the above example, the second confidence score of the log vector adjacent to the prediction window and located after the prediction window can be determined in the following way according to the prediction vector and the log vector adjacent to the prediction window and located after the prediction window:
[0063] Determine multiple distance metric values according to the prediction vector and the log vector adjacent to the prediction window and located after the prediction window;
[0064] Determine the second confidence score of the log vector adjacent to the prediction window and located after the prediction window according to the multiple distance metric values.
[0065] Specifically, different distance metric calculation methods are used to calculate different distance metrics between the predicted vector and the log vectors adjacent to and after the prediction window. For example, multiple distance metrics can be adopted, such as Euclidean distance, cosine similarity, Hamming distance, etc. The obtained multiple distance metrics are comprehensively processed, such as taking the mean, weighted average, etc., to obtain a comprehensive metric value. The comprehensive metric value is normalized to obtain the second confidence score of the log vector adjacent to and after the prediction window.
[0066] S150. Determine the abnormal status of each line of log text in the log to be analyzed according to the first confidence score corresponding to each log vector and at least two second confidence scores.
[0067] Specifically, for each line of log text, the abnormal status can be judged separately. If the line of log text only has the corresponding first confidence score, the abnormal status is directly judged according to the first confidence score. For example, it can be judged using a threshold, etc., to obtain the abnormal status of the line of log text; if the line of log text has the corresponding first confidence score and at least two corresponding second confidence scores, the first confidence score and each second confidence score are combined. For example, taking the mean, weighted average, or using a pre-trained model for calculation, etc., to obtain a comprehensive confidence score, and the abnormal status is judged according to the comprehensive confidence score. For example, it can be judged using a threshold, etc., to obtain the abnormal status of the line of log text. Thus, the abnormal status of each line of log text in the log to be analyzed can be obtained.
[0068] It should be noted that in this example, the number of pre-trained classical machine learning models is at least one, and the number of pre-trained end-to-end network models is at least two. The specific number can be determined according to the computing power that the executor can bear and is not limited here.
[0069] Based on the above example, the abnormal status of each line of log text in the log to be analyzed can be determined by the following method according to the first confidence score corresponding to each log vector and at least two second confidence scores:
[0070] For each log vector, when each second confidence score corresponding to the log vector is empty, determine the abnormal status of the log text corresponding to the log vector according to the first confidence score corresponding to the log vector and the first threshold;
[0071] For each log vector, when each second confidence score corresponding to the log vector is not empty, determine the comprehensive confidence score according to the first confidence score corresponding to the log vector, each second confidence score, and the pre-trained logistic regression model, and determine the abnormal status of the log text corresponding to the log vector according to the comprehensive confidence score corresponding to the log vector and the second threshold.
[0072] Among them, the first threshold is a pre-set threshold for determining whether the log vector is considered normal when using a classical machine learning model. The comprehensive confidence score is a value calculated by using a pre-trained logistic regression model for the first confidence score and each second confidence score. The second threshold is a pre-set threshold for determining whether the log vector corresponding to the comprehensive confidence score is normal.
[0073] Specifically, for each log vector, it is judged whether each second confidence score corresponding to the log vector is empty. If so, the abnormal state corresponding to the log vector can be determined only according to the corresponding first confidence score. The first confidence score is compared with the first threshold. If the first confidence score is greater than or equal to the first threshold, it is determined that the abnormal state corresponding to the log vector is abnormal; otherwise, it is determined that the abnormal state corresponding to the log vector is normal. If not, the first confidence score corresponding to the log vector and each second confidence score are input into the pre-trained logistic regression model, the comprehensive confidence score is output, and the comprehensive confidence score is compared with the second threshold. If the comprehensive confidence score is greater than or equal to the second threshold, it is determined that the abnormal state corresponding to the log vector is abnormal; otherwise, it is determined that the abnormal state corresponding to the log vector is normal.
[0074] Based on the above example, after determining the abnormal states of the log texts in each line of the log to be analyzed, the abnormal state of the entire log to be analyzed can be further judged. Specifically, it can be:
[0075] If there are at least a second number of log texts with abnormal states among at least a first number of consecutive log texts in the log to be analyzed, it is determined that the abnormal state of the log to be analyzed is abnormal;
[0076] If there are no at least a second number of log texts with abnormal states in each of the at least a first number of consecutive log texts in the log to be analyzed, it is determined that the abnormal state of the log to be analyzed is normal.
[0077] Among them, the first number of lines is a pre-set range of a certain number of lines for detection. The second number of lines is a pre-set number of lines for judging log anomalies.
[0078] Specifically, the log to be analyzed is grouped according to the first number of lines. If the number of lines of the log texts with abnormal states in each of the at least a first number of consecutive log texts is less than the second number of lines, it is determined that the abnormal state of the log to be analyzed is normal. If there is at least one group of at least a first number of consecutive log texts in which the number of lines of the log texts with abnormal states is greater than or equal to the second number of lines, it means that there are multiple lines of anomalies, and it can be determined that the abnormal state of the log to be analyzed is abnormal.
[0079] Exemplarily, first, data preprocessing is performed on the serial port logs during the startup process of the supercomputer operating system. Cleaning operations can be carried out, such as using regular expressions to remove redundant information in the brackets and their contents (such as "(0x0001)") in the serial port logs during the startup process of the supercomputer operating system. At the same time, lines with specific error codes (such as "0x0000000080000000") are skipped, and key feature lines are retained. For incomplete or short-line texts, the continuity and context consistency of the log content are ensured by appending them to the previous line. For duplicate log content, important information is filtered out through a difference comparison technique based on similarity calculation to reduce the processing burden of redundant data. After the cleaning operation, the logs to be grouped are obtained. For the log semantic grouping based on the pre-trained language model, the logs to be grouped can be grouped using the pre-trained language model, and the text features of the logs are extracted according to the understanding of the general semantics by the pre-trained language model. Furthermore, a linear classifier is used to divide the logs into multiple independent groups (logs to be analyzed), and each group implicitly refers to the serial port output of a stage of the system startup (such as kernel loading, hardware initialization, hardware self-check, etc.), that is, the category of the serial port logs during the startup process. During this process, the weights of some pre-trained language models are frozen to reduce the computational complexity and accelerate the inference process. Then, vector representation is performed on the logs to be analyzed respectively. An initial semantic vector corresponding to each line of log text in the logs to be analyzed is generated. Weight assignment also needs to be carried out, that is, an inverse document frequency (IDF) is assigned to each word, and then the IDF weight corresponding to each word can be determined: , where w i represents the i-th word, and n i represents the number of documents containing the word w iThe number of logs, the total number of N logs. The purpose of using the IDF weight is to enhance the importance of low-frequency events (such as abnormal logs). Further, combining the generated initial semantic vectors and the IDF weights, an embedding matrix is constructed. Each column is a feature vector obtained by weighting the initial semantic vector corresponding to a log text by the words it contains, that is, the log vector. Finally, anomaly detection based on heterogeneous network ensemble learning. It can be three parallel expert networks (two end-to-end network models and a classical machine learning model), processing the log vectors completed by embedding from the preprocessing stage. To avoid the built-in bias of a single expert network's structure fitting the training data, the expert networks adopt heterogeneous structures. For example, expert network 1 is a transformer decoder structure. To reduce computational overhead, it uses 1 self-attention and a feed-forward neural network layer; expert network 2 is a variational encoder structure; expert network 3 uses an isolation forest that is not a neural network and uses a tree structure to identify log anomalies. Both expert network 1 and expert network 2 adopt anomaly detection based on reconstruction error. A sliding window with a size of x rows of logs and a step size of 1 can be set (the prediction window corresponding to the preset prediction window parameters). The log vectors of each row within the window are used to predict the next row of log vector outside the window (the prediction vector) to capture the context relationship. The differences between the predicted vectors of the next row predicted by the transformer and the VAE and the log vectors corresponding to the actual current row of log text are compared respectively. Specifically, three distance metric values (such as Euclidean distance, cosine similarity, Hamming distance) can be used. Therefore, the results of integrating each distance metric value of expert network 1 and expert network 2 are respectively used as the second confidence score for the log text of the current row (normalized to the interval 0-1, where 1 indicates being confirmed as an anomaly). Expert network 3 detects anomalies by detecting the deviation of the log vector of the current row from the training data distribution, and uses the normalized value of the deviation gap as the confidence score for determining that the log text of the current row is abnormal, that is, the first confidence score. The output confidence scores of the three expert networks are used as the input of a logistic regression model to calculate the comprehensive confidence score to judge the anomaly status corresponding to each row of log text:
[0080]
[0081] where, f overall is the comprehensive confidence score, f IForest is the first confidence score, f trans and f VAE are the two second confidence scores, x i is the log vector of the i-th row, corresponding to the log text of the i-th row. α1, α2, and α3 respectively represent the weights for different expert networks, w is the weight term, b is the bias term. Among them, w, b, α1, α2, and α3 are numerically obtained through pre-training and can be used directly after being pre-configured during use without re-training.
[0082] Based on the detection of the abnormal status of each single-line log text, if the abnormal status of the log lines exceeding a certain threshold number (the second line number) within a certain range (the first line of the book) is abnormal, it is determined that the abnormal status of the corresponding log to be analyzed is abnormal. This is because there is a chain reaction in a single abnormal log text that appears during the startup process of the supercomputer system. If only one log text is abnormal, it may be a false alarm, and multiple abnormalities are required to analyze and judge the abnormal status of the log to be analyzed.
[0083] In the above example, by integrating multiple heterogeneous lightweight expert models, a balance between low false alarms and high precision is achieved. Through the semantic chunking strategy (log semantic grouping based on pre-trained language models), the data input scale is significantly reduced. At the same time, combined with the sliding window technique, the memory requirement is reduced. The dynamic attention mechanism and lightweight feature extraction optimize the computing efficiency, shortening the training time and detection time, which is suitable for resource-constrained environments. The log semantic grouping based on pre-trained language models can automatically parse the semantic features in the logs and adapt to various log formats, especially unstructured long text data (such as the startup logs of the supercomputer system serial port buffer). Through the max pooling and average pooling strategies of semantic chunking by pre-trained language models, different types of log data can be efficiently handled.
[0084] The present invention has the following technical effects: By determining the log vector corresponding to each log text line in the log to be analyzed according to the log to be analyzed, and determining the first confidence score corresponding to each log vector according to each log vector and a pre-trained classical machine learning model, to analyze in the first way, and according to each first confidence score, determining the startup prediction vector from each log vector to judge when to introduce the second way for analysis. The log vectors starting from the startup prediction vector are divided into each prediction window according to the preset prediction window parameters. For each prediction window, according to at least two pre-trained end-to-end network models, the abnormal status of each log text line in the prediction window, and the log vector, determining at least two second confidence scores corresponding to the log vector adjacent to and behind the prediction window for analysis in the second way. Using the sliding window method reduces resource occupancy. According to the first confidence score corresponding to each log vector and at least two second confidence scores, determining the abnormal status of each log text line in the log to be analyzed, and integrating multiple log analysis methods at the appropriate time, achieving the effects of improving the log analysis accuracy, reducing resource occupancy, and enhancing the wide applicability.
[0085] Embodiment 2
[0086] Figure 2 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 2 shown, the electronic device 200 includes one or more processors 201 and a memory 202.
[0087] The processor 201 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 200 to perform desired functions.
[0088] The memory 202 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 201 can run the program instructions to implement the serial port log analysis method of any embodiment of the present invention described above and / or other desired functions. Various contents such as initial external parameters and thresholds can also be stored in the computer-readable storage media.
[0089] In one example, the electronic device 200 can further include: an input device 203 and an output device 204, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). The input device 203 can include, for example, a keyboard, a mouse, etc. The output device 204 can output various information to the outside, including warning prompt information, braking force, etc. The output device 204 can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0090] Of course, for simplicity, Figure 2 only some of the components related to the present invention in the electronic device 200 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device 200 can further include any other appropriate components.
[0091] In addition to the above methods and devices, an embodiment of the present invention can also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps of the serial port log analysis method provided by any embodiment of the present invention.
[0092] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present invention. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0093] In addition, an embodiment of the present invention may also be a computer-readable storage medium storing computer program instructions, which, when run by a processor, cause the processor to execute the steps of the serial port log analysis method provided by any embodiment of the present invention.
[0094] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0095] It should be noted that the terms used in the present invention are only for describing specific embodiments and do not limit the scope of the present application. As shown in the specification of the present invention, unless the context clearly indicates an exception, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method or device comprising the element.
[0096] It should also be noted that the orientation or positional relationship indicated by terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. Unless otherwise clearly specified and defined, terms such as "installed", "connected", "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A serial port log analysis method, characterized in that: include: Determine, according to the log to be analyzed, a log vector corresponding to each line of log text in the log to be analyzed; Determine a first confidence score corresponding to each log vector according to each log vector and a pre-trained classic machine learning model, and determine a startup prediction vector from each log vector according to each first confidence score; Dividing each log vector starting from the startup prediction vector into prediction windows according to preset prediction window parameters; For each prediction window, determine at least two second confidence scores corresponding to the log vectors adjacent to and after the prediction window according to at least two pre-trained end-to-end network models, the abnormal state of each line of log text in the prediction window, and the log vector; The abnormal state of each line of log text in the log to be analyzed is determined according to the first confidence score corresponding to each log vector and at least two second confidence scores.
2. The method according to claim 1, characterized in that Before determining the log vector corresponding to each line of log text in the log to be analyzed according to the log to be analyzed, the method further includes: Obtaining a serial port log of a supercomputer operating system startup process, performing data cleaning on the serial port log of the supercomputer operating system startup process, and obtaining a log to be grouped; Inputting the logs to be grouped into a pre-trained language model to determine text features corresponding to the logs to be grouped; According to the text features corresponding to the logs to be grouped and the pre-trained classifier, the logs to be grouped are divided into logs to be analyzed corresponding to each startup process serial port log category.
3. The method according to claim 1, characterized in that The step of determining a first confidence score corresponding to each log vector according to each log vector and a pre-trained classical machine learning model, and determining a startup prediction vector from each log vector according to each first confidence score, includes: For each log vector, input the log vector into a pre-trained classical machine learning model, determine the deviation gap corresponding to the log vector, and normalize the deviation gap to obtain a first confidence score corresponding to the log vector; Taking the starting vector in each log vector as a candidate vector, taking the candidate vector as the starting point, determining a target window corresponding to the candidate vector according to a window length in a preset prediction window parameter, and judging whether the first confidence scores corresponding to each log vector in the target window are all less than a first threshold; If yes, the candidate vector is used as the start prediction vector; If not, the candidate vector is updated according to the moving step in the preset prediction window parameter, and the step of determining the target window corresponding to the candidate vector based on the window length in the preset prediction window parameter and taking the candidate vector as the starting point is returned to execute until the first confidence scores corresponding to the log vectors in the target window are all less than the first threshold or the number of log vectors in the target window is less than the window length.
4. The method according to claim 1, characterized in that: The determining, based on at least two pre-trained end-to-end network models, the abnormal state of each line of log text in the prediction window, and the log vector, at least two second confidence scores corresponding to the log vector adjacent to the prediction window and located after the prediction window includes: For each pre-trained end-to-end network model, according to the abnormal state of each line of log text in the prediction window, determine the to-be-input vector corresponding to each line of log text in the prediction window; Inputting each vector to be input into the end-to-end network model to determine a prediction vector corresponding to the prediction window; Determine a second confidence score of the log vector adjacent to the prediction window and located after the prediction window based on the prediction vector and the log vector adjacent to the prediction window and located after the prediction window; Among them, the first log vector is the abnormal state of each log vector in the prediction window of the start prediction vector, and all of them are normal.
5. The method according to claim 4, characterized in that The step of determining, according to the abnormal state of each line of log text in the prediction window, a to-be-input vector corresponding to each line of log text in the prediction window comprises: For each line of log text in the prediction window, when the abnormal state of the log text is normal, using the log vector corresponding to the log text as the to-be-input vector corresponding to the log text; For each line of log text in the prediction window, when the abnormal state of the log text is abnormal, the prediction vector corresponding to the log text is used as the to-be-input vector corresponding to the log text.
6. The method according to claim 4, characterized in that The step of determining a second confidence score of a log vector adjacent to and behind the prediction window according to the prediction vector and a log vector adjacent to and behind the prediction window comprises: Determining a plurality of distance metric values according to the prediction vector and a log vector adjacent to the prediction window and located after the prediction window; A second confidence score of the log vector adjacent to the prediction window and located after the prediction window is determined based on a plurality of distance metric values.
7. The method according to claim 1, characterized in that The determining, according to the first confidence score corresponding to each log vector and at least two second confidence scores, the abnormal state of each line of log text in the log to be analyzed includes: For each log vector, when each second confidence score corresponding to the log vector is empty, determining the abnormal state of the log text corresponding to the log vector according to the first confidence score corresponding to the log vector and a first threshold; For each log vector, when the second confidence scores corresponding to the log vector are not empty, a comprehensive confidence score is determined according to the first confidence score corresponding to the log vector, the second confidence scores and a pre-trained logistic regression model, and the abnormal state of the log text corresponding to the log vector is determined according to the comprehensive confidence score corresponding to the log vector and a second threshold.
8. The method according to claim 1, characterized in that After determining the abnormal state of each line of log text in the log to be analyzed, the method further includes: If, in at least one of the first consecutive lines of log text in the log to be analyzed, at least a second number of lines of log text have an abnormal state, determining that the abnormal state of the log to be analyzed is abnormal; If, in each of the first consecutive lines of log text in the log to be analyzed, there is no at least a second number of lines of log text whose abnormal state is abnormal, it is determined that the abnormal state of the log to be analyzed is normal.
9. An electronic device, characterized in that: The electronic device comprises: Processor and memory; The processor is used to execute the steps of the serial port log analysis method according to any one of claims 1 to 8 by calling the program or instruction stored in the memory.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program or instruction, and the program or instruction enables a computer to execute the steps of the serial port log analysis method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Abnormal behavior detection method and device, electronic equipment and readable storage medium
CN112149749A
Vehicle log anomaly detection method and device, electronic equipment and storage medium
CN117093443A
Service resource sharing method and service resource sharing system based on cloud computing
CN119052335A
Log data analysis method and device, equipment and storage medium
CN119292811A
Log error prediction method and electronic device based on random forest algorithm
CN119759716A