Serial Port Log Analysis Method, Device and Storage Medium
By using the combined analysis method of Word2vec model and multiple network models in the supercomputing system, the problem of high precision and low resource occupation of unstructured serial logs is solved, and efficient log analysis in the supercomputing system is realized.
Patent Information
- Application Number
- CN202510638296.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-19
AI Technical Summary
In supercomputing systems, it is difficult for the prior art to analyze unstructured serial port logs with limited computing resources, and there is a problem of high false alarm rate.
The Word2vec model is used for log vectorization, combining machine learning model and end-to-end network model, and through sliding window technology, a variety of analysis methods are combined to determine the abnormal state of the log text.
It realizes improving log analysis accuracy, reducing false positive rate under low resource usage, and improving the wide applicability of analysis.
Smart Images

Figure CN120162235B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of system operation and maintenance, and particularly to a serial port log analysis method, device, and storage medium. Background Art
[0002] In a large-scale supercomputer system, the serial port buffer startup logs of computing nodes are an important source of information for debugging and operation and maintenance. These logs record various types of information during the startup process of computing nodes in a form that does not intrude on the operating system resources of the computing nodes themselves. The serial port logs are stored in an unstructured long text format, and the complete logs of each node startup usually exceed 2000 lines, containing rich semantic information, which is the key to quickly locating system failures and abnormal behaviors.
[0003] Currently, log anomaly detection methods are mainly divided into: traditional machine learning methods and deep learning-based methods. However, machine learning methods perform well on structured or semi-structured log data, but due to the dependence on log parsers, they are still limited by the insufficient processing ability of the parsers for unstructured logs. Although deep learning methods have high detection accuracy, due to the need for large-scale data training and the dependence on high-performance hardware resources, it is difficult for them to be adapted and deployed on the board-level monitoring and management unit of a supercomputer system with limited computing resources. Therefore, for log data related to hardware such as serial port logs in a supercomputer system with unfixed word order, there is a problem of high false alarm rate, and it is difficult to capture cross-line abnormal patterns when processing complex long sequence data, resulting in limited detection accuracy.
[0004] In view of this, the present invention is specifically proposed. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a serial port log analysis method, device, and storage medium, which realizes high-precision and low false alarm analysis of serial port logs, reduces the resource occupation when analyzing serial port logs, and improves the wide applicability of serial port log analysis.
[0006] An embodiment of the present invention provides a serial port log analysis method, which includes:
[0007] Determine a log vector corresponding to each line of log text in the log to be analyzed according to the log to be analyzed;
[0008] Determine a first confidence score corresponding to each log vector according to each log vector and a pre-trained machine learning model, and determine a startup prediction vector from each log vector according to each first confidence score;
[0009] Divide each log vector starting from the startup prediction vector into each prediction window according to a preset prediction window parameter;
[0010] For each prediction window, at least two second confidence scores corresponding to the log vectors adjacent to and after the prediction window are determined according to at least two pre-trained end-to-end network models, the abnormal states of the log texts in each row within the prediction window, and the log vectors.
[0011] According to the first confidence scores corresponding to the respective log vectors and the at least two second confidence scores, the abnormal states of the log texts in each row of the log to be analyzed are determined.
[0012] An embodiment of the present invention provides an electronic device, which includes:
[0013] A processor and a memory;
[0014] The processor is configured to execute the steps of the serial port log analysis method according to any one of the embodiments by calling a program or instruction stored in the memory.
[0015] An embodiment of the present invention provides a computer-readable storage medium that stores a program or instruction, and the program or instruction causes a computer to execute the steps of the serial port log analysis method according to any one of the embodiments.
[0016] The embodiments of the present invention have the following technical effects:
[0017] By determining the log vector corresponding to each log text row in the log to be analyzed according to the log to be analyzed, determining the first confidence score corresponding to each log vector according to each log vector and a pre-trained machine learning model, and analyzing in the first way, and determining a start prediction vector from each log vector according to each first confidence score to determine when to introduce the second way for analysis, dividing the log vectors starting from the start prediction vector into each prediction window according to the preset prediction window parameters, for each prediction window, determining at least two second confidence scores corresponding to the log vectors adjacent to and after the prediction window according to at least two pre-trained end-to-end network models, the abnormal states of the log texts in each row within the prediction window, and the log vectors, and analyzing in the second way, using the sliding window method to reduce resource occupancy, and determining the abnormal states of the log texts in each row of the log to be analyzed according to the first confidence scores corresponding to the respective log vectors and the at least two second confidence scores, and integrating multiple log analysis methods at an appropriate time, achieving the effects of improving log analysis accuracy, reducing resource occupancy, and improving wide applicability. Description of the Drawings
[0018] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for the description of the specific embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0019] Figure 1 is a flowchart of a serial port log analysis method provided by an embodiment of the present invention;
[0020] Figure 2 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific Embodiments
[0021] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0022] The serial port log analysis method provided by the embodiments of the present invention is mainly applicable to the situation of analyzing and processing unstructured serial port logs without a fixed format in a supercomputer system with high precision and low resource consumption. The serial port log analysis method provided by the embodiments of the present invention can be executed by an electronic device.
[0023] Embodiment 1
[0024] Figure 1 is a flowchart of a serial port log analysis method provided by an embodiment of the present invention. Refer to Figure 1 , the serial port log analysis method specifically includes:
[0025] S110. Determine the log vector corresponding to each line of log text in the log to be analyzed according to the log to be analyzed.
[0026] Among them, the log to be analyzed is all or part of the serial port cache startup log of the supercomputer system. The log text is the unstructured long text data of each line in the log to be analyzed. The log vector is the result of vectorizing the log text.
[0027] Specifically, for each line of log text in the log to be analyzed, it is vectorized using Word2vec (Word to Vector, a related model for generating word vectors). First, an initial semantic vector is generated for each line of log text. An inverse document probability weight is assigned to each word in the log to be analyzed. An embedding matrix is constructed by combining the initial semantic vector corresponding to the log text and each inverse document probability weight. Each column is the feature vector obtained by weighting the words contained in a line of log text, that is, the log vector corresponding to that line of log text.
[0028] Based on the above example, before determining the log vector corresponding to each line of log text in the log to be analyzed according to the log to be analyzed, the serial port logs during the startup process of the supercomputer operating system can also be grouped according to different startup process types to obtain each log to be analyzed, which is convenient for subsequent grouped analysis. Specifically, it can be:
[0029] Obtain the serial port logs during the startup process of the supercomputer operating system, and perform data cleaning on the serial port logs during the startup process of the supercomputer operating system to obtain the logs to be grouped;
[0030] Input the logs to be grouped into a pre-trained language model to determine the text features corresponding to the logs to be grouped;
[0031] According to the text features corresponding to the logs to be grouped and a pre-trained classifier, the logs to be grouped are divided into the logs to be analyzed corresponding to each category of serial port logs in each startup process.
[0032] Among them, the serial port logs during the startup process of the supercomputer operating system are the original logs that need to be analyzed. The logs to be grouped are the logs obtained after data cleaning of the serial port logs during the startup process of the supercomputer operating system. The pre-trained language model is an artificial intelligence model for understanding natural language, usually obtaining knowledge through unsupervised learning on a large-scale text corpus. The pre-trained language model can be, for example, BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language model based on the Transformer architecture), etc. The text features are the features obtained by extracting text features from the logs to be grouped using the pre-trained language model. The classifier can be a linear classifier, etc., and is used to group the logs to be grouped according to the categories of serial port logs in each startup stage of the supercomputer system. The category of serial port logs during the startup process is the category of serial port output logs for the startup process of the supercomputer operating system, such as kernel loading, hardware initialization, hardware self-check, etc.
[0033] Specifically, obtain the serial port log during the startup process of the supercomputer operating system, clean the data of the serial port log during the startup process of the supercomputer operating system, and use the cleaned part as the log to be grouped. Input the log to be grouped into the pre-trained language model, and extract the text features of the log for the understanding of general semantics, so as to obtain the text features in the log to be grouped. Furthermore, input the text features corresponding to the log to be grouped into the pre-trained classifier, and group the log to be grouped according to the category of the serial port log during the startup process, so as to divide the log to be grouped into the logs to be analyzed corresponding to each category of the serial port log during the startup process.
[0034] Among them, in the above process, the calculation complexity can be reduced and the inference process can be accelerated by freezing some parameters in the pre-trained language model, such as some BERT weights.
[0035] Exemplarily, perform cleaning operations on the serial port log during the startup process of the supercomputer operating system. For example, use regular expressions to remove redundant information in the brackets and their contents in the log, skip the lines with specific error codes at the same time, and retain the log text lines with key features. For incomplete or short-line texts, ensure the continuity and context consistency of the log content by appending them to the previous line. For duplicate log contents, filter out important information through the difference comparison technology based on similarity calculation to reduce the processing burden of redundant data.
[0036] S120. According to each log vector and the pre-trained machine learning model, determine the first confidence score corresponding to each log vector, and determine the startup prediction vector from each log vector according to each first confidence score.
[0037] Among them, the machine learning model can be models such as Isolation Forest and Decision Tree that are not neural networks. The first confidence score is the analysis result evaluation value of the machine learning model corresponding to each row of log vectors. The startup prediction vector is the vector determined by the second confidence score when the first confidence scores of the log vectors of consecutive preset lines meet the normal requirements.
[0038] Specifically, input each log vector into the pre-trained machine learning model respectively, and the first confidence score corresponding to each log vector can be obtained. Analyze the first confidence scores of consecutive preset quantities in sequence according to the time sequence. If the first confidence scores of consecutive preset quantities all meet the normal requirements (for example, higher than a certain preset score), stop continuing to judge whether they all meet the normal requirements, and determine that the calculation of the subsequent second confidence score can be started. Therefore, the log vector corresponding to the first first confidence score among these consecutive preset quantities of first confidence scores is determined as the startup prediction vector.
[0039] Based on the above example, the following method can be used to determine the first confidence score corresponding to each log vector according to each log vector and a pre-trained machine learning model, and determine the startup prediction vector from each log vector according to each first confidence score:
[0040] For each log vector, input the log vector into the pre-trained machine learning model to determine the deviation gap corresponding to the log vector, and normalize the deviation gap to obtain the first confidence score corresponding to the log vector;
[0041] Use the starting vector in each log vector as the candidate vector. Starting from the candidate vector, determine the target window corresponding to the candidate vector according to the window length in the preset prediction window parameter, and judge whether the first confidence scores corresponding to each log vector in the target window are all less than the first threshold;
[0042] If so, use the candidate vector as the startup prediction vector;
[0043] If not, update the candidate vector according to the moving step length in the preset prediction window parameter, and return to execute the step of determining the target window corresponding to the candidate vector starting from the candidate vector according to the window length in the preset prediction window parameter until the first confidence scores corresponding to each log vector in the target window are all less than the first threshold or the number of log vectors in the target window is less than the window length.
[0044] Among them, the deviation gap can be the gap from normal logs. The preset prediction window parameter includes the window length and the moving step length. The window length is used to describe the amount of data that can be included in the window, and the moving step length is the distance when moving the window each time. The starting vector is the first vector in each log vector. The target window is the window for judging each first confidence score within the window. The candidate vector is the first vector in the target window. The first threshold is a preset threshold for determining whether the log vector is considered normal when using the machine learning model.
[0045] Specifically, for each log vector, the log vector is input into a pre-trained machine learning model, and the model can output the deviation gap corresponding to the log vector. Normalizing the deviation gap means normalizing it to between 0 and 1, where 1 represents an anomaly, and taking the value after normalization as the first confidence score corresponding to the log vector. Taking the starting vector in each log vector as the candidate vector, starting from the candidate vector, according to the window length in the preset prediction window parameter, obtaining the target window corresponding to the candidate vector, and determining whether the first confidence scores corresponding to the log vectors in the target window are all less than the first threshold. If so, it indicates that the processing of the second confidence score can be started. Therefore, the candidate vector can be used as the starting prediction vector. If not, it indicates that the processing of the second confidence score cannot be started yet, and the target window needs to be moved. That is, according to the moving step length in the preset prediction window parameter, move the window, take the first log vector in the moved target window as the new candidate vector, and return to execute the step of taking the candidate vector as the starting point and determining the target window corresponding to the candidate vector according to the window length in the preset prediction window parameter until the first confidence scores corresponding to the log vectors in the target window are all less than the first threshold or the number of log vectors in the target window is less than the window length.
[0046] S130. Divide each log vector starting from the starting prediction vector into each prediction window according to the preset prediction window parameter.
[0047] Among them, the prediction window is each window obtained by moving starting from the starting prediction vector according to the preset prediction window parameter.
[0048] Specifically, starting from the starting prediction vector, according to the preset prediction window parameter, perform window sliding, and multiple prediction windows can be obtained, or the log vectors within each prediction window can be determined.
[0049] S140. For each prediction window, determine at least two second confidence scores corresponding to the log vector adjacent to and after the prediction window according to at least two pre-trained end-to-end network models, the abnormal states of the log texts in each row of the prediction window, and the log vector.
[0050] Among them, the end-to-end network model is a deep learning network model. For example, it can be a Transformer prediction model (a deep neural network model based on the self-attention mechanism), a VAE (Variational Autoencoder), etc. The abnormal states include normal and abnormal. The second confidence score is the result evaluation value for analyzing the log vector calculated through the prediction results corresponding to each end-to-end network model.
[0051] Specifically, for each prediction window, similar processing is performed using each pre-trained end-to-end network model. Therefore, one of them can be taken as an example for illustration. Update the log vectors of each row within the prediction window according to the corresponding anomaly status. For each row of log vectors, if its anomaly status is abnormal, use the prediction vector of this end-to-end network model for this row of log vectors before. If its anomaly status is normal, directly use this log vector. Input the updated vectors in the prediction window into this pre-trained end-to-end network model to obtain the output result of this end-to-end network model, that is, the prediction vector corresponding to the log vector adjacent to and after the prediction window, calculate the distance metric value between the two, and perform normalization processing to obtain the second confidence score output by this end-to-end network model for the log vector adjacent to and after the prediction window.
[0052] Based on the above example, the following method can be used to determine at least two second confidence scores corresponding to the log vector adjacent to and after the prediction window according to at least two pre-trained end-to-end network models, the anomaly status of each row of log texts within the prediction window, and the log vectors:
[0053] For each pre-trained end-to-end network model, determine the vectors to be input corresponding to each row of log texts in the prediction window according to the anomaly status of each row of log texts in the prediction window;
[0054] Input each vector to be input into the end-to-end network model to determine the prediction vector corresponding to the prediction window;
[0055] Determine the second confidence score of the log vector adjacent to and after the prediction window according to the prediction vector and the log vector adjacent to and after the prediction window.
[0056] Among them, the anomaly status of each log vector in the prediction window that starts the prediction vector is normal. The vector to be input is the vector used for subsequent prediction by inputting into the end-to-end network model, which may be the corresponding log vector or the prediction vector corresponding to the corresponding log vector. The prediction vector is the result output by the same end-to-end network model for the log vector.
[0057] Specifically, for each pre-trained end-to-end network model, in combination with the abnormal status of each line of log text within the prediction window, it is determined that the input vector to be corresponding to each line of log text in the prediction window is either a log vector or a corresponding prediction vector. That is, if the abnormal status is abnormal, the input vector to be corresponding to this line of log text is the prediction vector corresponding to this line of log text output by the end-to-end network model; if the abnormal status is normal, the input vector to be corresponding to this line of log text is the log vector corresponding to this line of log text. Input the input vectors to be in the prediction window into the end-to-end network model for time series prediction, and the prediction vector corresponding to the prediction window can be obtained. Furthermore, calculate the distance metric value between the prediction vector and the log vector adjacent to and after the prediction window, and normalize the obtained distance metric to the interval 0-1. This normalized value is the second confidence score of the log vector adjacent to and after the prediction window.
[0058] Based on the above example, the input vectors to be corresponding to each line of log text in the prediction window can be determined in the following manner according to the abnormal status of each line of log text in the prediction window:
[0059] For each line of log text in the prediction window, when the abnormal status of the log text is normal, use the log vector corresponding to the log text as the input vector to be corresponding to the log text;
[0060] For each line of log text in the prediction window, when the abnormal status of the log text is abnormal, use the prediction vector corresponding to the log text as the input vector to be corresponding to the log text.
[0061] Specifically, for each line of log text in the prediction window, if the abnormal status of this line of log text is normal, the log vector corresponding to this line of log text can be directly used as the corresponding input vector to be; if the abnormal status of this line of log text is abnormal, when judging the abnormal status of this line of log text, the prediction vector output by the end-to-end network model can be used as the corresponding input vector to be.
[0062] Based on the above example, the second confidence score of the log vector adjacent to and after the prediction window can be determined in the following manner according to the prediction vector and the log vector adjacent to and after the prediction window:
[0063] Determine multiple distance metric values according to the prediction vector and the log vector adjacent to and after the prediction window;
[0064] Determine the second confidence score of the log vector adjacent to and after the prediction window according to the multiple distance metric values.
[0065] Specifically, different distance metric calculation methods are used to calculate different distance metric values between the prediction vector and the log vectors adjacent to and after the prediction window. For example, multiple distance metrics can be adopted, such as Euclidean distance, cosine similarity, Hamming distance, etc. The obtained multiple distance metric values are comprehensively processed, such as taking the mean, weighted average, etc., to obtain a comprehensive metric value. The comprehensive metric value is normalized to obtain the second confidence score of the log vector adjacent to and after the prediction window.
[0066] S150. Determine the abnormal status of each line of log text in the log to be analyzed according to the first confidence score corresponding to each log vector and at least two second confidence scores.
[0067] Specifically, for each line of log text, the judgment of the abnormal status can be performed separately. If the line of log text only has the corresponding first confidence score, the abnormal status judgment is directly performed according to the first confidence score. For example, judgment can be made using a threshold, etc., to obtain the abnormal status of the line of log text. If the line of log text has the corresponding first confidence score and at least two corresponding second confidence scores, the first confidence score and each second confidence score are combined. For example, taking the mean, weighted average, or calculating using a pre-trained model, etc., to obtain a comprehensive confidence score, and the abnormal status judgment is performed according to the comprehensive confidence score. For example, judgment can be made using a threshold, etc., to obtain the abnormal status of the line of log text. Thus, the abnormal status of each line of log text in the log to be analyzed can be obtained.
[0068] It should be noted that in this example, the number of pre-trained machine learning models is at least one, and the number of pre-trained end-to-end network models is at least two. The specific number can be determined according to the computing power that the executor can bear and is not limited here.
[0069] Based on the above example, the abnormal status of each line of log text in the log to be analyzed can be determined by the following method according to the first confidence score corresponding to each log vector and at least two second confidence scores:
[0070] For each log vector, when each of the second confidence scores corresponding to the log vector is empty, determine the abnormal status of the log text corresponding to the log vector according to the first confidence score corresponding to the log vector and the first threshold;
[0071] For each log vector, when each of the second confidence scores corresponding to the log vector is not empty, determine the comprehensive confidence score according to the first confidence score corresponding to the log vector, each second confidence score, and the pre-trained logistic regression model, and determine the abnormal status of the log text corresponding to the log vector according to the comprehensive confidence score corresponding to the log vector and the second threshold.
[0072] Among them, the first threshold is a pre-set threshold for determining whether the log vector is considered normal when using a machine learning model. The comprehensive confidence score is a value calculated by using a pre-trained logistic regression model for the first confidence score and each second confidence score. The second threshold is a pre-set threshold for determining whether the log vector corresponding to the comprehensive confidence score is normal.
[0073] Specifically, for each log vector, it is determined whether each second confidence score corresponding to the log vector is empty. If so, the abnormal state corresponding to the log vector can be determined only according to the corresponding first confidence score. The first confidence score is compared with the first threshold. If the first confidence score is greater than or equal to the first threshold, it is determined that the abnormal state corresponding to the log vector is abnormal; otherwise, it is determined that the abnormal state corresponding to the log vector is normal. If not, the first confidence score corresponding to the log vector and each second confidence score are input into the pre-trained logistic regression model to output the comprehensive confidence score, and the comprehensive confidence score is compared with the second threshold. If the comprehensive confidence score is greater than or equal to the second threshold, it is determined that the abnormal state corresponding to the log vector is abnormal; otherwise, it is determined that the abnormal state corresponding to the log vector is normal.
[0074] Based on the above example, after determining the abnormal state of each line of log text in the log to be analyzed, the abnormal state of the entire log to be analyzed can be further determined. Specifically, it can be:
[0075] If there are at least one log text with an abnormal state among at least the second number of lines in at least one consecutive first number of lines of the log to be analyzed, it is determined that the abnormal state of the log to be analyzed is abnormal;
[0076] If there is no log text with an abnormal state among at least the second number of lines in each consecutive first number of lines of the log to be analyzed, it is determined that the abnormal state of the log to be analyzed is normal.
[0077] Among them, the first number of lines is a pre-set range of a certain number of lines for detection. The second number of lines is a pre-set line number value for determining log anomalies.
[0078] Specifically, the log to be analyzed is grouped according to the first number of lines. If the number of lines of the log text with an abnormal state in each consecutive first number of lines is less than the second number of lines, it is determined that the abnormal state of the log to be analyzed is normal. If there is at least one group of consecutive first number of lines of log text where the number of lines of the log text with an abnormal state is greater than or equal to the second number of lines, it means that there are multiple lines of anomalies, and it can be determined that the abnormal state of the log to be analyzed is abnormal.
[0079] Exemplarily, first, data preprocessing is performed on the serial port logs during the startup process of the supercomputer operating system. For the cleaning operation, regular expressions can be used to remove redundant information in the brackets and their contents (such as "(0x0001)") in the serial port logs during the startup process of the supercomputer operating system. At the same time, lines with specific error codes (such as "0x0000000080000000") are skipped, and key feature lines are retained. For incomplete or short-line texts, the continuity and context consistency of the log content are ensured by appending them to the previous line. For duplicate log content, important information is filtered out through a difference comparison technique based on similarity calculation, reducing the burden of redundant data processing. After the cleaning operation, the logs to be grouped are obtained. For the log semantic grouping based on the pre-trained language model, the pre-trained language model can be used to group the logs to be grouped, and the text features of the logs are extracted according to the understanding of the general semantics by the pre-trained language model. Furthermore, a linear classifier is used to divide the logs into multiple independent groups (logs to be analyzed), and each group implicitly refers to the serial port output of a stage of the system startup (such as kernel loading, hardware initialization, hardware self-check, etc.), that is, the category of the serial port logs during the startup process. During this process, the weights of some pre-trained language models are frozen to reduce the computational complexity and accelerate the inference process. Then, vectorized representations are performed on the logs to be analyzed respectively. An initial semantic vector is generated for each line of log text in the logs to be analyzed. Weight assignment is also required, that is, an inverse document frequency (IDF) is assigned to each word, and then the IDF weight corresponding to each word can be determined: , where w i represents the i-th word, and n i represents the number of logs containing the word w iThe number of logs, the total number of N logs. The purpose of using IDF weights is to enhance the importance of low-frequency events (such as abnormal logs). Further, combining the generated initial semantic vectors and IDF weights, an embedding matrix is constructed. Each column is a feature vector obtained by weighting the initial semantic vector corresponding to a log text by the words it contains, that is, the log vector. Finally, anomaly detection based on heterogeneous network ensemble learning. It can be three parallel expert networks (two end-to-end network models and one machine learning model) to process the log vectors embedded in the preprocessing stage. To avoid the built-in bias of a single expert network's structure fitting the training data, the expert networks adopt heterogeneous structures. For example, expert network 1 is a transformer decoder structure. To reduce computational overhead, it uses 1 self-attention and forward neural network layer; expert network 2 is a variational encoder structure; expert network 3 uses an isolation forest that is not a neural network to identify log anomalies using a tree structure. Both expert network 1 and expert network 2 adopt anomaly detection based on reconstruction error. A sliding window with a size of x rows of logs and a step size of 1 can be set (the prediction window corresponding to the preset prediction window parameters). Predict the next row of log vector (prediction vector) outside the window through the log vectors of each row within the window to capture the context relationship. Compare the differences between the prediction vectors of the next row predicted by the transformer and the VAE respectively and the log vector corresponding to the actual current row of log text. Specifically, three distance metrics can be used (Euclidean distance, cosine similarity, Hamming distance). Therefore, the results of integrating each distance metric of expert network 1 and expert network 2 are used respectively as the second confidence score for the log text of the current row (normalized to the interval 0-1, 1 indicates confirmed as an anomaly). Expert network 3 detects anomalies by detecting the deviation of the current row of log vector from the training data distribution, and uses the normalized value of the deviation gap as the confidence score for determining the anomaly of the log text of the current row, that is, the first confidence score. The output confidence scores of the three expert networks are used as the input of a logistic regression model to calculate the comprehensive confidence score to judge the anomaly status corresponding to each row of log text:
[0080]
[0081] where f overall is the comprehensive confidence score, f IForest is the first confidence score, f trans and f VAE are the two second confidence scores, x i is the log vector of the i-th row, corresponding to the log text of the i-th row. α1, α2, and α3 respectively represent the weights for different expert networks, w is the weight term, and b is the bias term. Among them, w, b, α1, α2, and α3 are pre-trained numerical values, which can be pre-configured for use when needed without retraining.
[0082] Based on the detection of the abnormal status of each single-line log text, if the abnormal status of the log lines exceeding a certain threshold number (the second line number) within a certain range (the first line of text) is abnormal, then it is determined that the abnormal status of the corresponding log to be analyzed is abnormal. This is because there is a chain reaction in a single abnormal log text that appears during the startup process of the supercomputer system. If only one log text is abnormal, it may be a false alarm, and multiple abnormalities are required to analyze and judge the abnormal status of the log to be analyzed.
[0083] In the above example, by integrating multiple heterogeneous lightweight expert models, a balance between low false alarms and high precision is achieved. Through the semantic chunking strategy (log semantic grouping based on pre-trained language models), the data input scale is significantly reduced. At the same time, the sliding window technique is combined to reduce memory requirements, and the dynamic attention mechanism and lightweight feature extraction optimize the computing efficiency, shortening the training time and detection time, making it suitable for resource-constrained environments. The log semantic grouping based on pre-trained language models can automatically parse the semantic features in the log, adapt to various log formats, especially unstructured long text data (such as the startup log of the supercomputer system serial port cache). Through the max pooling and average pooling strategies of semantic chunking by the pre-trained language model, different types of log data can be efficiently handled.
[0084] The present invention has the following technical effects: By determining the log vector corresponding to each log text line in the log to be analyzed according to the log to be analyzed, and determining the first confidence score corresponding to each log vector according to each log vector and the pre-trained machine learning model, to analyze in the first way, and determining the startup prediction vector from each log vector according to each first confidence score to judge when to introduce the second way for analysis. The log vectors starting from the startup prediction vector are divided into each prediction window according to the preset prediction window parameters. For each prediction window, at least two second confidence scores corresponding to the log vectors adjacent to and after the prediction window are determined according to at least two pre-trained end-to-end network models, the abnormal status of each log text line in the prediction window, and the log vector, to analyze in the second way. Using the sliding window method reduces resource occupancy. According to the first confidence score corresponding to each log vector and at least two second confidence scores, the abnormal status of each log text line in the log to be analyzed is determined, and by comprehensively using multiple log analysis methods at the appropriate time, the effects of improving the log analysis accuracy, reducing resource occupancy, and enhancing the wide applicability are achieved.
[0085] Embodiment 2
[0086] Figure 2 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 2 shown, the electronic device 200 includes one or more processors 201 and a memory 202.
[0087] The processor 201 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 200 to perform desired functions.
[0088] The memory 202 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 201 can run the program instructions to implement the serial port log analysis method of any embodiment of the present invention described above and / or other desired functions. Various contents such as initial external parameters, thresholds, etc. can also be stored in the computer-readable storage media.
[0089] In one example, the electronic device 200 can further include: an input device 203 and an output device 204, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). The input device 203 can include, for example, a keyboard, a mouse, etc. The output device 204 can output various information to the outside, including warning prompt information, braking force, etc. The output device 204 can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0090] Of course, for simplicity, Figure 2 only some of the components related to the present invention in the electronic device 200 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 200 can further include any other appropriate components.
[0091] In addition to the above methods and devices, an embodiment of the present invention can also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps of the serial port log analysis method provided by any embodiment of the present invention.
[0092] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present invention. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0093] In addition, an embodiment of the present invention may also be a computer-readable storage medium having computer program instructions stored thereon, and the computer program instructions, when run by a processor, cause the processor to perform the steps of the serial port log analysis method provided by any embodiment of the present invention.
[0094] The computer-readable storage medium may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0095] It should be noted that the terms used in the present invention are only for describing specific embodiments and do not limit the scope of the present application. As shown in the specification of the present invention, unless the context clearly indicates an exceptional situation, words such as "a", "an", "one", and / or "the" do not specifically refer to the singular and may also include the plural. The term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, or device including the element.
[0096] It should also be noted that the orientation or positional relationship indicated by terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. Unless otherwise clearly specified and defined, terms such as "installed", "connected", "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A serial port log analysis method, characterized in that, Including: Determine the log vector corresponding to each line of log text in the log to be analyzed according to the log to be analyzed; For each log vector, input the log vector into a pre-trained machine learning model, determine the deviation gap corresponding to the log vector, and perform normalization processing on the deviation gap to obtain the first confidence score corresponding to the log vector; Use the starting vector in each log vector as the candidate vector, starting from the candidate vector, determine the target window corresponding to the candidate vector according to the window length in the preset prediction window parameter, and determine whether the first confidence scores corresponding to the log vectors in the target window are all less than the first threshold; If so, use the candidate vector as the start prediction vector; If not, update the candidate vector according to the moving step length in the preset prediction window parameter, and return to execute the step of determining the target window corresponding to the candidate vector starting from the candidate vector according to the window length in the preset prediction window parameter until the first confidence scores corresponding to the log vectors in the target window are all less than the first threshold or the number of log vectors in the target window is less than the window length; Divide the log vectors starting from the start prediction vector into each prediction window according to the preset prediction window parameter; For each prediction window, determine at least two second confidence scores corresponding to the log vector adjacent to and after the prediction window according to at least two pre-trained end-to-end network models, the abnormal state of each line of log text in the prediction window, and the log vector; Determine the abnormal state of each line of log text in the log to be analyzed according to the first confidence score corresponding to each log vector and at least two second confidence scores.
2. The method according to claim 1, wherein Before determining the log vector corresponding to each line of log text in the log to be analyzed, it further includes: Obtain the serial port log during the startup process of the supercomputer operating system, perform data cleaning on the serial port log during the startup process of the supercomputer operating system to obtain the log to be grouped; Input the log to be grouped into a pre-trained language model to determine the text feature corresponding to the log to be grouped; Divide the log to be grouped into the log to be analyzed corresponding to each startup process serial port log category according to the text feature corresponding to the log to be grouped and a pre-trained classifier.
3. The method according to claim 1, characterized in that, The determining at least two second confidence scores corresponding to the log vector adjacent to and after the prediction window according to at least two pre-trained end-to-end network models, the abnormal state of each line of log text in the prediction window, and the log vector includes: For each pre-trained end-to-end network model, determine the input vector to be input corresponding to each line of log text in the prediction window according to the abnormal state of each line of log text in the prediction window; Input each input vector to be input into the end-to-end network model to determine the prediction vector corresponding to the prediction window; Determine a second confidence score of a log vector that is adjacent to and after the prediction window according to the prediction vector and the log vector that is adjacent to and after the prediction window; Among them, the abnormal status of each log vector in the prediction window where the first log vector is the start prediction vector is normal.
4. The method according to claim 3, characterized in that, The determining the input vectors corresponding to the log texts in the prediction window according to the abnormal status of the log texts in each row of the prediction window includes: For each log text in the prediction window, when the abnormal status of the log text is normal, use the log vector corresponding to the log text as the input vector corresponding to the log text; For each log text in the prediction window, when the abnormal status of the log text is abnormal, use the prediction vector corresponding to the log text as the input vector corresponding to the log text.
5. The method according to claim 3, characterized in that, The determining the second confidence score of the log vector that is adjacent to and after the prediction window according to the prediction vector and the log vector that is adjacent to and after the prediction window includes: Determine a plurality of distance metric values according to the prediction vector and the log vector that is adjacent to and after the prediction window; Determine the second confidence score of the log vector that is adjacent to and after the prediction window according to the plurality of distance metric values.
6. The method according to claim 1, wherein The determining the abnormal status of the log texts in each row of the log to be analyzed according to the first confidence score corresponding to each log vector and at least two second confidence scores includes: For each log vector, when the second confidence scores corresponding to the log vector are empty, determine the abnormal status of the log text corresponding to the log vector according to the first confidence score corresponding to the log vector and the first threshold; For each log vector, when the second confidence scores corresponding to the log vector are not empty, determine a comprehensive confidence score according to the first confidence score corresponding to the log vector, the second confidence scores, and a pre-trained logistic regression model, and determine the abnormal status of the log text corresponding to the log vector according to the comprehensive confidence score corresponding to the log vector and the second threshold.
7. The method according to claim 1, characterized in that After determining the abnormal status of the log texts in each row of the log to be analyzed, it further includes: If there are at least the second number of log texts with abnormal status in at least one continuous first number of log texts in the log to be analyzed, determine that the abnormal status of the log to be analyzed is abnormal; If there are no at least the second number of log texts with abnormal status in each continuous first number of log texts in the log to be analyzed, determine that the abnormal status of the log to be analyzed is normal.
8. An electronic device, characterized in that, The electronic device includes: A processor and a memory; The processor is configured to execute the steps of the serial port log analysis method according to any one of claims 1 to 7 by calling the program or instructions stored in the memory.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instructions, and the program or instructions cause the computer to execute the steps of the serial port log analysis method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Vehicle log anomaly detection method and device, electronic equipment and storage medium
CN117093443A
Service resource sharing method and service resource sharing system based on cloud computing
CN119052335A