Context-driven fault solution suggestion generation method

By building the BERTomaly model and FP-Growth algorithm combined with the large language model, the real-time and accurate problems of fault detection in complex cross-business scenarios are solved, and the rapid identification and targeted resolution of faults are achieved, which improves the stability and business continuity of the information system.

CN120508425APending Publication Date: 2025-08-19SOUTHEAST UNIV +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510621959.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The prior art is difficult to achieve real-time, accurate and targeted fault detection in complex cross-business scenarios, and the lack of dynamic solutions, resulting in limited information system stability and business continuity.

Method used

Build a BERTomaly model and FP-Growth association analysis algorithm based on the multi-head cross attention mechanism, and combine it with a large language model to realize fault feature extraction, feature fusion and potential association mining to generate targeted solutions.

Benefits of technology

It improves the accuracy and timeliness of fault detection, can quickly identify and provide targeted solutions, reduce business interruption time, and improve system stability and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508425A_ABST
    Figure CN120508425A_ABST
Patent Text Reader

Abstract

The invention relates to a context-driven fault solution suggestion generation method, which specifically comprises the following steps of: firstly, collecting various log data generated in system operation, and preprocessing the data as a basis for subsequent analysis; then, a BERTomaly model based on a multi-head cross attention mechanism is constructed, feature extraction, feature fusion and template matching can be performed on the log data, anomaly detection of system faults can be realized, and then, based on the system operation log data, the system fault detection efficiency is improved. The method comprises the following steps of: firstly, constructing a fault solution library which is of a context-driven type and has a dynamic updating capability by using a large language model in combination with the output of a BERTomally model and the output of an FP-Growth association analysis algorithm, and carrying out real-time detection on each fault, so that the fault solution library has a dynamic updating capability. And a targeted solution is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of anomaly detection, and in particular to a context-driven fault resolution suggestion generation method. Background Art

[0002] With the continuous advancement of technology, the complexity of modern information systems is increasing, especially in complex cross-business scenarios. This complexity easily leads to the frequent occurrence of various failures during information system operation, and these failures are often hidden in the large amount of log data generated by the system. Although this log data contains rich fault information, its heterogeneity and diversity pose significant challenges to fault detection and resolution. Currently, the industry has developed a number of detection and analysis technologies for system faults, such as those based on statistical analysis, pattern matching, or machine learning models. These technologies can achieve fault analysis and detection to a certain extent, but they often fail to meet the real-time and accuracy requirements faced by the complex and changing fault modes across cross-business scenarios, and they also lack the ability to generate targeted solutions.

[0003] Although existing technologies have made some progress in the field of fault detection, many shortcomings still exist. First, most existing systems rely on predefined rules or templates, which make it difficult to adapt to dynamically changing fault modes, resulting in low accuracy of detection results. Second, facing cross-business scenarios, the log data generated by different modules in the system has diverse structures, and existing methods face significant challenges in data fusion and analysis. In addition, traditional technologies usually only focus on fault detection, ignoring the in-depth exploration of fault causes and the generation of subsequent solutions, making it impossible to provide targeted dynamic solutions. These problems greatly limit the applicability of existing technologies in complex scenarios.

[0004] To address these issues, we propose a context-driven fault resolution suggestion generation method that enables real-time fault detection and solution generation across business scenarios. Through feature extraction and pattern matching, this technology accurately identifies abnormal behavior of system faults and explores potential correlations between faults. Compared to traditional technologies, this solution innovates by introducing a dynamic update mechanism driven by a large language model. This not only improves the accuracy and timeliness of fault detection, but also generates targeted solutions. This intelligent, integrated technology enables information systems to more efficiently identify and resolve complex faults, improving system stability and business continuity. Summary of the Invention

[0005] The purpose of the present invention is to provide an intelligent solution for detecting and resolving system faults in complex cross-business scenarios, meeting the three core goals of real-time, high efficiency, and pertinence. By building a model that integrates fault identification, cause analysis, and solution generation, it is possible to achieve timely response and intelligent resolution to sudden system faults, reduce business interruption time caused by faults, and thus establish an adaptive, continuously optimized intelligent operation and maintenance environment, ultimately achieving stable business development and improved user experience.

[0006] The technical solution of the present invention is as follows: a context-driven fault resolution suggestion generation method, the method comprising the following steps:

[0007] Step A: Collect various log data generated by the power information system during operation and pre-process the data.

[0008] Step B: A BERTomaly model based on a multi-head cross-attention mechanism was constructed, which can extract and fuse features from log data and detect anomalies of system failures.

[0009] Step C: Based on the power information system operation log data, use the association analysis algorithm to further explore the potential associations between faults and fault solution decisions.

[0010] Step D: Use the public dataset in the field of anomaly detection, namely the large-scale cloud computing environment fault prediction dataset, to test many open source large models, and select the large model that can effectively understand and process anomaly detection data as the basic model.

[0011] Step E: Use the P-Tuning large model fine-tuning method to fine-tune the pre-trained basic large model and optimize it for professional terms and specific terms to adapt to the field of anomaly detection.

[0012] Step F: Use the big model to intelligently match faults with solutions. Based on this, the big model's powerful context generation capabilities are used to provide more targeted solutions for each fault.

[0013] Furthermore, various log data generated in the power information system are pre-processed, including data such as fault type, occurrence time, impact range, and treatment measures taken, and then integrated into a unified dataset named dataset M.

[0014] Furthermore, the data preprocessing includes:

[0015] Step 1: For problems such as outliers, missing values, inconsistent formats, etc. in the data, the data is cleaned by deletion, filling, standardization, etc.

[0016] Step 2: Normalize the data.

[0017]

[0018] Step 3: Normalize the data.

[0019]

[0020] Among them, μ is the mean, σ is the standard deviation, x represents the log data to be processed, and x′ represents the processed log data.

[0021] Step 4: For text data related to treatment measures, use regular expressions to clean noise from the text fields. Compared to traditional manual screening and rule matching methods, this solution can more accurately and efficiently process special characters, redundant symbols, and unstructured data in log or text data. Furthermore, this solution can quickly adapt to different types of text, improving data cleaning efficiency and making it suitable for large-scale text data cleaning and anomaly detection preprocessing.

[0022] Furthermore, for log data, a BERTomaly model based on a multi-head cross-attention mechanism is constructed, which can perform feature extraction, feature fusion and template matching on log data.

[0023] The BERTomaly model includes:

[0024] A sentence is treated as an arbitrary sequence of tokens, with a [CLS] token added at the beginning of the first sentence and a [SEP] token indicating the end of each sentence. First, the tokenizer converts the input sentence into tokens and computes token embeddings. These tokens are then fed into the segment and position embeddings. Finally, the tokenizer integrates all the embeddings and feeds the result into the model.

[0025] Furthermore, the BERTomaly model is used to segment and embed the log data. The first 128 tokens of each text are taken to save memory and speed up calculation. Then, the [CLS] tag of the BERTomaly model is used to generate a 768-dimensional embedding feature vector. Where N is the length of the text sequence, d w It is the text feature dimension.

[0026] The embedding generated by the BERTomaly model is concatenated with the time series data features to form a joint feature.

[0027] The joint features are divided into blocks, and multiple logs are divided into blocks according to the time series. Each block corresponds to a token. The shape after block division is Logs∈R B×N×D, where B is the batch size, N is the number of blocks, and D represents the global embedding dimension, which represents the complete dimension of the text embedding after being processed by the BERTomaly model. The BERTomaly model is used to process the block-based log sequence, and the dependency relationship between multiple logs is extracted based on the multi-head attention mechanism. It is divided into four steps: input embedding, multi-head attention mechanism layer, feedforward network, residual connection and normalization. The input embedding Input is expressed as:

[0028] Input=Concat(Time Embedding,Text Embedding)+Positional Encoding

[0029] Positional encoding refers to the use of sine and cosine functions to encode the position relationship in the sequence:

[0030]

[0031] Among them, pos refers to the position index in the sequence, i refers to the embedding dimension index, and d p Represents the embedding dimension used by the position encoding, which is used to calculate the index in the position encoding. In the present invention, its value is the same as the global embedding dimension D.

[0032] The embedding generated by the BERTomaly model is concatenated with the time series data features to form a joint feature.

[0033] Log Embedding=[Time Features; Text Embedding]

[0034] The joint features are divided into blocks, and multiple logs are divided into blocks according to the time series, and each block corresponds to a token. The shape after block division is Logs∈R B×N×D , where B is the batch size, N is the number of blocks, and D is the embedding dimension. The BERTomaly model is used to process the block-based log sequence and extract the dependencies between multiple logs based on the multi-head attention mechanism, specifically including:

[0035] Step 1: Provide input embeddings to the model.

[0036] The input embedding is represented as:

[0037] Input=Concat(Time Embedding,Text Embedding)+Positional Encoding

[0038] Positional encoding refers to the use of sine and cosine encoding sequences to encode positional relationships:

[0039]

[0040] Among them, pos refers to the position index in the sequence, i refers to the embedding dimension index,

[0041] Step 2: Transform the input embedding into query Q, key K, and value V through linear transformation

[0042] Q=XW Q , K=XW K , V=XW V

[0043] in, d×d k refers to the size of each matrix,

[0044] The similarity between each query and key is calculated by dot product and normalized by softmax:

[0045]

[0046] in, Refers to the scaling factor to prevent the dot product value of large-dimensional vectors from being too large, causing the gradient to disappear. T The transposed matrix K representing the query Q and key K T The dot product operation between

[0047] Divide Q, K, V into 8 heads and calculate attention separately:

[0048] Head i =Attention(Q i , K i , V i )

[0049] Among them, Q i , K i , V i is the submatrix obtained by slicing.

[0050] The outputs of each head are concatenated and then a linear transformation is performed to generate the final output of the multi-head attention:

[0051] MultiHead(Q,K,V)=Contact(Head1,...,Head h )W O

[0052] in, is the linear transformation matrix.

[0053] Step 3: Use a feedforward network (FFN) to apply a two-layer fully connected network and a nonlinear activation function to the output of the multi-head attention. Specifically:

[0054] FFN(Z)=ReLU(Zw1+b1)w2+b2

[0055] Where Z represents the input of the feedforward network FFN and the output of the multi-head attention mechanism. b1 refers to the bias of the first layer, and b2 refers to the bias of the second layer. Refers to the weight matrix of the fully connected layer, d ff refers to the hidden layer dimension,

[0056] Step 4: Add residual connections after the attention mechanism and feedforward network to prevent gradient disappearance. Specifically include:

[0057] Z′=X+MultiHead(Q,K,V)

[0058] Z″=Z′+FFN(Z′)

[0059] Where X represents the input embedding, Z′ represents the output of the multi-head attention, and contains residual connections. Z″ represents the output of the feedforward network FFN, and contains residual connections.

[0060] Step 5: After the residual connection, normalize the output of each layer. Specifically include:

[0061]

[0062] Where Z' represents the input vector, μ represents the mean of a single sample on the embedding dimension, and σ 2 represents the variance of a single sample on the embedding dimension, σ represents the standard deviation of a single sample on the embedding dimension, d L Indicates the embedding dimension used by LayerNorm normalization. In this invention, its value is the same as the global embedding dimension D. i Represents the value of a Token in the i-th dimension, and the final output is Z final ∈R N×D , which can be used for subsequent anomaly detection tasks.

[0063] Furthermore, the difference between the log embedding and the normal template is determined by taking the minimum Euclidean distance between the log embedding and the template set. Specifically, it includes:

[0064]

[0065] Among them, d ERepresents the embedding dimension used in similarity calculation, for the input embedding X and the template set {t1,......t n} similarity calculation, in this invention, its value is the same as the global embedding dimension D. This scheme uses the minimum Euclidean distance to measure the degree of deviation between the log embedding and the normal template. Compared with the existing technology, it has the advantages of high computational efficiency, strong adaptability, and easy implementation.

[0066] Furthermore, the FP-Growth association analysis algorithm is used to further explore the potential associations between faults and fault solution decisions.

[0067] Furthermore, the FP-Growth association analysis algorithm includes:

[0068] Association analysis is achieved through the FP-tree structure. For each indicator, the FP-tree is constructed by filtering and sorting the items in the transaction based on the frequent item information in the item header table.

[0069] Step 1: Map log events into specific item sets.

[0070] Step 2: After generating the transaction database, perform global frequency statistics on all items and sort them in descending order of frequency to generate an item header table.

[0071] Step 3: Build the FP-Tree by inserting transactions one by one. During insertion, arrange the items in the transaction according to the order in the item header table and update the node counts in the FP-Tree. If the path already exists, increment the count; if not, create a new node.

[0072] Step 4: After the FP-Tree is constructed, frequent itemsets are mined recursively. First, the conditional pattern base is extracted, i.e., the path containing the target item and its prefix. Then, a conditional FP-Tree is constructed based on the conditional pattern base, and frequent itemsets are mined recursively until no more conditional FP-Trees can be generated.

[0073] Furthermore, association rules between transactions are calculated using support and confidence. Support represents the frequency of a set of items appearing in a transaction, while confidence represents the probability of another item appearing under given conditions. Association rules that meet a set threshold are output as potential associations between log information. Specifically, this includes:

[0074]

[0075] Among them, A represents the event that occurs first, and B represents the event that occurs later.

[0076] Support(A) represents the probability of event A occurring.

[0077] Count(A) indicates the number of times event A occurs in the entire transaction database.

[0078] Total Transactions indicates the total number of transactions in the transaction database.

[0079] Support(A∪B) represents the probability of event A and event B occurring simultaneously.

[0080] Confidence(A→B) represents the probability that event B will occur if event A occurs.

[0081] This solution mines association rules between log transactions through support and confidence. It has the advantages of high computational efficiency, strong adaptability, good interpretability, and low computing resource consumption. It is particularly suitable for real-time log analysis, transaction pattern mining, and anomaly detection tasks.

[0082] Furthermore, the large model is used to intelligently match solutions to detected faults, providing intelligent solutions to system faults. Specifically, it includes:

[0083] Step 1: Use a publicly available dataset in the field of anomaly detection—a large-scale cloud computing environment fault prediction dataset—to test numerous open-source large models, selecting a large model that can effectively understand and process anomaly detection data as the base model.

[0084] Step 2: Use the P-Tuning large model fine-tuning method to fine-tune the pre-trained basic large model and optimize it for professional terms and specific terms to adapt to the field of anomaly detection.

[0085] Step 3: Use the big model to intelligently match faults with solutions. Based on this, the big model's powerful context generation capabilities are used to provide more targeted solutions for each fault.

[0086] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for generating context-driven fault resolution suggestions is implemented.

[0087] A computer-readable storage medium stores computer instructions, which, when executed by a processor, implement the method for generating context-driven fault resolution suggestions.

[0088] Compared to existing technologies, the present invention offers the following advantages: It improves fault detection efficiency while generating a large number of targeted solutions based on fault characteristics. In a specific application, the present invention rapidly detected a "sudden interruption of a microservice call chain" in a power information system in just 30 minutes and generated a fault report with targeted solutions, including "using circuit breakers (such as Netflix Hystrix) to limit calls to the faulty service," "reviewing call chain logs to locate the specific issue," and "using Kubernetes' Horizontal Pod Autoscaler (HPA) to dynamically increase service instances based on load." Compared to the traditional process of manually troubleshooting, analyzing the cause, and proposing solutions, the present invention shortens fault detection time by one hour and generates a large number of targeted solutions for each fault. Simulation results show that BERTomaly performs better than Drain and AEL on the complex BGL dataset, achieving a message accuracy rate of 89%, demonstrating that BERTomaly can better parse complex semantic tasks. Furthermore, on the complex BGL dataset, the BERTomaly model achieves an F1 score of 91.4%, significantly outperforming other models. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] Figure 1 Generate a method flow chart for a context-driven troubleshooting suggestion;

[0090] Figure 2 Build a schematic diagram for FP-tree;

[0091] Figure 3 Schematic diagram of the specific composition of the large-scale cloud computing environment fault prediction dataset. DETAILED DESCRIPTION

[0092] In order to deepen the understanding of the present invention, this embodiment is described in detail below with reference to the accompanying drawings.

[0093] Embodiment: The present invention is mainly aimed at fault detection and problem solving in power information systems, and proposes a model that integrates fault identification, cause analysis, and solution suggestion generation.

[0094] A context-driven approach to generating fault resolution suggestions, see Figure 1 , the method comprises the following steps:

[0095] First, various log data generated in the power information system are preprocessed.

[0096] Collect data including fault type, occurrence time, impact scope, and treatment measures taken, and integrate them into a unified dataset named dataset M.

[0097] For problems such as outliers, missing values, and inconsistent formats in the data, data cleaning is carried out by deletion, filling, and standardization. Specifically, the following methods are used:

[0098] Normalization:

[0099]

[0100] standardization:

[0101]

[0102] Among them, μ is the mean, σ is the standard deviation, x represents the log data to be processed, and x′ represents the processed log data.

[0103] For text data such as treatment actions, use regular expressions to clean up noise in text fields.

[0104] After preprocessing the log data, a BERTomaly model based on a multi-head cross-attention mechanism constructed by the present invention is used to perform feature extraction and feature fusion on the log data.

[0105] The steps to build the BERTomaly model are as follows:

[0106] A sentence is treated as an arbitrary sequence of tokens, with a [CLS] token added at the beginning of the first sentence and a [SEP] token indicating the end of each sentence. First, the tokenizer converts the input sentence into tokens and computes token embeddings. These tokens are then fed into the segment and position embeddings. Finally, the tokenizer integrates all the embeddings and feeds the result into BERT.

[0107] The BERTomaly model is used to segment and embed the log data. The first 128 tokens of each text are taken to save memory and speed up calculation. Then, the [CLS] tag of the BERTomaly model is used to generate a 768-dimensional embedding feature vector. Where N is the length of the text sequence, d w It is the text feature dimension.

[0108] The embedding generated by the BERTomaly model is concatenated with the time series data features to form a joint feature.

[0109] Log Embedding=[Time Features; Text Embedding]

[0110] The joint features are divided into blocks, and multiple logs are divided into blocks according to the time series, and each block corresponds to a token. The shape after block division is Logs∈R B×N×D , where B is the batch size, N is the number of chunks, and D represents the global embedding dimension, representing the complete dimension of the text embedding after processing by the BERTomaly model. The BERTomaly model processes the chunked log sequence and extracts dependencies between multiple logs using a multi-head attention mechanism. This process consists of four steps: input embedding, multi-head attention layer, feedforward network, residual connections, and normalization.

[0111] The input embedding X is represented as:

[0112] Input(X)=Concat(Time Embedding,Text Embedding)+Positional Encoding

[0113] Among them, time embedding is used to represent the time information of the input data. Text embedding is used to represent the text features in the input data. Positional encoding refers to the use of sine and cosine functions to encode the position relationship in the sequence:

[0114]

[0115] Among them, pos refers to the position index in the sequence, i refers to the embedding dimension index, and d p Represents the embedding dimension used by the position encoding, which is used to calculate the index in the position encoding. In this invention, its value is the same as the global embedding dimension D. 10000 in the position encoding 2i / d is the position scaling factor normalized by the embedding dimension.

[0116] The multi-head attention mechanism specifically includes:

[0117] Embed the input X through linear transformation to generate query Q, key K, value V

[0118] Q=XW Q , K=XW K , V=XW V

[0119] in, d×d k refers to the size of each matrix.

[0120] The similarity between each query and key is calculated by dot product and normalized by softmax:

[0121]

[0122] Among them, QK T The dot product of the query vector and the key vector is used to calculate the similarity between the query and the key. Refers to the scaling factor to prevent the dot product value of large-dimensional vectors from being too large, causing the gradient to disappear.

[0123] Divide Q, K, V into 8 heads and calculate attention separately:

[0124] Head i =Attention(Q i , K i , V i )

[0125] Among them, Q i , K i , V i is the submatrix obtained by slicing.

[0126] The outputs of each head are concatenated and then a linear transformation is performed to generate the final output of the multi-head attention:

[0127] MultiHead(Q,K,V)=Contact(Head1,...,Head h )W O

[0128] in, is the linear transformation matrix, and h represents the number of attention heads (8 in this invention).

[0129] The feedforward network specifically includes:

[0130] A feedforward network (FFN) is used to apply a two-layer fully connected network and a nonlinear activation function to the output of the multi-head attention MultiHead(Q, K, V).

[0131] FFN(Z)=ReLU(ZW1+b1)W2+b2

[0132] in, Refers to the weight matrix of the fully connected layer, d ff Refers to the hidden layer dimension, b1 and b2 refer to the bias vectors of the two fully connected layers in the feedforward network, and together with W1 and W2 define the linear transformation. The residual connection specifically includes:

[0133] Add residual connections after the attention mechanism and feedforward network to avoid gradient disappearance.

[0134] Z′=X+MultiHead(Q,K,V)

[0135] Z″=Z′+FFN(Z′)

[0136] Where Z represents the output of the multi-head attention as the input of the feedforward network. Z′ represents the result obtained after the residual connection. Z″ represents the output of the feedforward network FFN and includes the residual connection.

[0137] Normalization specifically includes:

[0138] After the residual connection, the output of each layer is normalized.

[0139]

[0140] Where Z' represents the input vector, μ represents the mean of a single sample on the embedding dimension, and σ 2 represents the variance of a single sample on the embedding dimension, σ represents the standard deviation of a single sample on the embedding dimension, d L Indicates the embedding dimension used by LayerNorm normalization. In this invention, its value is the same as the global embedding dimension D. The final output is Z final ∈R N×D , where N represents the length of the input sequence. This can be used for fault detection tasks. It mainly uses Euclidean similarity calculation to determine the difference between the log embedding and the normal template, including:

[0141]

[0142] Where X represents the input log embedding, t represents the normal template embedding, and d E Represents the embedding dimension used in similarity calculation, for the input embedding X and the template set {t1,......t n}, in the present invention, its value is the same as the global embedding dimension D.

[0143] For the input log embedding X and the template set {t1,......t n The similarity calculation of} is determined by taking the minimum Euclidean distance D:

[0144] d min =min(D1,......D n )

[0145]

[0146] The FP-Growth association analysis algorithm is used to further explore the potential correlation between faults and fault solution decisions for the feature values extracted by the BERTomaly model.

[0147] See Figure 2,FP-Growth association analysis algorithm mainly realizes association analysis through FP-tree structure, and the construction method is as follows:

[0148] First, log events are mapped into specific item sets.

[0149] After the transaction database is generated, global frequency statistics are performed on all items, and they are sorted in descending order of frequency to generate an item header table.

[0150] The FP-Tree is constructed by inserting transactions one by one. During insertion, the items in the transaction are sorted according to the order in the item header table and the node counts in the FP-Tree are updated. If the path already exists, the count is incremented; if not, a new node is created.

[0151] After the FP-Tree is constructed, frequent itemsets are mined recursively. First, a conditional pattern base is extracted, i.e., the path containing the target item and its prefix. Then, a conditional FP-Tree is constructed based on the conditional pattern base, and frequent itemsets are mined recursively until no more conditional FP-Trees can be generated.

[0152] After the FP-Tree is constructed, the association rules between transactions are calculated using support and confidence. The association rules that meet the set threshold are output as the potential association relationships between log information. This includes:

[0153]

[0154] Among them, A represents the event that occurs first, and B represents the event that occurs later.

[0155] Support(A) represents the probability of event A occurring.

[0156] Count(A) indicates the number of times event A occurs in the entire transaction database.

[0157] Total Transactions indicates the total number of transactions in the transaction database.

[0158] Support(A∪B) represents the probability of event A and event B occurring at the same time, and Confidence(A→B) represents the probability of event B occurring if event A occurs.

[0159] After performing anomaly detection and correlation analysis on faults, a large model is used to intelligently match solutions to detected faults, providing intelligent solutions to system faults.

[0160] First, see Figure 3, we use a public dataset in the field of anomaly detection - a large-scale cloud computing environment fault prediction dataset to test many open source large models, and select a large model that can effectively understand and process anomaly detection data as the basic model.

[0161] Secondly, the P-Tuning large model fine-tuning method is used to fine-tune the pre-trained basic large model and optimize it for professional terminology and specific terms to adapt to the field of anomaly detection.

[0162] Finally, the big model is used to intelligently match faults with solutions, and on this basis, the powerful context generation capability of the big model is used to provide more targeted solutions for each fault.

[0163] Through specific experiments and practical applications, the performance of the present invention in terms of log parsing and anomaly detection capabilities compared with other methods is verified.

[0164] Please refer to Table 1 and Table 2. Table 1 shows the performance of the three log parsing models on the HDFS dataset, and Table 2 shows the performance of the three log parsing models on the BGL dataset.

[0165] Table 1

[0166]

[0167] Table 2:

[0168]

[0169] The present invention is compared with other methods in terms of log parsing capabilities, including:

[0170] The experiment compares three log parsing methods: Drain, AEL and BERTomaly. The evaluation datasets used are: the log dataset HDFS collected from the Hadoop Distributed File System (HDFS) and the log dataset BGL collected from the BlueGene / L supercomputer system.

[0171] The evaluation uses three metrics: group accuracy, message-level accuracy, and edit distance. These metrics assess whether log messages can be correctly grouped, whether the parsed single log template matches the true template, and the minimum transformation distance between the log template structure and the true template.

[0172] The experimental results show that compared with Drain and AEL, BERTomaly performs better in the complex dataset BGL, with a message accuracy of 89%, indicating that BERTomaly can better parse complex semantic tasks.

[0173] Please refer to Table 3, which shows the performance of five anomaly detection models on the HDFS dataset and the BGL dataset respectively.

[0174] Table 3:

[0175]

[0176] The present invention is compared with other methods in terms of anomaly detection capabilities, including:

[0177] The experiment compares five anomaly detection models: DeepLog and LogAnomaly models based on deep learning, PCA and Lsolation Forest models based on traditional machine learning, and BERTomaly model. The evaluation datasets used are: HDFS, a log dataset collected from the Hadoop Distributed File System (HDFS), and BGL, a log dataset collected from the BlueGene / L supercomputer system.

[0178] The evaluation uses three indicators: precision, recall, and F1 score.

[0179] Experimental results show that on the complex BGL dataset, the BERTomaly model achieved an F1 score of 93.4%, significantly outperforming other models. Traditional methods such as PCA and Layout Forest, however, suffer from lower F1 scores due to their inability to fully utilize the semantic information in logs. While the deep learning-based DeepLog and LogAnomaly models have improved their use of semantic information, they still lag behind the BERTomaly model, particularly in terms of recall.

[0180] It should be noted that the above embodiments are not intended to limit the scope of protection of the present invention, and equivalent changes or substitutions made on the basis of the above technical solutions fall within the scope of protection of the claims of the present invention.

Claims

1. A context-driven fault resolution suggestion generation method, characterized in that: The method comprises the following steps: Step A: Collect various log data generated by the power information system during operation and pre-process the data. Step B: A BERTomaly model based on a multi-head cross-attention mechanism was constructed, which can extract and fuse features from log data and detect anomalies of system failures. Step C: Based on the power information system operation log data, use the association analysis algorithm to further explore the potential associations between faults and fault solution decisions. Step D: Use the public dataset in the field of anomaly detection, namely the large-scale cloud computing environment fault prediction dataset, to test many open source large models, and select the large model that can effectively understand and process anomaly detection data as the basic model. Step E: Use the P-Tuning large model fine-tuning method to fine-tune the pre-trained basic large model and optimize it for professional terms and specific terms to adapt to the field of anomaly detection. Step F: Use the big model to intelligently match faults with solutions. Based on this, the big model's powerful context generation capabilities are used to provide more targeted solutions for each fault.

2. A context-driven fault resolution suggestion generation method according to claim 1, characterized in that: In step A, it is necessary to pre-process various log data generated in the power information system. First, data including fault type, occurrence time, impact range, and treatment measures taken are collected and integrated into a unified data set, named data set M; For outliers, missing values, and inconsistent formats in the data, the data is cleaned by deletion, filling, and standardization, including: Normalization: standardization: Among them, μ is the mean of the entire data set, σ is the standard deviation of the entire data set, x represents the log data to be processed, and x′ represents the processed log data. For text data such as treatment actions, use regular expressions to clean up noise in text fields.

3. The context-driven fault resolution suggestion generation method according to claim 1, characterized in that: In step B, for log data, a BERTomaly model based on a multi-head cross attention mechanism is constructed to extract and fuse features from log data. Among them, the BERTomaly model includes: Treating sentences as arbitrary token sequences, a [CLS] tag is added at the beginning of the first sentence, and a [SEP] tag indicates the end of each sentence. First, the tagger converts the input sentence into tokens and then calculates the token embeddings. After that, these tokens are input into the segment and position embedding calculations. Finally, the tagger integrates all the embeddings and inputs the result into BERT. Use the BERTomaly model to segment and embed the log data. Take the first 128 tokens for each text to save memory and speed up calculations. Then, use the [CLS] tag of the BERTomaly model to generate a 768-dimensional embedding feature vector. Where N is the length of the text sequence, d w is the text feature dimension, The embedding generated by the BERTomaly model is concatenated with the time series data features to form a joint feature. Log Embedding=[Time Features; Text Embedding] The joint features are divided into blocks, and multiple logs are divided into blocks according to the time series. Each block corresponds to a token. The shape after block division is Logs∈R B×N×D , where B is the batch size, N is the number of blocks, and D represents the global embedding dimension, which represents the complete dimension of the text embedding after being processed by the BERTomaly model. The BERTomaly model is used to process the block-based log sequences, and the dependencies between multiple logs are extracted based on the multi-head attention mechanism. This process consists of four steps: input embedding, multi-head attention mechanism layer, feedforward network, residual connection, and normalization. The input embedding is expressed as: Input=Concat(Time Embedding,Text Embedding)+Positional Encoding, Positional encoding refers to the use of sine and cosine functions to encode the position relationship in the sequence: Among them, pos refers to the position index in the sequence, i refers to the embedding dimension index, and d p Represents the embedding dimension used by the position encoding, which is used to calculate the index in the position encoding. Its value is the same as the global embedding dimension D. The multi-head attention mechanism specifically includes: Embed the input into Input and generate query Q, key K, value V through linear transformation Q=XW Q ,K=XW K ,V=XW V Where X represents the input embedding, d×d k Refers to the size of each matrix, and the similarity between each query and key is calculated by dot product and normalized by softmax: in, Refers to the scaling factor to prevent the large-dimensional vector dot product value from being too large, causing the gradient to disappear. T The transposed matrix K representing the query Q and key K T The dot product operation between Divide Q, K, V into 8 heads and calculate attention separately: Head i =Attention(Q i ,K i ,V i ) Among them, Q i , K i , V i is the submatrix obtained by slicing, The outputs of each head are concatenated and then a linear transformation is performed to generate the final output of the multi-head attention: MultiHead(Q,K,V)=Contact(Head1,......,Head h )W O in, is the linear transformation matrix, Use the feedforward network FFN to apply a two-layer fully connected network and a nonlinear activation function to the output of the multi-head attention MultiHead (Q, K, V), specifically including: FFN(Z)=ReLU(Zw1+b1)w2+b2 Among them, Z represents the input of the feedforward network FFN, which comes from the output of the multi-head attention mechanism, b1 refers to the bias term Bias of the first layer, and b2 refers to the bias term Bias of the second layer. Refers to the weight matrix of the fully connected layer, d ff refers to the hidden layer dimension, Add residual connections after the attention mechanism and feedforward network to prevent gradient disappearance, including: Z′=X+MultiHead(Q,K,V) Z″=Z′+FFN(Z′) Among them, X represents the input embedding, Z′ represents the output result of the multi-head attention and contains residual connections, and Z″ represents the output result of the feedforward network FFN and contains residual connections. After the residual connection, the output of each layer is normalized, including: Among them, Z″ represents the normalized input, which comes from the output of the feedforward network FFN, μ represents the mean value of a single sample in the embedding dimension, σ 2 represents the variance of a single sample on the embedding dimension, σ represents the standard deviation of a single sample on the embedding dimension, d L Indicates the embedding dimension used by LayerNorm normalization. In this invention, its value is the same as the global embedding dimension D. i Represents the value of a Token in the i-th dimension, and the final output is Z final ∈R N×D , used for subsequent anomaly detection tasks, mainly uses Euclidean similarity calculation to judge the difference between log embeddings and normal templates, specifically including: Among them, d E Represents the embedding dimension used in similarity calculation, for the input embedding X and the template set {t1,......t n Its value is the same as the global embedding dimension D. By taking the minimum Euclidean distance D min To judge: D min =min(D1,......D n ) 4. The context-driven fault resolution suggestion generation method according to claim 1, characterized in that: In step C, the FP-Growth association analysis algorithm is used to further explore the potential associations between faults and fault solution decisions. The FP-Growth association analysis algorithm mainly implements association analysis through the FP-tree structure. For each indicator, the FP-tree is constructed by filtering and sorting the items in the transaction based on the frequent item information in the item header table. First, log events are mapped into specific item sets. After generating the transaction database, all items are globally counted and sorted in descending order of frequency to generate an item header table. FP-Tree is constructed by inserting transactions one by one. When inserting, the items in the transaction need to be arranged in the order in the item header table and the count of the nodes in the FP-Tree is updated. If the path already exists, the count is increased; if not, a new node is created. After the FP-Tree is constructed, frequent item sets are mined recursively. First, the conditional pattern base is extracted, that is, the path containing the target item and its prefix; then a conditional FP-Tree is constructed based on the conditional pattern base, and frequent item sets are mined recursively until no more conditional FP-Trees can be generated. The association rules between transactions are calculated through support and confidence. Support represents the frequency of a certain item set appearing in a transaction, and confidence represents the probability of another item appearing under given conditions. The association rules that meet the set threshold are output as the potential association relationship between log information. Among them, A represents the event that occurs first, and B represents the event that occurs later. Support(A) represents the probability of event A occurring. Count(A) indicates the number of times event A occurs in the entire transaction database. Total Transactions indicates the total number of transactions in the transaction database. Support(A∪B) represents the probability of event A and event B occurring simultaneously. Confidence(A→B) represents the probability that event B will occur if event A occurs.

5. The context-driven fault resolution suggestion generation method according to claim 1, characterized in that: In step D, we use a public dataset in the field of anomaly detection, namely a large-scale cloud computing environment fault prediction dataset, to test many open-source large models, and select a large model that can effectively understand and process anomaly detection data as the basic model. First, the large-scale cloud computing environment fault prediction dataset is preprocessed. The data is loaded into JSON format using Python's Pandas library. The data is then cleaned, missing and duplicate values are processed, invalid data is deleted, and finally the data is normalized. Then, the dataset is labeled, and the abnormal points are determined based on the abnormal event log number and downtime. The time point when no abnormality occurs is marked as 0, and the time point when an abnormal event occurs is marked as 1. The time window method is used to construct time series data to predict the status of the next data. Finally, we used the dataset to test a variety of open source large language models, and selected a large model that can effectively understand and process anomaly detection data as the basic model.

6. The context-driven fault resolution suggestion generation method according to claim 1, characterized in that: In step E, the P-Tuning large model fine-tuning method is used to fine-tune the pre-trained basic large model and optimize it for professional terminology and specific terms to adapt to the field of anomaly detection. Specifically, during the construction process, data cleaning and feature extraction are first performed on the large-scale cloud computing environment fault prediction dataset, and key log fields are extracted to build a professional terminology library to ensure that the model can accurately identify the fault type. During the model building phase, P-Tuning technology is used to add task-specific information to the pre-trained large model through soft prompts (trainable prefix vectors) to optimize its ability to understand abnormal logs and time series data. Specifically, a series of prompt templates are designed at the input layer. During the training process, cross entropy loss is used to optimize the classification task, and mean square error (MSE loss) is combined to predict the possibility of future anomalies. At the same time, incremental learning is used to regularly update the prompt vector so that the model can continuously adapt to new abnormal patterns.

7. The context-driven fault resolution suggestion generation method according to claim 1, characterized in that: In step F, the big model is used to intelligently match faults with solutions. On this basis, the big model's powerful context generation capability is used to provide more targeted solutions for each fault. Specifically, the big model is used to classify fault types based on the fault prediction results and the contents of the fault solution library. The pre-trained knowledge base is used to match existing fault solutions. Subsequently, the big model's context-driven mechanism is used to intelligently optimize and adjust the matched solutions based on the specific circumstances of the current fault, environmental variables, and historical solutions to make them more consistent with the characteristics of the current fault.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for generating context-driven fault resolution suggestions as described in any one of claims 1 to 7 is implemented.

9. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, a context-driven fault resolution suggestion generating method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Method for detecting log abnormity of power dispatching automation system

    CN120723588A

  • A method for detecting log anomalies in a power dispatch automation system

    CN120723588B

  • Flowmeter fault diagnosis method and system based on digital fusion large language model

    CN121255511A

  • A Method and System for Fault Diagnosis of Flow Meters Based on Data-Knowledge Fusion Large Language Model

    CN121255511B