Intelligent log analysis method based on DevOps assembly line
By building a structured knowledge base and a Transformer-based BERT deep learning model, combined with user feedback and an adaptive window algorithm, we solved the problems of insufficient flexibility and slow response speed of log analysis tools in the DevOps pipeline, and achieved intelligent log parsing and real-time problem location.
Patent Information
- Application Number
- CN202511285430.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-10-14
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing log analysis tools in DevOps pipelines lack flexibility and cannot adapt to changes in the system environment. They require a lot of manual intervention and have slow response speeds.
Build a structured knowledge base, adopt the BERT deep learning model based on the Transformer architecture, combine user feedback mechanism and adaptive window algorithm, dynamically optimize the log parsing model, optimize model parameters through incremental training and cross-entropy loss function, and achieve real-time problem location and solution recommendation.
It realizes intelligent log parsing, improves the flexibility and response speed of log analysis, reduces manual intervention, and supports real-time feedback and dynamic adjustment.
Smart Images

Figure CN120780558A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information technology, in particular to an intelligent log analysis method based on a DevOps pipeline. BACKGROUND
[0002] Currently, the technical solutions used by enterprises in the DevOps pipeline mainly include: Rule-based log analysis tools: using pre-set rules and patterns for log analysis. The tool monitors log files and matches exceptions according to rules. Common tools include ELK Stack (Elasticsearch, Logstash, Kibana), Splunk, etc.
[0003] Statistical methods: using statistical analysis techniques, through data collection and analysis, to identify frequently occurring error patterns. This method usually relies on data mining algorithms such as clustering analysis and anomaly detection.
[0004] Machine learning models: some companies have begun to use simple machine learning models to train classifiers based on historical data for log analysis. These models are mainly used for pattern recognition and anomaly detection, but often lack real-time feedback mechanisms and dynamic adjustment capabilities. These existing technologies, although they have improved the efficiency of log analysis, still face many problems, including: Rule limitations: rule-based tools rely on manually defined rules and are difficult to adapt to new or complex problems, lacking flexibility.
[0005] Static models: statistical methods and simple machine learning models often cannot adapt to changing system environments, leading to decreased model performance.
[0006] Manual intervention: although there is some automation, a large amount of manual intervention is still required, and the response speed is slow. SUMMARY
[0007] The purpose of the present application is to overcome the shortcomings of the prior art and provide an intelligent log analysis method based on a DevOps pipeline, comprising the following steps: Step one, build a structured knowledge base, which interacts with the DevOps pipeline platform in real time through an API interface, including a user information module, a log entry module, a solution module, a classification and label module, and a search and recommendation module; Step two, train a BERT deep learning model based on the Transformer architecture, which receives log data through the input layer, processes it through the embedding layer and the encoding layer, and outputs the matching solution through the decoding layer; Step three, dynamically extract context information based on error log features, determine the interception range through adaptive window algorithm, and dynamically adjust the window size according to the error type and historical data; Step four, introduce a user feedback mechanism to correct the solution probability by weighting the number of likes / dislikes, and dynamically optimize the model parameters through incremental training; Step five, form a closed-loop improvement system, feed back the unmatched problems to the knowledge base and analyze them, and update the model training data.
[0008] Further, the construction of the structured knowledge base, the knowledge base interacts with the DevOps pipeline platform in real time through the API interface, contains user information module, log entry module, solution module, classification and label module and search and recommendation module, including: Collect DevOps community Q&A data and store it by component type; Modular data structure, where the user information module records the feedback user identity and historical behavior; The log entry module stores the standardized log format and error feature label; The solution module associates error type and correction scheme; The classification and label module supports multi-dimensional retrieval; The search and recommendation module implements solution pushing based on semantic similarity.
[0009] Further, the BERT deep learning model based on the Transformer architecture is trained, the model receives log data through the input layer, processes it through the embedding layer and the encoding layer, and outputs the matched solution through the decoding layer, including: Data preprocessing stage: log segmentation, denoising and feature labeling; Use cross-entropy loss function to minimize the difference between predicted solution and real solution; Through the result feedback mechanism, unmatched cases are included in the training set; Deploy online learning module to support real-time reception of user feedback data and update model weights.
[0010] Further, the context information is dynamically extracted based on the error log features, and the interception range is determined through the adaptive window algorithm, wherein the window size is dynamically adjusted according to the error type and historical data, including: Initialize the window size based on the error type, where frequent errors use small windows and complex errors use large windows; Optimize the window calculation strategy through the feedback loop, the strategy calculates the historical error frequency with weighting; Solution matching degree is dynamically adjusted; The context interception range is adaptively expanded.
[0011] Further, the user feedback mechanism is introduced to correct the solution probability by weighting the number of likes / dislikes, and the model parameters are dynamically optimized through incremental training, including: Add a feedback count field in the solution record to record the number of likes and dislikes; Apply the weighted correction formula to adjust the solution probability: The correction probability = (original probability + number of likes * w1 - number of likes * w2) / total count Wherein: w1 and w2 are the weights of likes and dislikes; the total count is the sum of the number of likes and dislikes.
[0012] Further, the search and recommendation module adopts the BERT-wwm model to realize semantic retrieval, supports fuzzy matching and multi-condition combined query.
[0013] Further, the model output layer is connected with a classifier and a generator dual channel, wherein the classifier is used for error type determination, and the generator is used for solution text generation.
[0014] Further, the window size adjustment strategy contains a threshold judgment mechanism: when continuous N times of matching fail, the window range is automatically expanded to a preset maximum value.
[0015] Further, the difference between the predicted solution and the real solution is minimized by using a cross-entropy loss function, wherein the cross-entropy loss function is: Loss=− i =1∑ Nyi log( y ^ i ) Wherein yi is the real label, y ^ i is the probability output by the model.
[0016] Further, the window size is initialized based on the error type, wherein the window size is: W = n e + c × σ e Wherein: n e Indicates the average number of context lines, which represents the number of lines usually occupied by the error type; c × σ e Used to adjust the additional number of lines to ensure the integrity of the context, so as to ensure that important context information is not missed during parsing.
[0017] The beneficial effects of the present application are: through the introduction of AI technology and user feedback, intelligent log analysis and real-time problem positioning are realized. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 It is a flowchart of the intelligent log analysis method based on the DevOps pipeline; Figure 2This is a flowchart for the implementation of the intelligent log parsing method based on the DevOps pipeline. DETAILED DESCRIPTION
[0019] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the following.
[0020] The features and performance of the present invention are further described in detail below with reference to the embodiments.
[0021] like Figure 1 As shown in the figure, the intelligent log parsing method based on the DevOps pipeline includes the following steps: Step 1: Build a structured knowledge base that interacts with the DevOps pipeline platform in real time through an API interface. The knowledge base includes a user information module, a log entry module, a solution module, a classification and labeling module, and a search and recommendation module. Step 2: Train the BERT deep learning model based on the Transformer architecture. The model receives log data through the input layer, processes it through the embedding layer and encoding layer, and outputs the matching solution through the decoding layer. Step 3: Dynamically extract context information based on error log features and determine the interception range using an adaptive window algorithm, where the window size is dynamically adjusted based on the error type and historical data. Step 4: Introduce a user feedback mechanism to weight the solution probability based on the number of likes / dislikes, and dynamically optimize model parameters through incremental training. Step 5: Form a closed-loop improvement system to feed back unmatched problems to the knowledge base for analysis, and update the model training data at the same time.
[0022] The construction of the structured knowledge base, which interacts with the DevOps pipeline platform in real time through an API interface, includes a user information module, a log entry module, a solution module, a classification and labeling module, and a search and recommendation module, including: Collect DevOps community Q&A data and store it by component type; modular data structure, in which the user information module records the identity and historical behavior of the feedback user; the log entry module stores standardized log format and error feature labels; the solution module associates error types with correction solutions; the classification and labeling module supports multi-dimensional retrieval; the search and recommendation module implements solution push based on semantic similarity.
[0023] The training is based on the BERT deep learning model based on the Transformer architecture. The model receives log data through the input layer, processes it through the embedding layer and the encoding layer, and outputs a matching solution through the decoding layer, including: The data preprocessing stage performs word segmentation, denoising and feature labeling on the log; the cross-entropy loss function is used to minimize the difference between the predicted solution and the real solution; the result feedback mechanism is used to include the unmatched cases into the training set; the online learning module is deployed to support real-time reception of user feedback data and update the model weight.
[0024] The context information is dynamically extracted based on the error log features, and the interception range is determined by the adaptive window algorithm, wherein the window size is dynamically adjusted according to the error type and historical data, including: The window size is initialized based on the error type, wherein a small window is used for frequent errors and a large window is used for complex errors; the window calculation strategy is optimized through a feedback loop, the strategy is a historical error frequency weighted calculation; the solution matching degree is dynamically adjusted; the context interception range is adaptively expanded.
[0025] The user feedback mechanism is introduced to correct the solution probability by weighting the number of likes / dislikes, and the model parameters are dynamically optimized through incremental training, including: adding a feedback count field in the solution record to record the number of likes and dislikes; applying a weighted correction formula to adjust the solution probability: Corrected probability = (original probability + like number x w1 - dislike number x w2) / total count Where: w1 and w2 are the weights of likes and dislikes; total count is the sum of the number of likes and dislikes.
[0026] The search and recommendation module uses the BERT-wwm model to realize semantic retrieval, supporting fuzzy matching and multi-condition combined query.
[0027] The model output layer is connected to the classifier and the generator dual channel, wherein the classifier is used for error type determination, and the generator is used for solution text generation.
[0028] The window size adjustment strategy contains a threshold judgment mechanism: when the continuous N times of matching fails, the window range is automatically expanded to the preset maximum value.
[0029] The cross-entropy loss function is used to minimize the difference between the predicted solution and the real solution, wherein the cross-entropy loss function is: Loss=− i =1∑ Nyi log( y ^ i ) Where yi is the real label, y ^ i is the probability output by the model.
[0030] The window size is initialized based on the error type, wherein the window size is: W = n e + c × σ e wherein: n e represents the average number of context lines, reflecting the number of lines usually occupied by this error type; c × σ e for adjusting the number of additional lines to ensure the integrity of the context, ensuring that important context information is not missed during parsing.
[0031] Specifically, the present application proposes an intelligent log parsing method based on DevOps pipeline, aiming to help users quickly locate problems and find solutions. Specifically, the method includes the following steps, as shown in Figure 2 : Building a knowledge base Collecting question and answer data in the DevOps community to establish a question and answer knowledge base for each type of component. The knowledge base is connected to the DevOps pipeline platform through an API interface for real-time updating and expansion.
[0032] To ensure the effectiveness and scalability of the knowledge base, the structured design should include the following main modules: (1) User information module (2) Log entry module (3) Solution module (4) Classification and labeling module (5) Search and recommendation module Building an AI model The technical solution adopts a deep learning model based on the Transformer architecture: BERT (Bidirectional Encoder Representations from Transformers). The model details are as follows: (1) Model structure Input layer: The intercepted log context information is used as input. It can be multiple lines of text, which form word vectors after tokenization.
[0033] Embedding Layer: Utilize pre-trained word embeddings (e.g., Word2Vec or GloVe) to convert text into vector form.
[0034] Encoding Layer: Use multiple Transformer encoders to process the input context logs, capturing relevant features and dependencies in the context.
[0035] Decoding Layer: Output solution predictions using a softmax layer to generate a probability distribution for each solution.
[0036] (2) Training Process: During model training, the following data needs to be prepared: Input: Context logs (multiple lines of text).
[0037] Related error component types Related error labels (e.g., "database connection failed").
[0038] Output: A set of solutions and their corresponding probabilities (e.g., "check database configuration," "confirm network connection," etc.).
[0039] Data format examples are as follows: (3) Algorithm: 1. Data Preprocessing: Preprocess each log, including removing stop words and standardization.
[0040] Combine context logs and error types to form input features.
[0041] 2. Model Training: Use the cross-entropy loss function to train the model to minimize the difference between the true solutions and the model's predicted solutions.
[0042] Loss = − i =1∑ Nyi log( y ^ i ) where yi is the true label, y ^ i is the model's output probability.
[0043] (4) Model Evaluation: Use evaluation metrics such as accuracy, recall, and F1-score to verify the model's performance.
[0044] Based on the evaluation results, continuously adjust the model parameters to optimize the model's performance.
[0045] (5) Solution output: After obtaining the model output of the context log, the output of 5 solutions and their probabilities is sorted by probability value.
[0046] (6) Result feedback: If the probability of each solution does not exceed the set threshold, it indicates that no effective match is found, and the problem is recorded in the knowledge base and relevant technical personnel are notified for further analysis and solution.
[0047] The solved problems and related information are automatically included in the training set to enhance the learning ability and accuracy of the model.
[0048] Dynamic extraction of log context In this technical solution, the following algorithm strategy can be used to automatically determine the interception range by analyzing the characteristics of the error log and the corresponding context information.
[0049] 1. Determine the context window size: According to past experience and data, first define a basic window size parameter N , where N represents the number of lines before and after the error occurs.
[0050] 2. Window size automatic dynamic adjustment: dynamically adjust the window size according to the error type and historical data. The specific method is as follows: Determine the window size In the collected knowledge base sample data, count the error identifiers and calculate the frequency and parse the corresponding log context size.
[0051] Let: f e : frequency of error type E (happening times / total log number); n e : average log context line number corresponding to error type E ; σ e : standard deviation of log context line number corresponding to error type E ; c : safety factor (can be set according to demand, usually takes value between 1 and 2); then the appropriate window size W can be represented as: W = n e + c × σ e Where: ne Represents the average number of context lines, reflecting the number of lines usually occupied by this error type.
[0052] c × σ e Adjusts the number of additional lines to ensure context integrity, ensuring that important context information is not missed during parsing.
[0053] 1) Determine window size Small window for frequent errors: For errors with high frequency and simple log content, a relatively small window can be set to quickly respond and handle.
[0054] Large window for rare or complex errors: For errors with low frequency and complex log content, increase the window size to obtain more comprehensive context information.
[0055] Feedback loop: By continuously collecting new error cases and their effect feedback, adjust the above mean and standard deviation n e And σ e , so as to optimize the calculation of window size.
[0056] 3. Calculate limits: If the current line's log index is W , the context extraction formula is: Context start index = max(0, i − W ); Context end index = i + W; 4. Extract context logs: Use the start and end indices calculated above to extract context information from the log records.
[0057] User feedback mechanism Introduce the number of user likes as feedback to correct the model, the main steps are: Data structure update: Add like and dislike number fields in the solution record, such as: Weighted correction mechanism: Correction probability = (original probability + like number × w 1 - Like number ×w 2 ) / total count Where: w1 and w2 are the weights of likes and dislikes; total count is the sum of likes and dislikes, used for normalization.
[0058] Online learning: Periodically collect user feedback data for incremental training, dynamically adjust model parameters based on feedback.
[0059] Specific steps are as follows: 1、Feedback data collection: Collect user's like and dislike data regularly, and match it with the previous model output.
[0060] 2、Incremental training: Use the collected feedback data for incremental training, adjust the parameters of the model to enhance the recommendation of effective solutions and reduce the probability of inefficient solutions.
[0061] 3、Re-evaluate: Re-evaluate the performance of the model by increasing the effect of likes and reducing the effect of dislikes, and periodically learn from new feedback.
Claims
1. An intelligent log parsing method based on DevOps pipeline, characterized by: The steps include: Step 1: Build a structured knowledge base that interacts with the DevOps pipeline platform in real time through an API interface. The knowledge base includes a user information module, a log entry module, a solution module, a classification and labeling module, and a search and recommendation module. Step 2: Train the BERT deep learning model based on the Transformer architecture. The model receives log data through the input layer, processes it through the embedding layer and encoding layer, and outputs the matching solution through the decoding layer. Step 3: Dynamically extract context information based on error log features and determine the interception range using an adaptive window algorithm, where the window size is dynamically adjusted based on the error type and historical data. Step 4: Introduce a user feedback mechanism to weight the solution probability based on the number of likes / dislikes, and dynamically optimize model parameters through incremental training. Step 5: Form a closed-loop improvement system to feed back unmatched problems to the knowledge base for analysis, and update the model training data at the same time.
2. The intelligent log parsing method based on the DevOps pipeline according to claim 1 is characterized in that: The construction of the structured knowledge base, which interacts with the DevOps pipeline platform in real time through an API interface, includes a user information module, a log entry module, a solution module, a classification and labeling module, and a search and recommendation module, including: Collect DevOps community Q&A data and store it by component type; modular data structure, in which the user information module records the identity and historical behavior of the feedback user; the log entry module stores standardized log format and error feature labels; the solution module associates error types with correction solutions; the classification and labeling module supports multi-dimensional retrieval; the search and recommendation module implements solution push based on semantic similarity.
3. The intelligent log parsing method based on the DevOps pipeline according to claim 1 is characterized in that: The training is based on the BERT deep learning model based on the Transformer architecture. The model receives log data through the input layer, processes it through the embedding layer and the encoding layer, and outputs a matching solution through the decoding layer, including: During the data preprocessing phase, the logs are segmented, denoised, and feature labeled; the cross-entropy loss function is used to minimize the difference between the predicted solution and the actual solution; unmatched cases are included in the training set through a result feedback mechanism; and an online learning module is deployed to support real-time reception of user feedback data and update model weights.
4. The intelligent log parsing method based on the DevOps pipeline according to claim 1, characterized in that: The dynamic extraction of context information based on error log features and the determination of the interception range through an adaptive window algorithm, wherein the window size is dynamically adjusted according to the error type and historical data, include: The window size is initialized based on the error type, with small windows used for frequent errors and large windows for complex errors. The window calculation strategy is optimized through a feedback loop, with the strategy weighted by historical error frequency. The solution matching degree is dynamically adjusted, and the context interception range is adaptively expanded.
5. The intelligent log parsing method based on DevOps pipeline according to claim 1 is characterized in that: The user feedback mechanism described above weights the solution probability based on the number of likes and dislikes, and dynamically optimizes model parameters through incremental training. This includes: adding a feedback count field to the solution record to record the number of likes and dislikes; and applying a weighted correction formula to adjust the solution probability: Corrected probability = (original probability + number of likes × w1 - number of likes × w2) / total count Where: w1 and w2 are the weights of likes and dislikes; the total count is the sum of the number of likes and dislikes.
6. The intelligent log parsing method based on the DevOps pipeline according to claim 2 is characterized in that: The search and recommendation module uses the BERT-wwm model to implement semantic retrieval, supporting fuzzy matching and multi-condition combination queries.
7. The intelligent log parsing method based on DevOps pipeline according to claim 3 is characterized in that: The model output layer connects the classifier and the generator dual channels, where the classifier is used to determine the error type and the generator is used to generate the solution text.
8. The intelligent log parsing method based on DevOps pipeline according to claim 4 is characterized in that: The window size adjustment strategy includes a threshold judgment mechanism: when N consecutive matching failures occur, the window range is automatically expanded to a preset maximum value.
9. The intelligent log parsing method based on DevOps pipeline according to claim 3, characterized in that: The cross entropy loss function is used to minimize the difference between the predicted solution and the true solution, where the cross entropy loss function is: Loss=− i =1∑ Nyi log( y ^ i ) in yi is the true label, y ^ i is the probability of the model output.
10. The intelligent log parsing method based on DevOps pipeline according to claim 4, characterized in that: The window size is initialized based on the error type, where the window size is: W = n e + c × σ e in: n e Indicates the average number of context lines, reflecting the number of lines occupied by this error type; c × σ e Used to adjust the number of extra lines to ensure context completeness, ensuring that important context information is not missed during parsing.