Operation and maintenance log event association analysis method and system based on artificial intelligence

By performing feature extraction and correlation analysis on the operation and maintenance logs, and generating operation and maintenance event optimization strategies, the problem of inability to effectively understand the causes and propagation of system failures in the existing technology is solved, and the system operation and maintenance efficiency and stability are improved.

CN120276908AActive Publication Date: 2025-07-08SHANGHAI QINGCHUANG INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510758040.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The existing operation and maintenance log analysis methods cannot effectively mine the causal relationships and event propagation paths between logs, resulting in a lack of scientific basis for operation and maintenance decisions, which affects the stability and reliability of the system.

Method used

By obtaining the original operation and maintenance log collection, feature extraction, context semantic features and timeline behavior feature analysis, the pre-trained association analysis model is called to generate an association rule set, the operation and maintenance event optimization strategy is generated and feedback to the target system to trigger system optimization operations.

Benefits of technology

It realizes efficient analysis of operation and maintenance log events, accurately locates the root causes of system failures, improves the system operation and maintenance efficiency and stability, and reduces the losses caused by failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276908A_ABST
    Figure CN120276908A_ABST
Patent Text Reader

Abstract

The invention provides an operation and maintenance log event association analysis method and system based on artificial intelligence, and the method comprises the steps: firstly obtaining an original operation and maintenance log set of a target system, the original operation and maintenance log set comprising a plurality of operation and maintenance log fragments; secondly, performing feature extraction processing on the original operation and maintenance log set to obtain context semantic features and time sequence behavior features of each operation and maintenance log fragment, and then calling a pre-trained association analysis model to perform association rule matching processing on the features to generate an association rule set among the operation and maintenance log fragments; the association rule set comprises a causal relationship definition and an event propagation path description, generating an operation and maintenance event optimization strategy based on the association rule set, and finally feeding back the operation and maintenance event optimization strategy to an operation and maintenance management interface of a target system to trigger system optimization operation. And the system operation and maintenance efficiency and stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method and system for associative analysis of operation and maintenance log events based on artificial intelligence. Background Art

[0002] In modern complex information technology systems, operation and maintenance logs record various status and event information during the operation of the system, and are an important basis for system operation and maintenance and fault troubleshooting. However, with the continuous expansion of the system scale and the increasing complexity, the amount of operation and maintenance log data has increased explosively, and traditional operation and maintenance log analysis methods have gradually revealed many drawbacks.

[0003] Existing methods often only perform isolated analysis on individual operation and maintenance log segments, ignoring the potential association relationships between different log segments, resulting in an inability to comprehensively and accurately understand the causes and propagation processes of system failures. In the face of complex system failures, it is difficult for operation and maintenance personnel to quickly locate key information from a large amount of log data, and then formulate effective operation and maintenance strategies. In addition, traditional methods lack comprehensive consideration of the context semantics and temporal behavior of operation and maintenance logs, and cannot accurately mine the causal relationships and event propagation paths between logs, making operation and maintenance decisions lack a scientific basis, difficult to achieve timely optimization of the system and rapid recovery of faults, and seriously affecting the stability and reliability of the system. Summary of the Invention

[0004] In view of the above-mentioned problems, in combination with the first aspect of the present invention, embodiments of the present invention provide a method for associative analysis of operation and maintenance log events based on artificial intelligence, the method comprising: Obtain an original operation and maintenance log set of a target system, the original operation and maintenance log set including a plurality of operation and maintenance log segments; Perform feature extraction processing on the original operation and maintenance log set to obtain the context semantic features and temporal behavior features of each operation and maintenance log segment; Call a pre-trained association analysis model to perform association rule matching processing on the context semantic features and temporal behavior features, and generate an association rule set between the operation and maintenance log segments, the association rule set including the definition of causal relationships and the description of event propagation paths of the operation and maintenance log segments in the system failure link; Generate an operation and maintenance event optimization strategy based on the association rule set, the operation and maintenance event optimization strategy including the event response priority and system resource adjustment plan for each association rule in the association rule set; Feed back the operation and maintenance event optimization strategy to the operation and maintenance management interface of the target system to trigger system optimization operations.

[0005] In another aspect, an embodiment of the present invention further provides an operation and maintenance log event correlation analysis system based on artificial intelligence, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0006] Based on the above aspects, the embodiment of the present invention realizes the efficient analysis of the correlation of operation and maintenance log events. First, the original operation and maintenance log set is obtained and feature extraction is performed to obtain the context semantic features and temporal behavior features of each operation and maintenance log segment. Then, the pre-trained correlation analysis model is called to perform correlation rule matching processing on these features, and a set of correlation rules including the causal relationship definition and event propagation path description of the operation and maintenance log segments in the system fault link is generated, deeply revealing the internal connection between the log segments. Based on this set of correlation rules, an operation and maintenance event optimization strategy is generated, including the event response priority and system resource adjustment plan for each correlation rule. Finally, the operation and maintenance event optimization strategy is fed back to the operation and maintenance management interface of the target system to trigger system optimization operations, which can make full use of the information in the operation and maintenance logs, accurately locate the root cause of system failures, quickly formulate effective operation and maintenance strategies, significantly improve the operation and maintenance efficiency and stability of the system, and reduce the losses caused by system failures. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 is a schematic execution flow diagram of the operation and maintenance log event correlation analysis method based on artificial intelligence provided by the embodiment of the present invention.

[0008] Figure 2 is a schematic diagram of exemplary hardware and software components of the operation and maintenance log event correlation analysis system based on artificial intelligence provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0009] The present invention will be specifically described below in conjunction with the accompanying drawings of the specification. Figure 1 is a schematic flow diagram of the operation and maintenance log event correlation analysis method based on artificial intelligence provided by an embodiment of the present invention. The operation and maintenance log event correlation analysis method based on artificial intelligence will be introduced in detail below.

[0010] Step S110: Obtain the original operation and maintenance log set of the target system, and the original operation and maintenance log set includes a plurality of operation and maintenance log segments.

[0011] In this embodiment, the target system can be a financial trading system, which covers multiple key components such as the trading front-end, market data server, order processing engine, clearing system, and database cluster. These components generate rich and diverse operation and maintenance logs during operation. The trading front-end records users' trading requests, such as buy and sell instructions, the time, quantity, price, etc. of placing orders; the market data server records real-time market data, including stock prices, trading volumes, increases and decreases, etc.; the order processing engine records the order processing flow, such as order matching and execution status; the clearing system records transaction settlement information, such as fund transfers and securities deliveries; the database cluster records data read and write operations, backup situations, etc.

[0012] To collect these logs, the target system adopts a log collection architecture that combines centralized and distributed methods. Lightweight log collection agents are deployed on each component, and these agents send the logs generated by the components to the local log buffer. Then, the log forwarder centrally transmits the log data in the local buffer to the log storage center. The log storage center uses a distributed file system and database to store a large amount of log data, ensuring the security and scalability of the logs. Finally, the original operation and maintenance log set is obtained from the log storage center, and this set contains a large number of operation and maintenance log segments with different formats and contents.

[0013] Step S120: Perform feature extraction processing on the original operation and maintenance log set to obtain the context semantic features and temporal behavior features of each operation and maintenance log segment.

[0014] To mine valuable information from the original operation and maintenance log set, feature extraction is required. This step includes multiple sub-steps, which are detailed below.

[0015] Step S121: Perform syntactic structure parsing processing on the log statements in the operation and maintenance log segment, and split the log statements into multiple log word units. The log word units include system component identifiers, operation instruction keywords, and status descriptors.

[0016] Taking a log statement in the financial trading system as an example: "[2024-11-25 10:30:15] -[OrderProcessingEngine - Instance2] - [EXECUTE ORDER - BUY 100 SHARES OF ABCSTOCK AT $50] - [Status: PENDING] - [Error: Insufficient funds in account]".

[0017] When parsing the syntactic structure of this log statement, it is divided into different parts based on the delimiters in the log (such as square brackets, hyphens, etc.). In "[OrderProcessingEngine - Instance2]", "OrderProcessingEngine - Instance2" is the system component identifier, indicating that this log is generated by the second instance of the order processing engine. In "[EXECUTE ORDER - BUY 100 SHARES OF ABC STOCK AT $50]", "EXECUTEORDER" and "BUY" are operation instruction keywords, describing the specific trading operation. In "[Status:PENDING]", "PENDING" and "[Error: Insufficient funds in account]" are status descriptors. The former indicates the current processing status of the order, and the latter explains the problem encountered in order processing.

[0018] During the parsing process, considering the diversity of log formats, the methods of regular expression matching and pattern recognition are adopted. For example, through the regular expression "\[(.*?)\]", all the information enclosed by square brackets can be matched, and then the system component identifier, operation instruction keyword, and status descriptor are further distinguished according to the characteristics of different information. For ambiguous information, it is judged in combination with the context and predefined rules.

[0019] Step S122: Call the pre-trained feature encoder to perform context encoding processing on the multiple log word units to generate the semantic vector of the operation and maintenance log segment. The context encoding processing constructs a global semantic representation by capturing the dependency relationships between adjacent log word units.

[0020] After obtaining multiple log word units, they need to be converted into semantic vectors that can be understood and processed by the computer. Here, a pre-trained feature encoder is used. This feature encoder is trained based on a large amount of log data and has good semantic representation ability. This step further includes the following sub-steps: Step S1221: Input the log word units into the word embedding layer of the pre-trained feature encoder, and convert the discrete log word units into initial word vectors in the continuous space through distributed vector mapping.

[0021] Input the split log word units, such as "OrderProcessingEngine - Instance2", "EXECUTEORDER", "BUY", etc., into the word embedding layer of the pre-trained feature encoder. The word embedding layer uses the method of distributed vector mapping to map each discrete log word unit into a continuous vector space to form initial word vectors. For example, "OrderProcessingEngine - Instance2" will be mapped to a vector with a set dimension, and this vector contains the semantic information of this log word unit. The word embedding layer makes the log word units with similar semantics closer in the vector space by learning the lexical co-occurrence relationships in a large amount of log data.

[0022] Step S1222: Input the initial word vectors into the bidirectional attention layer of the feature encoder for context correlation analysis processing. The bidirectional attention layer calculates the semantic interaction weights between different log word units through the self-attention mechanism.

[0023] Input the obtained initial word vectors into the bidirectional attention layer of the feature encoder. The bidirectional attention layer uses the self-attention mechanism to analyze the semantic interaction relationships between different log word units. The self-attention mechanism determines the semantic interaction weights between them by calculating the similarity between each log word unit and other log word units. For example, there may be a strong semantic association between "EXECUTEORDER" and "BUY" because "BUY" is a specific form of the "EXECUTE ORDER" operation. The bidirectional attention layer will calculate a relatively high semantic interaction weight between these two log word units, indicating their close semantic connection. During the calculation process, the order and context information of the log word units will be taken into account to capture more accurate semantic interaction relationships.

[0024] Step S1223: Weightedly fuse the initial word vectors according to the semantic interaction weights to generate context-aware vectors for each log word unit. The context-aware vectors contain the context-dependent information of the log word unit in the entire log statement.

[0025] Based on the semantic interaction weights calculated by the bidirectional attention layer, the initial word vectors are weighted and fused. Specifically, for the initial word vector of each log word unit, a weighted sum is performed according to its semantic interaction weights with other log word units. For example, for the initial word vector of the log word unit "EXECUTE ORDER", it combines its semantic interaction weights with other log word units such as "BUY" and "100 SHARES OF ABC STOCK", and performs weighted fusion on the initial word vectors of these related log word units to obtain the context-aware vector of "EXECUTE ORDER". This context-aware vector contains the context-dependent information of "EXECUTE ORDER" in the entire log statement, that is, it considers its semantic relationship with other log word units.

[0026] Step S1224: Perform pooling on the context-aware vector, and generate a semantic vector with a unified dimension through a combined operation of max pooling and average pooling, where the window size and sliding step of the pooling operation are consistent for all operation and maintenance log segments, ensuring that the semantic vectors of different log segments have the same feature dimension.

[0027] Perform pooling on the obtained context-aware vector. A combined operation of max pooling and average pooling is used to generate a semantic vector with a unified dimension. Max pooling selects the maximum value within each pooling window as the output of the window, which can highlight the important features in the vector. Average pooling calculates the average value of the vector elements within each pooling window, which can smooth the features of the vector. For example, the context-aware vector is divided according to the set window size and sliding step, and the max pooling and average pooling operations are respectively performed on the vectors within each window, and then these two results are concatenated or combined in other ways to obtain a semantic vector with a unified dimension. To ensure that the semantic vectors of different operation and maintenance log segments have the same feature dimension, the window size and sliding step of the pooling operation are consistent for all operation and maintenance log segments.

[0028] Step S1225: Input the semantic vector into the output layer of the feature encoder for non-linear transformation to obtain a standardized semantic vector representation, and the standardization process eliminates the feature scale differences between different log segments.

[0029] Input the semantic vector obtained through pooling into the output layer of the feature encoder. The output layer will perform a non-linear transformation on the semantic vector, processing the vector elements through an activation function (such as ReLU, etc.) to enhance the non-linear expression ability of the model. At the same time, the output layer will also perform normalization processing to eliminate the feature scale differences between different log segments. For example, the semantic vectors of different log segments may have significant differences in the numerical range. Through normalization processing, the feature scales of these vectors can be unified into a suitable range, making subsequent analysis and processing more stable and accurate.

[0030] Step S123: Determine the context correlation degree between the operation and maintenance log segments based on the semantic vectors, and construct context semantic features according to the context correlation degree. The context correlation degree is used to quantify the logical relevance of different operation and maintenance log segments in the fault propagation path.

[0031] After obtaining the semantic vectors of the operation and maintenance log segments, in order to quantify the logical relevance of different operation and maintenance log segments in the fault propagation path, it is necessary to determine their context correlation degree and construct context semantic features based on this. This step includes the following sub-steps: Step S1231: Calculate the similarity score between the semantic vectors. The similarity score is obtained through vector inner product operation combined with normalization processing.

[0032] For each operation and maintenance log segment, its corresponding semantic vector has been obtained through the previous steps. Let the semantic vector of operation and maintenance log segment A be Va, and the semantic vector of operation and maintenance log segment B be Vb. In order to calculate the similarity score between these two semantic vectors, first perform normalization processing on the vectors. The purpose of normalization is to eliminate the influence of vector length on similarity calculation, making the similarity calculation focus more on the direction of the vectors.

[0033] For vector Va, the calculation method of its normalized vector Va' is: each element in Va' is equal to the corresponding element in Va divided by the modulus length of Va. Similarly, the normalized vector Vb' of Vb can be obtained.

[0034] Then, calculate the inner product of the normalized vectors Va' and Vb'. The vector inner product is to sum the products of the corresponding elements of the two vectors. The larger the result of the inner product, the closer the directions of the two vectors are, that is, the more similar the semantics are. The similarity score obtained in this way ranges from -1 to 1, and the value closer to 1 indicates that the semantics of the two operation and maintenance log segments are more similar.

[0035] Step S1232: Perform temporal sliding window smoothing processing on the similarity score to generate the smoothed context correlation degree.

[0036] Since the similarity score may be affected by noise or transient fluctuations, in order to obtain a more stable and reliable context correlation, it is necessary to perform a temporal sliding window smoothing process on the similarity score.

[0037] Set a temporal sliding window of a fixed size, assuming the window size is W. On the similarity score sequence, starting from the first score, sequentially select consecutive W similarity scores. For the scores within each window, calculate their average or weighted average. For example, for the similarity score at the i-th position, select a total of W scores before and after it as the center, and calculate the average of these scores as the smoothed context correlation at the i-th position.

[0038] By continuously moving the sliding window on the similarity score sequence and performing smoothing processing on each position, a smoothed context correlation sequence is finally obtained. This smoothing process can reduce the interference of noise and make the context correlation better reflect the long-term stable logical correlation between operation and maintenance log segments.

[0039] Step S1233: Concatenate the context correlation with the semantic vector, and map features of different dimensions to a unified semantic space through a fully connected network.

[0040] Concatenate the smoothed context correlation with the semantic vector. Assume the dimension of the semantic vector is D, and the context correlation is a one-dimensional value. When concatenating, add the context correlation as a new dimension to the end of the semantic vector to form a new feature vector with a dimension of D + 1.

[0041] Input the concatenated feature vector into the fully connected network. The fully connected network consists of multiple neuron layers, and each neuron is connected to all neurons in the previous layer. In the fully connected network, through a series of linear transformations and non-linear activation functions, features of different dimensions are combined and transformed, and finally mapped into a unified semantic space.

[0042] In this unified semantic space, the semantic relationships between different features are clearer and more comparable. For example, originally the semantic vector and the context correlation may represent information in different aspects. Through the mapping of the fully connected network, they can be better integrated in the new space, providing a more effective feature representation for subsequent clustering and pattern recognition.

[0043] Step S1234: In the unified semantic space, use a semi-supervised clustering algorithm to perform pattern partitioning on the concatenated features, where the cluster centers are initialized as the feature vectors corresponding to predefined system component role identifiers, identify the semantic grouping of the operation and maintenance log segments in the system failure scenario and associate them with specific role identifiers to obtain the semantic grouping result.

[0044] In the unified semantic space, a semi-supervised clustering algorithm is used to perform pattern partitioning on the spliced features. The semi-supervised clustering algorithm combines the advantages of supervised learning and unsupervised learning, and uses predefined information to guide the clustering process.

[0045] First, according to the architecture and component functions of the system, a series of system component role identifiers are predefined, such as "database server", "application server", "network device", etc., and corresponding feature vectors are defined for each role identifier. These feature vectors can be obtained through the analysis and statistics of historical operation and maintenance log data, or set by domain experts based on experience.

[0046] These predefined feature vectors are used as the initial values of the clustering centers. Then, the semi-supervised clustering algorithm will assign the operation and maintenance log segments to different clusters according to the distances between the spliced feature vectors and these initial clustering centers. During the assignment process, the algorithm will continuously adjust the positions of the clustering centers to make the feature vectors within each cluster more similar and the similarity between different clusters lower.

[0047] In this way, the semantic grouping of the operation and maintenance log segments in the system fault scenario is identified. Each semantic grouping corresponds to a specific system component role identifier, and the operation and maintenance log segments are associated with these role identifiers to obtain the semantic grouping result. For example, if the operation and maintenance log segments in a certain cluster are all related to the "database server", then these log segments are associated with the role identifier of "database server".

[0048] Step S1235: Construct the context semantic features based on the semantic grouping result, where the context semantic features include the role identifier and the semantic weight assignment result of the operation and maintenance log segments in the system operation context.

[0049] Construct the context semantic features according to the semantic grouping result. The context semantic features need to include the role identifier of the operation and maintenance log segments in the system operation context and the semantic weight assignment result.

[0050] For each operation and maintenance log segment, determine its corresponding system component role identifier according to the semantic grouping it belongs to. For example, if a certain log segment is assigned to the semantic grouping of "database server", then its role identifier is "database server".

[0051] The semantic weight assignment reflects the importance of the log segment in the semantic grouping it belongs to. The semantic weight can be determined according to the distance between the log segment and the clustering center. The closer the distance to the clustering center, the more similar the log segment is to the features of the semantic grouping, and the higher its semantic weight; conversely, the farther the distance from the clustering center, the lower the semantic weight.

[0052] When calculating the semantic weight specifically, a normalization method can be adopted. Let the distance between the log segment and the cluster center be d, the maximum distance between all log segments and the cluster center in this semantic group be Dmax, and the minimum distance be Dmin. Then the semantic weight w of this log segment can be calculated by the formula w = (Dmax - d) / (Dmax - Dmin), and the obtained semantic weight ranges from 0 to 1.

[0053] Combine the role identifier and semantic weight information to construct the context semantic feature. This context semantic feature can accurately reflect the role and semantic importance of the operation and maintenance log segment in the system operation context.

[0054] Step S124: Perform a temporal distribution analysis on the generation timestamp of the operation and maintenance log segment, extract the temporal interval feature and event trigger frequency feature of the operation and maintenance log segment, and fuse the temporal interval feature and event trigger frequency feature to generate the temporal behavior feature of the operation and maintenance log segment. The temporal interval feature reflects the continuity and suddenness patterns of event triggers, and the event trigger frequency feature reflects the event density within a preset time window. The fusion process eliminates redundant information through a feature dimensionality reduction algorithm and retains the key representations of the temporal pattern.

[0055] After obtaining the operation and maintenance log segment, its generation timestamp contains important temporal information. Conducting in-depth temporal distribution analysis on these timestamps can extract features that reflect the event trigger rules. This step includes the following sub-steps: Step S1241: Extract the set of generation timestamps of the operation and maintenance log segment, and sort the set of generation timestamps in chronological order.

[0056] From the original operation and maintenance log segment data, extract the generation timestamp corresponding to each log segment to form a set of generation timestamps. These timestamps record the specific moments when log events occur. Since subsequent analysis needs to be carried out in chronological order, the set of generation timestamps needs to be sorted.

[0057] Common sorting algorithms such as quicksort and mergesort can be used for the sorting process. Taking quicksort as an example, its basic idea is to select a reference timestamp, divide the set of timestamps into two parts, one part is less than the reference timestamp, and the other part is greater than the reference timestamp, and then recursively sort these two parts respectively. Finally, a set of generation timestamps arranged in chronological order is obtained. Such a sorted set can clearly show the occurrence order of log events and lay a foundation for subsequent temporal analysis.

[0058] Step S1242: Calculate the interval duration between adjacent generated timestamps in the sorted set of generated timestamps. The interval duration reflects the continuity and suddenness of event triggering.

[0059] After obtaining the set of generated timestamps arranged in chronological order, calculate the interval duration between adjacent generated timestamps in sequence. Let the sorted set of generated timestamps be T = {t1, t2, t3, …, tn}, then the set of interval durations of adjacent timestamps ΔT = {Δt1, Δt2, …, Δtn-1}, where Δti = ti+1 - ti (i = 1, 2, …, n - 1).

[0060] These interval durations contain information about the continuity and suddenness of event triggering. If the interval durations are relatively stable and short, it indicates that event triggering has strong continuity, and it may be that the system is in a stable operating mode, and events occur in sequence according to a certain rhythm; if the interval durations fluctuate greatly, with short and long intervals alternating, it indicates that event triggering has suddenness, and it may be that the system is affected by external interference or there are internal anomalies, resulting in irregular event occurrence times.

[0061] Step S1243: Perform distribution statistical processing on the interval durations, and extract the timing interval features characterizing the event triggering pattern. The timing interval features include the central tendency index and the dispersion degree index of the interval duration.

[0062] To describe the event triggering pattern more comprehensively, it is necessary to perform distribution statistical processing on the set of interval durations and extract the central tendency index and the dispersion degree index therein.

[0063] The central tendency index is used to describe the typical value of the interval duration. Common central tendency indexes include the average value and the median. The average value is obtained by adding all the interval durations and dividing by the number of interval durations, which reflects the overall average level of the interval duration. The median is the value located in the middle position after arranging the interval durations in ascending order (if the number is even, take the average of the two middle numbers), which is not affected by extreme values and can better represent the middle level of the data.

[0064] The dispersion degree index is used to describe the fluctuation of the interval duration. Common dispersion degree indexes include the standard deviation and the variance. The standard deviation is the square root of the variance, and the variance is the average of the squares of the differences between each interval duration and the average value. The larger the standard deviation and the variance, the greater the fluctuation of the interval duration and the worse the regularity of event triggering; conversely, the smaller the standard deviation and the variance, the more stable the interval duration and the more regular the event triggering.

[0065] Through these central tendency indexes and dispersion degree indexes, the timing interval features of event triggering can be comprehensively characterized.

[0066] Step S1244: Count the number of event triggering of the operation and maintenance log segment within a preset time window, and generate an event triggering frequency feature reflecting the event density.

[0067] Set a preset time window, such as in hours, days, etc. In this time window, count the number of event triggers of the operation and maintenance log fragment. Let the start time of the time window be t_start and the end time be t_end, traverse the sorted generation timestamp set T, and count the number of timestamps that satisfy t_start≤ti≤t_end. This number is the number of event triggers in this time window.

[0068] By changing the size and position of the time window, the number of event triggers in different time periods is counted to obtain a series of event trigger count values. These event trigger count values ​​constitute the event trigger frequency characteristics, which reflect the density of event occurrence in different time periods. If the number of event triggers in a certain time window is large, it means that the system activities are frequent in this time period, and there may be abnormal situations or business peaks; on the contrary, if the number of event triggers is small, it means that the system is in a relatively stable or idle state.

[0069] Step S1245: Normalize the time series interval characteristics and the event trigger frequency characteristics, and perform dimensionality reduction processing on the normalized time series interval characteristics and the event trigger frequency characteristics through the principal component analysis method to generate the time series behavior characteristics of the operation and maintenance log fragment, and the dimensionality reduction processing retains the feature component with the largest variance in the original time series data.

[0070] Since the time interval features and event trigger frequency features may have different dimensions and value ranges, they need to be normalized to make them comparable in subsequent analysis. Normalization can map the feature value to a set range. Common normalization methods include minimum-maximum normalization and z-score normalization.

[0071] Taking minimum-maximum normalization as an example, for a feature vector X={x1, x2, ..., xn}, its normalized vector X'={x1', x2', ..., xn'}, where xi'=(xi-min(X)) / (max(X)-min(X)), min(X) and max(X) are the minimum and maximum values ​​in the feature vector X. In this way, both the time interval feature and the event trigger frequency feature are normalized to the range of [0, 1].

[0072] After normalization, the principal component analysis (PCA) method is used to reduce the dimensionality of these two features. The basic idea of PCA is to find the directions with the largest variance in the original data and project the data onto these directions to achieve dimensionality reduction. The specific steps are as follows: First, calculate the covariance matrix of the normalized time interval features and event trigger frequency features. The covariance matrix reflects the correlation between each feature.

[0073] Then, perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and eigenvectors. The eigenvalue represents the variance size corresponding to each eigenvector, and the eigenvector represents the projection of the data in that direction.

[0074] Next, sort the eigenvectors in descending order of eigenvalues and select the top k eigenvectors, where k is the feature dimension to be retained after dimensionality reduction. The directions corresponding to these k eigenvectors are the directions with the largest variance.

[0075] Finally, project the normalized feature data onto these k eigenvectors to obtain the dimensionality-reduced feature data, that is, the time-series behavior features of the operation and maintenance log segments. Through this dimensionality reduction process, redundant information in the original features is eliminated, and the key representations of the time-series patterns are retained, making subsequent analysis more efficient and accurate.

[0076] Step S130: Invoke the pre-trained association analysis model to perform association rule matching on the context semantic features and time-series behavior features, and generate a set of association rules between the operation and maintenance log segments. The set of association rules includes the definition of the causal relationship and the description of the event propagation path of the operation and maintenance log segments in the system fault link.

[0077] After obtaining the context semantic features and time-series behavior features, it is necessary to invoke the pre-trained association analysis model to perform association rule matching to generate a set of association rules between the operation and maintenance log segments. This step includes the following sub-steps: Step S131: Input the context semantic features and time-series behavior features into the feature fusion layer of the pre-trained association analysis model to obtain the fused feature vectors. The feature fusion layer uses a gated recurrent unit to dynamically adjust the contribution ratio of different feature types.

[0078] Input the context semantic features and temporal behavior features into the feature fusion layer of the pre-trained correlation analysis model. The feature fusion layer uses a gated recurrent unit (GRU) to dynamically adjust the contribution ratios of different feature types. GRU is a type of recurrent neural network that contains an update gate and a reset gate, and can dynamically determine which information needs to be retained and which information needs to be updated according to the input features. For example, for context semantic features and temporal behavior features, GRU will dynamically adjust their weights in the fusion process according to their importance in the current task. Through the processing of GRU, the context semantic features and temporal behavior features are fused to obtain a fused feature vector.

[0079] Step S132: Perform a weight assignment process on the fused feature vector through the attention mechanism layer of the correlation analysis model to obtain a weight assignment result, and the weight assignment process captures the dependencies across log segments and the causal strength of event triggering.

[0080] Input the fused feature vector into the attention mechanism layer of the correlation analysis model. The role of the attention mechanism layer is to perform weight assignment on the fused feature vector to capture the dependencies across log segments and the causal strength of event triggering. This step further includes the following sub-steps: Step S1321: Perform a linear projection process on the fused feature vector to generate a query vector and a key vector, and the linear projection process realizes feature space transformation through a trainable parameter matrix.

[0081] In the attention mechanism layer of the correlation analysis model, for the input fused feature vector, a linear projection process is first performed. This linear projection is completed through a trainable parameter matrix. Assume that the fused feature vector is represented by the letter F, and there are two trainable parameter matrices, denoted as Wq and Wk respectively. Through matrix multiplication operation, multiplying the fused feature vector F by the parameter matrix Wq, the query vector Q is obtained, that is, Q = F×Wq; similarly, multiplying the fused feature vector F by the parameter matrix Wk, the key vector K is obtained, that is, K = F×Wk. The purpose of this linear projection process is to map the fused feature vector into different feature spaces for subsequent calculation of similarity scores.

[0082] Step S1322: Calculate the similarity score between the query vector and the key vector, and the similarity score is obtained through a scaled dot product operation.

[0083] After obtaining the query vector Q and the key vector K, it is necessary to calculate the similarity score between them. The scaled dot - product operation is used to achieve this calculation. Specifically, first calculate the dot - product of the query vector Q and the key vector K, that is, multiply each element in the query vector Q by the corresponding element in the key vector K, and then sum these products to obtain an unscaled dot - product result. To avoid the dot - product result being too large, it is also necessary to perform a scaling process. Assume the scaling factor is dk, and dk is usually the dimension of the key vector K. Divide the unscaled dot - product result by the square root of dk to obtain the final similarity score. This similarity score reflects the degree of similarity between the query vector and the key vector, and the higher the score, the more similar the two are.

[0084] Step S1323: Normalize the similarity score to convert the original score into attention weights in the form of a probability distribution.

[0085] After obtaining the similarity score, it is necessary to normalize it to convert the original similarity score into attention weights in the form of a probability distribution. The normalization process is implemented using the softmax function. The softmax function takes each similarity score as input and, through exponential operations and normalization operations, converts it into a value between 0 and 1, and the sum of all attention weights is 1. In this way, the similarity score is converted into a probability distribution form, and each attention weight represents the importance of the corresponding key vector in the subsequent weighted summation process.

[0086] Step S1324: Perform a weighted sum of the key vectors according to the attention weights to generate a context - enhanced feature representation.

[0087] Based on the normalized attention weights, perform a weighted sum operation on the key vectors. Specifically, multiply each key vector by the corresponding attention weight, and then sum these weighted key vectors to obtain a new vector, which is the context - enhanced feature representation. Through this weighted sum method, the information of the key vectors with higher similarity to the query vector can be highlighted, thereby enhancing the context relevance of the feature representation.

[0088] Step S1325: Perform a residual connection between the context - enhanced feature representation and the original fused feature vector, and capture the dependency patterns at different abstraction levels through a multi - layer attention stacking structure to generate a weight assignment result reflecting the potential relationship of the event chain.

[0089] Perform a residual connection between the context-enhanced feature representation and the original fused feature vector. The residual connection is to add the context-enhanced feature representation and the original fused feature vector to obtain a new vector. This connection method helps to alleviate the vanishing gradient problem, enabling the model to better learn the differences between features. Then, through a multi-layer attention stacking structure, that is, stacking multiple attention mechanism layers in sequence, further capture the dependency patterns at different abstraction levels. Each layer of the attention mechanism layer can learn feature information at different levels, and through multi-layer stacking, the potential relationships between event chains can be captured more comprehensively. Finally, after being processed by the multi-layer attention stacking structure, a weight assignment result reflecting the potential relationships between event chains is generated.

[0090] Step S133: Aggregate the fused feature vectors according to the weight assignment result to generate event-associated features. The aggregation process uses a graph neural network to mine the topological association structure between log events, where the graph nodes are defined as the semantic vectors of operation and maintenance log segments, the graph edge weights are jointly determined by the temporal interval features and context relevance between log segments, and the edge connection range is restricted within a preset time proximity threshold.

[0091] According to the weight assignment result obtained previously, aggregate the fused feature vectors. Here, a graph neural network is used to mine the topological association structure between log events. First, define the semantic vectors of operation and maintenance log segments as graph nodes. Each semantic vector represents a log event, and these nodes form the basic elements of the graph. The weights of the graph edges are jointly determined by the temporal interval features and context relevance between log segments. The temporal interval features reflect the time sequence relationship of event occurrences, and the context relevance reflects the semantic correlation between events. By combining these two features, the association strength between events can be described more accurately. At the same time, to avoid the graph being too complex, the edge connection range is restricted within a preset time proximity threshold. That is, only log events that are close in time will establish edge connections. In the graph neural network, through information transmission and aggregation operations between nodes, the feature representations of nodes are continuously updated, and finally event-associated features are generated. These event-associated features can reflect the topological association structure between log events.

[0092] Step S134: Invoke the rule generation layer of the association analysis model to perform rule matching processing on the event-associated features to obtain an initial set of association rules. The rule generation layer identifies association rules that meet the support threshold through a frequent pattern mining algorithm.

[0093] Input the generated event correlation features into the rule generation layer of the correlation analysis model, and perform rule matching processing to obtain an initial set of correlation rules. This step involves the application of frequent pattern mining algorithms, and during the mining process, it is necessary to carefully identify and screen the rules, which specifically includes the following sub-steps: Step S1341: Input the event correlation features into the fully connected network of the rule generation layer, and extract potential rule features through non-linear transformation.

[0094] The fully connected network of the rule generation layer consists of multiple neuron layers, and each neuron is connected to all neurons in the previous layer. When the event correlation features are input into the fully connected network, the network will perform a series of linear transformations and non-linear activation operations on them. The linear transformation is achieved by multiplying the weight matrix with the input feature vector and then adding a bias term, while the non-linear activation function (such as the ReLU function) introduces non-linearity to the network, enabling the network to learn more complex feature patterns. Through such processing, the fully connected network can extract potential rule features from the event correlation features. These potential rule features contain information about the possible correlation patterns and causal relationships between log events, but further processing is required to clarify the specific rules.

[0095] Step S1342: Perform multi-label classification processing on the potential rule features to identify the concurrent correlation categories and sequential correlation categories between the operation and maintenance log segments. When the same set of log segments simultaneously meets the concurrent and sequential correlation conditions, select the correlation category with a higher priority according to the confidence score of the causal relationship in the historical trigger records of the correlation rules.

[0096] After obtaining the potential rule features, it is necessary to perform multi-label classification processing on them to identify the concurrent correlation categories and sequential correlation categories between the operation and maintenance log segments. Multi-label classification means that a sample can belong to multiple categories simultaneously. In this scenario, there may be two situations between the operation and maintenance log segments: concurrent correlation (multiple events occur simultaneously) and sequential correlation (events occur in a set order).

[0097] During the classification process, a classifier (such as a support vector machine, neural network classifier, etc.) will be used to judge the potential rule features. The classifier will map the potential rule features to different correlation categories according to the trained model parameters. When the same set of log segments simultaneously meets the concurrent and sequential correlation conditions, it is necessary to select the correlation category with a higher priority according to the confidence score of the causal relationship in the historical trigger records of the correlation rules. The confidence score reflects the reliability of the causal relationship of the rule in historical events. The higher the score, the more credible the causal relationship of the rule. By comparing the confidence scores of the concurrent correlation and sequential correlation rules, select the correlation category with a higher score as the final correlation category for this set of log segments.

[0098] Step S1343: Determine the parallel constraint parameters in the event trigger condition according to the concurrent association category, where the parallel constraint parameters define the resource competition rules when multiple events are triggered simultaneously.

[0099] For the identified concurrent association category, it is necessary to determine the parallel constraint parameters in the event trigger condition. The parallel constraint parameters are used to define the resource competition rules when multiple events are triggered simultaneously. During the operation of the system, when multiple events occur simultaneously, it may cause competition for system resources (such as CPU, memory, network bandwidth, etc.). The parallel constraint parameters can specify the resource allocation method and limiting conditions in such a situation.

[0100] For example, if multiple events need to access database resources simultaneously, the parallel constraint parameters can specify the upper limit of the number of database connections that each event can occupy, or specify the allocation of database connections according to the priority of the events. By determining the parallel constraint parameters, it is possible to avoid a decline in system performance or errors caused by resource competition when multiple events are triggered simultaneously.

[0101] Step S1344: Determine the timing constraint parameters in the event trigger condition according to the sequential association category, where the timing constraint parameters define the sequential order and time window limit of event triggering.

[0102] For the sequential association category, it is necessary to determine the timing constraint parameters in the event trigger condition. The timing constraint parameters are used to define the sequential order and time window limit of event triggering. In sequential association, there is a clear sequential order relationship between events, and this order relationship usually needs to be satisfied within a set time range.

[0103] For example, event A must occur before event B, and after event A occurs, event B must be triggered within a specific time window. The timing constraint parameters can accurately describe this sequential order and time window requirement. By determining the timing constraint parameters, it can be ensured that events are triggered in the expected order and time requirements in sequence, thereby ensuring the normal operation of the system and the smooth execution of the business process.

[0104] Step S1345: Construct a trigger condition expression for the association rule based on the parallel constraint parameters and the timing constraint parameters.

[0105] After determining the parallel constraint parameters and the timing constraint parameters, it is necessary to construct a trigger condition expression for the association rule based on these parameters. The trigger condition expression is a logical expression that describes the conditions required for an event to be triggered.

[0106] For concurrent associations, the trigger condition expression will include a logical judgment of the parallel constraint parameters. For example, if the parallel constraint parameters specify the usage limit of a certain resource when multiple events are triggered simultaneously, the trigger condition expression will include a judgment of the usage of this resource. Only when the resource usage meets the conditions, the association rule will be triggered.

[0107] For sequential associations, the trigger condition expression will include a logical judgment of the timing constraint parameters. For example, if the timing constraint parameters specify that event B must be triggered within a certain time window after event A occurs, the trigger condition expression will include a judgment of the occurrence times of event A and event B. Only when both the time sequence and the time window requirements are met, the association rule will be triggered.

[0108] Step S1346: Perform logical simplification processing on the trigger condition expression, bind the simplified trigger condition expression with the event response actions, and generate an initial association rule set containing a complete definition of the causal relationship.

[0109] After obtaining the trigger condition expression of the association rule, it is necessary to perform logical simplification processing on it. The purpose of logical simplification is to remove the redundant parts in the expression to make the expression more concise and easy to understand. The rules and methods of Boolean algebra can be used to simplify the trigger condition expression, such as combining like terms and eliminating contradictory terms.

[0110] The simplified trigger condition expression is bound to the corresponding event response actions. Event response actions refer to the operations that the system needs to execute when the association rule is triggered, such as adjusting the system configuration, sending alarm information, etc. By binding the trigger condition expression with the event response actions, an initial association rule set containing a complete definition of the causal relationship is generated. Each rule in this initial association rule set clearly defines the conditions for event triggering and the response actions that the system should take.

[0111] Step S135: Perform pruning processing on the initial association rule set, remove redundant rules and conflicting rules, and generate a refined association rule set.

[0112] After obtaining the initial set of association rules, pruning processing needs to be performed on it to remove redundant rules and conflicting rules. Redundant rules refer to those rules that can be derived from other rules. Retaining these rules will increase the complexity of the rule set and may lead to duplicate processing. Conflicting rules refer to those rules that give different conclusions under the same conditions, which will lead to uncertainty when applying the rules. The process of pruning processing is to analyze and compare each rule in the initial set of association rules. For redundant rules, judgment is made through logical derivation and inclusion relationships between rules. If a rule can be derived from other rules, it is deleted from the set. For conflicting rules, evaluation is carried out according to indicators such as confidence and support of the rules. The more reliable rules are retained, and the unreliable rules are deleted. After pruning processing, a refined set of association rules is obtained, and the rules in this refined set of association rules are more concise and effective.

[0113] Step S136: Compare and verify the refined set of association rules with the historical event knowledge base. If any association rule conflicts with more than a preset proportion of the historical rules in the knowledge base in terms of trigger conditions and causal relationship definitions, then correct the abnormal trigger conditions and incorrect causal relationship definitions of this association rule according to the rule with the highest confidence in the knowledge base to generate the final set of association rules.

[0114] Compare and verify the refined set of association rules with the historical event knowledge base. The historical event knowledge base stores past events and their associated rules, and these rules have been verified in practice and have high reliability. During the comparison process, for each rule in the refined set of association rules, check whether its trigger conditions and causal relationship definitions conflict with the rules in the historical event knowledge base. If an association rule conflicts with more than a preset proportion of the historical rules in the knowledge base in terms of trigger conditions and causal relationship definitions, it indicates that this rule may be abnormal. At this time, correct the abnormal trigger conditions and incorrect causal relationship definitions of this association rule according to the rule with the highest confidence in the knowledge base. Confidence is an indicator to measure the reliability of a rule, and the higher the confidence, the more reliable the rule. After comparison, verification, and correction, the final set of association rules is generated. The association rules in this final set of association rules consider both the association patterns in the current log data and combine historical experience, and have higher accuracy and reliability.

[0115] Step S140: Generate an operation and maintenance event optimization strategy based on the set of association rules. The operation and maintenance event optimization strategy includes the event response priority and system resource adjustment plan for each association rule in the set of association rules.

[0116] After obtaining the final set of association rules, it is necessary to generate an operation and maintenance event optimization strategy based on these rules. This strategy aims to improve the operating efficiency and stability of the system and reduce the probability of failures. This step includes the following sub-steps: Step S141: Perform an impact degree evaluation process on each association rule in the association rule set to obtain a comprehensive scoring index.

[0117] Perform an impact degree evaluation on each association rule in the association rule set to determine its importance to the system. For example, this step further includes the following sub-steps: Step S1411: Extract the historical trigger record set of the association rule, and the historical trigger record set includes the event chain data and the system response log each time the rule is triggered.

[0118] Extract the historical trigger record set of each association rule from the system's log database. This historical trigger record set contains detailed information each time the rule is triggered. For example, the event chain data records the order and relationship of related events before and after the rule is triggered; and the system response log records the processing situation of the system when the rule is triggered, including resource usage, error information, etc. By analyzing these historical trigger records, information such as the trigger frequency and trigger consequences of the rule can be understood.

[0119] Step S1412: Count the number of triggers in the historical trigger record set, and the number of triggers reflects the activity level of the rule within the historical period.

[0120] Statistically analyze the extracted historical trigger record set and calculate the number of triggers for each association rule. The number of triggers reflects the activity level of the rule within the historical period. If a rule has a large number of triggers, it means that it is often triggered during the system operation and may have a greater impact on the stability and performance of the system; conversely, if the number of triggers is small, it means that the impact of this rule is relatively small.

[0121] Step S1413: Analyze the system processing results after each association rule is triggered, and calibrate the trigger consequence level according to the service recovery time and resource consumption in the system processing results.

[0122] Analyze the system processing results after each trigger of the association rule, with a focus on the service recovery time and resource consumption. The service recovery time refers to the time required for the system to recover from a failure or abnormal state caused by the rule trigger to the normal operating state, which reflects the impact of the rule trigger on the availability of system services. Resource consumption refers to the resources consumed by the system when processing the rule trigger event, such as CPU usage, memory usage, network bandwidth, etc., which reflects the occupancy of system resources by the rule trigger. Based on the service recovery time and resource consumption, the trigger consequence level can be calibrated. For example, if the service recovery time is long and the resource consumption is large, it indicates that the consequence of the rule trigger is relatively serious and the trigger consequence level is high; conversely, if the service recovery time is short and the resource consumption is small, it indicates that the consequence of the rule trigger is relatively light and the trigger consequence level is low.

[0123] Step S1414: Calculate the initial impact degree score based on the trigger times and the trigger consequence level, where the calculation process introduces a weight coefficient to balance the contribution ratios of frequency and severity.

[0124] Based on the statistically obtained trigger times and the calibrated trigger consequence level, calculate the initial impact degree score for each association rule. In the calculation process, a weight coefficient is introduced to balance the contribution ratios of the trigger times (frequency) and the trigger consequence level (severity). Assume that the trigger times are represented by the letter N, the trigger consequence level is represented by the letter S, the weight coefficients are w1 and w2 respectively, and w1 + w2 = 1. Then the initial impact degree score I can be calculated by the formula I = w1 × N + w2 × S. By adjusting the weight coefficient, the influence of frequency or severity on the impact degree score can be emphasized according to the actual situation.

[0125] Step S1415: Perform a timeliness attenuation process on the initial impact degree score to obtain the impact degree score after timeliness attenuation.

[0126] Considering the timeliness of historical data, perform a timeliness attenuation process on the initial impact degree score. As time goes by, the reference value of historical data for the current system state will gradually decrease. The timeliness attenuation process can be achieved through an attenuation function, such as an exponential attenuation function. Assume that the initial impact degree score is I, the time is t, and the attenuation coefficient is λ. Then the impact degree score I' after timeliness attenuation can be expressed as I' = I × exp(-λt). In this way, the contribution of the newer rule trigger records to the impact degree score is greater, and it can better reflect the actual situation of the current system.

[0127] Step S1416: Extract the matching degree parameter of the association rule in the current system environment, and the matching degree parameter is obtained by comparing the trigger condition of the association rule with the real-time system state.

[0128] Extract the matching parameters of each association rule in the current system environment. The matching parameters reflect the degree of consistency between the triggering conditions of the association rules and the real-time system status. The matching parameters are calculated by comparing the triggering conditions of the association rules with the real-time system status data, such as system resource usage, service operation status, etc. For example, if the triggering condition of the association rule is that the CPU usage exceeds a certain threshold, and the CPU usage of the current system is close to or exceeds the threshold, the matching parameter is high; conversely, if the CPU usage of the current system is far below the threshold, the matching parameter is low.

[0129] Step S1417: weighted fusion of the influence score after time-effect attenuation and the matching parameter to generate a final comprehensive score index.

[0130] The influence score after time decay is weighted and fused with the matching parameter to obtain the final comprehensive scoring index. Assume that the influence score after time decay is I', the matching parameter is M, the weight coefficients are w3 and w4 respectively, and w3+w4=1. Then the final comprehensive scoring index C can be calculated by the formula C=w3×I'+w4×M. This final comprehensive scoring index comprehensively considers the historical impact, timeliness and matching degree of the rule in the current system environment, and can more accurately reflect the impact of each association rule on the system.

[0131] Step S142: Prioritize the association rule set according to the comprehensive scoring index to generate a rule queue arranged in descending order of processing urgency.

[0132] According to the final comprehensive scoring index of each association rule, the association rule set is prioritized. The higher the scoring index, the greater the impact of the association rule on the system and the higher the processing urgency. The association rules are arranged from high to low according to the processing urgency to generate a rule queue. In this rule queue, the rules at the front need to be processed first to ensure that the system can respond to possible problems in a timely manner and improve the stability and reliability of the system.

[0133] Step S143: determining the event response priority of the target optimization operation based on the rule queue, wherein the event response priority is used to enable the event response action corresponding to the association rule with the highest ranking in the rule queue to obtain system resources first.

[0134] Based on the generated rule queue, determine the event response priority of the target optimization operation. The event response actions corresponding to the associated rules with a higher ranking should obtain system resources first to ensure that these important rules can be processed in a timely manner. For example, if the event response action corresponding to the first rule in the rule queue is to increase the number of instances of a certain service to improve system performance, then when performing the optimization operation, the required computing resources, memory resources, etc. should be allocated to this event response action first. By reasonably arranging the event response priority, the utilization efficiency of system resources can be improved, ensuring that the system can quickly and effectively respond to various operation and maintenance events.

[0135] Step S144: Configure a system resource adjustment plan according to the event response priority, where the system resource adjustment plan includes at least one of dynamically allocating computing resources, adjusting the number of service instances, and modifying the load balancing policy.

[0136] Configure a system resource adjustment plan according to the determined event response priority of the target optimization operation. The system resource adjustment plan can include various methods, such as dynamically allocating computing resources, flexibly adjusting the allocation of computing resources such as CPU and memory according to the requirements of different rules; adjusting the number of service instances, increasing the number of instances of a service when the load of a certain service is too high to improve processing capacity, and reducing the number of service instances to save resources when the load is low; modifying the load balancing policy, by adjusting the load balancing algorithm, more reasonably allocate requests to different service instances to improve the overall performance of the system. According to the event response priority and the specific requirements of the rules, select the appropriate resource adjustment method and formulate a detailed system resource adjustment plan.

[0137] Step S145: Package the event response priority and the system resource adjustment plan into an executable operation and maintenance event optimization strategy, and the packaging process includes adding a policy version identifier, an effective time range, and a rollback mechanism description.

[0138] Package the determined event response priority and the configured system resource adjustment plan into an executable operation and maintenance event optimization strategy. During the packaging process, it is necessary to add a policy version identifier to distinguish different versions of the policy for convenient management and tracking. At the same time, clarify the effective time range of the policy and specify the time period during which the policy is valid. In addition, it is also necessary to describe the rollback mechanism. When problems occur or the expected effects are not achieved during the execution of the policy, the system can be restored to the state before the policy execution according to the rollback mechanism. Through packaging, the event response priority and the system resource adjustment plan are integrated into a complete and manageable operation and maintenance event optimization strategy.

[0139] Step S146: Conduct a gray-box execution test on the optimized operation and maintenance event strategy. Split the optimized operation and maintenance event strategy into multiple sub-strategies and execute them in stages in the isolated sandbox environment of the target system according to the priority order. Detect the execution risks of each sub-strategy under limited resource constraints, and dynamically adjust the parameters of subsequent sub-strategies based on the feedback results of the sandbox environment to obtain the test results.

[0140] Conduct a gray-box execution test on the encapsulated optimized operation and maintenance event strategy. First, split the optimized operation and maintenance event strategy into multiple sub-strategies, where each sub-strategy corresponds to a part of the rules in the rule queue and the corresponding resource adjustment operations. Then, execute these sub-strategies in stages in the isolated sandbox environment of the target system according to the priority order. The isolated sandbox environment is a test environment isolated from the actual production environment, which can simulate the operation of the actual system without affecting the production environment. When executing each sub-strategy, detect its execution risks under limited resource constraints, such as insufficient resources and service interruptions. According to the feedback results of the sandbox environment, dynamically adjust the parameters of subsequent sub-strategies, such as adjusting the resource allocation ratio and modifying the number of service instances. Through the gray-box execution test, potential problems that may occur during the strategy execution can be discovered in advance, ensuring the security and effectiveness of the strategy in the actual production environment.

[0141] Step S147: Modify the parameter configuration in the optimized operation and maintenance event strategy according to the test results to generate the final version of the optimized operation and maintenance event strategy.

[0142] Modify the parameter configuration in the optimized operation and maintenance event strategy according to the results of the gray-box execution test. If it is found during the test that a certain sub-strategy has insufficient resources, the resource allocation amount corresponding to this sub-strategy can be increased; if it is found that the adjustment of a certain service instance leads to a decrease in system performance, the number of service instances can be adjusted or the load balancing strategy can be adjusted. By modifying the parameter configuration, optimize the optimized operation and maintenance event strategy to improve its adaptability and effectiveness in the actual production environment. The finally generated final version of the optimized operation and maintenance event strategy has been fully tested and optimized, and can better ensure the stable operation of the system.

[0143] Step S150: Feed back the optimized operation and maintenance event strategy to the operation and maintenance management interface of the target system to trigger system optimization operations.

[0144] Feed the optimized operation and maintenance event strategy in the final version back to the operation and maintenance management interface of the target system. The operation and maintenance management interface can receive the operation and maintenance event optimization strategy and convert it into specific system operation instructions. When the operation and maintenance event optimization strategy is received, the operation and maintenance management interface will parse and verify the strategy. The parsing process will convert each part of the strategy, such as the event response priority and the system resource adjustment plan, into instructions that the system can understand and execute. The verification process will check the legality and rationality of the strategy to ensure that the strategy will not damage the system.

[0145] After passing the verification, the operation and maintenance management interface will trigger corresponding system optimization operations in sequence according to the event response priority in the strategy. For the operation of dynamically allocating computing resources in the system resource adjustment plan, the operation and maintenance management interface will interact with the resource management module of the system. The resource management module is responsible for monitoring the resource usage of each component in the system, such as CPU, memory, storage, etc. When it is necessary to allocate computing resources for the event response action corresponding to a certain rule, the operation and maintenance management interface will send an instruction to the resource management module, and the resource management module will allocate the corresponding resources from the idle resource pool to the specified component or service according to the instruction. For example, if the strategy requires increasing the CPU resources for a certain key service, the resource management module will adjust the CPU scheduling policy of the server where the service is located and allocate more CPU time slices to the service.

[0146] For the operation of adjusting the number of service instances, the operation and maintenance management interface will communicate with the container orchestration tool or service management platform of the system. Container orchestration tools such as Kubernetes can automatically manage and schedule containerized service instances. When the strategy requires increasing or decreasing the number of service instances, the operation and maintenance management interface will send corresponding instructions to the container orchestration tool. If it is to increase the service instances, the container orchestration tool will start new service instances on appropriate nodes according to the preset images and configuration information; if it is to decrease the service instances, the container orchestration tool will select appropriate instances to shut down and release the relevant resources.

[0147] For the operation of modifying the load balancing strategy, the operation and maintenance management interface will interact with the load balancer of the system. The load balancer is responsible for distributing client requests to multiple backend service instances to achieve balanced processing of requests and high availability. The operation and maintenance management interface will send new configuration information to the load balancer to modify the load balancing algorithm or adjust the weights of the backend service instances. For example, if the strategy requires distributing more requests to service instances with better performance, the operation and maintenance management interface will adjust the weight distribution of the load balancer so that these service instances with better performance can receive more requests.

[0148] During the execution of system optimization operations, the operation and maintenance management interface will monitor the execution status of the operations and the feedback information of the system in real time. If the operations are executed successfully, the performance metrics and running status of the system will gradually improve, such as shorter service response time and higher resource utilization. The operation and maintenance management interface will record this feedback information for subsequent evaluation and analysis.

[0149] However, if abnormal situations occur during the operation execution, such as resource allocation failure or service instance startup failure, the operation and maintenance management interface will handle them according to the preset error handling mechanism. The error handling mechanism may include retry operations, rollback operations, or sending alarm notifications to the operation and maintenance personnel. For example, if a service instance fails to start, the operation and maintenance management interface will attempt to restart the instance a set number of times; if the attempts still fail after multiple tries, a rollback operation will be triggered to restore the system to the state before the operation execution and an alarm will be sent to notify the operation and maintenance personnel for further investigation and handling.

[0150] After the execution of system optimization operations is completed, the operation and maintenance management interface will summarize and evaluate the entire process. The evaluation content includes the execution effect of the strategy, the improvement of system performance, and the utilization efficiency of resources. By analyzing these evaluation results, the effectiveness and deficiencies of the operation and maintenance event optimization strategy can be understood.

[0151] In addition, to ensure the security of the system and the privacy of data, strict privacy protection and anti-disclosure technical means are adopted for data collection and processing throughout the process. For example, when obtaining the original operation and maintenance log set, the log data is encrypted, and only authorized modules and personnel can access and process this data. During the training and use of the correlation analysis model, technologies such as differential privacy are also adopted to process the input and output data of the model to prevent the leakage of sensitive information. At the same time, the system also sets up a strict access control mechanism to assign different permissions to users and modules with different roles, ensuring that only personnel and modules with corresponding permissions can collect, process, and use data.

[0152] During the long-term operation of the system, as the system continuously changes and develops, such as the increase in business requirements and the adjustment of system architecture, the content and characteristics of the original operation and maintenance log set will also change. Therefore, it is necessary to regularly update and optimize the correlation analysis model. The update process includes collecting new operation and maintenance log data, re-performing feature extraction, and model training. When re-training the model, it is necessary to adjust the parameters and structure of the model to adapt to the new data characteristics and system changes. At the same time, the operation and maintenance event optimization strategy also needs to be adjusted accordingly to ensure that the strategy can continuously and effectively improve the performance and stability of the system.

[0153] In the training of the association analysis model, its training process is a complex and crucial link. First, a large amount of labeled training data needs to be prepared. This data is extracted from historical operation and maintenance logs and has been manually annotated. The annotation content includes the association relationship, causal relationship, etc. between log fragments. The training data is divided into a training set and a validation set. The training set is used for model training, and the validation set is used to evaluate the performance and generalization ability of the association analysis model.

[0154] The structure of the association analysis model contains multiple key modules, such as a feature fusion layer, an attention mechanism layer, a rule generation layer, etc. During the training process, optimization algorithms such as stochastic gradient descent are used to adjust the parameters of the association analysis model. First, the feature data in the training set is input into the association analysis model. The context semantic features and temporal behavior features are fused through the feature fusion layer. In the feature fusion layer, the gated recurrent unit dynamically adjusts the contribution ratio of different feature types according to the input features, enabling the association analysis model to better capture the association information between different features.

[0155] Then, the fused feature vector enters the attention mechanism layer. The attention mechanism layer calculates the semantic interaction weights between different log fragments through the self-attention mechanism, capturing the dependency relationships across log fragments and the causal intensity of event triggers. During this process, query vectors and key vectors are generated through linear projection, the similarity scores between them are calculated, and normalization processing is performed to obtain the attention weights. The key vectors are weighted and summed according to the attention weights to generate a context-enhanced feature representation. Through a multi-layer attention stacking structure, the dependency patterns at different abstraction levels are continuously captured, enabling the association analysis model to more comprehensively understand the association information in the log data.

[0156] Next, the feature vector processed by the attention mechanism layer enters the rule generation layer. The rule generation layer uses a frequent pattern mining algorithm to identify association rules that meet the support threshold. During the training process, the association analysis model continuously adjusts its parameters to make the generated association rules more accurate and reliable. At the same time, pruning is performed on the generated rules to remove redundant rules and conflicting rules, improving the quality and effectiveness of the rules.

[0157] During the training process, the validation set is used to evaluate the association analysis performance of the association analysis model. The evaluation metrics include the accuracy, recall rate, F1 value, etc. of the rules. By continuously adjusting the parameters and structure of the association analysis model, the performance of the association analysis model on the validation set is optimized. When the performance of the association analysis model on the validation set is stable and meets the preset index requirements, it is considered that the model training is completed.

[0158] To further improve the performance and adaptability of the association analysis model, techniques such as transfer learning can also be adopted. Transfer learning can utilize the model parameters trained in other related fields or systems as the initial parameters of the current model, and then fine-tune them on the current operation and maintenance log data. This can reduce the training time and data requirements of the model, while improving the generalization ability and performance of the association analysis model.

[0159] In practical applications, the association analysis model and the operation and maintenance event optimization strategy are closely integrated with large financial trading systems. The input data of the association analysis model is the original operation and maintenance log set collected from the system, and after feature extraction, context semantic features and temporal behavior features are obtained. The output data is the set of association rules between operation and maintenance log segments and the corresponding operation and maintenance event optimization strategy. In this way, the association analysis model and algorithm are combined with the specific financial trading system scenario, realizing the effective analysis and optimization of system operation and maintenance events, improving the stability and reliability of the system, and ensuring the smooth progress of financial transactions.

[0160] Figure 2 FIG. shows a schematic diagram of exemplary hardware and software components of an operation and maintenance log event association analysis system 100 based on artificial intelligence that can implement the idea of the present application provided by some embodiments of the present application. For example, the processor 120 can be used on the operation and maintenance log event association analysis system 100 based on artificial intelligence and is used to execute the functions in the present application.

[0161] The operation and maintenance log event association analysis system 100 based on artificial intelligence can be a general-purpose server or a special-purpose server, both of which can be used to implement the operation and maintenance log event association analysis method based on artificial intelligence of the present application. Although only one server is shown in the present application, for convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0162] For example, the operation and maintenance log event association analysis system 100 based on artificial intelligence can include a network port 110 connected to the network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROM, or RAM, or any combination thereof. Exemplarily, the operation and maintenance log event association analysis system 100 based on artificial intelligence can also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. According to these program instructions, the method of the present application can be implemented. The operation and maintenance log event association analysis system 100 based on artificial intelligence also includes an I / O interface 150 between the computer and other input and output devices.

[0163] For the sake of convenience in description, only one processor is described in the artificial intelligence-based operation and maintenance log event correlation analysis system 100. However, it should be noted that the artificial intelligence-based operation and maintenance log event correlation analysis system 100 in the present application may also include multiple processors. Therefore, the steps executed by one processor described in the present application may also be jointly executed or separately executed by multiple processors. For example, if the processor of the artificial intelligence-based operation and maintenance log event correlation analysis system 100 executes step A and step B, it should be understood that step A and step B may also be jointly executed by two different processors or separately executed in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.

[0164] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above-mentioned artificial intelligence-based operation and maintenance log event correlation analysis method is implemented.

[0165] It should be noted that, in order to simplify the expression of the disclosure of the present invention and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are merged into one embodiment, drawing, or description thereof.

Claims

1. An operation and maintenance log event correlation analysis method based on artificial intelligence, characterized in that, The method includes: Obtaining an original operation and maintenance log set of the target system, where the original operation and maintenance log set includes multiple operation and maintenance log segments; Performing feature extraction processing on the original operation and maintenance log set to obtain the context semantic features and temporal behavior features of each operation and maintenance log segment; Invoking a pre-trained association analysis model to perform association rule matching processing on the context semantic features and temporal behavior features, generating an association rule set between the operation and maintenance log segments, where the association rule set includes the definition of the causal relationship and the description of the event propagation path in the system failure link for the operation and maintenance log segments; Generating an operation and maintenance event optimization strategy based on the association rule set, where the operation and maintenance event optimization strategy includes the event response priority and the system resource adjustment plan for each association rule in the association rule set; Feeding back the operation and maintenance event optimization strategy to the operation and maintenance management interface of the target system to trigger system optimization operations.

2. The method for analyzing the correlation of operation and maintenance log events based on artificial intelligence according to claim 1, wherein The performing feature extraction processing on the original operation and maintenance log set to obtain the context semantic features and temporal behavior features of each operation and maintenance log segment includes: Performing syntactic structure parsing processing on the log statements in the operation and maintenance log segments, splitting the log statements into multiple log word units, where the log word units include system component identifiers, operation instruction keywords, and status descriptors; Invoking a pre-trained feature encoder to perform context encoding processing on the multiple log word units to generate a semantic vector of the operation and maintenance log segment, where the context encoding processing constructs a global semantic representation by capturing the dependency relationships between adjacent log word units; Determining the context correlation degree between the operation and maintenance log segments based on the semantic vector, and constructing context semantic features according to the context correlation degree, where the context correlation degree is used to quantify the logical relevance of different operation and maintenance log segments in the failure propagation path; Performing temporal distribution analysis processing on the generation timestamps of the operation and maintenance log segments, extracting the temporal interval features and event trigger frequency features of the operation and maintenance log segments, and fusing the temporal interval features and event trigger frequency features to generate the temporal behavior features of the operation and maintenance log segments, where the temporal interval features reflect the continuity and suddenness patterns of event triggers, the event trigger frequency features reflect the event density within a preset time window, and the fusion process eliminates redundant information through a feature dimensionality reduction algorithm and retains the key representations of the temporal patterns.

3. The method for analyzing the correlation of operation and maintenance log events based on artificial intelligence according to claim 2, wherein The invoking a pre-trained feature encoder to perform context encoding processing on the multiple log word units to generate a semantic vector of the operation and maintenance log segment includes: Inputting the log word units into the word embedding layer of the pre-trained feature encoder, and converting the discrete log word units into initial word vectors in a continuous space through distributed vector mapping; Inputting the initial word vectors into the bidirectional attention layer of the feature encoder for context association analysis processing, where the bidirectional attention layer calculates the semantic interaction weights between different log word units through a self-attention mechanism; Weightedly fuse the initial word vectors according to the semantic interaction weights to generate context-aware vectors for each log word unit, where the context-aware vectors contain the context-dependent information of the log word unit in the entire log statement; Perform pooling on the context-aware vectors, and generate semantic vectors of a unified dimension through a combined operation of max pooling and average pooling, where the window size and sliding step of the pooling operation are consistent for all operation and maintenance log segments, ensuring that the semantic vectors of different log segments have the same feature dimension; Input the semantic vectors into the output layer of the feature encoder for non-linear transformation to obtain a standardized semantic vector representation, and the standardization process eliminates the feature scale differences between different log segments.

4. The method for analyzing the correlation of operation and maintenance log events based on artificial intelligence according to claim 2, wherein Determine the context correlation between the operation and maintenance log segments based on the semantic vectors, and construct context semantic features according to the context correlation, including: Calculate the similarity scores between the semantic vectors, and the similarity scores are obtained through vector inner product operation combined with normalization processing; Perform time-series sliding window smoothing processing on the similarity scores to generate smoothed context correlation; Perform splicing processing on the context correlation and the semantic vectors, and map features of different dimensions to a unified semantic space through a fully connected network; In the unified semantic space, use a semi-supervised clustering algorithm to perform pattern division on the spliced features, where the cluster centers are initialized as the feature vectors corresponding to predefined system component role identifiers, identify the semantic grouping of the operation and maintenance log segments in the system failure scenario and associate them with specific role identifiers to obtain a semantic grouping result; Construct the context semantic features based on the semantic grouping result, and the context semantic features include the role identifiers and semantic weight allocation results of the operation and maintenance log segments in the system operation context.

5. The method for analyzing the correlation of operation and maintenance log events based on artificial intelligence according to claim 2, wherein Perform time-series distribution analysis processing on the generation timestamps of the operation and maintenance log segments, extract the time-series interval features and event trigger frequency features of the operation and maintenance log segments, and fuse the time-series interval features and event trigger frequency features to generate the time-series behavior features of the operation and maintenance log segments, including: Extract the set of generation timestamps of the operation and maintenance log segments, and sort the set of generation timestamps in chronological order; Calculate the interval duration between adjacent generation timestamps in the sorted set of generation timestamps, and the interval duration reflects the continuity and suddenness of event triggers; Perform distribution statistical processing on the interval duration, and extract the time-series interval features characterizing the event trigger pattern, where the time-series interval features include the central tendency index and the dispersion degree index of the interval duration; Count the number of event triggers of the operation and maintenance log segments within a preset time window to generate an event trigger frequency feature reflecting the event density; Perform normalization processing on the time-series interval features and the event trigger frequency features, and perform dimensionality reduction processing on the normalized time-series interval features and the event trigger frequency features through the principal component analysis method to generate the time-series behavior features of the operation and maintenance log segments, and the dimensionality reduction processing retains the feature components with the largest variance in the original time-series data.

6. The method for associative analysis of operation and maintenance log events based on artificial intelligence according to claim 1, wherein Invoking the pre-trained association analysis model to perform association rule matching processing on the context semantic features and temporal behavior features to generate an association rule set between the operation and maintenance log segments, including: Inputting the context semantic features and temporal behavior features into the feature fusion layer of the pre-trained association analysis model to obtain a fused feature vector, and the feature fusion layer dynamically adjusts the contribution ratio of different feature types using a gated recurrent unit; Performing weight assignment processing on the fused feature vector through the attention mechanism layer of the association analysis model to obtain a weight assignment result, and the weight assignment processing captures the dependency relationship across log segments and the causal strength of event triggering; Performing aggregation processing on the fused feature vector according to the weight assignment result to generate event association features, and the aggregation processing uses a graph neural network to mine the topological association structure between log events, where the graph nodes are defined as the semantic vectors of operation and maintenance log segments, and the graph edge weights are jointly determined by the temporal interval features and context relevance between log segments, and the edge connection range is restricted within a preset time proximity threshold; Invoking the rule generation layer of the association analysis model to perform rule matching processing on the event association features to obtain an initial association rule set, and the rule generation layer identifies association rules that meet the support threshold through a frequent pattern mining algorithm; Performing pruning processing on the initial association rule set to remove redundant rules and conflicting rules to generate a refined association rule set; Comparing and validating the refined association rule set with the historical event knowledge base. If any association rule conflicts with more than a preset proportion of historical rules in the knowledge base in terms of trigger conditions and causal relationship definitions, then correct the abnormal trigger conditions and incorrect causal relationship definitions of this association rule according to the rule with the highest confidence in the knowledge base to generate a final association rule set.

7. The method for analyzing the association of operation and maintenance log events based on artificial intelligence according to claim 6, wherein The performing weight assignment processing on the fused feature vector through the attention mechanism layer of the association analysis model to obtain a weight assignment result includes: Performing linear projection processing on the fused feature vector to generate a query vector and a key vector, and the linear projection processing realizes feature space transformation through a trainable parameter matrix; Calculating the similarity score between the query vector and the key vector, and the similarity score is obtained through a scaled dot product operation; Performing normalization processing on the similarity score to convert the original score into an attention weight in the form of a probability distribution; Performing weighted summation on the key vector according to the attention weight to generate a context-enhanced feature representation; Performing a residual connection between the context-enhanced feature representation and the original fused feature vector, and capturing the dependency patterns at different abstraction levels through a multi-layer attention stacking structure to generate a weight assignment result reflecting the potential relationship of the event chain.

8. The method for correlative analysis of operation and maintenance log events based on artificial intelligence according to claim 6, characterized in that The invoking the rule generation layer of the association analysis model to perform rule matching processing on the event association features to obtain an initial association rule set includes: Inputting the event association features into the fully connected network of the rule generation layer to extract potential rule features through non-linear transformation; Perform multi-label classification processing on the potential rule features to identify concurrent association categories and sequential association categories between the operation and maintenance log fragments. When the same group of log fragments simultaneously meets concurrent and sequential association conditions, select an association category with a higher priority according to the confidence score of the causal relationship in the historical trigger record of the association rule; Determine a parallel constraint parameter in an event triggering condition according to the concurrent association category, wherein the parallel constraint parameter defines a resource competition rule when multiple events are triggered simultaneously; Determine the timing constraint parameters in the event triggering condition according to the sequence association category, wherein the timing constraint parameters define the sequence of event triggering and the time window limit; Constructing a trigger condition expression of the association rule based on the parallel constraint parameter and the timing constraint parameter; The trigger condition expression is logically simplified, and the simplified trigger condition expression is bound to the event response action to generate an initial association rule set containing a complete causal relationship definition.

9. The method for analyzing the correlation of operation and maintenance log events based on artificial intelligence according to claim 1, characterized in that The generating of the operation and maintenance event optimization strategy based on the association rule set includes: Performing an impact evaluation process on each association rule in the association rule set to obtain a comprehensive scoring index; Prioritize the association rule set according to the comprehensive scoring index to generate a rule queue arranged in descending order of processing urgency; Determine an event response priority of a target optimization operation based on the rule queue, wherein the event response priority is used to enable an event response action corresponding to an association rule ranked first in the rule queue to obtain system resources first; Configure a system resource adjustment plan according to the event response priority, wherein the system resource adjustment plan includes at least one of dynamically allocating computing resources, adjusting the number of service instances, and modifying the load balancing strategy; Encapsulating the event response priority and system resource adjustment plan into an executable operation and maintenance event optimization strategy, wherein the encapsulation process includes adding a strategy version identifier, an effective time range, and a rollback mechanism description; Perform grayscale execution test on the operation and maintenance event optimization strategy, split the operation and maintenance event optimization strategy into multiple sub-strategies and execute them in stages in the isolated sandbox environment of the target system in order of priority, detect the execution risk of each sub-strategy under limited resource constraints, and dynamically adjust the parameters of subsequent sub-strategies according to the feedback results of the sandbox environment to obtain test results; The parameter configuration in the operation and maintenance event optimization strategy is modified according to the test results to generate a final version of the operation and maintenance event optimization strategy.

10. An operation and maintenance log event correlation analysis system based on artificial intelligence, characterized in that, It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the artificial intelligence-based operation and maintenance log event correlation analysis method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Operation and maintenance log security detection method and device based on big data

    CN108667678A

  • Artificial intelligence driven adaptive firewall rule optimization method and system

    CN119484148A

  • Business data analysis method and device based on big data, equipment and storage medium

    CN119669309A

Cited By

  • Charging pile quality inspection data management method and system based on intelligent processing

    CN120494638A

  • Building facility digital operation and maintenance method and system combined with BIM (Building Information Modeling)

    CN120563116A

  • Multi-dimensional knowledge extraction construction method and system applied to operation and maintenance technology service

    CN120803879A

  • Knowledge graph driven intelligent log assertion reasoning method and system

    CN120822620A

  • Intelligent log assertion reasoning method and system driven by knowledge graph

    CN120822620B