Automatic operation and maintenance event processing method and system based on AI intelligent agent

By constructing an event correlation graph and combining it with time series information, and using graph neural networks and adaptive differential evolution algorithm optimization strategies, the problems of insufficient adaptability and prediction accuracy in traditional operation and maintenance event processing are solved, and efficient and intelligent operation and maintenance event processing is achieved.

CN120654736AInactive Publication Date: 2025-09-16HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511003429.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-09-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional operation and maintenance event processing methods are difficult to cope with complex scenarios with high frequency, strong correlation, and multiple variables. Existing algorithms lack adaptability and the ability to mine dependencies between events, resulting in inaccurate predictions and poor generalization of strategy optimization, and lack of closed-loop self-learning capabilities.

Method used

Combining graph neural networks with adaptive differential evolution algorithms, an event association graph is constructed. The graph neural network is used to analyze the dependencies between devices and make predictions based on time series information. The adaptive differential evolution algorithm is used to optimize the processing strategy to form a closed-loop self-learning system.

Benefits of technology

It achieves accurate prediction and strategy optimization of operation and maintenance events, improves processing efficiency and decision-making accuracy, has high generalization and adaptability, reduces manual intervention costs, and improves operation and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654736A_ABST
    Figure CN120654736A_ABST
Patent Text Reader

Abstract

The invention discloses an operation and maintenance event automatic processing method and system based on an AI intelligent agent, and the system comprises a data collection module which is used for collecting various types of data in an operation and maintenance system and carrying out the preprocessing of the data; the event association analysis module is used for generating operation and maintenance event types and association mode information; the event time prediction module is used for predicting the evolution trend of the operation and maintenance event; the event processing strategy optimization module is used for adjusting the priority of operation and maintenance events and a resource scheduling scheme; the event automatic processing module is used for executing automatic processing of operation and maintenance events; the processing effect evaluation module is used for evaluating the processing effect of the operation and maintenance event; and the strategy dynamic adjustment module is used for realizing dynamic optimization of the strategy. The invention relates to the technical field of intelligent operation and maintenance, aims to solve the problems of slow response and stiff strategy in traditional operation and maintenance event processing, can provide an efficient and scientific optimization scheme in operation and maintenance event automatic processing, and brings remarkable technical value for practical application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent operation and maintenance, and in particular to a method and system for automatically processing operation and maintenance events based on an AI agent. Background Art

[0002] With the continuous expansion and complexity of information technology infrastructure, the number of events generated in operation and maintenance systems has increased rapidly. Traditional operation and maintenance event processing methods rely mostly on manual judgment and static rule configuration, which makes it difficult to cope with complex event scenarios with high frequency, strong correlation, and multiple variables. At present, some automated operation and maintenance solutions based on rule engines or simple machine learning models have achieved event classification and processing to a certain extent, but there are still problems such as the lack of adaptability of processing strategies, insufficient ability to mine dependencies between events, and poor fusion of multi-source heterogeneous data.

[0003] In existing technologies, algorithms that combine graph structures with time series find it difficult to accurately capture the complex topological relationships between devices and the temporal evolution of events, resulting in inaccurate event predictions and poor generalization of policy optimization. At the same time, existing optimization algorithms are mostly based on static parameter tuning, making it difficult to dynamically adjust strategies based on real-time feedback and lacking closed-loop self-learning capabilities. Therefore, there is an urgent need for a new intelligent agent method that combines graph neural networks with advanced optimization algorithms to achieve efficient and automated processing of the entire process of operation and maintenance events, from collection, analysis, prediction to policy optimization and execution. Summary of the Invention

[0004] One purpose of the present invention is to propose an automated processing method and system for operation and maintenance events based on AI intelligent agents. The present invention combines graph neural networks with adaptive differential evolution algorithms, uses the topological relationship between devices and event time series information to construct an event association graph, and classifies events, performs correlation analysis and trend prediction. At the same time, the processing strategy is adaptively adjusted through a dynamic optimization mechanism to achieve comprehensive optimization of event priority, resource scheduling and fault recovery. This method drives strategy evolution through real-time feedback to form a closed-loop processing system, effectively improving the intelligence level, processing efficiency and decision-making accuracy of operation and maintenance event response.

[0005] According to an embodiment of the present invention, an AI-agent-based automated operation and maintenance event processing method and system includes: The data acquisition module is used to collect and pre-process various operation and maintenance event data in the operation and maintenance system, including equipment status monitoring data, equipment operation logs, and network topology information; The event correlation analysis module is used to perform graph neural network analysis based on the event correlation graph, identify the dependencies between devices, classify and predict operation and maintenance events, and generate operation and maintenance event types and correlation pattern information; The event time prediction module is used to predict the evolution trend of operation and maintenance events by combining time series information and generate time prediction information for operation and maintenance events; The event processing strategy optimization module is used to optimize the operation and maintenance event processing strategy based on the operation and maintenance event type, correlation pattern and time prediction information using the adaptive differential evolution algorithm, and adjust the operation and maintenance event priority, resource scheduling plan and fault recovery strategy; The event automation processing module is used to perform automated processing operations for operation and maintenance events according to the optimized policy set, including fault detection, automatic repair, resource allocation and event response; The processing effect evaluation module is used to evaluate the processing effect of each operation and maintenance event based on the feedback information of the event processing results; The policy dynamic adjustment module is used to monitor the operation and maintenance event processing process in real time, adjust the processing strategy based on feedback, and realize dynamic optimization and adaptive adjustment of the strategy.

[0006] Optionally, modules can be connected using the following methods: S1. Collect device status monitoring data, device operation logs, network topology information, and operation and maintenance event data from the operation and maintenance system, pre-process the collected data, construct graph structure data based on the relationship between devices and events, and generate an event correlation graph; S2. Analyze the event association graph using a graph neural network to identify dependencies between devices, perform operation and maintenance event classification and association analysis, and generate operation and maintenance event type and association pattern information; S3. Introducing time series information into graph neural networks, using graph convolution combined with temporal dependencies to analyze the evolution trend of operation and maintenance events, predict the state changes of future events, and generate time prediction information for operation and maintenance events. S4. Based on the generated event types, correlation pattern information, and time prediction information of operation and maintenance events, an adaptive differential evolution algorithm is applied to optimize the event processing strategy, adjust the operation and maintenance event processing priority, resource scheduling plan, and fault recovery strategy; S5. Execute automated O&M event processing operations based on the optimized event handling strategy, including fault detection, automatic repair, resource allocation, and event response. S6. Based on the feedback information of the event processing results, the event processing strategy is dynamically adjusted using the adaptive differential evolution algorithm, and the event processing strategy is continuously optimized.

[0007] Optionally, S1 includes the following specific steps: S11. Collect device status monitoring data and network topology information from the operation and maintenance system, and collect operation and maintenance event data from different data sources, and perform unified processing based on the timestamps of the operation and maintenance events; S12. De-noising the collected equipment status monitoring data by smoothing the data using a sliding window technique to filter out short-term fluctuations; S13. Align the timestamps of the collected device operation logs; S14. Construct an event association graph based on the relationship between devices and the timing information of operation and maintenance events. The nodes of the event association graph represent devices and operation and maintenance events, and the edges represent the communication relationship between devices and the causal relationship between operation and maintenance events.

[0008] Optionally, S2 includes the following specific steps: S21. Extract features from each node based on the event association graph, and use the graph convolution layer in the graph neural network to perform convolution operations on the event association graph to update the feature representation of each node, namely: ; in, Representation node In the Feature representation after layer graph convolution, For the The weight matrix of the layer graph convolution, For nodes The set of neighbor nodes of For nodes With neighboring nodes The connection weights between is the activation function, Representation node In the Feature representation after layer graph convolution; S22. Perform maximum pooling on the node features after the graph convolution operation to obtain the global features of the graph; S23. Classify and predict events based on the global features of the graph, and generate type information and association pattern information of operation and maintenance events.

[0009] Optionally, S3 includes the following specific steps: S31. Based on the graph neural network model, time series information is introduced to model the temporal dependency of operation and maintenance events. Each operation and maintenance event is associated with timestamp information to generate operation and maintenance event node features with time series information. S32. Perform graph convolution processing on the operation and maintenance event node features with time series information, update the node feature representation, and combine the device status, the order and time interval of the operation and maintenance events to obtain the time evolution characteristics of the operation and maintenance events, namely: ; in, For nodes In the Feature representation after layer graph convolution, For nodes In the Feature representation after layer graph convolution, For the The weight matrix of the layer graph convolution, For nodes The set of neighbor nodes of For nodes With neighboring nodes The connection weights between For nodes Timestamp information, for Activation function; The node Timestamp information Expressed as: ; in, Represents an event or device node The actual occurrence timestamp, Indicates the time interval between the node event and its previous associated event. represents the temporal context encoding vector of node v; S33. Extract temporal dependencies through a multi-layer graph convolutional network and perform time series modeling to obtain future state predictions for operation and maintenance events: ; in, Operation and maintenance events The future state prediction value, is the prediction function, For events Neighbor node features, For events Timestamp information.

[0010] Optionally, S4 includes the following specific steps: S41. Construct an event processing optimization model based on the generated operation and maintenance event type information and associated pattern information, combined with the future state prediction value of the operation and maintenance event, and define an optimization objective function; S42. Apply the adaptive differential evolution algorithm to solve the optimization objective function and initialize a set of candidate operation and maintenance event processing strategies. ,in, Process the decision vector for each operation and maintenance event, expressed as: ; in, The priority of the operation and maintenance event. For resource allocation plan, For fault recovery solutions; S43. Calculate the fitness value of each candidate strategy according to the standard process of the adaptive differential evolution algorithm: ; in, For strategy The fitness value of For strategy The corresponding optimization target value; S44. Based on the adaptive differential evolution algorithm, select the strategy with the best fitness value for crossover and mutation operations, and generate a new candidate strategy set : ; in, For crossover operation, Generate a new set of candidate strategies for mutation operations; S45, through repeated iterative crossover, mutation and selection operations, gradually approach the optimal processing strategy, and finally output the optimal processing strategy set : ; in, is the optimal event processing strategy set, is the corresponding optimization target value.

[0011] Optionally, the calculation formula of the event processing optimization model in S41 is: ; in, Indicates the processing strategy output of event e, represents the ReLU activation function, The node feature vector representing event e, Representing an event The time state prediction value of Representing an event Characteristics of the association relationship with other events, is the weight matrix, is the bias term; The event processing optimization model is optimized by minimizing the following objective function: ; in, is the overall optimization goal, is a collection of events, For events The weight of Operation and maintenance events Type information and future state predictions The loss function between For association mode information Operation and maintenance events The impact measure, is the event processing strategy set optimized by the adaptive differential evolution algorithm and event time predictions The resource adjustment cost between and is the weight coefficient.

[0012] Optionally, S5 includes the following specific steps: S51. Assign a corresponding processing priority to each operation and maintenance event based on the generated optimal event processing strategy set; S52. Based on the operation and maintenance event priority and resource availability information, a resource scheduling algorithm is used to allocate resources to the processing strategy and generate a resource allocation plan for each operation and maintenance event; S53. Generate a fault recovery plan for each operation and maintenance event, and select a recovery plan based on the event type and its associated mode; S54. Execute automated processing of maintenance events based on their priorities, resource allocation plans, and fault recovery plans, including fault detection, automatic repair, resource scheduling, and event response. S55. During the operation and maintenance event processing, the processing effect is monitored in real time, the processed operation and maintenance events are dynamically evaluated, and the processing strategy is adjusted according to the evaluation results.

[0013] Optionally, S6 includes the following specific steps: S61. Evaluate the processing effect of each operation and maintenance event based on the operation and maintenance event processing result and generate a processing effect evaluation value; S62. Calculate the overall effect evaluation value of the optimal event processing strategy set based on the processing effect evaluation values ​​of all operation and maintenance events; S63. Based on the overall effect evaluation value, the event processing strategy is dynamically adjusted using the adaptive differential evolution algorithm; S64. Apply the adjusted strategy set to the operation and maintenance event processing process, continuously collect feedback information on the processing results through the real-time monitoring system and feedback mechanism, and further optimize the event processing strategy based on the new feedback data.

[0014] The beneficial effects of the present invention are: First, the present invention realizes structured modeling, precise prediction and strategy optimization of complex operation and maintenance events by introducing the collaborative mechanism of graph neural network and adaptive differential evolution algorithm, overcoming the problems of difficulty in extracting correlation relationships between events and static and rigid processing strategies in the existing technology. By constructing an event correlation graph and integrating time series information, the system can identify the topological dependencies between devices and the temporal evolution trends of event development, thereby realizing early warning and intervention of potential risks. Compared with traditional solutions that rely on rule engines or single machine learning models, the present invention has higher generalization and adaptability.

[0015] In addition, the introduction of the adaptive differential evolution algorithm enables event processing strategies to be dynamically updated based on real-time feedback, achieving optimal decisions for resource scheduling and fault recovery, and forming a closed loop of strategy generation and execution. When facing multi-source heterogeneous data, the system has strong robustness and scalability, and can automatically adapt to different operation and maintenance scenarios, significantly improving the accuracy, response speed and system stability of event processing, thereby effectively reducing the cost of manual intervention and improving overall operation and maintenance efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is a functional block diagram of the automatic processing system for operation and maintenance events in the present invention; Figure 2 This is a flow chart of the method for automatically processing operation and maintenance events in the present invention. DETAILED DESCRIPTION

[0017] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0018] Example 1 The current operation and maintenance event automation processing platform often uses a combination of graph structure and time series for analysis when dealing with operation and maintenance events. The analytical ability of the combination of the two is weak and cannot accurately capture the complex topological relationship between devices and the time evolution law of events, resulting in inaccurate event prediction and poor generalization ability of strategy optimization. Based on this, this embodiment proposes an operation and maintenance event automation processing method based on AI intelligent body, which can identify the topological dependency between devices and the time evolution trend of event development, thereby realizing early warning and intervention of potential risks. Compared with traditional solutions that rely on rule engines or single machine learning models, the present invention has higher generalization and adaptability. Figure 1 and Figure 2 As shown, the method includes the following steps: S1. Collect device status monitoring data, device operation logs, network topology information, and operation and maintenance event data from the operation and maintenance system, pre-process the collected data, construct graph structure data based on the relationship between devices and events, and generate an event correlation graph. The specific steps include the following: S11. Collect device status monitoring data and network topology information from the operation and maintenance system, and collect operation and maintenance event data from different data sources, and perform unified processing based on the timestamps of the operation and maintenance events; S12. De-noising the collected equipment status monitoring data by smoothing the data using a sliding window technique to filter out short-term fluctuations; S13. Align the timestamps of the collected device operation logs; S14. Construct an event association graph based on the relationship between devices and the timing information of operation and maintenance events. The nodes of the event association graph represent devices and operation and maintenance events, and the edges represent the communication relationship between devices and the causal relationship between operation and maintenance events.

[0019] The present invention uses graph neural networks to analyze event correlation graphs, which can effectively mine the dependencies between devices and the potential correlation patterns between events, and achieve more accurate event classification and causal identification. This process avoids the limitations of traditional rule systems in their ability to handle complex relationships, and significantly improves the accuracy and intelligence level of event identification and structured modeling.

[0020] S2. Use graph neural networks to analyze the event correlation graph, identify the dependencies between devices, perform operation and maintenance event classification and correlation analysis, and generate operation and maintenance event type and correlation pattern information. This includes the following steps: S21. Extract features from each node based on the event association graph, and use the graph convolution layer in the graph neural network to perform convolution operations on the event association graph to update the feature representation of each node, namely: ; in, Representation node In the Feature representation after layer graph convolution, For the The weight matrix of the layer graph convolution, For nodes The set of neighbor nodes of For nodes With neighboring nodes The connection weights between is the activation function, Representation node In the Feature representation after layer graph convolution; S22. Perform maximum pooling on the node features after the graph convolution operation to obtain the global features of the graph; S23. Classify and predict events based on the global features of the graph, and generate type information and association pattern information of operation and maintenance events.

[0021] By introducing time series features and embedding time-dependent modeling capabilities in graph convolution, the present invention enables the system to accurately capture the dynamic evolution process of events and achieve forward-looking predictions of event development trends. This predictive capability provides a basis for early intervention and risk avoidance, and improves the system's early warning and proactive processing capabilities.

[0022] S3. Introduce time series information into the graph neural network, analyze the evolution trend of operation and maintenance events through graph convolution combined with time dependency, predict the state changes of future events, and generate time prediction information for operation and maintenance events. This includes the following steps: S31. Based on the graph neural network model, time series information is introduced to model the temporal dependency of operation and maintenance events. Each operation and maintenance event is associated with timestamp information to generate operation and maintenance event node features with time series information. S32. Perform graph convolution processing on the operation and maintenance event node features with time series information, update the node feature representation, and combine the device status, the order and time interval of the operation and maintenance events to obtain the time evolution characteristics of the operation and maintenance events, namely: ; in, For nodes In the Feature representation after layer graph convolution, For nodes In the Feature representation after layer graph convolution, For the The weight matrix of the layer graph convolution, For nodes The set of neighbor nodes of For nodes With neighboring nodes The connection weights between For nodes Timestamp information, for Activation function; S33. Extract temporal dependencies through a multi-layer graph convolutional network and perform time series modeling to obtain future state predictions for operation and maintenance events: ; in, Operation and maintenance events The future state prediction value, is the prediction function, For events Neighbor node features, For events Timestamp information.

[0023] The present invention dynamically optimizes event processing strategies through an adaptive differential evolution algorithm. The system can automatically generate processing priorities, resource scheduling, and recovery plans based on a comprehensive consideration of event types, correlations, and development trends, thereby improving the optimality and scenario adaptability of strategies, reducing the burden of manual configuration, and enhancing the system's autonomous decision-making capabilities.

[0024] S4: Based on the generated event types, associated pattern information, and time prediction information of operation and maintenance events, an adaptive differential evolution algorithm is applied to optimize the event processing strategy, adjust the operation and maintenance event processing priority, resource scheduling scheme, and fault recovery strategy. This includes the following steps: S41. Construct an event processing optimization model based on the generated operation and maintenance event type information and associated pattern information, combined with the future state prediction value of the operation and maintenance event, and define an optimization objective function; The calculation formula of the event processing optimization model is: ; in, Indicates the processing strategy output of event e, represents the ReLU activation function, The node feature vector representing event e, Representing an event The time state prediction value of Representing an event Characteristics of the association relationship with other events, is the weight matrix, is the bias term; The event processing optimization model is optimized by minimizing the following objective function: ; in, is the overall optimization goal, is a collection of events, For events The weight of Operation and maintenance events Type information and future state predictions The loss function between For association mode information Operation and maintenance events The impact measure, is the event processing strategy set optimized by the adaptive differential evolution algorithm and event time predictions The resource adjustment cost between and is the weight coefficient.

[0025] S42. Apply the adaptive differential evolution algorithm to solve the optimization objective function and initialize a set of candidate operation and maintenance event processing strategies. ,in, Process the decision vector for each operation and maintenance event, expressed as: ; in, The priority of the operation and maintenance event. For resource allocation plan, For fault recovery solutions; S43. Calculate the fitness value of each candidate strategy according to the standard process of the adaptive differential evolution algorithm: ; in, For strategy The fitness value of For strategy The corresponding optimization target value; S44. Based on the adaptive differential evolution algorithm, select the strategy with the best fitness value for crossover and mutation operations, and generate a new candidate strategy set : ; in, For crossover operation, Generate a new set of candidate strategies for mutation operations; S45, through repeated iterative crossover, mutation and selection operations, gradually approach the optimal processing strategy, and finally output the optimal processing strategy set : ; in, is the optimal event processing strategy set, is the corresponding optimization target value.

[0026] The present invention executes event processing operations according to the optimized strategy, realizing a full-process automated response from fault detection to automatic repair and resource allocation. This processing method can significantly reduce event response time, improve system fault handling efficiency, and at the same time ensure the rationality of resource utilization and the continuity of system operation.

[0027] S5. Based on the optimized event handling strategy, perform automated operation and maintenance event handling operations, including fault detection, automatic repair, resource allocation, and event response. This includes the following steps: S51. Assign a corresponding processing priority to each operation and maintenance event based on the generated optimal event processing strategy set; S52. Based on the operation and maintenance event priority and resource availability information, a resource scheduling algorithm is used to allocate resources to the processing strategy and generate a resource allocation plan for each operation and maintenance event; S53. Generate a fault recovery plan for each operation and maintenance event, and select a recovery plan based on the event type and its associated mode; S54. Execute automated processing of maintenance events based on their priorities, resource allocation plans, and fault recovery plans, including fault detection, automatic repair, resource scheduling, and event response. S55. During the operation and maintenance event processing, the processing effect is monitored in real time, the processed operation and maintenance events are dynamically evaluated, and the processing strategy is adjusted according to the evaluation results.

[0028] The present invention constructs a self-learning closed-loop mechanism through real-time feedback and evaluation of processing results, combined with the dynamic adjustment strategy of the adaptive differential evolution algorithm. The system can continuously iterate and optimize event processing solutions, improve the stability and robustness of processing effects, and ensure that the system continues to maintain efficient operation and intelligent evolution capabilities in a changing environment.

[0029] S6. Based on the feedback from the event processing results, the adaptive differential evolution algorithm is used to dynamically adjust the event processing strategy and continuously optimize the event processing strategy to improve the efficiency and accuracy of event processing. This includes the following steps: S61. Evaluate the processing effect of each operation and maintenance event based on the operation and maintenance event processing result and generate a processing effect evaluation value; S62. Calculate the overall effect evaluation value of the optimal event processing strategy set based on the processing effect evaluation values ​​of all operation and maintenance events; S63. Based on the overall effect evaluation value, the event processing strategy is dynamically adjusted using the adaptive differential evolution algorithm; S64. Apply the adjusted strategy set to the operation and maintenance event processing process, continuously collect feedback information on the processing results through the real-time monitoring system and feedback mechanism, and further optimize the event processing strategy based on the new feedback data.

[0030] Example 2 Based on the automatic processing method of operation and maintenance events proposed in Example 1, this embodiment provides an automatic processing system for operation and maintenance events based on AI intelligent agent, which consists of a data acquisition module, an event correlation analysis module, an event time prediction module, an event processing strategy optimization module, an event automatic processing module, a processing effect evaluation module and a strategy dynamic adjustment module.

[0031] The data acquisition module is used to collect various types of operation and maintenance event data in the operation and maintenance system, including equipment status monitoring data, equipment operation logs, and network topology information, and perform preprocessing; the event correlation analysis module is used to perform graph neural network analysis based on the event correlation graph, identify the dependencies between devices, classify and predict operation and maintenance events, and generate operation and maintenance event types and correlation pattern information; the event time prediction module is used to predict the evolution trend of operation and maintenance events based on time series information and generate time prediction information for operation and maintenance events.

[0032] The event processing strategy optimization module is used to optimize the operation and maintenance event processing strategy based on the operation and maintenance event type, correlation pattern and time prediction information, using the adaptive differential evolution algorithm to adjust the operation and maintenance event priority, resource scheduling plan and fault recovery strategy; the event automation processing module is used to perform automated processing operations of operation and maintenance events according to the optimized strategy set, including fault detection, automatic repair, resource allocation and event response; the processing effect evaluation module is used to evaluate the processing effect of each operation and maintenance event based on the feedback information of the event processing results; the strategy dynamic adjustment module is used to monitor the operation and maintenance event processing process in real time, adjust the processing strategy according to the feedback, and realize dynamic optimization and adaptive adjustment of the strategy.

[0033] Example 3 This example is based on a large provincial-level financial data center that serves over 7,000 branches and nearly 10 million users. Its IT infrastructure includes 3,500 servers, 820 network switching devices, and 1,200 business application modules. During system operation, the center generates an average of approximately 120,000 pieces of device status data, 280,000 system logs, and approximately 3,000 fault and abnormal operation and maintenance events daily. Due to the large number of operation and maintenance events, the center has long faced problems such as inefficient manual response, rigid operation and maintenance strategies, and weak fault prediction capabilities.

[0034] To solve the above problems, the center introduced the AI ​​intelligent agent-based automatic processing method for operation and maintenance events of the present invention and deployed a complete system. The system was launched in November 2024. Comparative tests were conducted before and after operation, mainly including fault response time, resource utilization efficiency, fault prediction accuracy and other dimensions.

[0035] During the application process, the system was first deployed in a database cluster and selected load balancing devices in the core financial district. Device operational status (such as CPU load, memory usage, and I / O latency), system logs, and network topology information were collected and preprocessed in real time. An event correlation graph was dynamically constructed, and graph neural network modeling was used to analyze dependency paths between database nodes, as well as the high-frequency correlation between master-slave synchronization failures and network jitter.

[0036] Within two weeks, the system automatically identified 387 hidden anomalies across five categories of events. By introducing a time series prediction model, it predicted that 234 of these would evolve into high-impact failures within 72 hours. The adaptive differential evolution algorithm generated the optimal processing strategy based on event importance, time prediction trends, and device criticality, automatically scheduled backup services, and preemptively performed lightweight cache cleanup and hot data migration operations on 26 database servers with potential write delays.

[0037] After receiving system predictions and generating strategies, the event processing module no longer relies on on-duty engineers to manually execute commands. Instead, it implements a closed-loop disposal mechanism that automatically issues network isolation policies, adjusts resource pools, and links application degradation with alarm notifications. The strategy optimization engine conducts real-time evaluation and feedback adjustments based on each processing result, completing 94 rounds of strategy self-learning and weight updates in less than 10 days.

[0038] Compared with traditional static O&M methods, after deploying the new system, fault response time was reduced from an average of 8.7 minutes to 1.3 minutes, the predictive repair rate increased to 82.6%, average resource scheduling efficiency increased by 21%, and the recurrence rate of abnormal events decreased by nearly 50%. The following table shows the key comparative data collected: Table 1 Comparison of event processing efficiency before and after system introduction

[0039] Table 2 Abnormal event prediction accuracy and early processing effectiveness

[0040] To sum up, this embodiment clearly demonstrates the application effect of the present invention in an actual operation and maintenance environment. By constructing an event correlation graph, time series prediction and strategy optimization linkage processing model, the system not only greatly improves the automation and intelligence of event processing, but also significantly shortens the response time and reduces operation and maintenance costs, verifying that the present invention has obvious technological advancement and practical value.

[0041] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. An automated processing method for operation and maintenance events based on AI agents, characterized in that: The steps include: S1. Collect device status monitoring data, device operation logs, network topology information, and operation and maintenance event data from the operation and maintenance system, pre-process the collected data, construct graph structure data based on the relationship between devices and events, and generate an event correlation graph; S2. Analyze the event association graph using a graph neural network to identify dependencies between devices, perform operation and maintenance event classification and association analysis, and generate operation and maintenance event type and association pattern information; S3. Introducing time series information into graph neural networks, using graph convolution combined with temporal dependencies to analyze the evolution trend of operation and maintenance events, predict the state changes of future events, and generate time prediction information for operation and maintenance events. S4. Based on the generated event types, correlation pattern information, and time prediction information of operation and maintenance events, an adaptive differential evolution algorithm is applied to optimize the event processing strategy, adjust the operation and maintenance event processing priority, resource scheduling plan, and fault recovery strategy; S5. Execute automated O&M event processing operations based on the optimized event handling strategy, including fault detection, automatic repair, resource allocation, and event response. S6. Based on the feedback information of the event processing results, the event processing strategy is dynamically adjusted using the adaptive differential evolution algorithm, and the event processing strategy is continuously optimized.

2. The method for automatically processing operation and maintenance events according to claim 1, characterized in that: The S1 includes the following specific steps: S11. Collect device status monitoring data and network topology information from the operation and maintenance system, and collect operation and maintenance event data from different data sources, and perform unified processing based on the timestamps of the operation and maintenance event data; S12. De-noising the collected equipment status monitoring data by smoothing the data using a sliding window technique to filter out short-term fluctuations; S13. Align the timestamps of the collected device operation logs; S14. Construct an event association graph based on the relationship between devices and the timing information of operation and maintenance events. The nodes of the event association graph represent devices and operation and maintenance events, and the edges represent the communication relationship between devices and the causal relationship between operation and maintenance events.

3. The method for automatically processing operation and maintenance events according to claim 1, characterized in that: The S2 includes the following specific steps: S21. Extract features from each node based on the event association graph, and use the graph convolution layer in the graph neural network to perform convolution operations on the event association graph to update the feature representation of each node, namely: ; in, Representation node In the Feature representation after layer graph convolution, For the The weight matrix of the layer graph convolution, For nodes The set of neighbor nodes of For nodes With neighboring nodes The connection weights between is the activation function, Representation node In the Feature representation after layer graph convolution; S22. Perform maximum pooling on the node features after the graph convolution operation to obtain the global features of the graph; S23. Classify and predict events based on the global features of the graph, and generate type information and association pattern information of operation and maintenance events.

4. The method for automatically processing operation and maintenance events according to claim 1, characterized in that: The S3 includes the following specific steps: S31. Based on the graph neural network model, time series information is introduced to model the temporal dependency of operation and maintenance events. Each operation and maintenance event is associated with timestamp information to generate operation and maintenance event node features with time series information. S32. Perform graph convolution processing on the operation and maintenance event node features with time series information, update the node feature representation, and combine the device status, the order and time interval of the operation and maintenance events to obtain the time evolution characteristics of the operation and maintenance events, namely: ; in, For nodes In the Feature representation after layer graph convolution, For nodes In the Feature representation after layer graph convolution, For nodes In the Feature representation after layer graph convolution, For the The weight matrix of the layer graph convolution, For nodes The set of neighbor nodes of For nodes With neighboring nodes The connection weights between For nodes Timestamp information, for Activation function; S33. Extract time dependencies through multi-layer graph convolutional networks and perform time series modeling to obtain future state predictions of operation and maintenance events.

5. The method for automatically processing operation and maintenance events according to claim 1, characterized in that: The S4 includes the following specific steps: S41. Construct an event processing optimization model based on the generated operation and maintenance event type information and associated pattern information, combined with the future state prediction value of the operation and maintenance event, and define an optimization objective function; S42. Apply the adaptive differential evolution algorithm to solve the optimization objective function and initialize a set of candidate operation and maintenance event processing strategies. ,in, Process the decision vector for each operation and maintenance event, expressed as: ; in, The priority of the operation and maintenance event. For resource allocation plan, For fault recovery solutions; S43. Calculate the fitness value of each candidate strategy according to the standard process of the adaptive differential evolution algorithm: ; in, For strategy The fitness value of For strategy The corresponding optimization target value; S44. Based on the adaptive differential evolution algorithm, select the strategy with the best fitness value for crossover and mutation operations, and generate a new candidate strategy set : ; in, For crossover operation, Generate a new set of candidate strategies for mutation operations; S45, through repeated iterative crossover, mutation and selection operations, gradually approach the optimal processing strategy, and finally output the optimal processing strategy set : ; in, is the optimal event processing strategy set, is the corresponding optimization target value.

6. The method for automatically processing operation and maintenance events according to claim 5, characterized in that: The calculation formula of the event processing optimization model of S41 is: ; in, Indicates the processing strategy output of event e, represents the ReLU activation function, The node feature vector representing event e, Representing an event The time state prediction value of Representing an event Characteristics of the association relationship with other events, is the weight matrix, is the bias term; The event processing optimization model is optimized by minimizing the following objective function: ; in, is the overall optimization goal, is a collection of events, For events The weight of Operation and maintenance events Type information and future state predictions The loss function between For association mode information Operation and maintenance events The impact measure, is the event processing strategy set optimized by the adaptive differential evolution algorithm and event time predictions The resource adjustment cost between and is the weight coefficient.

7. The method for automatically processing operation and maintenance events according to claim 1, characterized in that: The S5 includes the following specific steps: S51. Assign a corresponding processing priority to each operation and maintenance event based on the generated optimal event processing strategy set; S52. Based on the operation and maintenance event priority and resource availability information, a resource scheduling algorithm is used to allocate resources to the processing strategy and generate a resource allocation plan for each operation and maintenance event; S53. Generate a fault recovery plan for each operation and maintenance event, and select a recovery plan based on the event type and its associated mode; S54. Execute automated processing of maintenance events based on their priorities, resource allocation plans, and fault recovery plans, including fault detection, automatic repair, resource scheduling, and event response. S55. During the operation and maintenance event processing, the processing effect is monitored in real time, the processed operation and maintenance events are dynamically evaluated, and the processing strategy is adjusted according to the evaluation results.

8. The method for automatically processing operation and maintenance events according to claim 1, characterized in that: The S6 includes the following specific steps: S61. Evaluate the processing effect of each operation and maintenance event based on the operation and maintenance event processing result and generate a processing effect evaluation value; S62. Calculate the overall effect evaluation value of the optimal event processing strategy set based on the processing effect evaluation values ​​of all operation and maintenance events; S63. Based on the overall effect evaluation value, the event processing strategy is dynamically adjusted using the adaptive differential evolution algorithm; S64. Apply the adjusted strategy set to the operation and maintenance event processing process, continuously collect feedback information on the processing results through the real-time monitoring system and feedback mechanism, and further optimize the event processing strategy based on the new feedback data.

9. An automated operation and maintenance event processing system, executing the automated operation and maintenance event processing method according to any one of claims 1 to 8, characterized in that: Includes the following modules: The data acquisition module is used to collect and pre-process various operation and maintenance event data in the operation and maintenance system, including equipment status monitoring data, equipment operation logs, and network topology information; The event correlation analysis module is used to perform graph neural network analysis based on the event correlation graph, identify the dependencies between devices, classify and predict operation and maintenance events, and generate operation and maintenance event types and correlation pattern information; The event time prediction module is used to predict the evolution trend of operation and maintenance events by combining time series information and generate time prediction information for operation and maintenance events; The event processing strategy optimization module is used to optimize the operation and maintenance event processing strategy based on the operation and maintenance event type, correlation pattern and time prediction information using the adaptive differential evolution algorithm, and adjust the operation and maintenance event priority, resource scheduling plan and fault recovery strategy; The event automation processing module is used to perform automated processing operations for operation and maintenance events according to the optimized policy set, including fault detection, automatic repair, resource allocation and event response; The processing effect evaluation module is used to evaluate the processing effect of each operation and maintenance event based on the feedback information of the event processing results; The policy dynamic adjustment module is used to monitor the operation and maintenance event processing process in real time, adjust the processing strategy based on feedback, and realize dynamic optimization and adaptive adjustment of the strategy.

Citation Information

Cited By

  • Equipment linkage strategy generation method and system based on space-time event

    CN121151433A

  • Unified operation and maintenance management and control method

    CN121580790A