Intelligent operation and maintenance system based on AI Agent

Through an intelligent operation and maintenance system based on AI Agent, machine learning and graph neural network decomposition and optimization of operation and maintenance tasks are used to solve the problems of low efficiency, error prone and lack of forward-looking in the traditional operation and maintenance model, and efficient and reliable intelligent operation and maintenance are achieved.

CN120336116AActive Publication Date: 2025-07-18XIAN DUNXUN INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510394663.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The traditional operation and maintenance model is inefficient, prone to errors, lacks forward-lookingness and low degree of automation, making it difficult to meet the needs of efficient and stable operation of modern complex systems.

Method used

The intelligent operation and maintenance system based on AI Agent is adopted to monitor the operation data of the operation and maintenance system through the content collection module, decompose tasks by the decision-making and planning module, execute monitoring module calling tools, feedback optimization module optimization and collaboration, and use machine learning and graph neural network for feature extraction and task scheduling to achieve intelligent operation and maintenance.

Benefits of technology

It improves the collaboration efficiency and reliability of the operation and maintenance system, can quickly locate problems, optimize task execution processes, reduce human errors, and improve the automation and foresight of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336116A_ABST
    Figure CN120336116A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent operation and maintenance system based on an AI Agent, and the system is characterized in that the system comprises a content collection module which obtains abnormal information in operation data and an operation and maintenance problem submitted by a user, and determines a target operation and maintenance task; the decision planning module is used for decomposing the target operation and maintenance task into a plurality of sub-tasks, determining a closest solution based on each sub-task, and determining required principles and tools according to the solutions; the execution monitoring module is used for calling and splicing tools, monitoring the completion state of each subtask and summarizing and memorizing each cooperation of the principle and the tools by taking completion of the subtasks as a target through the principle; and the feedback optimization module is used for reasoning the summarized and memorized content and summarizing the cooperation of the principle and the tool into a solution based on the target operation and maintenance task. The operation and maintenance system is perceived, understood, decided, executed and fed back through the AI Agent, and intelligent operation and maintenance are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of operation and maintenance systems, and particularly to an intelligent operation and maintenance system based on an AI Agent. Background Art

[0002] In the traditional operation and maintenance mode, operation and maintenance personnel mainly rely on manual operations and personal experience to handle various tasks, such as system monitoring, fault troubleshooting, and configuration management. This mode has many deficiencies: on the one hand, in the face of complex systems and a large number of repetitive tasks, operation and maintenance personnel rely solely on personal experience and professional knowledge to handle manually, which is not only inefficient but also error-prone. Especially when the system fails, it often takes a long time to locate and solve the problem, resulting in a serious impact on business operations. On the other hand, the traditional operation and maintenance mode lacks foresight and mainly focuses on the current running state of the system, making it difficult to predict potential faults, such as performance bottlenecks or hardware failures. Usually, it can only be processed after the problem occurs, increasing the risk of the system. In addition, the automation degree of the traditional operation and maintenance mode is relatively low, and many operations need to be completed manually, such as manually uploading installation packages, configuring parameters, and restarting services during software updates. This not only further reduces efficiency but may also cause new system failures due to human errors. In short, the traditional operation and maintenance mode has obvious defects in terms of efficiency, cost, foresight, and automation degree, and it is difficult to meet the requirements of the efficient and stable operation of modern complex systems. Summary of the Invention

[0003] The purpose of the present invention is to provide an intelligent operation and maintenance system based on an AI Agent, which realizes intelligent operation and maintenance through the perception, understanding, decision-making, execution, and feedback of the AI Agent on the operation and maintenance system.

[0004] To achieve the above invention purpose, the present invention provides an intelligent operation and maintenance system based on an AI Agent, and the system includes:

[0005] A content collection module, which is used to monitor the operation data of the operation and maintenance system, obtain the abnormal information in the operation data and the operation and maintenance problems submitted by users to determine the target operation and maintenance tasks;

[0006] A decision-making and planning module, which is used to decompose the target operation and maintenance tasks into several subtasks, determine the closest solution based on each subtask, and determine the required principles and tools according to the solution;

[0007] An execution and monitoring module, which is used to call and splice tools based on the principles with the goal of completing the subtasks, monitor the completion status of each subtask. If the subtask is successfully completed, summarize and remember the cooperation between the principles and tools. If the subtask fails to complete, optimize the cooperation between the tools and principles until the target operation and maintenance tasks are completed, and summarize and remember each cooperation;

[0008] A feedback optimization module, which is used to reason about the summarized and memorized content, and summarize the collaboration of principles and tools into solutions based on the target operation and maintenance tasks.

[0009] Furthermore, determining the target operation and maintenance tasks specifically includes:

[0010] Preprocess the obtained abnormal information, extract features from the preprocessed abnormal information through machine learning to obtain abnormal features, pre-train an abnormal detection model, and determine the type and severity of the abnormality through the abnormal detection model using the abnormal features;

[0011] Perform natural language processing on the operation and maintenance problems submitted by users to obtain semantic features, pre-train a problem recognition model, and determine the user's intention through the problem recognition model using the semantic features;

[0012] Fuse the abnormal features and semantic features through a multi-modal fusion algorithm to obtain fusion features, pre-train a task generation model, and determine the target operation and maintenance tasks by inputting the fusion features into the task generation model.

[0013] Furthermore, decompose the target operation and maintenance tasks into several subtasks through an AI Agent, specifically including:

[0014] Decompose the target operation and maintenance tasks into several subtasks according to the requirements of the target operation and maintenance tasks through a task decomposition algorithm;

[0015] Determine the dependency relationships between the subtasks, and set corresponding evaluation indicators according to the dependency relationships between the subtasks.

[0016] Furthermore, call and splice the tools with the goal of completing the subtasks through principles, specifically including:

[0017] Determine the dependency relationships between the subtasks, determine the tools required for the subtasks with dependency relationships, and determine the splicing order of the tools through principles;

[0018] Determine whether the subtasks without dependency relationships need to call the same tool, determine whether the same tool called can be called simultaneously, and if not, determine the splicing order of the tools through principles;

[0019] Merge the subtasks with dependency relationships into a collective, determine whether the tools required for the collective and the remaining subtasks can be called simultaneously, and if not, determine the splicing order of the tools through principles.

[0020] Furthermore, summarize and memorize the collaboration of principles and tools, specifically including:

[0021] Represent the collaborative relationship between tools and principles as a graph structure, where nodes represent tools and principles, and edges represent the collaborative relationship between tools and principles;

[0022] Learn the feature vectors of nodes and edges in the graph structure through a graph neural network, and classify the collaboration between tools and principles by passing the extracted feature vectors through a classification algorithm;

[0023] Summarize and memorize the collaborative classification data through a deep classification learning model to obtain a collaborative strategy model.

[0024] Furthermore, optimize the collaboration between tools and principles, specifically including:

[0025] Consider the task execution status, the dependencies between tasks, and the overall status of the operation and maintenance system when all subtasks fail;

[0026] Construct the collaborative dependency relationship between tools and principles among failed subtasks as a directed acyclic graph, and transform the dependencies, status of subtasks during task execution, and the overall status of the operation and maintenance system into weight vectors and embed them into the directed acyclic graph;

[0027] Calculate the attention coefficients between each task and its neighboring tasks in the remaining tasks of the directed acyclic graph, and learn the high-dimensional state features of each task in the remaining tasks through a graph neural network;

[0028] Use the high-dimensional state features as reference information for the task scheduling algorithm to regenerate the scheduling scheme for the collaboration between tools and principles.

[0029] Furthermore, reason about the summarized and memorized content, specifically including:

[0030] Collect the collaborative data of tools and principles in operation and maintenance tasks, including successful cases and failure cases. Regard the constructed graph structure as a successful case and the constructed directed acyclic graph as a failure case;

[0031] Learn the first inference feature vectors of nodes and edges in the graph structure and the directed acyclic graph through a graph neural network;

[0032] Input the graph structure and the directed acyclic graph into the collaborative strategy model to obtain the second inference feature vector;

[0033] Map the first inference feature vector and the second inference feature vector to a vector space through knowledge graph embedding technology, and predict new principle and tool collaboration relationships by calculating the similarity between vectors.

[0034] Compared with the prior art, the beneficial effects of the present invention are:

[0035] An intelligent operation and maintenance system based on AI Agent provided by the present invention monitors the operation data of the operation and maintenance system through a content collection module, obtains abnormal information and problems submitted by users, and determines the target operation and maintenance tasks. The decision-making and planning module decomposes the target operation and maintenance tasks into several subtasks, determines the closest solutions based on each subtask, and clarifies the required principles and tools to reasonably arrange the task execution order. The execution monitoring module calls and splices tools with the goal of completing subtasks, represents the collaboration relationship between tools and principles as a graph structure, learns feature vectors through a graph neural network, performs similarity calculations using knowledge graph embedding technology, predicts new collaboration relationships, and continuously optimizes the collaboration between tools and principles; by monitoring the completion status of each subtask and optimizing and adjusting the failed subtasks, it realizes the intelligent management and execution of the target operation and maintenance tasks. The feedback optimization module reasons about the summarized and memorized content, generalizes the collaboration between principles and tools into solutions, accumulates operation and maintenance experience through summarized memory and knowledge graph embedding technology, provides reasoning references for subsequent new operation and maintenance tasks, and improves the collaboration efficiency of the operation and maintenance system. The present invention realizes the efficient execution of operation and maintenance tasks and improves the reliability of the operation and maintenance system based on AI Agent. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0037] Figure 1 It is a schematic structural diagram of an intelligent operation and maintenance system based on AI Agent provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described here are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the drawings.

[0039] Refer to Figure 1 , this embodiment provides an intelligent operation and maintenance system based on AI Agent, and the system includes:

[0040] A content collection module, configured to monitor the operation data of the operation and maintenance system, obtain abnormal information in the operation data, and determine the target operation and maintenance tasks for the operation and maintenance problems submitted by users.

[0041] In this embodiment, obtaining the operation data in the operation and maintenance system includes, but is not limited to, system resource data, application performance data, log data, and business metric data. Specifically, the system resource data reflects the load situation and resource consumption status of the server, which are important indicators for evaluating system performance and stability, such as the CPU usage rate, memory utilization rate, disk space, network traffic, etc. of the server. The application performance data includes the response time, throughput, error rate, etc. of the application program. For example, if the response time of the payment interface of an e-commerce website is too long, it may affect the user experience and even lead to transaction failures. A large amount of logs will be generated during the operation of the system and application programs, which may contain records of abnormal situations such as error messages and warning messages. By analyzing the logs, the root cause of the problem can be traced. Obtaining the metrics related to the enterprise business helps the AI Agent evaluate the operation status of the system from the business perspective and promptly discover problems that may affect the business, such as the order volume and user activity.

[0042] Optionally, through existing operation and maintenance monitoring tools and platforms, the operation data is collected and analyzed in real time to automatically detect abnormal situations and promptly send an alarm to the AI Agent, such as Keynote Listening Cloud, the operation and maintenance center of Alibaba Cloud, etc. The AI Agent determines whether to identify the alerted information as abnormal information according to the alarm information, and performs a re-abnormality determination on the obtained operation data through a preset threshold. When the metric exceeds the set judgment threshold, it is considered abnormal. For example, the CPU usage rate threshold of the server is set to 80%. When the usage rate exceeds this value, the usage rate is regarded as abnormal information. After the system has been running for a period of time, the historical data is analyzed. When there is a large deviation between the newly collected data and the model, it is judged that an abnormality may occur. For example, machine learning algorithms are used to model the usage of system resources and predict future usage trends. If the actual value deviates too much from the predicted value, the operation data with too large a deviation is determined as abnormal information.

[0043] Problems encountered by users during the use of the system can reflect the abnormal conditions of the system from the perspective of actual use. By combining the problems submitted by users and the abnormal information in the operation data, the operation status of the system can be understood more comprehensively, and the operation and maintenance tasks with greater impact on the business can be determined and processed preferentially. For example, if a user reports that a certain business function cannot be used normally, and the operation data shows abnormal network traffic of the relevant server, the AI Agent can take repairing the network problem corresponding to the business function as the target operation and maintenance task to quickly restore the normal operation of the business. Specifically, the abnormal information in the operation data often means that there may be potential faults or performance bottlenecks in the system. By analyzing these abnormal information and the feedback from users, the location of the problem can be quickly determined, so as to determine the tasks that need to be carried out for operation and maintenance. For example, if it is found that the disk space of a certain server is full, the AI Agent needs to clean up the disk space or expand the capacity in time to prevent the system from malfunctioning due to insufficient disk space.

[0044] Determine the target operation and maintenance tasks, specifically including:

[0045] Preprocess the obtained abnormal information, extract features from the preprocessed abnormal information through machine learning to obtain abnormal features, pre-train the abnormal detection model, and determine the type and severity of the abnormality through the abnormal detection model with the abnormal features.

[0046] Perform natural language processing on the operation and maintenance problems submitted by users to obtain semantic features, pre-train the problem recognition model, and determine the user's intention through the problem recognition model with the semantic features.

[0047] Fuse the abnormal features and semantic features through a multi-modal fusion algorithm to obtain fusion features, pre-train the task generation model, and determine the target operation and maintenance tasks by inputting the fusion features into the task generation model.

[0048] In this embodiment, by preprocessing and feature extraction of the abnormal information, the key features related to the abnormality can be screened out from a large amount of complex data, which can accurately point to the specific abnormal situation in the system. For example, through the analysis of the machine learning model, it is accurately identified that certain specific system resource usage patterns indicate that the server is about to be overloaded, so as to determine the operation and maintenance tasks such as resource allocation or capacity expansion for the server. Specifically, in the face of a large number of servers and complex systems in a large data center, the abnormal detection model can quickly locate the problematic server or module, enabling the AI Agent to process it in time and reducing the impact of system failures on the business.

[0049] Performing natural language processing on the operation and maintenance problems submitted by users can convert the natural language descriptions of users into semantic features that can be understood by computers, and then accurately judge the real needs of users through a problem recognition model. For example, when a user describes that a certain function cannot be used normally, semantic analysis can determine whether it is a network connection problem, a server response problem, or a problem with the application itself, so as to provide a clear direction for subsequent task determination.

[0050] The multi-modal fusion algorithm fuses abnormal features and semantic features, comprehensively considers abnormal information and user feedback, and avoids the one-sidedness that may be caused by a single information source. The fused features can more accurately reflect the actual operation status of the system and the needs of users, so as to determine the target operation and maintenance tasks. For example, the system operation data may show a performance decline of a certain module, while the user feedback may focus on the usage problems of this module. The fused information can clearly indicate that operation and maintenance tasks such as performance optimization or fault troubleshooting need to be carried out on this module.

[0051] The decision-making and planning module is used to decompose the target operation and maintenance tasks into several subtasks, determine the closest solution based on each subtask, and determine the required principles and tools according to the solution.

[0052] In this embodiment, the principles and tools involved in the solution are principles with similar fundamentality and generality in the operation and maintenance field, such as the principles of automated operation and maintenance, fault diagnosis, and machine learning. The solution refers to the specific steps, strategies, or plans taken to complete a certain subtask or achieve a certain goal. It is a relatively high-level concept that describes the overall idea and process of how to solve problems. The solution usually combines the use of multiple principles and tools to achieve the best solution effect for the target operation and maintenance tasks. For example, through data analysis tools and machine learning algorithms, abnormal patterns and bottlenecks in resource usage are identified.

[0053] The principle refers to the basic laws and theories of science, mathematics, or engineering on which the solution is based. It is the theoretical basis of the solution and explains why a certain method can effectively solve the problem. The principle usually has universality and abstraction and is a general guiding principle for solving similar problems. For example, using machine learning algorithms to analyze and model operation and maintenance data to achieve intelligent operation and maintenance functions such as anomaly detection and fault prediction, the principle is based on mathematical theories such as statistics, probability theory, and linear algebra, and learns the laws and patterns in the data by constructing models.

[0054] A tool refers to a specific software, hardware, or technical means used to implement a solution method. A tool is the actual carrier of the solution method and principle. By operating the tool, specific operation and maintenance tasks can be completed. Tools usually have practicality and pertinence, and can be directly applied to actual work scenarios to improve work efficiency and quality. For example, TensorFlow, PyTorch, Scikit-learn, etc. are used to build and train machine learning models, mine and analyze operation and maintenance data, realize intelligent operation and maintenance decision-making, and provide rich algorithm libraries and model building frameworks to facilitate users to perform data processing and model training.

[0055] Specifically, the principle of automated operation and maintenance realizes the automated execution of operation and maintenance tasks through scripts or tools, improving efficiency and accuracy. For example, using automated operation and maintenance tools such as Ansible, repetitive operation and maintenance operations are written into reusable task scripts, and the AI Agent only needs to directly call them to complete complex and repetitive daily operation and maintenance work. The principle of fault diagnosis is based on system operation data and user feedback information, and uses methods such as logical reasoning and data analysis to determine the location and cause of the fault. For example, when analyzing the performance bottleneck of a server, principles such as queuing theory can be used. By analyzing data such as the usage of system resources and the length of the task queue, it can be judged whether there are problems such as resource competition or task blocking. The principle of machine learning uses machine learning algorithms to analyze and model operation and maintenance data to realize intelligent operation and maintenance functions such as anomaly detection and fault prediction. For example, through clustering algorithms, the system operation data is clustered and analyzed, and similar operation states are classified into one category, so as to quickly discover abnormal operation patterns.

[0056] A tool is a specific implementation form of a principle. Its design and function are based on the corresponding principle. The process of using a tool is the application process of the principle in an actual scenario. During the use of the tool, new problems and requirements may be encountered, which prompts the further development and innovation of the principle. For example, with the continuous increase in data volume and complexity, traditional data analysis principles may not be able to meet the requirements, thus promoting the innovation of big data analysis principles and algorithms. The solution method is formulated under the guidance of the principle. The principle provides the theoretical basis and guiding principles for the solution method. The tool is the actual executor of the solution method, and various tasks in the solution method are completed by operating the tool. The solution method, principle, and tool cooperate with each other to jointly form a complete solution. The principle provides theoretical support for the solution method, the solution method guides the use of the tool, and the tool in turn verifies and optimizes the principle. The three interact with each other to jointly promote the effective completion of operation and maintenance tasks.

[0057] Exemplarily, the target operation and maintenance task is that the server disk space is insufficient. The AI Agent obtains the target operation and maintenance task and splits it into subtasks, and calls the principle and tool through matching the solution method to solve the target operation and maintenance task:

[0058] Subtask 1: Monitor disk usage

[0059] Solution: Select appropriate monitoring tools such as Zabbix, Prometheus, etc.; Configure the monitoring tools, set monitoring metrics (such as disk usage rate of each partition) and monitoring frequency; Run the monitoring tools to collect server disk usage data in real time.

[0060] Principle applied: System monitoring principle, using the interfaces and data collection technologies provided by the operating system to obtain key metrics such as disk usage in real time. For example, Zabbix collects system data through agent programs, and Prometheus obtains metric data from target servers through the pull mode; Data collection and transmission principle, organizing and transmitting the collected disk usage data to ensure the accuracy and timeliness of the data.

[0061] Tools applied: System commands such as df - h to view disk usage, du - sh to view the size of a specified directory, etc.; Monitoring tools such as Zabbix, Prometheus, etc., which can monitor the disk usage of the server in real time, set alarm thresholds, and promptly detect problems of insufficient disk space.

[0062] Subtask 2: Analyze the reasons for disk space occupation

[0063] Solution: Use system commands (such as df - h, du - sh, etc.) to initially view the disk usage of each partition; Deeply analyze the types of files and directories that occupy a large amount of space to determine whether they are temporary files, log files, business data, etc.; Combine business requirements and data growth trends to analyze the root cause of insufficient disk space.

[0064] Principle applied: Data management principle, classifying, cleaning, and optimizing the data on the disk, following the theory of data life cycle management and storage optimization, reasonably allocating and utilizing storage resources to improve the utilization rate of disk space. For example, based on attributes such as the last access time and modification time of files, determine which files have not been used for a long time and can be cleaned or transferred; Data analysis principle, deeply analyzing the disk space usage through methods such as statistical analysis and trend prediction. For example, calculate the growth rate of disk usage in each partition and predict possible future space shortages.

[0065] Tools applied: System commands such as ls - l to list detailed information of files and directories, find to search for specific types of files, etc.; Data analysis tools such as Excel, the Pandas library in Python, etc., which are used to organize, analyze, and visualize the collected disk usage data to help more intuitively understand the disk space occupation situation.

[0066] Subtask 3: Clean and Optimize Disk Space

[0067] Solution: Clean unnecessary temporary files and cache files, and use tools (such as tmpreaper) to automatically clean expired temporary files; manage log files, use log rotation tools (such as logrotate) to control the size and quantity of log files, and delete useless log files; transfer or delete long-unused business data, back it up to external storage devices or delete expired business data.

[0068] Principle Applied: Data cleaning principle, based on the importance and usage frequency of data, clean unnecessary data to reduce disk space occupation; storage optimization principle, by adjusting data storage methods and structures, improve the utilization rate of disk space, such as merging small files for storage, adopting more efficient compression algorithms, etc.

[0069] Tools Used: Data cleaning tools, such as tmpreaper for cleaning temporary files, logrotate for managing the size and quantity of log files to prevent log files from overly occupying disk space; backup tools, such as rsync, tar, etc., for backing up long-unused business data to external storage devices to ensure data security and recoverability.

[0070] Subtask 4: Optimize Storage Strategy

[0071] Solution: Evaluate the current storage architecture to determine whether to increase disk capacity or adopt a more efficient storage architecture; plan data storage methods, and reasonably allocate storage resources according to the access frequency and importance of data; implement storage optimization measures, such as adding new disks, adjusting storage partitions, etc., and verify the optimization effect.

[0072] Principle Applied: Storage architecture design principle, according to business requirements and data characteristics, select a suitable storage architecture, such as RAID, distributed storage, etc., to improve the performance, reliability, and scalability of the storage system; resource allocation and scheduling principle, reasonably allocate storage resources, and perform dynamic resource scheduling according to the access pattern and priority of data to ensure the efficient operation of the system.

[0073] Tools Used: Storage management tools, such as LVM (Logical Volume Management) for adjusting storage partitions and disk capacity, RAID management tools for configuring and managing RAID arrays, etc.; performance testing tools, such as Iozone for testing disk I / O performance, helping to evaluate the effect of storage optimization measures and providing a basis for subsequent optimization and adjustment.

[0074] The target operation and maintenance tasks are decomposed into several subtasks by the AI Agent, specifically including:

[0075] According to the requirements of the target operation and maintenance task, the target operation and maintenance task is decomposed into several subtasks through a task decomposition algorithm.

[0076] Determine the dependency relationships between the subtasks, and set corresponding evaluation indicators according to the dependency relationships between the subtasks.

[0077] In this embodiment, in the operation and maintenance task decomposition, the target operation and maintenance task is decomposed into smaller work units to make complex tasks easier to manage and track. According to the requirements of the target operation and maintenance task, starting from the overall target operation and maintenance task, it is gradually decomposed into more specific subtasks. For example, for a target operation and maintenance task of system fault repair, it is first decomposed into subtasks such as "determine the fault scope", "analyze the fault cause", "formulate a repair plan", "execute the repair operation", "verify the repair effect", etc.; it is decomposed from dimensions such as the type, impact range, and processing steps of the operation and maintenance task. For example, for a target operation and maintenance task of server performance optimization, it can be decomposed into subtasks such as "monitor resource usage", "analyze performance bottlenecks", "optimize configuration parameters", "test the optimization effect", etc.; considering existing similar historical operation and maintenance tasks or ready-made templates, the decomposition method can be referred to to quickly determine the subtasks of the current task. For example, for a common server update task, referring to the previous update process template, the task can be decomposed into subtasks such as "backup data", "download the update package", "install the update", "restart the service", "verify the update success", etc.

[0078] According to the above-defined subtasks, determine the dependency relationships between the subtasks. According to the actual execution process of the operation and maintenance task, analyze which subtasks must be completed before other subtasks. For example, in the case of "insufficient server disk space", "monitor disk usage" (subtask 1) must be completed first before "analyze the reason for disk space occupation" (subtask 2) can be carried out based on its results. Determine which subtask outputs are inputs to other subtasks. For example, "analyze the reason for disk space occupation" (subtask 2) requires the disk usage data provided by "monitor disk usage" (subtask 1). Consider whether there is competition or dependence on the same resource between subtasks. For example, if multiple subtasks need to access the same server or the same set of storage devices, their execution order needs to be reasonably arranged to avoid resource conflicts.

[0079] Set corresponding evaluation metrics according to the dependencies of subtasks. In the "Monitor Disk Usage" (subtask 1), evaluate whether the monitoring data can timely reflect the real-time usage status of the server disk and whether it can issue an alarm in time when the disk space is approaching exhaustion. In the "Analyze the Reasons for Disk Space Occupancy" (subtask 2), evaluate whether the analysis results accurately identify the key factors causing insufficient disk space. In the "Clean and Optimize Disk Space" (subtask 3), evaluate whether the cleaning operation reasonably reduces the disk space occupancy without causing unnecessary resource consumption or performance impact on the normal operation of the server. In the "Optimize Storage Strategy" (subtask 4), evaluate whether the new storage strategy can effectively prevent the problem of insufficient disk space from occurring again in the long term.

[0080] Execute the monitoring module, which is used to call and splice tools based on principles with the goal of completing subtasks, monitor the completion status of each subtask. If a subtask is successfully completed, summarize and remember the collaboration between the principle and the tool. If a subtask fails to complete, optimize the collaboration between the tool and the principle until the target operation and maintenance task is completed, and summarize and remember each collaboration.

[0081] In this embodiment, by summarizing and remembering the collaboration, it is avoided to repeat a large amount of effort in researching and testing different combinations of principles and tools for the same or similar tasks. The AI Agent can directly select a proven and suitable solution based on the existing memory, and invest more time and energy in other important and innovative operation and maintenance work.

[0082] Call and splice tools based on principles with the goal of completing subtasks, specifically including:

[0083] Determine the dependencies between subtasks, determine the tools required for the subtasks with dependencies, and determine the splicing order of the tools based on principles.

[0084] Determine whether the subtasks without dependencies need to call the same tool, determine whether the same tool called can be called simultaneously. If it cannot be called simultaneously, determine the splicing order of the tools based on principles.

[0085] Merge the subtasks with dependencies into a group, determine whether the tools required by the group and the remaining subtasks can be called simultaneously. If it cannot be called simultaneously, determine the splicing order of the tools based on principles.

[0086] In this embodiment, the execution processes and data dependencies of each subtask are analyzed to determine the sequence among them. For example, "monitoring disk usage" (subtask 1) is the basis of the entire task and must be completed first to provide data support for the subsequent "analyzing the reasons for disk space occupation" (subtask 2). According to the dependency relationships of the subtasks, the required tools to be called and their splicing order are determined. For example, in subtask 1, the system command df - h is first used to initially view the disk usage, and then the monitoring tool Zabbix is used for real - time monitoring and data collection. For subtasks without dependency relationships, it is judged whether the same tool needs to be called. For example, "cleaning and optimizing disk space" (subtask 3) and "optimizing storage policies" (subtask 4) have no direct logical dependency, but both may need to use the storage management tool LVM. It is determined whether the same tool called can be called simultaneously. If it can, such as different functional modules of LVM can operate on different storage partitions simultaneously, then they can be executed in parallel; if they cannot be called simultaneously, such as some tools can only perform exclusive operations on specific resources at the same time, then the splicing order of the tools needs to be determined according to the principle, for example, performing storage optimization operations in sequence according to the importance and urgency of the data. A reasonable splicing order of tool calls can ensure that the most appropriate principles and tools are applied when processing each subtask, improving the pertinence and effectiveness of problem - solving.

[0087] Subtasks with dependency relationships are combined into a collective. For example, subtask 1 and subtask 2 are combined because subtask 2 depends on the output of subtask 1. It is determined whether the tools required to be called by the collective and the remaining subtasks can be called simultaneously. If they can, they can be executed in parallel, such as the monitoring tool and the data analysis tool can run on different system resources simultaneously; if they cannot be called simultaneously, then the splicing order of the tools needs to be determined through the principle, such as some tools may occupy the same system resources resulting in conflicts. Considering subtasks with dependency relationships as a collective can grasp the logical structure and execution process of the operation and maintenance tasks as a whole, making the solution more systematic and comprehensive, and avoiding missing key steps or ignoring potential influencing factors.

[0088] During the operation and maintenance process, new situations or requirement changes may occur. By flexibly adjusting the subtasks and the tool call order, these changes can be quickly adapted to, and corresponding measures can be taken in a timely manner to ensure the stable operation of the system. For example, if new space occupation problems or business requirement changes are found during the disk space optimization process, the optimization strategy and operation steps can be quickly adjusted according to the established principles and tool call logic.

[0089] Summarize and memorize the collaboration between the principle and the tool, specifically including:

[0090] Represent the collaboration relationship between the tool and the principle as a graph structure, where the nodes represent the tool and the principle, and the edges represent the collaboration relationship between the tool and the principle.

[0091] In this embodiment, taking the operation and maintenance task of "server disk space shortage" as an example, tools (such as df - h, Zabbix, tmpreaper, LVM, etc.) and principles (such as system monitoring principle, data management principle, storage optimization principle, etc.) are used as nodes, and their collaboration relationships (such as using the df - h command to view disk usage based on the system monitoring principle) are used as edges to construct a graph structure. Attribute information such as the type of tools, the field of principles, and the frequency of collaboration is added to the nodes and edges to more comprehensively describe their characteristics and relationships.

[0092] The feature vectors of the nodes and edges in the graph structure are learned through a graph neural network, and the extracted feature vectors are used to classify the collaboration between tools and principles through a classification algorithm.

[0093] In this embodiment, a suitable graph neural network model, such as a graph convolutional network (GCN) or a graph attention network (GAT), is selected to construct a network architecture. The constructed graph structure is input into the graph neural network, and through the process of forward propagation and backward propagation, the feature vectors of the nodes and edges are learned to capture the key features and patterns of the collaboration between tools and principles.

[0094] The collaboration classification data is summarized and memorized through a deep classification learning model to obtain a collaboration strategy model.

[0095] In this embodiment, the extracted feature vectors are input into a classification algorithm, such as a support vector machine (SVM), a random forest (RF), or a fully connected neural network in deep learning, etc., to classify the collaboration between tools and principles. For example, the collaboration is classified into two categories: success and failure, or multi - category classification is performed according to the effect of the collaboration. A deep classification learning model, such as a multi - layer perceptron (MLP) or a convolutional neural network (CNN), is constructed, and the classification results are used as supervision information to train the model. By continuously adjusting the parameters of the model, it can accurately classify and predict new collaboration cases, and finally obtain a collaboration strategy model. When facing complex operation and maintenance problems, the collaboration strategy model can comprehensively consider the collaboration of multiple principles and tools, provide a comprehensive and systematic solution, assist the AI Agent to consider the whole, avoid missing key factors, and improve the problem - solving ability.

[0096] Optimize the collaboration between tools and principles, specifically including:

[0097] Simultaneously consider the status of the task when all subtasks fail to complete, the dependency relationships between tasks, and the overall status of the operation and maintenance system. Construct the collaboration dependency relationships between tools and principles among the subtasks that fail to complete as a directed acyclic graph, and transform the dependency relationships, status, and the overall status of the operation and maintenance system of the subtasks during task execution into weight vectors and embed them into the directed acyclic graph.

[0098] In this embodiment, the collaborative dependency relationships of tools and principles between the sub-tasks that have failed (such as "analyze the reasons for disk space occupancy" and "clean and optimize disk space") are represented as a directed acyclic graph (DAG). For example, if "analyze the reasons for disk space occupancy" needs to be executed prior to "clean and optimize disk space", a directed edge from "analyze" to "clean" is added in the DAG. The dependency relationships, status (such as failure reasons, data volume, etc.) of the sub-tasks during execution and the overall status of the operation and maintenance system (such as the current disk usage rate, system load, etc.) are transformed into weight vectors and embedded into the DAG. For example, weights are assigned to the dependency edges to represent their tightness; weights are assigned to the nodes (sub-tasks) to represent their execution priorities or difficulties.

[0099] Calculate the attention coefficients between each task and its neighboring tasks among the remaining tasks in the directed acyclic graph, and learn the high-dimensional state features of each task in the remaining tasks through a graph neural network.

[0100] In this embodiment, for each task of the remaining tasks in the DAG, the attention coefficients between it and its neighboring tasks are calculated by measuring the correlation or dependency strength between tasks. For example, the cosine similarity is used to calculate the similarity between task feature vectors as the attention coefficient. The DAG is encoded using a graph neural network (GNN) to learn the high-dimensional state features of each task. The GNN can aggregate the features of the task itself and its neighboring tasks to capture the complex relationships between tasks. The graph convolutional network (GCN) fuses the features of each task with the features of its neighboring tasks through multiple convolutional operations to generate richer high-dimensional representations.

[0101] Use the high-dimensional state features as reference information for the task scheduling algorithm to regenerate the scheduling scheme for the collaboration of tools and principles.

[0102] In this embodiment, the learned high-dimensional state features are used as reference information for the task scheduling algorithm. Based on the scheduling algorithm of reinforcement learning, the high-dimensional state features are used as state inputs, and the task scheduling actions are output through a policy network to select the next task to be executed and its corresponding combination of tools and principles. According to the new scheduling scheme, the collaboration order of tools and principles is rearranged to optimize the task execution process. If it is found that the "clean and optimize disk space" task can be executed earlier in some cases to reduce the impact on subsequent tasks, its priority in the scheduling is adjusted.

[0103] A feedback optimization module for reasoning about the content of the summary memory and generalizing the collaboration of principles and tools into solutions based on the target operation and maintenance tasks.

[0104] Reasoning about the content of the summary memory specifically includes:

[0105] Collect the collaborative data of tools and principles in the operation and maintenance tasks, including successful cases and failure cases. Consider the constructed graph structure as a successful case and the constructed directed acyclic graph as a failure case.

[0106] In this embodiment, represent the collaborative relationship between tools and principles in the successful case as a graph structure, where nodes represent tools and principles, and edges represent their collaborative relationship. Represent the collaborative relationship between tools and principles in the failure case as a directed acyclic graph (DAG), where nodes represent tools and principles, and edges represent their collaborative relationship, and there are no cycles in the graph.

[0107] Learn the first inference feature vectors of nodes and edges in the graph structure and the directed acyclic graph through a graph neural network.

[0108] In this embodiment, use a graph neural network to encode the graph structure and the directed acyclic graph, learn the feature vectors of nodes and edges to capture the complex relationships between nodes and edges, extract high-dimensional feature representations, and obtain the first inference feature vectors. The first inference feature vectors, as preliminary feature representations, provide a basis for subsequent inference tasks. When learning the high-dimensional feature representations of nodes and edges, GNN can retain the structural information of the graph and more accurately capture the complex relationships in the graph structure, making the inference process of the model more interpretable and thus better generalizing to unseen graph data.

[0109] Input the graph structure and the directed acyclic graph into the collaborative policy model to obtain the second inference feature vectors.

[0110] In this embodiment, input the graph structure and the directed acyclic graph into a pre-trained collaborative policy model, and output the second inference feature vectors of each node and edge. The second inference feature vectors further integrate the global information of the graph, the relationships between nodes at greater distances, and high-level semantic information, and more deeply characterize the roles and functions of nodes and edges in the graph.

[0111] Map the first inference feature vectors and the second inference feature vectors to a vector space through knowledge graph embedding technology, and predict new principle and tool collaboration relationships by calculating the similarity between vectors.

[0112] In this embodiment, use knowledge graph embedding technology to map the first inference feature vectors and the second inference feature vectors to the same vector space, so that feature vectors from different sources can be compared and calculated in the same space. In the vector space, predict new principle and tool collaboration relationships by calculating the similarity between vectors (such as cosine similarity, Euclidean distance, etc.). Vectors with high similarity mean that their collaboration relationships may be similar, discover new entity combinations or relationships similar to existing patterns, and thus predict new principle and tool collaboration relationships.

[0113] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intelligent operation and maintenance system based on an AI Agent, characterized in that, The system includes: A content collection module, which is used to monitor the operation data of the operation and maintenance system, obtain the abnormal information in the operation data and the operation and maintenance problems submitted by users to determine the target operation and maintenance tasks; A decision-making and planning module, which is used to decompose the target operation and maintenance tasks into several subtasks, determine the closest solution based on each subtask, and determine the required principles and tools according to the solution; An execution monitoring module, which is used to call and splice tools based on the principle with the goal of completing subtasks, monitor the completion status of each subtask. If a subtask is successfully completed, summarize and remember the collaboration between the principle and the tool. If a subtask fails to complete, optimize the collaboration between the tool and the principle until the target operation and maintenance tasks are completed, and summarize and remember each collaboration; A feedback and optimization module, which is used to reason about the summarized and remembered content, and summarize the collaboration between the principle and the tool as a solution based on the target operation and maintenance tasks.

2. The intelligent operation and maintenance system based on the AI Agent according to claim 1, wherein Determining the target operation and maintenance tasks specifically includes: Preprocess the obtained abnormal information, extract features from the preprocessed abnormal information through machine learning to obtain abnormal features, pre-train an abnormal detection model, and determine the type and severity of the abnormality through the abnormal detection model using the abnormal features; Perform natural language processing on the operation and maintenance problems submitted by users to obtain semantic features, pre-train a problem recognition model, and determine the user's intention through the problem recognition model using the semantic features; Fuse the abnormal features and semantic features through a multi-modal fusion algorithm to obtain fusion features, pre-train a task generation model, and determine the target operation and maintenance tasks by inputting the fusion features into the task generation model.

3. The intelligent operation and maintenance system based on the AI Agent according to claim 1, characterized in that, Decompose the target operation and maintenance tasks into several subtasks through an AI Agent, specifically including: According to the requirements of the target operation and maintenance tasks, decompose the target operation and maintenance tasks into several subtasks through a task decomposition algorithm; Determine the dependency relationships between the subtasks, and set corresponding evaluation indicators according to the dependency relationships between the subtasks.

4. The intelligent operation and maintenance system based on the AI Agent according to claim 3, wherein, Call and splice tools based on the principle with the goal of completing subtasks, specifically including: Determine the dependency relationships between subtasks, determine the tools required for the subtasks with dependency relationships, and determine the splicing order of the tools through the principle; Determine whether the subtasks without dependency relationships need to call the same tool, determine whether the same tool called can be called simultaneously. If it cannot be called simultaneously, determine the splicing order of the tools through the principle; Merge the subtasks with dependency relationships into a group, determine whether the tools required by the group and the other subtasks can be called simultaneously. If it cannot be called simultaneously, determine the splicing order of the tools through the principle.

5. The intelligent operation and maintenance system based on the AI Agent according to claim 1, wherein Summarize and remember the collaboration between the principle and the tool, specifically including: Represent the collaboration relationship between the tool and the principle as a graph structure, where the nodes represent the tool and the principle, and the edges represent the collaboration relationship between the tool and the principle; Learn the feature vectors of the nodes and edges in the graph structure through a graph neural network, and classify the collaboration between the tool and the principle by using the extracted feature vectors through a classification algorithm; Summarize and remember the collaboration classification data through a deep classification learning model to obtain a collaboration strategy model.

6. The intelligent operation and maintenance system based on the AI Agent according to claim 5, wherein Optimize the collaboration between the tool and the principle, specifically including: Consider the status of tasks when all subtasks fail to complete, the dependencies between tasks, and the overall status of the operation and maintenance system simultaneously; Construct the collaborative dependencies of tools and principles between subtasks that have failed to complete as a directed acyclic graph, and transform the dependencies, status, and overall status of the operation and maintenance system of subtasks during task execution into weight vectors and embed them into the directed acyclic graph respectively; Calculate the attention coefficients between each task and its neighboring tasks among the remaining tasks in the directed acyclic graph, and learn the high-dimensional status features of each task in the remaining tasks through a graph neural network; Use the high-dimensional status features as reference information for the task scheduling algorithm to regenerate the scheduling scheme for the collaboration of tools and principles.

7. The intelligent operation and maintenance system based on the AI Agent according to claim 6, wherein Infer the content of the summary memory, specifically including: Collect the collaborative data of tools and principles in the operation and maintenance tasks, including successful cases and failure cases. Regard the constructed graph structure as a successful case and the constructed directed acyclic graph as a failure case; Learn the first inference feature vectors of nodes and edges in the graph structure and the directed acyclic graph through a graph neural network; Input the graph structure and the directed acyclic graph into the collaboration strategy model to obtain the second inference feature vector; Map the first inference feature vector and the second inference feature vector to the vector space through knowledge graph embedding technology, and predict new principle and tool collaboration relationships by calculating the similarity between vectors.

Citation Information

Patent Citations

  • Intelligent operation and maintenance system based on artificial intelligence

    CN115858226A

  • Petrochemical production process anomaly diagnosis and optimization method and system integrated with knowledge graph

    CN119668245A

  • Document generation method and device based on artificial intelligence

    CN119719354A

  • Platform-enabled orchestration and optimization of digital workflows

    WO2025064639A1

Cited By

  • Intelligent operation and maintenance method and system based on MCP protocol

    CN120614263A

  • Large model system efficient operation maintenance method and device

    CN120929337A

  • Automatic operation and maintenance method and system suitable for closed system

    CN121563467A

  • Belt conveyor multi-mode operation and maintenance fusion system and method based on three-tower type cognitive model

    CN121959466A

  • RPA, AI and LLM-based server product operation and maintenance method, device and equipment for realizing Agent, and storage medium

    CN121967152A