Intelligent inspection assistant system based on AI Agent multi-agent

Through the intelligent inspection assistant system based on AI Agent multi-agent, inspection tasks and report generation are automatically completed, solving the problems of manual time and relying on manual analysis in the existing technology, and an efficient and automated inspection process is achieved.

CN119988149AInactive Publication Date: 2025-05-13SHENZHEN SUNLINE TECH CO LTD

Patent Information

Application Number
CN202510473397.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

During the inspection process, the existing technology has problems such as manual time-consuming, unified time-based inspections are not suitable for the differences between different subsystem nodes, time-consuming reporting and sorting, and relying on manual analysis and repair suggestions.

Method used

The intelligent inspection assistant system based on AI Agent multi-agent is adopted to realize self-reflection and learning mechanisms, automatically complete inspection tasks, and generate reports and repair suggestions to reduce manual intervention.

Benefits of technology

Greatly save labor costs, improve inspection work efficiency, reduce inspection time, and realize automated inspection task creation, execution, report generation and repair suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988149A_ABST
    Figure CN119988149A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent inspection assistant system based on AI Agent multi-agents, and the system comprises a main AI Agent module which is used for receiving an inspection instruction and carrying out the intention recognition based on the inspection instruction; the guide AI Agent module is used for performing task allocation based on the intention recognition result of the main AI Agent module; the inspection task generation AI Agent module is used for generating an inspection task; the inspection task execution AI Agent module is used for executing an inspection task; the interior of the inspection task generation AI Agent module and the interior of the inspection task execution AI Agent module adopt workflow to define main execution processes and logics; and the RAG knowledge base is used for storing the inspection execution report and the inspection problem investigation report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent inspection assistant system based on AI Agent multi-agents. Background Art

[0002] Inspection is a kind of operation and maintenance method, which is usually used to regularly check and maintain IT infrastructure such as servers and network equipment. Generally speaking, inspection focuses on the business continuity of the production system, ensuring that the status of each component of the entire system is normal and healthy at key time nodes (such as marketing activities, day changes, etc.).

[0003] Inspection is mainly carried out in the form of inspection tasks. In the past, inspection work was mainly carried out manually. The working mode of manual inspection is to organize an inspection team to first determine the inspection content items of the inspection: such as computing resource utilization, disk file occupancy, file handles, JVM gc times and time consumption, etc. Then, each machine node is executed through a pre-set Shell command to check the health of the inspected content. And finally a report is formed for reporting. However, manual inspection has the following defects: (1) A single inspection takes several hours, and generally occurs multiple times a day; (2) The inspection time points of different subsystem nodes may be different, and the current manual practice is to conduct inspections at a unified time; (3) The inspection report compilation work is relatively time-consuming; (4) If some inspection items are not met during the inspection, it is necessary to manually record the unmet inspection items first, and analyze the problem content before giving repair suggestions. This decision-making time is relatively long.

[0004] Although there are currently automated inspection tools, such as inspection scripting tools and RPA robot process automation inspection, they all have shortcomings. Inspection scripting tools are used to compile a list of all machines to be inspected and all the inspection instructions that need to be executed on each machine. The system inspection task is completed by calling each functional script with a general script, and the items that do not meet the inspection requirements are output to a specific file. After the inspection is completed, the items that do not meet the inspection requirements are extracted, and the root cause and repair plan are obtained through manual analysis. The repair plan is then run by a dedicated operation and maintenance personnel. Although inspection scripting tools can use a "one-click script" method to launch inspection tasks for each component in the system, they have several obvious disadvantages: (1) Writing scripts is time-consuming and labor-intensive; (2) The scripts need to be deployed to each machine one by one to run. For a large system, the number of machine nodes is very large. If containerized deployment is used, the number of pods will be even greater. (3) After each script is run, there will be a corresponding running record. The administrator needs to combine these running records to compile a report and identify the inspection items that failed the inspection. (4) For inspection items that failed during the inspection, you need to manually check the failed items and give repair suggestions based on your own experience and repair them. If there are many failed items, you need to give repair suggestions one by one.

[0005] The inspection of RPA robot process automation is mainly combined with RPA technology. RPA technology can automate the execution of inspection scripts so that they can be executed on specific nodes. Finally, by collecting the execution results of each RPA robot, the items that do not meet the inspection requirements are summarized and obtained. Finally, the root cause and repair plan are obtained through manual analysis, and the repair plan is run by dedicated operation and maintenance personnel. The inspection of RPA robot process automation is actually a supplement to the inspection script tool. It can save the work related to script writing and script deployment, but it relies heavily on the professional experience of relevant personnel in terms of collating reports based on operation records, identifying inspection items that failed inspections, and repair suggestions. Summary of the invention

[0006] In order to solve the above problems, the purpose of the present invention is to provide an intelligent inspection assistant system based on AI Agent multi-agent, which can have self-reflection and learning mechanisms, complete all corresponding tasks independently, and no human intervention is required throughout the process, thereby greatly saving labor costs. Humans can be freed from cumbersome configurations and can issue inspection-related instructions in human language.

[0007] The present invention provides an intelligent inspection assistant system based on AI Agent multi-agents, the system comprising: The main AI Agent module is used to receive inspection instructions and perform intent recognition based on the inspection instructions; A guiding AI Agent module, connected to the main AI Agent module, for performing task allocation based on the intention recognition result of the main AI Agent module; Inspection task generation AI Agent module, used to generate inspection tasks; Inspection task execution AI Agent module, used to execute inspection tasks; the inspection task generation AI Agent module and the inspection task execution AI Agent module use workflow to define the main execution process and logic; RAG knowledge base is used to store inspection execution reports and inspection problem troubleshooting reports.

[0008] Optionally, when receiving the call instruction of the guiding AI Agent module, the inspection task generating AI Agent module extracts the inspection information based on the inspection instruction and assembles the inspection task page; Call the inspection task configuration tool to create an inspection task.

[0009] Optionally, the inspection task includes an inspection scenario, multiple inspection rules, and an inspection task execution plan; The inspection scenario is a description of the inspection task and applicable scenario, and also summarizes the scope and boundaries of the inspection; Inspection rules are specific execution rules determined based on inspection scenarios, including business capacity rules, configuration funnel rules, and service governance rules; The inspection task execution plan is an execution plan for the inspection task configured in combination with the inspection rules, including rule-triggered execution, scheduled task-triggered execution, or manual triggering execution.

[0010] Optionally, the inspection rule includes one or more inspection plug-ins, and the inspection plug-in obtains a customized inspection plug-in template from the inspection plug-in library and configures it; Each inspection plug-in includes one or more checkpoints. Each checkpoint represents a content to be checked. Each checkpoint contains a checkpoint assertion and one or more execution scripts. The checkpoint assertion is an assertion expression with only two values: True and False. The execution script is a Shell script that supports checkpoint assertions.

[0011] Optionally, the inspection task generates an AI Agent module for intelligently predicting inspection points and rules based on specific inspection scenarios, and intelligently generating corresponding inspection points and inspection rules based on the description of activity events and environment switching events.

[0012] Optionally, the system further comprises: A repair multi-solution recommendation module is used to search for relevant problem solving reports in the RAG knowledge base when unhealthy inspection items are detected during inspection, so as to provide at least one repair solution, as well as repair steps and dependency analysis descriptions corresponding to each repair solution; The multi-repair scheme recommendation module is provided with a repair effect prediction model, which is used to simulate the repair scenario of each of the repair schemes to evaluate the repair results corresponding to different repair schemes.

[0013] Optionally, the inspection execution AI Agent module has different adaptation anti-corrosion layers to adapt to different inspection objects; The inspection objects include traditional applications, containerized applications, database nodes, and middleware. The middleware includes traditional deployment, containerized CRD deployment, and StatefulSet deployment.

[0014] Optionally, the main AI Agent module is also used for a dynamic collaboration mechanism of intelligent agents based on task complexity to achieve intelligent decomposition and collaborative execution of inspection tasks, including: Task complexity assessment, intelligent task decomposition, dependency analysis, complex directed acyclic graph construction, and collaborative execution strategy.

[0015] Optionally, the task complexity analysis includes performing a complexity assessment on the received inspection tasks, and the assessment dimensions include: the number of inspection rules, the number of inspection plug-ins, the total number of checkpoints, the number of inspected nodes, the estimated time consumption of executing the script, and the complexity of the dependency relationship between nodes; The intelligent task decomposition includes: Rule-level decomposition, treating different inspection rules as independent task units; Plug-in level decomposition, further decomposing the plug-ins under each rule; Node-level decomposition, grouped by server group or node type; Parallelism evaluation, evaluates the maximum parallel execution degree based on the system load capacity.

[0016] Optionally, the dependency relationship includes: inspection rule dependency relationship, checkpoint dependency relationship and node dependency relationship; When a complex directed acyclic graph is constructed, a DAG graph is constructed based on the dependency relationship.

[0017] The intelligent inspection assistant system based on AI Agent multi-agents of the present invention allows inspectors to express inspection requirements in human language through the collaboration of multiple AI Agents. The system automatically collaborates with the large model to complete a series of tasks such as inspection task creation, inspection execution, inspection report generation, and inspection repair suggestions, thereby greatly improving the overall efficiency of the inspection work.

[0018] Based on the following detailed description of specific embodiments of the present invention in conjunction with the accompanying drawings, those skilled in the art will become more aware of the above and other objects, advantages and features of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings: Figure 1 It is a schematic diagram of the overall architecture of the intelligent inspection assistant system based on AI Agent multi-agents according to an embodiment of the present invention; Figure 2 is a conceptual schematic diagram of inspection according to an embodiment of the present invention; Figure 3 is a schematic diagram of a role-playing conceptual model of inspection according to an embodiment of the present invention; Figure 4 is a schematic diagram of inspection rule dependency in an embodiment of the present invention; Figure 5 is a schematic diagram of checkpoint dependency relationships in an embodiment of the present invention; Figure 6 is a schematic diagram of node dependency relationships in an embodiment of the present invention; Figure 7 It is a schematic diagram of a complex DAG diagram formed by three dependency relationships in an embodiment of the present invention; Figure 8 It is a schematic diagram of DAG construction in a promotion activity inspection example of an embodiment of the present invention; Fig. 9 The inspection configuration administrator of the embodiment of the present invention generates the inspection task diagram using human language; Fig.10 The inspection task execution AI Agent of the embodiment of the present invention automatically performs inspection according to the timer and generates an execution result report and a repair suggestion diagram; Fig.11 The inspection configuration administrator of the embodiment of the present invention uses human language to immediately initiate inspection execution and generate an execution result report and a repair suggestion flow chart; Fig.12 It is a blueprint diagram of the functional modules of the intelligent inspection assistant system based on AI Agent multi-agents in an embodiment of the present invention. DETAILED DESCRIPTION

[0020] The embodiments of the present invention are described below in conjunction with the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the present invention and are not limiting.

[0021] The embodiment of the present invention provides an intelligent inspection assistant system based on AI Agent multi-agents. The intelligent inspection assistant system based on AI Agent multi-agents in this embodiment includes a main AI Agent module, which is used to receive inspection instructions and perform intent recognition based on the inspection instructions; a guiding AI Agent module, which is connected to the main AI Agent module and is used to perform task allocation based on the intent recognition result of the main AI Agent module; an inspection task generation AI Agent module, which is used to generate inspection tasks; an inspection task execution AI Agent module, which is used to execute inspection tasks; the inspection task generation AI Agent module and the inspection task execution AI Agent module use workflows to define the main execution process and logic; and a RAG knowledge base, which is used to store inspection execution reports and inspection problem troubleshooting reports. In addition, the system of this embodiment can also be set to generate a large model, and the large model is coupled with each module in the system, such as the main AI Agent module, the guiding AI Agent module, the inspection task generation AI Agent module, the inspection task execution AI Agent module, and the RAG knowledge base, to assist each model in working.

[0022] In practical applications, deepseek-r1 or other models can be used to generate large models. The generated large models have extremely strong reasoning capabilities and can perform reasoning-related tasks such as graph recognition, report summary generation, and intelligent repair suggestions based on inspection results.

[0023] The AI ​​Agent multi-agent intelligent inspection assistant system of this embodiment realizes the intelligent decomposition and collaborative execution of inspection tasks based on the dynamic cooperation mechanism of agents according to task complexity; the secure communication protocol between agents ensures the confidentiality and integrity of inspection instructions and results; the working memory mechanism of agents enables them to maintain context consistency in long-term inspection tasks; the self-evaluation and optimization mechanism of agents automatically adjusts the inspection strategy according to the historical execution results. The following is a detailed description of the different modules.

[0024] 1. Main AI Agent Module The main AI Agent module is used to receive inspection instructions and perform intent recognition based on the inspection instructions. The working process and details of the main AI Agent module are as follows: 1. Access layer design 1. Multi-channel command reception: supports multiple access methods such as natural language input, structured API calls, chat interface, command line, etc.; 2. Instruction preprocessing: Perform preliminary cleaning and standardization on inspection instructions to remove redundant information and special characters.

[0025] 2. Large Models Prompt Engineering Architecture 1. Prompt template system: Design a multi-level prompt template library, including basic intent recognition templates, domain-specific inspection templates, and complex scenario processing templates; Dynamic template assembly mechanism automatically selects and combines appropriate prompt templates according to input instruction characteristics.

[0026] 2. Context Management: Working memory mechanisms, which preserve recent interaction history to maintain contextual coherence; Sliding window memory compression algorithm automatically condenses non-critical context information and reduces token consumption; Key information marking and retention mechanisms ensure that important instructions are not forgotten.

[0027] 3. Tip injection technology: Role definition injection: "You are a professional inspection system expert, responsible for understanding the inspection needs of users"; Thinking chain injection: "Please think in steps: 1. What type of inspection instruction is this? 2. What is the specific goal of the instruction?"; Example injection: Provides sample pairs of typical inspection instructions and corresponding intents.

[0028] 3. Intent Recognition Process 1. Preliminary intent classification: The instructions are initially classified into the following main categories: creating inspection tasks, executing inspection tasks, querying inspection results, modifying inspection configuration, etc. A multi-round reasoning strategy is adopted, first coarse classification and then fine classification, to improve recognition accuracy.

[0029] 2. Intent parameter extraction: Through structured output design, the large model is guided to extract key parameters such as inspection objects, inspection scope, threshold settings, and time requirements; The JSON response format is used to ensure that the extraction results can be directly processed by the system.

[0030] 3. Dealing with ambiguity and clarification: Identify ambiguous or contradictory statements in instructions and generate clarifying questions; A multiple-option suggestion mechanism provides choices when there are multiple possible interpretations.

[0031] 4. Intent confirmation and error correction: The recognition results are self-verified, and the rationality of the extracted intention is checked through secondary reasoning; Low confidence intent tagging, proactively requesting confirmation when recognition is uncertain.

[0032] 2. Guide AI Agent Module The guiding AI Agent module is connected to the main AI Agent module to perform task allocation based on the intention recognition result of the main AI Agent module.

[0033] There are two task AI Agents. The first is the AI ​​Agent for patrol task generation, which is mainly responsible for patrol task generation; the other is the AI ​​Agent for patrol task execution, which is mainly responsible for patrol task execution. Both patrol task AI Agents use workflows to define the main execution process and logic. AI Agents can call tools to complete the corresponding work.

[0034] 3. Inspection Task Generation AI Agent Module The inspection task generation AI Agent module is used to generate inspection tasks. Specifically, when the inspection task generation AI Agent module receives the call instruction of the guiding AI Agent module, it extracts the inspection information based on the inspection instruction, assembles the inspection task page, and calls the inspection task configuration tool to create the inspection task.

[0035] The working process of the inspection task generation AI Agent module in the embodiment of the present invention is as follows: 1. Receive the call instruction of the guide AI Agent module: When the inspection configuration administrator sends the instruction to create an inspection task to the main AI Agent, the main AI Agent module recognizes the intention and determines it as "create an inspection task", and then passes it to the guide AI Agent module. The guide AI Agent module calls the inspection task generation AI Agent through the `@InspectionTaskGen-Agent` method.

[0036] 2. Extract inspection information: The inspection task generates an AI Agent module to start the workflow and use LLM to extract necessary information from the context, including: Inspection rules (such as disk space, CPU utilization, memory utilization, number of file handles, etc.); Checkpoint assertion configuration in the inspection plug-in of each rule (such as space not exceeding 10G, CPU not exceeding 70%, etc.); Execution plan (e.g., once at midnight and once at 12 o'clock every day); List of inspected servers (application servers, database servers, gateway servers, etc.).

[0037] 3. Assemble the inspection task page: After extracting the necessary information, the AI ​​Agent assembles the information into the inspection task page and generates the inspection task details. This page can be used for the inspection configuration administrator to confirm or directly execute.

[0038] 4. Call the inspection task configuration tool: The workflow calls the inspection task configuration tool (Tool Call), which completes the following operations: Call the data access object (DAO) layer to add inspection task records and corresponding details in the database; Through the pre-configured WebHook, relevant personnel are notified by email or SMS; the newly created inspection tasks are displayed on the inspection dashboard.

[0039] Through this process, the system realizes the automatic conversion from human language instructions to structured inspection tasks, eliminating the need for manual writing of inspection scripts, greatly improving the efficiency of inspection task creation.

[0040] Inspection tasks are issued in natural language, which may be overly colloquial or unclear. The inspection task generation AI Agent module can introduce a matching and clarification mechanism for fuzzy natural language requirements through the intention recognition of the large model, thereby improving the accuracy of inspection task creation.

[0041] The inspection task generation AI Agent module has the following functions: (1) The adaptive inspection rule generation algorithm for the microservice architecture can automatically generate the optimal inspection path based on the service topology. When generating inspection tasks, if it is for the nodes of the microservice architecture, there is actually a topological relationship between different microservice nodes. There is a corresponding adaptive inspection rule generation algorithm that can determine the inspection priority based on the service topology and automatically generate the optimal inspection path, so as to carry out inspections efficiently.

[0042] (2) Intelligent inspection priority determination mechanism based on service dependency graph; (3) Fuzzy matching and clarification mechanism of natural language requirements to improve the accuracy of inspection task creation; (4) Automatic iteration and optimization system of task templates based on historical inspection effect feedback.

[0043] The results of each inspection will be stored in the data and uploaded to the RAG knowledge base. The inspection effect can be gradually improved as the agent learns and improves itself. It can automatically iterate and optimize the system based on the task templates of historical inspection effect feedback.

[0044] 4. Inspection Task Execution AI Agent Module The inspection task execution AI Agent module is used to execute inspection tasks; the inspection task generation AI Agent module and the inspection task execution AI Agent module use workflow to define the main execution process and logic. The inspection task execution AI Agent module has a distributed inspection load balancing mechanism to avoid performance impact on the production system; adaptive retry and fault recovery mechanism to ensure that the inspection task can still be completed under abnormal conditions such as network fluctuations. Of course, in actual applications, external workflows (not the workflows within the AI ​​Agent intelligent body) can also be used to semi-automate the entire process.

[0045] The inspection system is the production system, and during the inspection process, the inspection execution AI Agent will generate high instantaneous traffic to each node of the production system. This traffic is mainly used to distribute inspection tasks and transmit inspection task execution scripts, which may cause performance impact on the production system. Therefore, a distributed inspection load balancing mechanism is introduced to load balance the same type of inspection work (for example, all server computing resource inspections) to avoid performance impact on the production system.

[0046] The distributed inspection load balancing mechanism is designed to solve the performance impact that may be caused to the production system during the inspection process. During the inspection execution process, the inspection execution AI Agent will simultaneously distribute inspection tasks and execute scripts to multiple nodes of the production system, generating high instantaneous traffic, which may affect the normal operation of the production system. The execution process is as follows: 1. Inspection task classification: The system first classifies inspection tasks by type (such as computing resource inspection, disk space inspection, JVM inspection, etc.); 2. Load evaluation: Evaluate the resource consumption of each type of inspection task; 3. Time staggering: stagger the execution of similar inspection tasks at different times; 4. Batch grouping: Divide a large number of nodes into multiple batches to avoid simultaneous inspection; 5. Priority sorting: sorting according to business importance and system load; 6. Real-time monitoring and adjustment: dynamically adjust the inspection rate according to the system load; Execution example: Assume that there is a large e-commerce system consisting of 100 application servers, 20 database servers, and 10 gateway servers, which requires a comprehensive inspection.

[0047] 1. Inspection task breakdown: Computing resource inspection: Check CPU usage (<70%) and memory usage (<60%); Disk space inspection: check the root directory space (<10G) and database file space (<50G); File handle inspection: Check the number of file handles (<60000).

[0048] 2. Batch execution: First batch: First inspect 25 application servers, 5 database servers, and 3 gateway servers; After an interval of 5 minutes, inspect the second batch... and so on.

[0049] 3. Type peak shifting: For the first batch of servers, perform computing resource inspection first; After an interval of 2 minutes, perform a disk space inspection; After another 2 minutes, perform a file handle inspection.

[0050] 4. Dynamic adjustment: If the monitoring finds that the load of a server has reached 85%, the inspection of the server will be automatically postponed; It is found that the system load is low at night, so the number of concurrent inspections is automatically increased.

[0051] 5. Exception handling: If a node inspection fails, the system will first complete the inspection of other nodes and then retry the failed node; In case of network fluctuations, an adaptive retry mechanism is used to ensure inspection completion; Through this distributed load balancing mechanism, the system can efficiently complete all inspection tasks without affecting the normal operation of the production environment.

[0052] The inspection process involves network communication. If a node being inspected is healthy, but the network between the AI ​​Agent executing the inspection and the inspected node fluctuates abnormally, it may cause accidental damage due to failure to obtain the correct inspection results. Therefore, an adaptive retry mechanism and a fault recovery mechanism are introduced to ensure that the inspection task can be completed smoothly even in the case of network fluctuations.

[0053] During the execution of the adaptive retry mechanism and the fault recovery mechanism, the core principle is: the adaptive retry mechanism dynamically adjusts the retry strategy according to the network status and error type, and the fault recovery mechanism ensures that the inspection task can be correctly restored after a network failure.

[0054] Execution process example: 1. Initial connection detection The inspection execution AI Agent sends an inspection task to the target node 192.168.0.1; network fluctuations are detected and the connection times out.

[0055] 2. Failure classification and analysis System record failure type: connection timeout; Analyzing historical data, it was found that the node had been running stably; It is judged to be a temporary network fluctuation rather than a node failure.

[0056] 3. Adaptive retry strategy formulation First retry: wait 2 seconds and try again; If it fails again: the waiting time is increased to 5 seconds; Third failure: the waiting time is increased to 10 seconds; Set the maximum number of retries to 5.

[0057] 4. Execute retry First retry: The connection still times out; Second retry: The connection is established successfully but data transmission is interrupted; The system identifies different types of faults and adjusts the waiting time to 8 seconds; Third retry: Successfully connected and completed inspection.

[0058] 5. Recovery and Verification After obtaining the inspection data, the system performs integrity verification; Record the network anomaly in the node history; Update the network status characteristic model of the node.

[0059] 6. Learning and Optimization According to the characteristics of this network fluctuation, the system updates the retry parameters; Optimize and adjust the retry strategy of the IP segment node; Record successful recovery methods in the knowledge base.

[0060] Measures to prevent accidental injuries: If the checkpoint shows that the disk usage is above the threshold, but the data is incomplete due to network fluctuations: The system will mark the data as "suspicious data"; After the retry is successful, the two data are compared. If the difference is large, the new data is used; If it is still impossible to obtain complete data, the result will be clearly marked as "data acquisition exception" instead of "check failure"; This mechanism ensures that even when the network is unstable, the inspection results can reflect the true status of the node to the greatest extent, effectively avoiding accidental damage.

[0061] In actual applications, the systems to be inspected are not only applications and databases, but also various middleware. Some cannot directly log in to the corresponding nodes to execute the inspection script, while some can only obtain the inspection results through monitoring systems such as Prometheus. Therefore, the embodiment of the present invention can also put forward the requirement of unified inspection interface adaptation layer for the inspection execution AI Agent. First, the inspection tasks are classified and several situations are abstracted, such as whether the inspected application is a traditional application (usually an application developed with a monolithic architecture, all functional modules and components are packaged into an independent application, sharing the same database and code base), a containerized application (the application and its dependencies are packaged into a standardized, portable unit, which is called a container), a database node or a middleware. The middleware is divided into traditional deployment, containerized CRD deployment, and StatefulSet deployment. From the perspective of access mode, it can also be divided into direct communication access by process, access through openapi, etc. Therefore, according to different modes, there are different adaptation anti-corrosion layers, so that different inspection objects can be adapted to solve the problem of differentiated adaptation.

[0062] 5. RAG Knowledge Base RAG knowledge base is used to store inspection execution reports and inspection problem troubleshooting reports.

[0063] The AI ​​Agent multi-agent based intelligent inspection assistant system of the embodiment of the present invention is a multi-AI Agent intelligent body architecture, wherein the main AI Agent is responsible for intention recognition, and the guiding AI Agent is responsible for assigning the current work to the corresponding task agent.

[0064] The AI ​​Agent multi-agent-based intelligent inspection assistant system of this embodiment also refers to the RAG knowledge base, which mainly stores two types of information. One type is the inspection execution report, which is incrementally added to the knowledge base through the knowledge base content construction engine; the other type is the inspection problem troubleshooting report, which is actively uploaded to the knowledge base by the knowledge base maintainer.

[0065] The vector model of the RAG knowledge base mainly uses bge-m3 (multi-functional text embedding model). The vector model candidates are generally bge-m3 and bm25. For semantic search, the essence is the distance between points in high-dimensional space, so it needs to be high-dimensional, so bge-m3 is more suitable. BM25 is a sparse embedding (most of it is 0, and only a few tokens have values). It is more suitable for finding relevance but cannot understand the context. It is an improvement on tf-idf, so it is suitable for keyword search. For AI inspection scenarios, a strong semantic relevance interpretation is required, so bge-m3 is used.

[0066] The RAG knowledge base contains two types of documents, one is various inspection execution reports, and the other is inspection problem repair solution reports that are manually improved. Therefore, the problem of knowledge fragmentation will exist in the RAG knowledge base. The content of different inspection items at different nodes is not only relevant but also specific, and is also related to the site during the inspection. Therefore, this embodiment can also introduce an automatic association and integration mechanism based on the big model, that is, how to associate the inspection report with the corresponding inspection solution. First, let the big model extract keywords and summaries for each document, and compare them based on keywords and summaries to achieve the association between the inspection report and the corresponding inspection solution, thereby improving the coherence and integrity of the knowledge base.

[0067] For inspection problem repair reports, different people have different writing levels, which will also lead to different values ​​and details of each problem repair report. For the same problem, there may be multiple inspection problem repair reports, but the quality of each is different. The corresponding weights can be assigned and dynamically adjusted based on the frequency of inspection problem repair reports being recalled and the actual solution effect after adopting this inspection problem repair report. The assistant system of this embodiment can use weights and dynamic adjustment algorithms, that is, if the inspection problem repair report is repeatedly adopted, it will be recognized as a high-quality report and will be given more weight.

[0068] At the same time, the assistant system of this embodiment has an "anti-starvation" mechanism. That is, if there is already a repair report A for a certain inspection problem, and it has been cited many times, but now the administrator uploads another repair report B, and the quality and detail of repair report B are much higher than those of repair report A. Under normal circumstances, if only the frequency of citation and recall is considered, the new repair report B will have a low score because it has not been cited before, so repair report A will eventually be selected, and repair report B will be starved to death. At this time, the solution is to give the new report a very high first weight (but it disappears the second time), that is, to give the new report a chance. If the new repair report B is finally selected through large model analysis, then repair report B will retain this extremely high weight, so that it will be adopted first in similar inspection problems. Otherwise, its weight will be reduced to the same weight as the report that was previously introduced the most times. In this way, starvation can be prevented.

[0069] In practical applications, a knowledge timeliness management mechanism can also be introduced to ensure that outdated information is not used to generate repair suggestions.

[0070] The RAG knowledge base has the following functions: (1) an automatic association and integration mechanism for knowledge fragments to improve the coherence and completeness of the knowledge base; (2) an algorithm for dynamically adjusting knowledge weights based on recall frequency and resolution effect; (3) an anti-starvation mechanism for high-quality inspection and repair reports; and (4) a knowledge timeliness management mechanism to ensure that outdated information is not used to generate repair suggestions.

[0071] In actual applications, when an unhealthy inspection item appears during the inspection, the inspection is required to execute the AI ​​Agent to give repair suggestions. The repair suggestions are formed by searching for similar problem-solving reports in the RAG knowledge base. For the same problem, there may be multiple repair methods, such as killing the process, modifying the operating system kernel parameters, and even restarting. Therefore, the embodiment of the present invention also introduces a repair multi-solution recommendation module. That is, when one or more repair solutions are recalled from RAG, they are not directly used as recommended repair suggestions, but are combined with the actual situation of the current system (such as the degree of exceeding the threshold, the healthy operation time of the system, and the upstream and downstream conditions of the system) to comprehensively give the first recommended repair solution, the second recommended repair solution, the third recommended repair solution, etc.

[0072] The repair plan given is not just a command, but often requires more detailed information, such as repair steps and dependency analysis, to ensure that the current repair operation is safe and will not affect the current system.

[0073] To this end, a repair effect prediction model is introduced to evaluate the potential results of different repair solutions. Specifically, if a repair suggestion is given, a scenario can be exercised to determine the series of impacts that may be caused to the current node after the suggestion is executed, or whether a cascading impact may occur. For example, for file handles, if the repair suggestion is to blindly increase the number, when a large number of file handles are sockets, a large amount of memory will be occupied, and these memories are not exchangeable; in addition, if a container is Out of Memory and is discovered by the inspection. If the inspection only recommends increasing the maximum memory of the current container, it may affect other Pods scheduled to the Node. Other Pods may be OOMKilled due to insufficient memory, which would not have been OOMKilled originally. Therefore, the introduction of the repair effect prediction model can simulate scenarios during inspection and repair to prevent the situation where the current problem is fixed but other problems are introduced at the same time.

[0074] When generating repair suggestions, a multi-scheme recommendation module based on the risk assessment of the repair operation, automatic sorting of repair steps and dependency analysis mechanism are used to ensure the safety of the repair operation; a repair effect prediction model is built to evaluate the potential results of different repair schemes.

[0075] There are two key decision points here: First, the inspection and repair operation is somewhat dangerous. If the repair suggestion is wrong and executed automatically, it will have a catastrophic impact. Therefore, AI only gives repair suggestions, but will not automatically execute the repair action through a dedicated AI Agent; second, the repair suggestions automatically generated by AI will not be directly put into the RAG knowledge base, because these have not been reviewed by experts. If the suggestions are wrong and enter the RAG knowledge base, they may be repeatedly recalled by other similar scenarios, thus cascading to more scenarios with wrong repair suggestions. Therefore, only inspection problem repair reports that have been reviewed by human experts and carefully written can be added to the RAG knowledge base as high-quality assets.

[0076] The intelligent inspection assistant system based on AI Agent multi-agent of this embodiment can abstract the entire inspection scene based on the observation, thinking, action and memory mechanism of multi-agent. Without human intervention, it can efficiently complete a series of tasks such as creating inspection tasks, executing inspection actions, analyzing inspection results, and giving optimization suggestions. Greatly improve work efficiency. Furthermore, the multi-agent architecture of the system architecture of the embodiment has a clear division of labor, can automatically recognize intent based on human language instructions, and automatically call the AI ​​Agent that can complete the current task. For the access and deployment methods of differentiated systems, components, and middleware, a unified inspection interface is adopted to adapt to anti-corrosion.

[0077] In an optional embodiment of the present invention, Figure 2 As shown, the inspection task includes an inspection scenario, multiple inspection rules, and an inspection task execution plan. The inspection scenario is a description of the inspection task and the applicable scenario, and summarizes the scope and boundary of the inspection.

[0078] Inspection rules are specific execution rules determined based on inspection scenarios, as well as certain rule subdivisions, such as business capacity rules, configuration funnel rules, and service governance rules. Inspection rules include one or more inspection plug-ins, which can obtain customized inspection plug-in templates from the inspection plug-in library and configure them; each inspection plug-in includes one or more checkpoints, each checkpoint represents an item to be checked, and each checkpoint contains a checkpoint assertion and one or more execution scripts; a checkpoint assertion is an assertion expression with only two values, True and False; and an execution script is a Shell script that supports checkpoint assertions.

[0079] Optionally, the system of this embodiment can perform intelligent inspection point and rule prediction based on specific inspection scenarios, and can intelligently generate corresponding inspection points and inspection rules based on the description of activity events and environment switching events.

[0080] Inspection is not only a routine inspection of server resources, but also provides inspection for specific events. Events can be divided into two categories. One is activity events, such as promotions, which have certain requirements for resource stability and high concurrency. Another category is environment switching events, such as disaster recovery switching, system expansion, grayscale release, etc., which also have requirements for resource capacity, switching stability, etc. The optional embodiment of the present invention can intelligently generate corresponding inspection points and inspection rules based on the description of events. Taking promotion as an example, in the description of the promotion scene, there will be corresponding content such as duration, the number of high concurrency that the system must support, etc., so a series of inspection items can be automatically generated based on this information, and these inspection items will be very different from conventional inspections. In addition, rule prejudgment must be performed. For example, there will be fuse downgrade and current limiting protection during the promotion activity, so the corresponding content will also be inspected during the inspection, such as whether the downgraded service can run normally through the inspection.

[0081] The inspection task execution plan is an execution plan for the inspection task configured in combination with the inspection rules. It can be triggered by rules, scheduled tasks, or manually.

[0082] The execution of the inspection task will generate a corresponding inspection task execution report and archive it in the inspection task report library. In the current design, if the inspection task encounters a problem, it will be solved through a specific processing script, and the solution process and results will also be incorporated into the inspection task report.

[0083] Take a case study of an application health check case for a large system cluster as an example. Figure 3 As shown in the figure, the conceptual model of inspection is as follows: create an inspection task and assign a unique inspection task ID.

[0084] 1. Set the inspection task scenario to system application health check.

[0085] 2. Define inspection rules: Inspection rule 1: Disk space usage inspection rule. It contains two plug-ins: Plug-in 1 is for log space usage inspection, and plug-in 2 is for dbfile space usage inspection.

[0086] a. Inspection plug-in 1 contains one checkpoint - log space. Its checkpoint assertion statement is Assert (log directory space size occupied < threshold).

[0087] b. Inspection plug-in 2 contains one checkpoint - dbfile space. Its checkpoint assertion statement is Assert (database file dbfile space size occupied < threshold).

[0088] Inspection Rule 2: JVM Inspection Rule. It contains 3 plugins: Plugin 1 is for gc times and duration check, Plugin 2 is for young generation check, and Plugin 3 is for old generation check.

[0089] a. Inspection Plugin 1 contains 4 checkpoints - ygc times, ygc duration, fgc times, fgc duration. The corresponding checkpoint assertion statements are Assert(ygc times < threshold), Assert(ygc duration < threshold), Assert(fgc times < threshold), Assert(fgc duration < threshold) b. Inspection Plugin 2 contains 1 checkpoint. Its checkpoint assertion statement is Assert(young generation size < threshold) c. Inspection Plugin 3 contains 1 checkpoint. Its checkpoint assertion statement is Assert(old generation size < threshold) Inspection Rule 3: Computing Resource Utilization Inspection Rule. It contains 2 plugins: Plugin 1 is for CPU utilization check, and Plugin 2 is for memory utilization check.

[0090] a. Inspection Plugin 1 contains 1 checkpoint. Its checkpoint assertion statement is Assert(CPU% < threshold) b. Inspection Plugin 2 contains 1 checkpoint. Its checkpoint assertion statement is Assert (Memory% < threshold) Inspection Rule 4: File Handle Inspection Rule. It contains 1 plugin: the plugin is for file handle check.

[0091] a. Inspection Plugin 1 contains 1 checkpoint. Its checkpoint assertion statement is Assert (number of file handles < threshold) 3. Configure the inspection task execution plan.

[0092] Inspection Plugin 1: Scope = {192.168.0.1, 192.168.0.2, 192.168.0.3}, Triggering method = triggered by scheduled task <cronjob expression> Inspection Plugin 2: Scope = {172.19.3.35, 172.19.3.36, 172.19.3.37, 172.19.3.38}, Triggering method = triggered by scheduled task <cronjob expression> + manual trigger Inspection Plugin 3: Scope = {xx.xx.xx.xx, xx.xx.xx.xx.....} Triggering method = triggered by scheduled task <cronjob expression> + triggered by Chinese festivals (Spring Festival | Lantern Festival | Tomb-Sweeping Day | Dragon Boat Festival...) Patrol Inspection Plugin 4: Scope of action = {yy.yy.yy.yy, yy.yy.yy.yy.....}, Trigger method = Triggered by scheduled tasks <cronjob expression> + Triggered before various marketing activities.

[0093] The intelligent patrol inspection assistant system based on AI Agent multi - agents and the intelligent agent dynamic cooperation mechanism based on task complexity can realize the intelligent decomposition and collaborative execution of patrol inspection tasks. Because if all patrol inspection tasks are executed serially, it will take a relatively long time. Therefore, patrol inspection tasks, especially complex ones, can be intelligently decomposed, and the logical relationships and dependency trees among them can be clarified to form a DAG (Directed Acyclic Graph), and collaborative execution can be carried out based on the dependency relationships.

[0094] As introduced above, the intelligent patrol inspection assistant system based on AI Agent multi - agents in the embodiments of the present invention can realize the intelligent decomposition and collaborative execution of patrol inspection tasks based on the intelligent agent dynamic cooperation mechanism based on task complexity. The following details the working process of the intelligent agent dynamic cooperation mechanism based on task complexity.

[0095] I. Task Complexity Assessment The intelligent patrol inspection system first conducts a complexity assessment on the received patrol inspection tasks. The assessment dimensions include: the number of patrol inspection rules, the number of patrol inspection plugins, the total number of checkpoints, the number of nodes to be inspected, the estimated time consumption of the execution script, and the complexity of the dependency relationships between nodes.

[0096] II. Task Intelligent Decomposition Process Rule - level decomposition: Different patrol inspection rules (such as disk space, JVM, computing resources, file handles, etc.) are used as independent task units.

[0097] Plugin - level decomposition: The plugins under each rule (such as log space, dbfile space) are further decomposed.

[0098] Node - level decomposition: Grouping is carried out according to server groups or node types (application servers, database servers, gateway servers).

[0099] Parallelism assessment: Based on the system load capacity, the maximum parallel execution degree is evaluated.

[0100] III. Dependency Relationship Analysis 1. Patrol Inspection Rule Dependency Relationship Some patrol inspection rules must be executed before other rules, forming a pre - dependency relationship. For example, it is necessary to check the rules at the infrastructure level first and then the rules at the application level. As Figure 4 shown, Disk space check (A) → JVM check (B). Only when the disk space is sufficient is it necessary to conduct a JVM check.

[0101] 2. Checkpoint Dependency Relationship Checkpoints are the specific execution units of inspection rules. There are dependencies between checkpoints. The result of a checkpoint will directly affect the judgment logic of subsequent checkpoints.

[0102] like Figure 5 As shown, database connection check (X) → database performance check (Y). The performance can be checked only when the connection is normal.

[0103] 3. Node Dependencies In a distributed system environment, there are dependencies between different nodes, such as master-slave nodes, upstream and downstream service nodes, etc. These dependencies determine the execution order of inspections.

[0104] like Figure 6 As shown, database master node (M) → database slave node (N), the health of the master node is the prerequisite for the normal operation of the slave node.

[0105] When inspecting the master-slave database: first check the connection status of the master database, then check the master-slave synchronization delay, and finally check the data consistency of the slave database.

[0106] 4. Complex Directed Acyclic Graph Construction The system builds a DAG graph based on dependencies: 1. Each checkpoint is a node in the graph 2. Dependencies as directed edges 3. Topological sorting determines the execution order 4. Identify branches that can be executed in parallel In actual inspection tasks, the above three dependencies together form a complex DAG graph, based on which the intelligent system performs task decomposition and scheduling optimization. Figure 7 shown.

[0107] V. Collaborative Execution Strategy 1. Dynamic resource allocation: Dynamically allocate execution resources based on task priority and system load 2. Task scheduling optimization: Prioritize tasks on the critical path; predict task duration based on historical execution data; dynamically adjust execution parallelism.

[0108] 3. Load balancing mechanism: avoid performance impact on production systems.

[0109] Fault recovery strategy: Automatically retry or adjust the strategy when a node fails to execute.

[0110] 6. Example: Inspection of promotional activities The inspection of e-commerce systems before a large-scale promotion is a typical complex inspection scenario. At this time, the inspection task needs to ensure that all components from the network infrastructure to the application layer are in optimal condition, and that the traffic limit and fuse mechanism are correctly configured to cope with the upcoming high concurrent access.

[0111] It is now necessary to conduct a pre-promotion inspection for the e-commerce system: 1. Complexity assessment: 100 application servers; 20 database servers, 5 inspection rules, 15 plug-ins, 30 checkpoints; the assessment complexity is "high".

[0112] 2. Intelligent decomposition: The application servers are divided into 10 groups for parallel inspection; the database servers are divided into the master database group and the slave database group; the JVM inspection and the disk inspection are performed in parallel.

[0113] DAG construction, such as Figure 8 shown.

[0114] 3. Collaborative Execution: Assign 3 execution agents to handle application server checks; assign 1 dedicated agent to handle database checks; and reserve 1 agent to handle emergencies and generate repair suggestions.

[0115] 4. Agent collaborative execution process: Agent-1: Responsible for infrastructure layer checks (network connectivity, load balancing).

[0116] Agent-2: responsible for data storage layer checks (database connection, cache service).

[0117] Agent-3: responsible for application service layer inspection (application service, message queue).

[0118] Agent-4: responsible for protection mechanism checks (flow limits, circuit breaker parameters, capacity assessment) The main coordinating agent dynamically arranges the order of task execution based on the DAG graph. Agent-1 and Agent-2 can be started in parallel. When they complete their first check items, Agent-3 and Agent-4 can start working. This collaborative mechanism ensures efficient execution while also following the dependencies between check items.

[0119] 5. Merge execution results: The inspection results of each intelligent agent are summarized and the main AI Agent generates a final report through comprehensive analysis.

[0120] This mechanism reduces inspection time from 4 hours to 40 minutes, greatly improving efficiency.

[0121] In a multi-agent architecture, different agents are different nodes, and there is a possibility of man-in-the-middle attacks between them, leading to information tampering. In addition, the execution of commands from the inspection execution agent to each node in the system may also be attacked by man-in-the-middle attacks. Therefore, a secure communication protocol is introduced to encrypt the communication channel to ensure the confidentiality and integrity of inspection instructions and results.

[0122] Inspection tasks may take a long time, especially for large-scale systems. At this time, the working memory mechanism of the agent is introduced to enable it to maintain context consistency during long-term inspection tasks.

[0123] At the beginning of the inspection, there are no inspection results and repair suggestions as a reference. However, with multiple inspections, the intelligent agent can also automatically adjust the inspection strategy according to the historical execution results through self-evaluation and optimization mechanism, so that it can complete the overall inspection work most efficiently.

[0124] The system of this embodiment can be compatible with the differentiated configuration of inspection time points of different subsystem nodes, automatically generate inspection reports in a specific format, and identify possible unsatisfied items. It can automatically and intelligently provide correct repair suggestions for current inspection problems through a large model combined with the RAG knowledge base.

[0125] [Scenario 1]: The inspection configuration administrator uses human language to generate inspection tasks, such as Fig. 9 shown.

[0126] Step 1: The configuration administrator responsible for inspection issues instructions to the main AI Agent. The reference instructions are as follows: Please help me create an inspection task, mainly to check the health of system applications.

[0127] First, check the disk space of the root directory of all components and middleware, which must not exceed 10G. Check the disk space of the database db file, which must not exceed 50G. Then check the CPU utilization of all nodes, which must not be higher than 70%. The memory utilization of all nodes must not be higher than 60%. Then check the number of file handles of all nodes, which must not exceed 60,000.

[0128] The inspection task is performed twice a day, once at midnight and once at noon.

[0129] The list of application servers being inspected is:<app_ip1> ,<app_ip> ..., the list of inspected database servers is<db_ip1> ,<db_ip2> ..., the list of gateway servers inspected is xxxx,xxx.

[0130] Step 2: The main AI Agent calls LLM to perform intent recognition, and the recognized intent is "create inspection task".

[0131] Step 3: The main AI Agent passes the task to the guiding AI Agent to assign tasks.

[0132] Step 4: Guide the Agent to identify the intent and find that it is to create an inspection task. Therefore, @InspectionTaskGen-Agent is used to let the AI ​​Agent responsible for generating inspection tasks to generate the inspection task.

[0133] Step 5: The inspection task generates an AI Agent to start a workflow. First, LLM is used to extract the necessary information from the context, including the definition of three inspection rules, the configuration of the checkpoint assertion in the inspection plug-in of each rule, the replacement of the IP in the execution script, etc., and then the execution plan task is associated.

[0134] Step 6: After extracting the necessary information, assemble it into an inspection task page, accompanied by inspection task details. On this page, the inspection configuration administrator can click Confirm (or execute directly without confirmation).

[0135] Step 7: The workflow will call the inspection task configuration tool to create an inspection task. This is a Tool Call.

[0136] Step 8: The tool for creating inspection tasks calls the data access object (dao) layer and adds inspection task records and corresponding details in the corresponding database.

[0137] Step 9: The tool that creates the inspection task will also notify relevant personnel via email or SMS notifications through the pre-configured WebHook.

[0138] Step 10: You can see the corresponding inspection tasks on the inspection dashboard, and you can also view the inspection tasks that were created previously.

[0139] [Scenario 2]: Inspection task execution AI Agent automatically performs inspections according to the timer and generates execution result reports and repair suggestions, such as Fig.10 shown.

[0140] Step 1: Inspection task execution AI Agent automatically performs inspections according to the timer configuration. It will pull up a workflow and call the inspection task execution tool, which is a Tool Call.

[0141] Step 2: The inspection task execution tool obtains the corresponding inspection task and related configuration details from the inspection task configuration database.

[0142] Step 3: The inspection task execution tool traverses each inspection plug-in and calls the corresponding execution script (for example, to check resource consumption on different servers), then summarizes each result and generates the final result of the plug-in assertion statement.

[0143] Step 4: The workflow calls the next node: the inspection result generation tool, which is also a Tool Call.

[0144] Step 5: The inspection result generation tool calls the data access object Dao layer to record the inspection results in the inspection result execution database.

[0145] Step 6: The inspection result generation tool calls LLM to organize the report generated by the results, such as how many inspection rules there are, which inspection plug-ins are executed, whether the assertion of each inspection plug-in passes, and on which machines the execution fails, etc.

[0146] Step 7: Output the inspection results in the required format and display the corresponding inspection result report, including the details of the inspection rules that failed.

[0147] Step 8: If the inspection result is completely healthy, the relevant personnel will be notified through the configuration of WebHook, and the workflow ends; if the inspection result is not completely healthy, this step of WebHook will not be executed first, and the workflow will continue.

[0148] Step 9: [Called only when the inspection result is not completely healthy] The workflow calls the next node: Repair Suggestion Tool. This is also a Tool Call.

[0149] Step 10: The repair suggestion tool first recalls the document fragments corresponding to the inspection execution report and the inspection problem troubleshooting report from the RAG knowledge base.

[0150] Step 11: The repair suggestion tool puts the document fragments recalled from RAG into the LLM context, and based on the LLM-assisted reasoning, gives corresponding repair suggestions for the current inspection problem.

[0151] Step 12: After organizing the language, the repair suggestion tool generates inspection issues and corresponding repair suggestions and displays the content.

[0152] Step 13: Notify relevant personnel through the WebHook configuration and the workflow ends.

[0153] Step 14: [Asynchronous operation] The knowledge base content building engine will automatically add the inspection result content to the RAG knowledge base in an incremental manner. At the same time, the knowledge base maintenance personnel will also regularly add the inspection problem troubleshooting report to the RAG knowledge base.

[0154] [Scene 3]: The inspection configuration administrator uses human language to immediately initiate inspection execution and generate execution result reports and repair suggestions, such as Fig.11 shown.

[0155] Step 1: The configuration administrator in charge of inspection issues a command to the main AI Agent: Please help me execute the previously generated inspection task with inspection task number XXX immediately.

[0156] Step 2: The main AI Agent calls LLM to perform intent recognition, and the recognized intent is "perform inspection tasks."

[0157] Step 3: The main AI Agent passes the task to the guiding AI Agent to assign tasks.

[0158] Step 4: Guide the Agent to identify the intent and find that the inspection task is to be executed. Therefore, @InspectionTaskExec-Agent is used to let the AI ​​Agent responsible for the inspection task to execute the inspection task.

[0159] Step 5: Inspection task execution AI Agent starts a workflow and calls the inspection task execution tool to specifically execute the inspection task. This is a Tool Call.

[0160] Step 6: The inspection task execution tool obtains the corresponding inspection task and related configuration details from the inspection task configuration database based on the inspection task number.

[0161] All subsequent steps are similar to those in the previous scenario and will not be repeated here.

[0162] In general, the AI ​​Agent multi-agent intelligent inspection assistant system of the embodiment of the present invention has multiple functions. Fig.12 It is a blueprint diagram of the functional module.

[0163] Microservice Architecture Containerized deployment is adopted, with each agent acting as an independent microservice; service mesh is used to manage communication between agents; and loosely coupled component interaction is achieved based on event-driven architecture.

[0164]

AI Technology Stack

Collaboration Mechanism between Agents

[0165] 2. Consensus decision-making mechanism: Use multi-agent voting, evidence weight and confidence assessment to reach consensus on complex inspection assertions and improve the accuracy of results.

[0166] 3. Collaborative communication protocol: defines standardized message formats and communication processes, supports multiple modes such as request-response, publish-subscribe, and broadcast, and ensures efficient information transmission.

[0167] 4. Task decomposition strategy: Automatically decompose inspection tasks based on complexity and dependencies, create optimal execution paths, and balance parallel execution and resource utilization.

[0168] 5. Conflict resolution mechanism: Implement mechanisms such as priority sorting, evidence evaluation and arbitration agent to handle judgment conflicts among multiple agents and ensure consistency of results.

[0169] 6. Collective knowledge update: Maintain a shared knowledge base, update shared knowledge through inspection experience accumulation, and achieve continuous optimization of collective intelligence.

[0170] Monitoring and Observability 1. Distributed tracing system monitors cross-agent calls 2. Time series database storage performance indicators 3. Visual dashboard showing inspection status

Deployment and expansion

[0171] 2. Based on the large model and RAG knowledge base, it is possible to correctly associate existing inspection problems with corresponding repair solutions, and explainability is also provided based on RAG, that is, explaining why this repair solution is feasible.

[0172] 3. It can intelligently generate inspection reports in a specific format and identify possible unsatisfactory items. However, both manual sorting and rule-based sorting methods lack efficiency and flexibility.

[0173] 4. It can be compatible with the differentiated configuration of inspection time points of different subsystem nodes. The dynamic collaboration mechanism of intelligent agents based on task complexity can realize the intelligent decomposition and collaborative execution of inspection tasks, which greatly adapts to the complexity of inspection tasks.

[0174] 5. Secure communication protocol between agents to ensure the confidentiality and integrity of inspection instructions and results, and prevent man-in-the-middle attacks.

[0175] 6. The working memory mechanism of the intelligent agent enables it to maintain contextual consistency during long-term inspection tasks, thereby reducing token consumption, that is, operating costs.

[0176] 7. The intelligent agent self-evaluation and optimization mechanism automatically adjusts the inspection strategy according to the historical execution results, thereby automatically improving the results.

[0177] 8. The adaptive inspection rule generation algorithm for the microservice architecture can automatically generate the optimal inspection path according to the service topology, thereby greatly improving the operational efficiency of the entire inspection.

[0178] 9. The intelligent inspection priority determination mechanism based on the service dependency graph can handle the dependency of inspection objects well, such as the situation where the accuracy of the inspection assertion of a certain inspection item depends on the inspection assertion result of the inspection item of the previous node.

[0179] 10. The fuzzy matching and clarification mechanism of natural language requirements can improve the accuracy of inspection task creation.

[0180] 11. The task template automatic iteration optimization system based on historical inspection effect feedback can be continuously optimized based on historical inspection results.

[0181] 12. Based on the description of activity events and environment switching events, the corresponding inspection points and inspection rules can be intelligently generated, thus eliminating the trouble of manually creating various complex inspection points and inspection rules.

[0182] 13. Distributed inspection load balancing mechanism to avoid performance impact on production systems.

[0183] 14. Based on the adaptive retry and fault recovery mechanism, it ensures that the inspection task can still be completed under abnormal conditions such as network fluctuations, thereby greatly reducing the false alarm rate caused by network fluctuations.

[0184] 15. The multi-plan recommendation system based on the risk assessment of the repair operation ensures that there will be multiple plans for the inspection and repair executors to choose from.

[0185] 16. Automatic sorting of repair steps and dependency analysis mechanism ensure the safety of repair operations and prevent the introduction of new inspection problems during inspection and repair.

[0186] 17. The repair effect prediction model is used to evaluate the potential results of different repair plans and prevent a repair suggestion from causing damage to the original system.

[0187] 18. Introduce a mechanism for automatic association and integration of knowledge fragments to improve the coherence and integrity of the knowledge base.

[0188] 19. Introduce a knowledge weight dynamic adjustment algorithm based on recall frequency and resolution effect to more effectively find the best repair suggestions.

[0189] 20. Introduce a mechanism to prevent starvation of high-quality inspection and repair reports to prevent high-quality new inspection and repair reports from not being given priority due to insufficient citations.

[0190] 21. Introduce a knowledge timeliness management mechanism to ensure that outdated information is not used to generate repair suggestions.

[0191] 22. Introduce access and deployment methods for differentiated systems, components, and middleware, and use a unified inspection interface to adapt the anti-corrosion layer, so that it can be compatible with various differentiated component inspection scenarios.

[0192] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. An intelligent inspection assistant system based on AI Agent multi-agent, characterized in that: The system comprises: The main AI Agent module is used to receive inspection instructions and perform intent recognition based on the inspection instructions; A guiding AI Agent module, connected to the main AI Agent module, for performing task allocation based on the intention recognition result of the main AI Agent module; Inspection task generation AI Agent module, used to generate inspection tasks; Inspection task execution AI Agent module, used to execute inspection tasks; the inspection task generation AI Agent module and the inspection task execution AI Agent module use workflow to define the main execution process and logic; RAG knowledge base is used to store inspection execution reports and inspection problem troubleshooting reports.

2. The system according to claim 1, characterized in that When the inspection task generation AI Agent module receives the call instruction of the guiding AI Agent module, it extracts the inspection information based on the inspection instruction and assembles the inspection task page; Call the inspection task configuration tool to create an inspection task.

3. The system according to claim 1, characterized in that Inspection tasks include inspection scenarios, multiple inspection rules, and inspection task execution plans; The inspection scenario is a description of the inspection task and applicable scenario, and also summarizes the scope and boundaries of the inspection; Inspection rules are specific execution rules determined based on inspection scenarios, including business capacity rules, configuration funnel rules, and service governance rules; The inspection task execution plan is an execution plan for the inspection task configured in combination with the inspection rules, including rule-triggered execution, scheduled task-triggered execution, or manual triggering execution.

4. The system according to claim 3, characterized in that The inspection rules include one or more inspection plug-ins. The inspection plug-ins obtain customized inspection plug-in templates from the inspection plug-in library and configure them. Each inspection plug-in includes one or more checkpoints. Each checkpoint represents a content to be checked. Each checkpoint contains a checkpoint assertion and one or more execution scripts. A checkpoint assertion is an assertion expression with only two possible values: True and False. The execution script is a Shell script that supports checkpoint assertions.

5. The system according to claim 4, characterized in that The inspection task generation AI Agent module is used to intelligently predict inspection points and rules based on specific inspection scenarios, and intelligently generate corresponding inspection points and inspection rules based on the description of activity events and environment switching events.

6. The system according to any one of claims 1 to 5, characterized in that: The system further comprises: A repair multi-solution recommendation module is used to search for relevant problem solving reports in the RAG knowledge base when unhealthy inspection items are detected during inspection, so as to provide at least one repair solution, as well as repair steps and dependency analysis descriptions corresponding to each repair solution; The multi-repair scheme recommendation module is provided with a repair effect prediction model, which is used to simulate the repair scenario of each of the repair schemes to evaluate the repair results corresponding to different repair schemes.

7. The system according to any one of claims 1 to 5, characterized in that: The inspection execution AI Agent module has different adaptation anti-corrosion layers to adapt to different inspection objects; The inspection objects include traditional applications, containerized applications, database nodes, and middleware. The middleware includes traditional deployment, containerized CRD deployment, and StatefulSet deployment.

8. The system according to any one of claims 1 to 5, characterized in that: The main AI Agent module is also used for the dynamic collaboration mechanism of intelligent agents based on task complexity to achieve intelligent decomposition and collaborative execution of inspection tasks, including: Task complexity assessment, intelligent task decomposition, dependency analysis, complex directed acyclic graph construction, and collaborative execution strategy.

9. The system according to claim 8, characterized in that The task complexity analysis includes a complexity assessment of the received inspection tasks, and the assessment dimensions include: the number of inspection rules, the number of inspection plug-ins, the total number of checkpoints, the number of inspected nodes, the estimated time consumption of executing the script, and the complexity of the dependency relationship between nodes; The intelligent decomposition of tasks includes: Rule-level decomposition, treating different inspection rules as independent task units; Plug-in level decomposition, further decomposing the plug-ins under each rule; Node-level decomposition, grouped by server group or node type; Parallelism evaluation, evaluates the maximum parallel execution degree based on the system load capacity.

10. The system according to claim 8, characterized in that The dependencies include: inspection rule dependencies, checkpoint dependencies, and node dependencies; When a complex directed acyclic graph is constructed, a DAG graph is constructed based on the dependency relationship.

Citation Information

Patent Citations

  • Multi-agent cooperative electric power inspection method, system and device and storage medium

    CN119106101A

  • Unmanned aerial vehicle inspection system and method based on multi-agent cooperation

    CN119717889A

  • Multi-agent system and control method therefor

    WO2022001120A1

Cited By

  • Financial text logicality evaluation method and system based on AI Agent

    CN120542431A

  • Book data processing and intelligent service system based on artificial intelligence

    CN120653760A

  • Home monitoring method and system based on multi-agent large model

    CN120751089A

  • Intelligent gray release method and system and computer equipment

    CN120849236A

  • Machine learning model quality inspection system and method based on multi-agent cooperation

    CN120875093A