Dynamic problem troubleshooting method and device, electronic equipment and readable storage medium

CN121681206BActive Publication Date: 2026-09-22BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511913664.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-09-22
Estimated Expiration
2045-12-17

AI Technical Summary

Technical Problem

[0003]这套模式在面对诸如“服务突然变慢”、“CPU利用率异常飙升”这类动态排查性问题时,就会明显力不从心

Benefits of technology

[0010]本公开所提供的动态排查问题处理方案,通过获取需要通过至少两种排查方式来确定真实原因的待排查问题,并利用以知识工程领域专用语言预先构建的结构化知识来确定对应的排查方案,明确了基于不同排查方式和特定排查顺序的任务规划,其中结构化知识通过逻辑分析节点、数据需求和工具接口以结合的方式呈现,得以将抽象的排查逻辑转化为可具体执行的标准化任务描述,然后通过将方案中各排查方式按照既定顺序下发给基于数据需求与能力匹配的智能体来逐步执行,实现了对复杂异常状况的协同式、步骤化排查,直至定位真实原因。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121681206B_ABST
    Figure CN121681206B_ABST
Patent Text Reader

Abstract

The present disclosure provides a dynamic problem troubleshooting method and device, electronic equipment and readable storage medium, which are related to the fields of artificial intelligence such as knowledge engineering, structured information, intelligent agent, external tool calling and visualization. The method comprises: obtaining a to-be-troubleshooted problem initiated for an abnormal condition; determining a troubleshooting scheme corresponding to the to-be-troubleshooted problem by using structured knowledge constructed in advance according to a knowledge engineering domain specific language (KDSL), wherein the troubleshooting scheme is determined based on different troubleshooting methods and troubleshooting sequences; and issuing each troubleshooting method to a matched intelligent agent according to the troubleshooting sequence for step-by-step troubleshooting until the real cause of the abnormal condition is determined. The method can significantly improve the coverage and solving capability of systematic diagnosis of complex problems without clear intention or requiring multi-step reasoning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, specifically to artificial intelligence technologies such as knowledge engineering, structured information, intelligent agents, external tool invocation, and visualization, and particularly to a dynamic problem-solving method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] In the current operation and maintenance of large-scale computer infrastructure, although intelligent customer service has been applied to handle basic inquiries and operational guidance, its core relies on "intent recognition" plus "fixed standard operating procedures (SOPs)".

[0003] This model falls short when dealing with dynamic, troubleshooting issues such as "sudden service slowdown" or "abnormal spikes in CPU utilization." This is because these problems often lack a single, clear intent and require a multi-step, iterative analysis and reasoning process, much like a detective solving a case, combining real-time system data. Summary of the Invention

[0004] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for dynamically troubleshooting problems.

[0005] In a first aspect, this disclosure proposes a dynamic problem-solving method, comprising: acquiring a problem to be investigated in response to an abnormal situation; wherein the problem to be investigated requires at least two investigation methods to determine the true cause of the abnormal situation; determining an investigation plan corresponding to the problem to be investigated using structured knowledge pre-constructed according to the knowledge engineering domain-specific language KDSL; wherein the investigation plan is determined based on different investigation methods and investigation order, the structured knowledge includes logical analysis nodes, data requirements, and tool interfaces, the logical analysis nodes correspond to the investigation methods, the tool interfaces are used to describe the tools that can be called when the corresponding logical analysis node performs the corresponding investigation task, and the data requirements are used to describe the structure of the input / output data of the corresponding logical analysis node and the called tools when performing the corresponding investigation task; distributing each investigation method to the matched intelligent agent in the investigation order to investigate step by step until the true cause of the abnormal situation is determined; wherein the matching between the investigation method and the intelligent agent is determined based on the data requirements of the corresponding task and the capabilities of different intelligent agents.

[0006] Secondly, this disclosure proposes a dynamic problem-solving apparatus, comprising: a problem-to-be-solved-problem acquisition unit, configured to acquire a problem to be solved in response to an abnormal situation; wherein the problem to be solved requires at least two investigation methods to determine the true cause of the abnormal situation; an investigation scheme determination unit, configured to determine an investigation scheme corresponding to the problem to be solved using structured knowledge pre-constructed according to the knowledge engineering domain-specific language KDSL; wherein the investigation scheme is determined based on different investigation methods and investigation order, the structured knowledge includes logical analysis nodes, data requirements, and tool interfaces, the logical analysis nodes correspond to the investigation methods, the tool interfaces describe the tools that can be called when the corresponding logical analysis node performs the corresponding investigation task, and the data requirements describe the structure of the input / output data of the corresponding logical analysis node and the called tools when performing the corresponding investigation task; and a step-by-step investigation unit, configured to distribute each investigation method to a matching agent in the investigation order for step-by-step investigation until the true cause of the abnormal situation is determined; wherein the matching between the investigation method and the agent is determined based on the data requirements of the corresponding task and the capabilities of different agents.

[0007] Thirdly, embodiments of this disclosure provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the dynamic troubleshooting method as described in the first aspect.

[0008] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions that enable a computer to implement the dynamic troubleshooting method as described in the first aspect when executed.

[0009] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, can implement the steps of the dynamic troubleshooting method as described in the first aspect.

[0010] The dynamic problem-solving solution provided in this disclosure acquires problems that require at least two investigation methods to determine the root cause, and utilizes pre-built structured knowledge in a knowledge engineering-specific language to determine the corresponding investigation plan. It clarifies the task planning based on different investigation methods and specific investigation sequences. The structured knowledge is presented in a combined manner through logical analysis nodes, data requirements, and tool interfaces, which transforms abstract investigation logic into a standardized task description that can be specifically executed. Then, by distributing the various investigation methods in the solution to intelligent agents based on matching data requirements and capabilities in a predetermined order, it achieves collaborative and step-by-step investigation of complex anomalies until the root cause is located.

[0011] This solution transforms the dynamic investigation process, which originally relied on human experience and was difficult to systematize, into a standardized process driven by structured knowledge and executed collaboratively by multiple specialized intelligent agents. This significantly improves the coverage and problem-solving capabilities for systematically diagnosing complex problems without clear intent or requiring multi-step reasoning. At the same time, through task decomposition and precise matching of capabilities among intelligent agents, it effectively reduces the cognitive load and error risk when a single model handles a complex full-process, enhancing the reliability, interpretability, and reproducibility of the entire investigation process.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is an exemplary system architecture to which this disclosure can be applied; Figure 2 A flowchart of a dynamic problem-solving method provided in this embodiment of the disclosure; Figure 3 A flowchart of a step-by-step troubleshooting method provided in this embodiment of the disclosure; Figure 4 A branch diagram illustrating two different methods for determining the next screening step provided in embodiments of this disclosure; Figure 5 A flowchart illustrating a task assignment and execution method provided in this embodiment of the disclosure; Figure 6-1 , Figure 6-2 , Figure 6-3 Figure 7 A structural block diagram of a dynamic problem-solving device provided in this embodiment of the present disclosure; Figure 8 This is a schematic diagram of the structure of an electronic device suitable for performing a dynamic troubleshooting method, provided as an embodiment of the present disclosure. Detailed Implementation

[0014] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0015] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0016] Figure 1 An exemplary system architecture 100 is shown, in which embodiments of the dynamic troubleshooting methods, apparatuses, electronic devices, and computer-readable storage media of this disclosure can be applied.

[0017] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0018] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed. These applications include question-and-answer interaction applications, troubleshooting applications, and instant messaging applications.

[0019] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.

[0020] Server 105 can provide various services through its built-in applications. Taking a question-and-answer interactive application that provides troubleshooting services as an example, when running this application, server 105 can achieve the following: First, it receives user-initiated troubleshooting questions from terminal devices 101, 102, and 103 via network 104. These questions require at least two troubleshooting methods to determine the true cause of the anomaly. Next, using structured knowledge pre-built according to the knowledge engineering domain-specific language KDSL, it determines a troubleshooting plan corresponding to the question. This plan is based on different troubleshooting methods... Once the investigation methods and order are determined, the structured knowledge includes logical analysis nodes, data requirements, and tool interfaces. Each logical analysis node corresponds to a specific investigation method, the tool interface describes the tools that can be invoked by the corresponding logical analysis node when performing the corresponding investigation task, and the data requirements describe the structure of the input / output data of the corresponding logical analysis node and the invoked tools when performing the corresponding investigation task. Then, each investigation method is distributed to the matching agent according to the investigation order to investigate step by step until the true cause of the anomaly is determined. The matching between the investigation method and the agent is determined based on the data requirements of the corresponding task and the capabilities of different agents.

[0021] It should be noted that, in addition to being obtained in real time from terminal devices 101, 102, and 103 via network 104, the issues to be investigated can also be pre-stored locally on server 105 through various means. Therefore, when server 105 detects that this data is already stored locally (for example, when starting to process previously reserved investigation tasks), it can choose to directly retrieve this data from locally. In this case, the exemplary system architecture 100 may also exclude terminal devices 101, 102, and 103 and network 104.

[0022] Because dynamic troubleshooting of complex problems requires significant computing resources and power, the dynamic troubleshooting methods provided in the subsequent embodiments of this disclosure are generally executed by a server 105 with strong computing power and abundant computing resources. Correspondingly, the dynamic troubleshooting device is also generally located within the server 105. However, it should also be noted that when terminal devices 101, 102, and 103 also possess sufficient computing power and resources, they can also perform the aforementioned calculations performed by the server 105 through their installed question-and-answer interactive applications, thereby outputting the same results as the server 105. Especially when multiple terminal devices with different computing capabilities exist simultaneously, but the question-and-answer interactive application determines that its terminal device has strong computing power and abundant remaining computing resources, it can allow the terminal device to perform the aforementioned calculations, thereby appropriately reducing the computing pressure on the server 105. Accordingly, the dynamic troubleshooting device can also be located within the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude the server 105 and the network 104.

[0023] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0024] Please refer to Figure 2 , Figure 2 A flowchart of a dynamic troubleshooting method provided in this disclosure embodiment is provided, wherein process 200 includes the following steps: Step 201: Obtain the issues to be investigated in response to the abnormal situation that occurred; This step aims to address the issue by the entity responsible for dynamically identifying and handling the problem (e.g., [the entity]). Figure 1 The question-and-answer interactive application running on server 105 shown receives interactive requests from user terminals (such as workstations of maintenance personnel), and obtains the questions to be investigated by the user in response to the abnormal situation by interpreting and identifying the interactive request.

[0025] The term "abnormal condition" here is a broad technical concept. It refers to any negative event or phenomenon in the system, network, or application environment monitored or served by the server that deviates from the expected normal state or performance indicators. Examples include abnormally long service response times, abnormally high CPU usage on the server, or exhaustion of the database connection pool. User requests, such as "My service is slow, please check the cause" or "Machine A's CPU is experiencing a hotspot," are preliminary and often superficial descriptions of this type of abnormal condition.

[0026] To accurately identify such problems, it's crucial to consider their inherent complexity: they "require at least two investigation methods to determine the true cause of the anomaly," distinguishing them from simple question-and-answer queries. Technically, "at least two investigation methods" implies that the underlying cause of the anomaly is not fixed or singular, or that its root cause is hidden beneath multiple layers of appearance, making it impossible to reach a conclusion directly by querying a single knowledge point or performing an isolated check (e.g., simply checking the current CPU load). It suggests that locating the root cause must rely on a chain of reasoning involving multiple logical steps, potentially involving different data sources and investigation tools. For example, regarding the problem of "service slowdown," possible causes include backend application code logic defects, database query performance bottlenecks, network congestion, and underlying host resource contention, among others. Any one or a combination of these could be the "true cause." Therefore, when the aforementioned implementing entity obtains such questions, it can perform intent analysis and complexity assessment on the question text based on natural language processing models or predefined rules to determine whether it belongs to a complex question category that requires initiating the multi-step, multi-agent collaborative investigation process of this solution, rather than a simple question for which the answer can be directly retrieved from the knowledge base.

[0027] In practice, the aforementioned execution entity can obtain the problem to be investigated in various ways. The most common method is receiving natural language input from the user through an integrated dialog interface; alternatively, it may receive structured alarm information from other monitoring systems or automated scripts via application programming interfaces (APIs), which encapsulates a description of the anomaly. Regardless of the input format, the execution entity can internally standardize it into a data structure containing a problem description, potentially accompanying contextual information (such as the time of occurrence and the host identifiers involved), and an implicit "requires complex investigation" label. This data structure can then serve as the input object for all subsequent processing flows.

[0028] Step 202: Using the structured knowledge pre-constructed using a language specific to the knowledge engineering field, determine the investigation plan corresponding to the problem to be investigated; Building upon step 201, this step aims to have the aforementioned implementing entity utilize structured knowledge pre-constructed using Knowledge-Domain Specific Language (KDSL or KDSL) to determine an investigation plan corresponding to the problem to be investigated. This investigation plan is determined based on different investigation methods and sequences.

[0029] The "structured knowledge pre-built using a knowledge engineering domain-specific language" is a special knowledge base stored and maintained on the server side. Essentially, it's the result of encoding domain experts' troubleshooting experience, failure modes, and operational procedures using a machine-readable and logically rigorous formal language—the knowledge engineering domain-specific language. This structured knowledge is not a simple question-and-answer pair or document index, but rather a network of interconnected components designed to enable the server to perform logical reasoning and task planning. It primarily comprises three core components: logical analysis nodes, data requirements, and tool interfaces.

[0030] Among them, the logical analysis node is the most basic reasoning unit in structured knowledge. Each node encapsulates an independent, purposeful investigation logic or judgment rule, such as "checking host network connectivity", "analyzing CPU utilization trends" or "verifying the status of a specific service port". A logical analysis node corresponds to one of the aforementioned troubleshooting methods; while the tool interface is a description of the execution capabilities attached to the logical analysis node. It defines the specifications of the specific software tools, functions, or application programming interfaces that the server can call to complete the troubleshooting task represented by the node. For example, a network diagnostic tool interface called "async_ping" or a data interface called "get_metric_data" for querying monitoring metrics. In other words, the tool interface binds abstract troubleshooting actions with specific, executable operations. The data requirements precisely describe the format, type, and structure of the input data required to execute the troubleshooting task represented by a logical analysis node, as well as the expected output data specifications after the task is executed. For example, for the "check host network connectivity" node, its data requirements may explicitly specify that the input must include the target host's IP address (string type), and the output must be a structured result containing latency and packet loss rate. The data requirements also constrain the input and output of the tool interface, ensuring the compatibility and clarity of data when flowing between logical nodes and tools.

[0031] At the practical level, when the aforementioned implementing entity needs to determine a troubleshooting plan for a specific problem, it can try to launch a plan generation engine. This engine takes the problem description as input and retrieves, matches, and logically combines data from its internally stored structured knowledge base. The process can be based on the troubleshooting capabilities represented by the logical analysis nodes and the pre-defined or dynamically derived logical relationships (such as sequence, branching, and dependencies) between nodes. For example, facing the problem of "slow service response," the knowledge base may contain multiple related logical analysis nodes such as "check network latency," "check database query performance," and "check application server load." The plan generation engine analyzes the dependencies between these nodes (e.g., usually checking the network first, then the application) and the context of the current problem, organizing them in a reasonable "troubleshooting order" to form a preliminary, step-by-step "troubleshooting plan." This plan clearly lists which troubleshooting methods need to be performed (i.e., which logical analysis nodes need to be activated) and their approximate execution order. When the aforementioned execution entity performs this scheme combination, it will simultaneously consider the data requirements of each node to confirm whether the scheme has an executable data foundation. However, at this time, it does not actually call tools or assign intelligent agents. That is, the final output of this step is a structured scheme object, which can guide the subsequent distributed investigation steps.

[0032] Furthermore, to deepen the understanding of how to construct this structured knowledge base, one possible implementation method, including but not limited to: domain experts or system administrators can use dedicated editing tools to define logical analysis nodes, declare their data requirements, and bind existing tool interfaces in a relatively natural way. The management service on the aforementioned execution entity will compile and persistently store these definitions as knowledge entries in KDSL format. In addition, the knowledge base can support version management and incremental updates, allowing the logical nodes and relationships to be continuously enriched and corrected as operational experience accumulates, enabling the troubleshooting solutions developed by the server to continuously evolve, becoming more accurate and efficient.

[0033] Step 203: Distribute each investigation method to the matching agent in the order of investigation to investigate step by step until the real cause of the abnormal situation is determined.

[0034] Building upon step 202, this step aims to transform the abstract investigation plan into specific concurrent or sequential execution tasks by the aforementioned execution entity acting as a central scheduler. A resource matching mechanism ensures that each task is undertaken by a suitable execution unit (i.e., an agent) until the final investigation goal is achieved. The matching between the investigation method and the agent is determined based on the data requirements of the corresponding task and the capabilities of different agents.

[0035] From a technical perspective, this investigation sequence is the logical thread of this step, derived from the investigation plan determined in the previous step. It defines the execution order and potential dependencies between various "investigation methods" (i.e., the specific tasks represented by logical analysis nodes). The aforementioned execution entity can maintain an execution state machine or task queue to strictly follow this sequence to schedule tasks. An intelligent agent refers to a software module or service instance with a certain degree of autonomous execution capability (e.g., built on a tool, machine learning algorithm, deep neural network, or large model as a foundation) within or managed by the server. Each intelligent agent is designed to specialize in performing a specific type of operation, such as calling a monitoring API to obtain metrics, executing a diagnostic command, or analyzing a log file of a certain format. Its "capabilities" can be formally described as a set of data structures it can process, a set of tool functions it can call, and its execution constraints. The matching process is essentially a dynamic resource scheduling and task allocation problem, with the core basis being the compatibility between the data requirements of the corresponding task and the capabilities of different intelligent agents. The aforementioned implementing entities need to compare the clearly defined input / output data structures (data requirements) associated with each investigation method in the plan with the data formats and ranges that each agent declares it can process, thereby selecting one or more agents that can meet the execution conditions of the task.

[0036] In practical terms, the execution process described in this step can include several coherent operational steps: First, the executing entity can sequentially retrieve the currently pending investigation method and its complete context description (including its data requirements) from its task queue; then, it initiates a matching query, for example, by pre-maintaining an agent capability registry, which records the capability profiles of all available agents in a structured manner (e.g., using a descriptive language or attribute key-value pairs). For example, agent A's capability description is "can execute SSH commands, with input being the host IP (string) and the command (string), and output being a text result"; agent B's capability description is "can call the Metrics API, with input being the metric name (string) and the time range, and output being time series data in JSON format." The executing entity uses the data requirements of the current investigation task as the query condition, performs matching calculations with the registry, and filters out all theoretically capable agent candidate sets. Then, the server selects one or more target agents from the candidate set according to a certain strategy (such as being the first to be idle, having the highest historical success rate, or having the lowest load). Once selected, the server generates a formatted task instruction, which encapsulates the specific operation request, necessary input parameters, and task identifier. This instruction is then "issued" to the agent via a predefined communication mechanism (such as remote procedure call, message queue, or event bus). After receiving and executing the task, the agent returns the execution result (i.e., the "troubleshooting result") to the server. The server collects and analyzes the results. If the true cause of the anomaly can be determined, the entire process ends. Otherwise, based on the solution logic and the current results, the server determines the next troubleshooting method to be executed and repeats the matching and issuing process, forming a closed loop of "step-by-step troubleshooting" until the root cause of the problem is located.

[0037] Furthermore, additional scheduling strategies can be introduced when implementing the matching and task assignment mechanism. For example, agent capability registration can be dynamic, allowing agents to proactively register or update their capability descriptions with the server upon startup or capability changes, thereby improving system scalability. In addition, when assigning tasks, besides static capability matching, the server can also incorporate real-time health status checks to avoid assigning tasks to faulty or overloaded agents, thus improving the robustness and success rate of the entire troubleshooting process.

[0038] The dynamic troubleshooting method provided in this disclosure acquires the problem to be investigated, which requires at least two investigation methods to determine the true cause. It then uses structured knowledge pre-built in a knowledge engineering-specific language to determine the corresponding investigation plan, clarifying the task planning based on different investigation methods and specific investigation sequences. The structured knowledge is presented in a combined manner through logical analysis nodes, data requirements, and tool interfaces, which transforms the abstract investigation logic into a standardized task description that can be specifically executed. Then, by distributing each investigation method in the plan to an intelligent agent based on matching data requirements and capabilities in a predetermined order, the method is executed step by step, realizing a collaborative and step-by-step investigation of complex abnormal situations until the true cause is located.

[0039] This solution transforms the dynamic investigation process, which originally relied on human experience and was difficult to systematize, into a standardized process driven by structured knowledge and executed collaboratively by multiple specialized intelligent agents. This significantly improves the coverage and problem-solving capabilities for systematically diagnosing complex problems without clear intent or requiring multi-step reasoning. At the same time, through task decomposition and precise matching of capabilities among intelligent agents, it effectively reduces the cognitive load and error risk when a single model handles a complex full-process, enhancing the reliability, interpretability, and reproducibility of the entire investigation process.

[0040] Based on the above embodiments, in order to improve the accuracy and efficiency of generating investigation solutions as much as possible, it is also possible to try to retrieve relevant knowledge for generating investigation solutions for the problem to be investigated from structured knowledge. Combining semantic and vectorization techniques, one implementation method, including but not limited to, is as follows: First, determine the target semantic vector corresponding to the problem to be investigated. Then, retrieve the target knowledge that has at least vector similarity to the target semantic vector from the structured knowledge pre-constructed according to KDSL. Then, the investigation solution corresponding to the problem to be investigated can be determined based on the target knowledge.

[0041] The solution provided in this embodiment aims to intelligently retrieve and combine investigation solutions adapted to the current problem from a pre-built structured knowledge base. The core idea is to associate and match problems described in natural language, which may have diverse wording, with formally defined structured knowledge through the medium of "semantic vectors." Here, the target semantic vector is a mathematical representation, typically mapping the textual description of the problem to be investigated to a specific point in a high-dimensional numerical vector space through a pre-trained text embedding model (e.g., a language model based on the Transformer architecture). A key characteristic of this vector is that, in this vector space, the vectors corresponding to semantically similar or topically related text fragments are spatially close (e.g., they have high similarity scores when measured by cosine similarity). Therefore, the vector can serve as a computable and comparable digital "fingerprint" of text semantics. The aforementioned execution entity needs to determine this vector, i.e., complete a forward computation process: inputting the problem text into the model and obtaining its corresponding fixed-dimensional output vector. Because structured knowledge not only preserves its inherent logical analysis nodes, data requirements, and tool interfaces during storage, each knowledge unit (e.g., each logical analysis node and its description) also has its corresponding semantic vector pre-calculated and stored using the same text embedding model. This results in the entire knowledge base forming a series of "coordinate points" in the vector space. Therefore, the retrieval process involves the server calculating the similarity between the target semantic vector and all knowledge semantic vectors in the knowledge base (e.g., calculating cosine similarity), and then selecting the Top-K most similar results based on a preset threshold to find the knowledge units (i.e., the target knowledge) that are semantically most relevant to the current problem.

[0042] In practical terms, the process executed by the aforementioned entity can be as follows: First, it calls its integrated text vectorization service or library, passing in the string representing the problem to be investigated. This service encapsulates a lightweight offline or online text embedding model specifically designed to generate high-quality sentence-level vectors. After the model finishes processing, the server holds the target semantic vector representing the current problem in its memory. Next, the retrieval process is initiated, which can be achieved using a dedicated vector retrieval index engine (e.g., an index built based on an approximate nearest neighbor search algorithm). The server submits the target semantic vector to the index engine, which quickly returns a set of knowledge entry IDs with the highest similarity scores and their scores. These entries may correspond to one or more independent logical analysis nodes, or they may correspond to certain predefined troubleshooting patterns containing multiple nodes. Finally, the aforementioned execution entity determines the troubleshooting plan corresponding to the problem to be investigated based on the target knowledge: by reading the complete structured definition of these target knowledge entries, analyzing the logical relationships between them (for example, there are data dependencies between input and output between certain nodes, or certain typical node execution orders are defined in the knowledge base), and then by parsing these relationships and the context of the current problem, integrating and sorting these discrete target knowledge, and planning an executable troubleshooting plan that includes specific "troubleshooting methods" and "troubleshooting orders". For example, for the "database access timeout" problem, the retrieved target knowledge may include three logical nodes: "check network links", "verify database connection pool status", and "analyze slow query logs". The server can assemble them into a specific troubleshooting plan in sequence according to the dependencies defined in the knowledge base (usually checking the network first, then checking the connections, and finally analyzing the logs).

[0043] Please refer to Figure 3 , Figure 3 A flowchart of a step-by-step troubleshooting method provided in this disclosure embodiment, wherein process 300 includes the following steps: Step 301: Among the various investigation methods constituting the current investigation plan, determine the current investigation method as the starting point according to the investigation order, and execute the following investigation steps: Step 302: Distribute the investigation task corresponding to the current investigation method to at least one matched agent to perform the corresponding investigation operation and obtain the current investigation result; One specific implementation method is to: distribute the investigation task corresponding to the current investigation method to at least one matched intelligent agent, and then control at least one intelligent agent to call the target external tool through the model context protocol to execute at least one corresponding investigation operation step.

[0044] The Model Context Protocol (MCP) is a predefined interface protocol used to standardize communication and data exchange between an agent (as the requester) and an external tool (as the service provider). This protocol specifies the format of requests (e.g., it must include a tool identifier, operation commands, and a list of parameters), the structure of responses (e.g., it must include a status code and result data), and possible error handling methods. The server controls the agent by issuing standardized instructions conforming to the protocol, and the agent, as a faithful client of the protocol, is responsible for assembling requests according to the protocol specifications and sending them to the correct target. An external tool refers to any software entity or service that exists outside the server environment but can be accessed through a network API, command-line interface, or specific driver, such as a standalone monitoring system data query interface, a cloud platform management API, a command-line toolset for performing remote diagnostics, or a database query service. The server and the agent are not concerned with the internal implementation of these tools, only with their exposed calling methods that conform to the Model Context Protocol.

[0045] At the practical level, the server first matches the agent and issues tasks. The instructions include a "tool invocation blueprint" encapsulated according to the protocol format. For example, for the investigation task of "checking host CPU utilization," the issued instruction will explicitly specify that an external tool named "get_cpu_metric" needs to be invoked through the model context protocol, with the required parameters: {"host_id": "server-01", "time_range": "last_5_minutes"}. After receiving this instruction, the agent's internal protocol client module is activated. This module constructs a rigorously formatted request message according to the "tool invocation blueprint." Subsequently, the agent locates the service instance of the "get_cpu_metric" tool through a pre-configured network endpoint or service discovery mechanism and sends the request. After processing the request, the external tool encapsulates the time-series CPU utilization data according to the JSON format specified by the protocol and returns it to the agent. Upon receiving the response, the agent does not perform complex business logic processing but instead sends the raw response or the slightly formatted result back to the server as the output of the "investigation operation." Throughout the process, although the server does not directly invoke the tool, it indirectly but effectively controls and supervises "at least one investigation operation step" by issuing standardized protocol instructions and receiving and listening to the results returned by the intelligent agent.

[0046] 303: In response to the inability to determine the true cause based on the current investigation results, determine the next investigation method based on the current investigation results and the current investigation plan; Step 304: Take the next investigation method as the new current investigation method, and repeat the investigation steps described in steps 302-303 until the real cause of the abnormal situation is determined based on the current investigation results.

[0047] Step 301 marks the initialization of the loop, where the execution entity parses the determined troubleshooting plan: based on the defined troubleshooting order, it logically deduces or directly reads from multiple troubleshooting methods to determine the current troubleshooting method to be executed first. This is similar to setting a starting node for a workflow engine. The loop then begins to run, and step 302 is the execution phase within the loop: the specific troubleshooting task corresponding to the current troubleshooting method (the content of which is fully defined by the associated logical analysis nodes, data requirements, and tool interfaces) is assigned to a suitable agent through a matching mechanism. The agent, as an independent execution unit, completes the actual operation (e.g., running diagnostic scripts, querying the database) and encapsulates the raw execution data into a structured current troubleshooting result, returning it to the server. Further, in step 303, the critical decision gating phase begins: at this point, the server parses and evaluates the obtained current troubleshooting result. Its core logic is to determine whether the result is sufficient to directly or indirectly "determine the true cause," which can be done based on predefined rules or models, such as whether the result contains a clear error code, whether it exceeds a certain critical threshold, or whether it completely matches the characteristics of a known fault mode. If the judgment is "yes", the loop terminates; if the judgment is "no", it means "the real cause cannot be determined based on the current investigation results", and further planning of subsequent paths is required.

[0048] The latter half of step 303, together with step 304, constitutes the loop's progression mechanism: when the true cause cannot be determined, the aforementioned executing entity must decide what to do next, i.e., "determine the next investigation method." This process requires integrating information from both the "current investigation results" and the "current investigation plan." In principle, the investigation results provide new context: they may rule out certain potential causes, thus narrowing the hypothesis space; they may also trigger new clues, requiring adjustments to the original path. Based on the logical relationships defined in the plan (e.g., branch conditions, dependencies) and dynamic reasoning rules, the server calculates the most appropriate investigation method to execute next. Subsequently, in step 304, the aforementioned executing entity updates its state, setting this newly determined "next investigation method" as the "current investigation method" for the next round of the loop. This assignment operation triggers the loop iteration, and the control flow jumps back to step 302, starting a new round of the "task issuance - result acquisition - analysis and decision-making" process. This process repeats until, in step 303 of a certain cycle, the result analyzer determines that the current investigation results are sufficient to determine the true cause. At this point, the cycle terminates, and the entire dynamic investigation process is successfully completed.

[0049] In this embodiment, after obtaining a troubleshooting plan with a clear troubleshooting order, a result-oriented dynamic execution loop is started and driven through steps 301-304. This loop process is the central logic of the server to transform the static plan into dynamic troubleshooting behavior. Its essence is a controlled, iterative "execution-evaluation-decision" loop until the predetermined termination condition is met—that is, the true cause of the abnormal situation is located.

[0050] Furthermore, the loop mechanism can be further adjusted. For example, the server can record not only the results in each iteration but also the time and resource costs consumed. This allows for cost-benefit analysis when determining the next investigation method, prioritizing high-probability, low-cost investigation paths. Additionally, the loop can be designed to support a breadth-first exploration strategy, simultaneously distributing investigation methods corresponding to multiple hypotheses to different agents for parallel execution, thereby accelerating root cause localization. The loop termination condition can also be configured, including not only determining the true cause but also conditions such as "all possible investigation methods have been exhausted" or "the cumulative time budget has been exhausted," ensuring the process always has an exit point. This more clearly defines how the aforementioned executing entities transform static knowledge into dynamic diagnostic actions through an automatically advancing closed loop, fully demonstrating their intelligent scheduling and decision-making capabilities.

[0051] In the loop described in the previous embodiment where the agent performs investigation operations and obtains results, a quality control and adaptive optimization mechanism can be further introduced. This mechanism proactively performs in-depth evaluation of each returned "investigation result" and dynamically influences the decision-making of subsequent investigation paths based on the insights generated by the evaluation, thereby improving the accuracy and intelligence level of the entire process. One implementation method, including but not limited to, is as follows: the effectiveness of the investigation results obtained from each step of the investigation operation performed by the agent is evaluated; then, investigation correction suggestions are generated based on the evaluation results; and these suggestions are used to determine the next investigation method.

[0052] The effectiveness evaluation refers to the quantitative or qualitative judgment process by which the executing entity assesses the reliability, information content, and logical rationality of the output of a single investigation operation. The "effectiveness" of this evaluation does not refer to whether the task itself was successfully executed, but rather to the value and credibility of the execution result in advancing the ultimate goal of root cause analysis. The evaluation can be based on a pre-established set of criteria or models. Its inputs are specific investigation result data and their context (such as the corresponding investigation method and input parameters), and the output is an "evaluation result." This result can be a simple Boolean value (valid / invalid), a score (e.g., a confidence score from 0 to 1), or a multi-dimensional set of labels (e.g., "data complete," "indicator anomaly," "consistent with historical patterns"). Furthermore, the knowledge sources upon which the evaluation is based are diverse, including: universally accepted correct rules within the domain (e.g., "ping successful" means normal network layer connectivity), patterns verified in historical cases, corroborating evidence from other monitoring data in the current system environment, or samples pre-labeled by domain experts. The "investigation correction suggestions" are guiding opinions on the direction of subsequent actions, generated by the server's analytical logic based on the effectiveness evaluation results. It can be a strategic input, such as: "The current result has low confidence, it is recommended to re-verify using backup tool B", "The result strongly suggests direction X, it is recommended to prioritize the subsequent investigation method Y related to direction X", "The result is abnormal but does not match any known pattern, it is recommended to record it as an unknown case and transfer it to manual review".

[0053] In practical terms, the implementation of the solution provided in this embodiment should be embedded within each round of investigation. That is, after the server receives the "current investigation result" from an agent, it will synchronously or asynchronously start the evaluation process. An independent "evaluation engine" module will be invoked. This engine loads the corresponding evaluation rule set or model and calculates the result data. For example, for an investigation result of "checking CPU utilization", the evaluation engine may perform multiple checks simultaneously: verifying whether the returned JSON data structure is complete; determining whether the CPU utilization value is within a reasonable physical range (e.g., 0%-100%); and comparing the value with historical data from the same host one minute ago to check for any sudden anomalies. The combined conclusions of these checks constitute the current evaluation result. Next, the server's "suggestion generator" module works based on the evaluation result. If the evaluation result shows that the result is credible and clearly pointed, the generated correction suggestion may be "continue according to the original solution order"; if the evaluation finds that the data is incomplete, the suggestion may be "supplement with memory metrics from the last 5 minutes to assist in the judgment"; if the evaluation believes that the result deviates significantly from expectations, the suggestion may be "trigger a check on the health status of the monitoring data source itself".

[0054] "Incorporating the suggested corrections into determining the next investigation method" means that at subsequent logical decision points (such as the step of determining the next investigation method), the server's decision algorithm will use this suggested correction as an important input variable. The decision algorithm will comprehensively consider the original investigation plan ("investigation order") and the dynamically generated suggested corrections, and may choose to follow the suggestions, partially adopt the suggestions, or arbitrate among multiple suggestions, thereby ultimately determining a better, real-time corrected "next investigation method".

[0055] Furthermore, this effectiveness evaluation system can be further designed as follows: the aforementioned implementing entity records the inputs, outputs, and final verification conclusions (whether the root cause has truly been found) of each evaluation in a log, forming training data. Then, by periodically analyzing this data, the parameters of the evaluation model can be automatically optimized or the weights of the evaluation rules can be adjusted, making the evaluation criteria increasingly accurate and forming a "data flywheel" that drives the continuous improvement of the overall system capability. In addition, the evaluation dimensions can not be limited to single-step results, but can be extended to the long-term evaluation of the agent's performance itself. For example, recording the historical effectiveness of a specific agent on a specific type of task can introduce a reputation mechanism in future task matching.

[0056] For a deeper understanding of how to determine the next screening method, please refer to [link / reference needed]. Figure 4 , Figure 4 This document provides branch diagrams illustrating two different methods for determining the next investigation method in embodiments of this disclosure. For the description of the higher-level solution for determining the next investigation method based on the current investigation result and the current investigation plan, two different implementation methods are given for two different scenarios: Scenario 1: If it is determined that the current investigation plan needs to be adjusted based on the current investigation results, the current investigation plan is adjusted according to the current investigation results to obtain the adjusted investigation plan. Among the various investigation methods that constitute the adjusted investigation plan, the next investigation method is determined according to the adjusted investigation order. Scenario 2: If it is determined from the current investigation results that no adjustment to the current investigation plan is needed, then the next investigation method shall be determined according to the current investigation order among the various investigation methods that constitute the current investigation plan.

[0057] The fundamental difference between the two scenarios lies in the server's internal judgment of the Boolean state regarding "whether adjustment is needed." This judgment can be made by the server's "solution evaluator" module, whose inputs are the "current troubleshooting results" and the "current troubleshooting solution," and whose output is a decision signal. The evaluation logic can be based on various rules: for example, whether the current result fully matches the expected output of this step in the original solution (such as returning a "normal" state). If it does, no adjustment may be needed; conversely, if the result is abnormal (such as returning "error code 404" or "CPU utilization 100%), and this abnormal pattern triggers the predefined "solution adjustment rules" in the knowledge base (e.g., "when error code 404 is detected, network proxies should be investigated first, rather than application services"), then it is determined that adjustment is needed.

[0058] In scenario one, the server determines that adjustment is needed and activates the adjustment engine. Based on new information from the current investigation results (e.g., the problem has been confirmed to be at the network layer), this engine modifies the current investigation plan in real time according to pre-defined adjustment logic (such as replacing, inserting, deleting, or reordering logical analysis nodes), generating an adjusted investigation plan that may differ from the original plan in subsequent steps. Subsequently, the server determines the next investigation method from the logical starting point of this new plan or based on its new adjusted investigation order.

[0059] In scenario two, the server determines that no adjustment is needed. This means that the current investigation results do not provide enough disruptive information to change the original plan. Therefore, the server maintains the current investigation plan and simply selects the next node of the currently executed investigation method as the next investigation method, according to the established current investigation order.

[0060] Specifically, after obtaining the current investigation results, the aforementioned executing entity can query an adjustment rule table associated with the structured knowledge base. This table defines which result patterns should trigger which adjustment operations. For example, a rule might specify "If the result of 'Check Host Connectivity' is 'Timeout,' then the adjustment plan is: insert the 'Check Firewall Rule' node in the next hop." If a rule is matched, then proceed to Case 1. The adjustment engine can obtain the internal representation of the original plan (which may be a graph structure or a task list), apply the graph transformation operations defined in the rules (such as adding nodes or changing edge connections), and generate and persist a new plan version. Then, the next node to be executed is parsed from this new version. If no adjustment rule is matched, or the stability threshold is not triggered (e.g., the confidence level of the result is higher than a certain value), then proceed to Case 2. In this case, the server only needs to maintain a pointer to the current plan execution progress; moving this pointer one position forward allows it to read the next scheduled task from the original plan's data structure.

[0061] During the dynamic investigation loop driven by the server execution entity, this embodiment defines in detail the two logical branches that the server relies on at this decision point. Essentially, these are two differentiated path selection strategies adopted by the server after dynamically evaluating the original plan (i.e., the current investigation scheme) based on the latest obtained investigation results. The aim is to give the investigation process the highest possible flexibility and adaptability through this mechanism, so that it can optimize the execution path based on real-time feedback, rather than rigidly following the initial plan.

[0062] To further understand how to assign investigation tasks to matching agents for execution, based on any of the above embodiments, please refer to [link to relevant documentation]. Figure 5 , Figure 5 A flowchart of a task assignment and execution method provided in this embodiment of the disclosure includes the following steps: Step 501: Determine the investigation task corresponding to the current investigation method; In this step, the "troubleshooting task" is an instantiation of the "troubleshooting method" within a specific context. A troubleshooting method (such as "check host connectivity") is a general logical template, while the troubleshooting task is populated with specific execution parameters (such as the target host IP address being "192.168.1.100"). The server generates a task object containing all necessary operation instructions and input parameters by parsing the logical analysis node definition associated with the current troubleshooting method and combining it with the specific context information in the session (such as the faulty machine specified by the user).

[0063] Step 502: Determine the data requirements corresponding to the investigation task, and determine the complete capability requirements based on the data requirements; The core of this step is to transform the task into capability requirements for the executor. This data requirement explicitly defines the data type and structure of the input data needed to complete the task (e.g., requiring the target host's IP address string) and also specifies the format of the task's output data (e.g., returning a JSON object containing latency and packet loss rate fields). Based on this data requirement, the executor derives the complete capability requirements needed to perform the task. This set of requirements is a structured description of the agent's capabilities; it not only includes the ability to process specific input / output data formats but may also implicitly contain constraints on the execution environment, permissions, or the type of encapsulated tools. For example, a data requirement to "execute SSH commands and parse the returned text" corresponds to a complete capability requirement that the agent possess both "SSH protocol operation capabilities" and "text parsing capabilities."

[0064] Step 503: Based on the preset mapping table, determine the set of target intelligent agents that jointly meet the requirements for complete capabilities; The mapping table records the correspondence between different data requirements and intelligent agents with different capabilities, and the target intelligent agent set includes at least one target intelligent agent.

[0065] This step is the core of resource matching, where the server accesses a "pre-defined mapping table." This mapping table is essentially a capability directory or service registry, recording all agent instances managed by the server and their declared capability attributes. The mapping relationship can be direct, such as associating the capability of "being able to call the async_ping tool" with a specific agent ID; or it can be conditional, such as declaring that a certain agent can handle a type of data requirement where "input is an IP address, output is a numerical latency".

[0066] The aforementioned executing entity uses the complete capability requirements obtained in step 502 as query conditions to search, compare, and filter in the mapping table, ultimately identifying one or more target agents that can jointly meet all capability requirements, thus forming a target agent set. This set may contain only one agent (when a single agent has comprehensive capabilities), or it may contain multiple agents that need to cooperate (for example, one responsible for data collection and another responsible for data analysis).

[0067] Step 504: Distribute the investigation task to the target intelligent entity set to perform the corresponding investigation operation and obtain the current investigation results.

[0068] This step is the final delivery of the task. The aforementioned executing entity distributes the encapsulated investigation task instructions to each member of the target agent set through internal communication mechanisms (such as message queues or remote procedure calls). After receiving the instructions, the agents independently or collaboratively execute their assigned investigation operations (such as running tools or querying interfaces), and format the output data after execution, returning it to the server as the current investigation result, thus completing the closed loop of task distribution and result collection for this round.

[0069] At the practical level, the aforementioned execution entities may read parameters from a shared session context and populate them into logical analysis node templates loaded from a knowledge base to quickly instantiate tasks. The derivation from data requirements to complete capability requirements may be achieved through a set of predefined transformation rules, which are built into the server's task planner. In actual systems, the mapping table may be manifested as a microservice registry or a dedicated capability graph database, supporting flexible queries. Matching algorithms can run on this basis, and in addition to capability matching, simple load balancing strategies may be added to select the best from multiple candidate agents that meet the capability requirements. Task distribution can usually adopt an asynchronous communication mode, starting a timeout timer after issuing the instruction and waiting for the agent's callback or polling result. For tasks that require collaboration among multiple agents, the aforementioned execution entities can also coordinate their execution order or data transfer methods.

[0070] Furthermore, this mapping table can also support dynamic updates, allowing agents to actively register or deregister when starting up, going offline, or upgrading capabilities, thus adapting to elastic deployment environments. The description of capability requirements can also try to introduce version numbers or performance indicators, so that the server can prioritize agents with updated versions or faster historical responses among those that meet basic functional requirements. In addition, for complex capability requirements that require multi-step collaboration to meet, the server can also add preliminary task decomposition capabilities, breaking down a large task into several sub-tasks, matching an agent to each sub-task, and attaching coordination instructions when issuing them, thereby realizing more complex distributed troubleshooting scenarios.

[0071] In the process of finding and assigning suitable executors to the currently selected investigation method, this embodiment details how the server transforms the abstract "investigation method" into a concrete, executable "investigation task," and precisely schedules execution resources (agents) to complete the task through a capability-based matching mechanism, thereby obtaining the crucial "current investigation result." This process ensures alignment between task requirements and executor expertise, forming the foundation for efficient and reliable collaborative investigation.

[0072] Based on any of the above embodiments, the remote desktop system provided by Virtual Network Computing (VNC) technology can also be used to present the process of gradually investigating by distributing various investigation methods to the matching intelligent agents in the order of investigation in a live broadcast.

[0073] This step, while driving the core investigation process, also provides a transparent observation channel for users. Its core lies in showcasing the process to users within a structured framework using real-time streaming technology similar to live streaming. This "live streaming" does not refer to video streaming, but rather to the server continuously feeding back the dynamic changes of its internal processes to the user interface with extremely low latency.

[0074] VNC is a graphical desktop sharing system based on the RFB (Remote Frame Buffer) protocol. Its core principle is to encode and compress changes to the framebuffer of a graphical desktop (or application window) on the server side—that is, updates to the screen pixel content—and transmit them over the network to the client. Simultaneously, it can transmit client input events (such as mouse clicks and keyboard input) back to the server. In this solution, the server does not share a real physical desktop but instead initializes a virtual graphical desktop session specifically for this investigation task. Within this session, the server can run various graphical monitoring tools, command-line terminals, or web browsers. When the server controls the agent to perform investigation operations (for example, the agent "simulates" clicking a button in a browser to query monitoring charts), these operations are actually executed in this virtual desktop environment, producing corresponding graphical interface changes. The VNC server component continuously captures these pixel changes and converts them into VNC protocol data streams. Users can connect to the specific port provided by the server using any standard VNC client software to receive this continuous desktop image stream, thus watching the entire process of the investigation task being issued to the agent in a "live" format, triggering a coherent graphical interface operation.

[0075] In practical terms, when a live-streamed investigation process is initiated, the scheduler on the server commands an isolated container (such as a Docker container) to start. This container is pre-installed with an operating system containing a graphical interface and necessary tools, and automatically runs a VNC server. Subsequently, the server's main control program establishes a control connection with the environment within this container. It sends a series of simulated operation commands to the container, driving applications in the graphical interface (e.g., automatically filling in login information and opening a network monitoring web page). All these visual feedbacks occurring within the container's virtual desktop are captured and generated as data streams by the VNC server within the container. The server itself is responsible for securely providing the access endpoint of this VNC stream (such as a WebSocket proxy address) to the front-end user interface. Users can connect to this endpoint via an integrated VNC web client or a standalone VNC viewer to view a dynamically updated remote desktop window in real time. The window clearly displays each step of the investigation process, showing the agent "operating" various graphical tools, such as the generation of graphs, the scrolling of log files, or the navigation of configuration pages.

[0076] Furthermore, this VNC-based live streaming solution can also attempt to provide an "operation highlighting" function, which overlays a transparent graphic annotation on the video stream, using conspicuous boxes or arrows to indicate the specific interface elements currently being operated by the intelligent agent, helping users to more clearly understand the operation intent. In addition, the entire VNC session can be fully recorded and timestamped with all machine-readable logs and events during the investigation process. Afterwards, users can not only replay the video but also click on any point on the video timeline to directly view the corresponding structured data (such as the precise command being executed at that time and the returned raw data), achieving deep integration and auditing of audio-visual recordings and machine data. This approach combines highly intuitive visual presentation with traceable technical details, greatly improving the transparency and credibility of complex investigation processes.

[0077] To deepen understanding, this disclosure also addresses the prior art problems mentioned in the background section by providing a specific intelligent customer service system for complex problems. Its core lies in structuring the inherent logic of complex domain knowledge into KDSL-based knowledge. Simultaneously, in engineering, KDSL is mapped to multiple agents, supplemented by MCP resource injection, tool injection, multi-agent isolation, and StepScore (a step-based scoring mechanism, i.e., the effectiveness evaluation mentioned in the above embodiments) scoring to achieve a low-illusion, traceable, and easily iterative closed loop. Then, live streaming is used to broadcast the operations implemented by the large model, improving the credibility of the content and conclusions, and enhancing the user experience.

[0078] The system has the following architectural components (modules): 1. Knowledge-DSL Layer (hereinafter referred to as KDSL) The structured semantics describing domain knowledge include Analyzer, Data, and Tool. Analyzer corresponds to a logical node that can be analyzed independently (one or more steps), containing a natural language description of the logical node, the inputs required to complete the analysis, and the conclusions that can be drawn. Analyzer-Data is used to make the focus explicit, that is, to clarify what data is needed to complete the analysis and what data conclusions can be drawn. Analyzer-Analyzer connects the next steps of analysis through the "obtainable data conclusions." Data refers to the data itself, used to clarify the inputs and outputs required by Analyzer, as well as the inputs and outputs required by Tool. Tool is responsible for informing the model of the tools that can be invoked.

[0079] This functional layer, through reasonable abstraction of knowledge in the domain of large-scale computer infrastructure, provides expert teams with a shareable and iterative knowledge persistence language framework; for models, it provides the core logic and panoramic view of domain knowledge. Compared to process abstraction, which focuses more on the implementation of fragmented knowledge rather than knowledge itself; and compared to knowledge graphs, which primarily describe the relationships between entities in objective reality, KDSL describes the logic and relationships within a knowledge domain (see [link to relevant documentation]). Figure 6-1 The structured knowledge record page shown clearly defines the input / output data structure during the analysis process in the left column, such as "IP / hostname", "ping result", and "SSH banner information". The middle column contains specific callable tool interfaces such as "async_ping", "check_ssh_banner", and "get_matrix_page_status". The logic flowchart in the right column is a visualization of the analysis logic. Each circular node represents an independent Analyzer, such as "host connectivity diagnosis" and "third-party monitoring availability check". The arrows between nodes indicate the logical flow.

[0080] 2. Multi-agent mapping

[0081] Knowledge conforming to the KDSL structure is mapped to multi-agent systems using customized rules.

[0082] Function: To provide a reasonable mapping of knowledge, which can specifically include the following points: a) Divide and conquer: Breaking down large tasks into smaller ones allows the large model to focus on the scheduling, execution, and summarization of subtasks without interference from other analytical contexts. This reduces the probability of illusions, interference, and errors in each small problem. b) Enhance the autonomy and robustness of the system. Traditional single-model systems are merely question-and-answer sessions, while multi-agent systems can autonomously propose supplementary data, autonomously decide on the next step of analysis (based on KDSL's Analyzer-Analyzer chain or choosing a new starting point based on understanding), and autonomously invoke tools for verification and error correction. This forms a self-reflective, self-playing, and self-converging capability. Compared to single-model multi-turn dialogues, it is more stable, less prone to illusions, better covers long-link analysis and reasoning, and is more capable of error correction and backtracking across multiple links.

[0083] c) Enabling knowledge to possess both "executability" and "evolutionizability." KDSL is responsible for transforming the set of processes into a set of logic. Multi-agent systems are responsible for transforming knowledge into an executable engineering structure. This engineering structure ensures the following during runtime: knowledge is executable, meaning it is no longer just part of a structural diagram, but rather an individual executable agent; and it is traceable and evolutionary, as each problem's execution generates a series of steps, which can serve as the basis for knowledge evolution.

[0084] 3. Vector Recall Layer

[0085] By vectorizing user requests, key points in KDSL knowledge, and semantics of the knowledge base, relevant knowledge content is retrieved and injected into multiple agents as prompts to provide more suggestions and improve performance.

[0086] Function: After vectorizing user requests into word segments, it suggests the relevance of certain knowledge blocks in the model. It assists in global intelligent decision-making, uncovers the possibilities between seemingly unrelated knowledge blocks, and integrates multiple knowledge blocks into atypical questions to generate new processes and improve answer performance.

[0087] 4. MCP injection

[0088] Tools are integrated into the orchestrator team in the form of MCPs (Multi-Choice Programming) for human-machine collaboration. When writing logical knowledge, the MCP can display a list of tools and resources, allowing for targeted assembly and iterative addition of missing tools as needed.

[0089] Functions: Real-time state acquisition and real-time execution of agent operations. MCP resource provides the real-time state of relevant resources. MCP tool provides execution capabilities. MCP prompt provides customized prompt word assembly. Together, they reduce the illusion caused by the model relying on only a few samples.

[0090] 5. StepScore

[0091] For each request returned to the user, a score is assigned for each step taken by the multi-agent team. The scoring criteria include comparison with existing knowledge and human evaluation. The aim is to distinguish whether the steps are meaningful, consistent with existing knowledge, and complete in the context of the problem.

[0092] Function: To score each step of the model's response. This clarifies whether the knowledge documentation is insufficiently clear or if there are issues with the model's execution and invocation. It optimizes the knowledge system and model input at each step. Based on the scoring, it creates a data flywheel between knowledge and question answering, achieving low-illusion responses to user questions with high-quality knowledge.

[0093] 6. VNC Live Streaming Layer

[0094] At the start of the task, the VNC container is initialized. Every use of tools and access to external resources is connected to the VNC container via MCP, allowing the model's real-time operations to be displayed and archived. A token is used to describe the mapping between user requests and the VNC container. Simultaneously, on the front end, changes to internal VNC operations are displayed to the user in a live broadcast manner.

[0095] Purpose: To broadcast the model's actions in real time using technological means. This reduces the need for users to verify the source of information when they have questions. It makes the answer steps more intuitive and credible. Ultimately, it enhances product usability.

[0096] Please see Figure 6-3 It details the specific components of the VNC live streaming layer, with its core components including: Chrome Container: Runs the Chrome browser, simulating human operations (such as accessing matrix pages, checking agent status); X11VNC: Converts the X11 desktop screen of the Chrome browser into a VNC stream (remote desktop protocol); Websockify Container: Converts the VNC stream into a WebSocket stream (a browser-supported protocol), presenting it to the user through a live streaming interface; ChromeControl MCP: A concrete implementation based on the MCP protocol, responsible for the interaction between LLM and Chrome (such as LLM sending "click" or "scroll" commands, and Chrome Control MCP executing the operations and returning the results); User interaction flow: The user sends a task through the control interface (such as "check for machine hotspots") → LLM generates operation commands → Chrome Control MCP controls the Chrome browser to execute operations (such as accessing matrix pages) → The operation screen is displayed to the user in real time through the live streaming interface.

[0097] The above-mentioned functional components / modules can be operated according to the following steps: 1) User initiates a session (text / alarm trigger / API); 2) VNC container initialization, token mapping into the pool; 3) The system uses vector matching to match relevant knowledge blocks based on user requests; 4) Obtain resources through MCP; 5) Initialize the multi-agent group using resources and related knowledge blocks; 6) Obtain the tool list through MCP and inject it into the multi-agent group; 7) The Orchestrator, based on past knowledge, plans the next step and assigns sub-tasks; 8) The sub-Agent executes the sub-task and provides a response, while simultaneously operating the VNC container to implement each step of the operation; 9) The Orchestrator determines the current progress, whether there is an infinite loop, and whether the user's needs have been met.

[0098] Then repeat the above steps until an answer is returned to the user.

[0099] Based on the above solutions, new interaction designs can be designed for the two roles separately: 1. Regular users / on-duty engineers: only ask questions, view steps, and accept results; 2. Customer service knowledge provider (expert team): Only need to write logic in natural language and fill the tools into the MCP template. No need to understand workflow / intent recognition / prompt stacking.

[0100] The differences and interaction methods between the solution provided in this embodiment and the prior art are explained from these two perspectives respectively: Role 1 Perspective: User Perspective (Engineer / Business User) Input method: Whether it is a web page or a client, you only need to describe the problem.

[0101] System interaction method: Visible, transparent, and traceable multi-agent step display. View the action container and verify the model's action processing.

[0102] Final output: Structured diagnostic report (see attached) Figure 6-2 The above is a troubleshooting report for a machine with hotspots.

[0103] Role Two Perspective: Customer Service Knowledge Provider (Expert Team) Perspective

[0104] You only need to write the logic in natural language (in the form of a knowledge framework in KDSL format), without having to consider how to identify intent, how to jump between workflow nodes, or how to adjust prompt words.

[0105] Because multi-agent systems perform real-time analysis by referencing the logical links of KDSL, the logic of the large executable model modules of knowledge mapping has been hidden through engineering methods. Users only need to focus on the logic itself.

[0106] Furthermore, you only need to fill the tool into the MCP server template, implement the tool's input / output, and the system will automatically inject it into the Agent for use. There is no need to write large model calls, complex tool bindings, or maintain the link.

[0107] By applying the solution provided in this embodiment, the following technical effects can be achieved compared to the prior art: 1) Eliminates problems caused by intent analysis: In customer service scenarios, a large proportion of users asking questions have no intent, making intent recognition even more confusing. This solution eliminates the confusion caused by intent analysis from the perspectives of macro-decision making and dynamic analysis. While being compatible with questions with clear intent, it is also compatible with questions without intent, thus expanding the scope of capabilities.

[0108] 2) Low illusion: The operability of KDSL's structured knowledge and the divide-and-conquer approach to the problem by multiple agents keep the illusion of the model's answer at a low level; 3) Iterable: Building on the low illusion, the multi-agent steps combined with StepScore allow each imperfect answer to feed back into knowledge iteration and model optimization, further reducing the illusion. 4) Enhanced Practicality: Building upon real-time, low-illusion Q&A, live streaming further optimizes product practicality. Users can see the model's actions in real time and understand the source of its answers, eliminating the need for secondary verification and thus improving product usability.

[0109] Further reference Figure 7 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a dynamic problem-solving device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0110] like Figure 7As shown, the dynamic problem-solving device 700 of this embodiment may include: a problem acquisition unit 701, a problem-solving plan determination unit 702, and a step-by-step problem-solving unit 703. The system includes a problem acquisition unit 701, configured to acquire problems to be investigated in response to an anomaly; each problem requires at least two investigation methods to determine the true cause of the anomaly; an investigation scheme determination unit 702, configured to use structured knowledge pre-built according to the knowledge engineering domain-specific language KDSL to determine an investigation scheme corresponding to the problem; the investigation scheme is determined based on different investigation methods and sequences; the structured knowledge includes logical analysis nodes, data requirements, and tool interfaces; logical analysis nodes correspond to investigation methods; tool interfaces describe the tools that can be called when the corresponding logical analysis node performs the corresponding investigation task; and data requirements describe the structure of the input / output data of the corresponding logical analysis node and the called tools when performing the corresponding investigation task; and a step-by-step investigation unit 703, configured to distribute each investigation method to the matching agent in the investigation order for step-by-step investigation until the true cause of the anomaly is determined; the matching between the investigation method and the agent is determined based on the data requirements of the corresponding task and the capabilities of different agents.

[0111] In this embodiment, the specific processing and technical effects of the following components in the dynamic problem investigation and handling device 700—namely, the problem acquisition unit 701, the investigation plan determination unit 702, and the step-by-step investigation unit 703—can be referred to separately. Figure 2 The relevant descriptions of steps 201-203 in the corresponding embodiments will not be repeated here.

[0112] In some other alternative implementations of this example, the step-by-step investigation unit 703 includes: The current investigation method determination and investigation subunit are configured to determine the current investigation method as the starting point according to the investigation order among the various investigation methods constituting the current investigation plan, and execute the following investigation steps: distribute the investigation task corresponding to the current investigation method to at least one matched intelligent agent to perform the corresponding investigation operation and obtain the current investigation result; in response to the inability to determine the true cause based on the current investigation result, determine the next investigation method based on the current investigation result and the current investigation plan; The troubleshooting step cyclic execution subunit is configured to take the next troubleshooting method as the new current troubleshooting method and cyclically execute the troubleshooting steps until the real cause of the abnormal situation is determined based on the current troubleshooting results.

[0113] In some other alternative implementations of this example, the step-by-step investigation unit 703 also includes: The effectiveness evaluation subunit is configured to evaluate the effectiveness of the investigation results obtained from each step of the investigation operation performed by the agent. The correction suggestion generation sub-unit is configured to generate investigation correction suggestions based on the evaluation results of the effectiveness evaluation, and to participate in determining the next investigation method.

[0114] In some other optional implementations of this example, the current investigation method determination and investigation subunit include a next investigation method determination module configured to determine the next investigation method based on the current investigation results and the current investigation plan. The next investigation method determination module is further configured to: In response to the determination that the current investigation plan needs to be adjusted based on the current investigation results, the current investigation plan is adjusted accordingly to obtain the adjusted investigation plan; Among the various investigation methods that constitute the adjusted investigation plan, the next investigation method is determined according to the adjusted investigation order.

[0115] In some other optional implementations of this example, the current investigation method determination and investigation subunit include a next investigation method determination module configured to determine the next investigation method based on the current investigation results and the current investigation plan. The next investigation method determination module is further configured to: In response to the determination that no adjustment to the current investigation plan is needed based on the current investigation results, the next investigation method is determined according to the current investigation order among the various investigation methods that constitute the current investigation plan.

[0116] In some other optional implementations of this example, the current investigation method determination and investigation subunit include a next investigation method determination module configured to determine the next investigation method based on the current investigation results and the current investigation plan. The next investigation method determination module is further configured to: Determine the investigation tasks corresponding to the current investigation methods; Identify the data requirements corresponding to the investigation tasks, and determine the complete capability requirements based on the data requirements; Based on a pre-defined mapping table, a set of target intelligent agents that collectively meet the requirements for complete capabilities is determined; wherein, the mapping table records the correspondence between different data requirements and intelligent agents with different capabilities, and the set of target intelligent agents includes at least one target intelligent agent; The investigation task is distributed to the target intelligent entity set to perform the corresponding investigation operation and obtain the current investigation results.

[0117] In some other optional implementations of this example, the current investigation method determination and investigation subunit includes a current investigation subunit configured to distribute the investigation task corresponding to the current investigation method to at least one matched agent to perform the corresponding investigation operation and obtain the current investigation result. The current investigation subunit includes: The investigation task corresponding to the current investigation method is sent to at least one matched intelligent agent; Control at least one intelligent agent to invoke the target external tool through the model context protocol to perform at least one corresponding investigation operation.

[0118] In some other alternative implementations of this example, the troubleshooting scheme determination unit 702 is further configured to: Determine the target semantic vector corresponding to the problem to be investigated; From the structured knowledge pre-constructed according to KDSL, target knowledge that has at least vector similarity to the target semantic vector is retrieved; Based on the target knowledge, determine the investigation plan corresponding to the problem to be investigated.

[0119] In some other alternative implementations of this example, the dynamic troubleshooting and handling device 700 may also include: The live streaming unit is configured to use a preset Model-View-Controller (MVC) framework to present the process of gradually investigating by distributing various investigation methods to the matching agents in the order of investigation in a live streaming format.

[0120] This embodiment exists as a device embodiment corresponding to the above method embodiment. The dynamic troubleshooting device provided in this embodiment acquires the problem to be investigated, which requires at least two investigation methods to determine the true cause, and uses structured knowledge pre-built in a language specific to the knowledge engineering field to determine the corresponding investigation plan. It clarifies the task planning based on different investigation methods and specific investigation sequences. The structured knowledge is presented in a combined manner through logical analysis nodes, data requirements, and tool interfaces, which can transform the abstract investigation logic into a standardized task description that can be specifically executed. Then, by distributing each investigation method in the plan to an intelligent agent based on data requirements and capability matching in a predetermined order, it is executed step by step, realizing collaborative and step-by-step investigation of complex abnormal situations until the true cause is located.

[0121] This device can transform the dynamic investigation process, which originally relied on human experience and was difficult to systematize, into a standardized process driven by structured knowledge and executed collaboratively by multiple specialized intelligent agents. This significantly improves the coverage and problem-solving capabilities for systematically diagnosing complex problems without clear intent or requiring multi-step reasoning. At the same time, through task decomposition and precise matching of capabilities between intelligent agents, it effectively reduces the cognitive load and error risk when a single model handles a complex full process, enhancing the reliability, interpretability, and reproducibility of the entire investigation process.

[0122] According to embodiments of this disclosure, this disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to implement the dynamic troubleshooting method described in any of the above embodiments.

[0123] According to embodiments of this disclosure, this disclosure also provides a readable storage medium storing computer instructions that enable a computer to implement the dynamic troubleshooting method described in any of the above embodiments when executed.

[0124] According to embodiments of this disclosure, this disclosure also provides a computer program product that, when executed by a processor, can implement the dynamic troubleshooting method described in any of the above embodiments.

[0125] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0126] like Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0127] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0128] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as dynamic troubleshooting methods. For example, in some embodiments, the dynamic troubleshooting method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the dynamic troubleshooting method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform dynamic troubleshooting methods by any other suitable means (e.g., by means of firmware).

[0129] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0130] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0131] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0133] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0134] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0135] According to the technical solution of this disclosure, by acquiring the problem to be investigated that requires at least two investigation methods to determine the true cause, and by using structured knowledge pre-built in a language specific to the knowledge engineering field to determine the corresponding investigation plan, the task planning based on different investigation methods and specific investigation order is clarified. The structured knowledge is presented in a combined manner through logical analysis nodes, data requirements and tool interfaces, so as to transform the abstract investigation logic into a standardized task description that can be specifically executed. Then, by distributing each investigation method in the plan to an intelligent agent based on data requirements and capability matching in a predetermined order, the collaborative and step-by-step investigation of complex abnormal situations is realized until the true cause is located. This solution transforms the dynamic investigation process, which originally relied on human experience and was difficult to systematize, into a standardized process driven by structured knowledge and executed collaboratively by multiple specialized intelligent agents. This significantly improves the coverage and problem-solving capabilities for systematically diagnosing complex problems without clear intent or requiring multi-step reasoning. At the same time, through task decomposition and precise matching of capabilities among intelligent agents, it effectively reduces the cognitive load and error risk when a single model handles a complex full-process, enhancing the reliability, interpretability, and reproducibility of the entire investigation process.

[0136] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0137] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A dynamic problem-solving method, comprising: The system acquires issues to be investigated in response to an abnormal situation; wherein the issues to be investigated require at least two investigation methods to determine the true cause of the abnormal situation. Using structured knowledge pre-constructed according to KDSL, a specialized language for knowledge engineering, a troubleshooting plan corresponding to the problem to be investigated is determined. The troubleshooting plan is determined based on different troubleshooting methods and sequences. The structured knowledge includes logical analysis nodes, data requirements, and tool interfaces. The logical analysis nodes correspond to the troubleshooting methods, the tool interfaces describe the tools that can be called when the corresponding logical analysis node performs the corresponding troubleshooting task, and the data requirements describe the structure of the input / output data of the corresponding logical analysis node and the called tools when performing the corresponding troubleshooting task. In the various investigation methods constituting the current investigation plan, the current investigation method is determined as the starting point according to the investigation order, and the following investigation steps are executed: the investigation task corresponding to the current investigation method is issued to at least one matched intelligent agent to perform the corresponding investigation operation to obtain the current investigation result; in response to the inability to determine the true cause based on the current investigation result, the next investigation method is determined based on the current investigation result and the current investigation plan; the next investigation method is taken as the new current investigation method, and the investigation steps are executed cyclically until the true cause of the abnormal situation is determined based on the obtained current investigation result; wherein, the matching between the investigation method and the intelligent agent is determined based on the data requirements of the corresponding task and the capabilities of different intelligent agents.

2. The method according to claim 1, further comprising: The effectiveness of the investigation results obtained from each step of the investigation operation performed by the intelligent agent is evaluated. Based on the evaluation results of the effectiveness evaluation, a rectification suggestion is generated, and the rectification suggestion is used to determine the next rectification method.

3. The method according to claim 1, wherein, The step of determining the next investigation method based on the current investigation results and the current investigation plan includes: In response to determining that the current investigation plan needs to be adjusted based on the current investigation results, the current investigation plan is adjusted according to the current investigation results to obtain the adjusted investigation plan; Among the various investigation methods that constitute the adjusted investigation plan, the next investigation method is determined according to the adjusted investigation order.

4. The method according to claim 1, wherein, The step of determining the next investigation method based on the current investigation results and the current investigation plan includes: In response to the determination based on the current investigation results that no adjustment to the current investigation plan is required, the next investigation method is determined according to the current investigation order among the various investigation methods constituting the current investigation plan.

5. The method according to claim 1, wherein, The step of issuing the investigation task corresponding to the current investigation method to at least one matched intelligent agent to perform the corresponding investigation operation and obtain the current investigation result includes: Determine the investigation task corresponding to the current investigation method; Determine the data requirements corresponding to the investigation task, and determine the complete capability requirements based on the data requirements; According to a preset mapping table, a set of target intelligent agents that jointly meet the complete capability requirements is determined; wherein, the mapping table records the correspondence between different data requirements and intelligent agents with different capabilities, and the set of target intelligent agents includes at least one target intelligent agent; The investigation task is sent to the target intelligent entity set to perform the corresponding investigation operation and obtain the current investigation result.

6. The method according to claim 1, wherein, The step of issuing the investigation task corresponding to the current investigation method to at least one matched intelligent agent to perform the corresponding investigation operation and obtain the current investigation result includes: The investigation task corresponding to the current investigation method is sent to at least one matched intelligent agent; The control system calls the target external tool through the model context protocol to perform at least one corresponding investigation operation.

7. The method according to claim 1, wherein, The step of determining the investigation plan corresponding to the problem to be investigated by utilizing structured knowledge pre-constructed according to the knowledge engineering domain-specific language KDSL includes: Determine the target semantic vector corresponding to the problem to be investigated; From the structured knowledge pre-constructed according to the KDSL, target knowledge that has at least vector similarity to the target semantic vector is retrieved; Based on the target knowledge, determine the investigation plan corresponding to the problem to be investigated.

8. The method according to any one of claims 1-7, further comprising: The process of gradually investigating by distributing the various investigation methods to the matched intelligent agents in the order of investigation using a remote desktop system provided by virtual network computing technology is presented in the form of a live broadcast.

9. A dynamic problem-solving device, comprising: The problem to be investigated unit is configured to acquire problems to be investigated in response to an abnormal situation; wherein, the problem to be investigated needs to be determined by at least two investigation methods to determine the true cause of the abnormal situation; The investigation plan determination unit is configured to use structured knowledge pre-constructed according to the knowledge engineering domain-specific language KDSL to determine the investigation plan corresponding to the problem to be investigated; wherein, the investigation plan is determined based on different investigation methods and investigation order, the structured knowledge includes logical analysis nodes, data requirements and tool interfaces, the logical analysis nodes correspond to the investigation methods, the tool interfaces are used to describe the tools that can be called when the corresponding logical analysis node performs the corresponding investigation task, and the data requirements are used to describe the structure of the input / output data of the corresponding logical analysis node and the called tools when performing the corresponding investigation task; A step-by-step investigation unit is configured to determine the current investigation method as the starting point according to the investigation order among the various investigation methods constituting the current investigation plan, and execute the following investigation steps: distribute the investigation task corresponding to the current investigation method to at least one matched agent to perform the corresponding investigation operation, and obtain the current investigation result; in response to the inability to determine the true cause based on the current investigation result, determine the next investigation method based on the current investigation result and the current investigation plan; use the next investigation method as the new current investigation method, and repeatedly execute the investigation steps until the true cause of the abnormal situation is determined based on the obtained current investigation result; wherein, the matching between the investigation method and the agent is determined based on the data requirements of the corresponding task and the capabilities of different agents.

10. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the dynamic troubleshooting method according to any one of claims 1-8.

11. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the dynamic troubleshooting method according to any one of claims 1-8.

12. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the dynamic troubleshooting method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Problem checking method and device, electronic equipment and storage medium

    CN115222371A

  • Database anomaly diagnosis method and system based on large language model

    CN120104385A