Data center moving ring fault root cause tracing method and system

By constructing a multimodal unified view and multi-type intelligent agent collaboration, combined with large language models and digital twin technology, the problems of alarm silos and low root cause tracing efficiency in data center environmental monitoring systems are solved, achieving efficient and secure fault root cause localization and repair plan generation.

CN121743087APending Publication Date: 2026-03-27INST OF AUTOMATION CHINESE ACAD OF SCI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511626076.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing data center environmental monitoring systems suffer from problems such as isolated alarms, low efficiency in root cause tracing, insufficient knowledge accumulation, and a single intelligent operation and maintenance agent, resulting in low operation and maintenance efficiency and high security risks.

Method used

By constructing a multimodal unified view of the raw data stream, multiple types of intelligent agents are used to perform operations such as abnormal event summary generation, directed acyclic graph diagnostic workflow, fault root cause identification, and root cause pre-implementation verification. Combined with domain knowledge graph and large language model fine-tuning mechanism, a fault repair plan is generated and pre-implemented in a digital twin environment.

Benefits of technology

It enables cross-domain collaborative root cause tracing, improves operational efficiency, provides a comprehensive and consistent data foundation, ensures the effectiveness and security of fault repair plans, and avoids the risks associated with direct application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121743087A_ABST
    Figure CN121743087A_ABST
Patent Text Reader

Abstract

The invention provides a data center moving ring fault root cause tracing method and system, which can be applied to the technical field of data center infrastructure intelligent operation and maintenance. The data center moving ring fault root cause tracing method comprises the following steps: constructing an original data stream with a multi-modal unified view by utilizing a configuration management data stream and a moving ring data stream stored in a time sequence data lake; based on the original data stream, executing an abnormal event abstract generation operation, a diagnosis workflow generation operation of a directed acyclic graph, a fault root cause confirmation operation and a root cause rehearsal verification operation by calling a multi-type agent to obtain a fault repair plan scheme; optimizing a fault root cause confirmation operation and a root cause rehearsal verification operation by utilizing a domain knowledge graph and a large language model fine tuning mechanism; and executing a digital twin rehearsal operation on the fault recovery plan scheme, and executing an issuing operation and a storage operation on a control script of the fault recovery plan scheme based on a result of the digital twin rehearsal operation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data center infrastructure intelligent operation and maintenance, in particular to a data center dynamic environment fault root cause tracing method and system. BACKGROUND

[0002] The data center dynamic environment (power environment) monitoring system is a core intelligent management platform for ensuring stable operation of the data center. By monitoring power equipment and environmental parameters in real time, the data center dynamic environment monitoring system realizes unattended operation and efficient operation and maintenance of the computer room. The data center dynamic environment monitoring system usually includes four subsystems of power supply, heating and ventilation, security and IT (Information Technology). However, the existing data center dynamic environment monitoring system has the following problems: alarm island, inefficient root cause positioning, insufficient knowledge sedimentation and single operation and maintenance intelligent agent. SUMMARY

[0003] In view of the above problems, the present application provides a data center dynamic environment fault root cause tracing method and system.

[0004] According to a first aspect of the present application, a data center dynamic environment fault root cause tracing method is provided, comprising: collecting dynamic environment data streams of the data center in real time, and constructing original data streams with a multi-modal unified view by using configuration management data streams and dynamic environment data streams stored in a time series data lake; based on the original data streams, calling multiple types of intelligent agents to execute abnormal event summary generation operations, directed acyclic graph diagnosis workflow generation operations, fault root cause confirmation operations and root cause pre-play verification operations, to obtain a fault repair plan scheme; using a domain knowledge graph and a large language model fine-tuning mechanism to optimize the fault root cause confirmation operations and the root cause pre-play verification operations; performing a digital twin pre-play operation on the fault repair plan scheme, and based on the result of the digital twin pre-play operation, performing a control script execution issuing operation and a storage operation on the fault repair plan scheme.

[0005] According to the embodiment of the application, the real-time acquisition of the dynamic environment data stream of the data center and the construction of the original data stream with the multi-modal unified view by using the configuration management data stream and the dynamic environment data stream stored in the time series data lake include: real-time acquisition of the dynamic environment data stream of the data center through a multi-type communication protocol, wherein the dynamic environment data stream includes power data stream, heating data stream, security information stream and software and hardware information stream, and the multi-type communication includes Internet of Things communication protocol, industrial automation communication protocol, battery management protocol and network management protocol; writing the dynamic environment data stream into the time series data lake, and synchronously acquiring the configuration management data stream associated with the dynamic environment data stream in the process of real-time acquisition of the dynamic environment data stream, wherein the configuration management data stream includes topology information, configuration information and historical alarm information of the management database; based on the unified metadata specification, performing real-time association operation, data fusion operation and visualization mapping operation on the dynamic environment data stream and the configuration management data stream to obtain the original data stream with the multi-modal unified view.

[0006] According to the embodiment of the application, the above-mentioned fault repair plan scheme is obtained based on the original data stream by calling multi-type agents to perform abnormal event summary generation operation, directed acyclic graph diagnosis workflow generation operation, fault root cause confirmation operation and root cause pre-play verification operation, which includes: performing data stream compression operation on the original data stream by using a perception agent, and performing abnormal event detection operation on the compressed original data stream based on an unsupervised anomaly detection mechanism to generate an abnormal event summary, wherein the abnormal event summary includes time information, event object, event feature information and event abnormality degree.

[0007] According to the embodiment of the application, the above-mentioned fault repair plan scheme is obtained based on the original data stream by calling multi-type agents to perform abnormal event summary generation operation, directed acyclic graph diagnosis workflow generation operation, fault root cause confirmation operation and root cause pre-play verification operation, which further includes: based on the thinking chain technology, performing big language model-based analysis on the abnormal event summary by using an arrangement agent to obtain a candidate fault hypothesis set, wherein the big language model performs real-time parameter fine-tuning through a low-rank adaptive fine-tuning mechanism; dynamically instantiating multi-modal expert agents by using the candidate fault hypothesis set, hypothesis activity and field dependence, and dynamically generating a directed acyclic graph diagnosis workflow by using the multi-modal expert agents.

[0008] According to the embodiment of the present application, the above-mentioned fault repair plan scheme is obtained by calling multiple types of agents to perform the abnormal event summary generation operation, the directed acyclic graph-based diagnostic workflow generation operation, the fault root cause confirmation operation, and the root cause pre-play verification operation based on the original data stream, and further includes: based on the domain knowledge graph, performing extension processing on the local knowledge graph by using the multi-modal expert agent, and performing simulation operation on the multi-modal system of the data center based on the extended local knowledge graph, wherein the multi-modal expert agent corresponds to the multi-modal system one by one; during the simulation operation, the fault root cause set of the data center and the confidence of the fault root cause set are diagnosed in parallel by using the multi-modal expert agent based on the directed acyclic graph-based diagnostic workflow; based on the multi-agent consensus mechanism, the fault root cause set and the confidence of the fault root cause set are interactively processed by using the orchestration agent and the multi-modal expert agent to obtain the fault target root cause.

[0009] According to the embodiment of the present application, the above-mentioned interactive processing of the fault root cause set and the confidence of the fault root cause set by using the orchestration agent and the multi-modal expert agent based on the multi-agent consensus mechanism to obtain the fault target root cause includes: based on the confidence of the fault root cause set, voting the fault root cause set based on the local consensus mechanism by using the multi-modal expert agent to obtain the fault initial root cause set; optimizing the low-rank adaptive fine-tuning mechanism and the extended local knowledge graph by using the fault initial root cause set stored in the experience replay pool; updating the parameters of the large language model by using the optimized low-rank adaptive fine-tuning mechanism, and optimizing the simulation operation of the multi-modal expert agent by using the optimized local knowledge graph; optimizing the orchestration agent by using the large language model after parameter updating, and performing the Bayesian aggregation operation on the fault initial root cause set by using the optimized orchestration agent to obtain the fault target root cause.

[0010] According to the embodiment of the present application, the above-mentioned fault repair plan scheme is obtained by calling multiple types of agents to perform the abnormal event summary generation operation, the directed acyclic graph-based diagnostic workflow generation operation, the fault root cause confirmation operation, and the root cause pre-play verification operation based on the original data stream, and further includes: generating an initial fault repair plan scheme based on the fault target root cause obtained by the fault root cause confirmation operation, and performing the root cause pre-play verification operation on the fault repair plan scheme to obtain the fault repair plan scheme.

[0011] According to the embodiment of the present application, the digital twin pre-visualization operation is performed on the fault repair plan scheme, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-visualization operation, and the control script of the fault repair plan scheme is executed based on the result of the fault repair plan scheme.

[0012] According to a second aspect of the present application, a data center dynamic environment fault root cause tracing system is provided, comprising: a data acquisition module for acquiring dynamic environment data streams of a data center in real time, and constructing original data streams with a multi-modal unified view by using configuration management data streams and dynamic environment data streams stored in a time series data lake; a multi-type agent module for performing abnormal event summary generation operation, directed acyclic graph diagnosis workflow generation operation, fault root cause confirmation operation and root cause pre-visualization verification operation by calling multi-type agents based on the original data streams, to obtain a fault repair plan scheme; a domain knowledge base module for optimizing the fault root cause confirmation operation and the root cause pre-visualization verification operation by using a domain knowledge graph and a large language model fine-tuning mechanism; a dynamic environment digital twin module for performing a digital twin pre-visualization operation on the fault repair plan scheme, and executing a control script of the fault repair plan scheme based on the result of the digital twin pre-visualization operation.

[0013] According to the embodiment of the present application, the neighborhood knowledge base module comprises a dynamic environment fault tree unit, a treatment rule unit and an experience replay pool; the dynamic environment digital twin module comprises a state coding unit, a state evolution unit and a closed-loop verification unit; and the multi-type agent module comprises a perception agent unit, an arrangement agent unit, a multi-modal expert agent unit and an action and verification unit.

[0014] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method.

[0015] The fourth aspect of the present application also provides a computer readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions are executed by a processor to implement the steps of the method.

[0016] The data center dynamic environment fault root cause tracing method provided by the application integrates real-time dynamic environment data streams and configuration management data streams, constructs a unified data view covering multiple dimensions and multiple modes of the dynamic environment of the data center, solves the data island problem faced in the traditional data center dynamic environment operation and maintenance process, and provides a comprehensive and consistent data basis for root cause analysis and tracing; meanwhile, the data center dynamic environment fault root cause tracing method provided by the application adopts a multi-type intelligent agent division of labor and cooperation mode, greatly improves the efficiency of root cause tracing and positioning in a closed-loop processing mode; in addition, the data center dynamic environment fault root cause tracing method provided by the application combines a large language model fine-tuning mechanism, performs semantic enhancement and logic optimization on the root cause analysis and tracing process, and preforms the obtained fault repair plan scheme in a digital twin environment, ensures the effectiveness of the control script, and avoids the security risks brought by directly applying the fault repair plan scheme to the dynamic environment of the data center. BRIEF DESCRIPTION OF DRAWINGS

[0017] The above and other objects, features and advantages of the present application will become more apparent from the following description of embodiments of the present application, taken in conjunction with the accompanying drawings, in which:

[0018] Figure 1 An application scenario diagram of the data center dynamic environment fault root cause tracing method according to an embodiment of the application is schematically shown.

[0019] Figure 2 A flowchart of the data center dynamic environment fault root cause tracing method according to an embodiment of the application is schematically shown.

[0020] Figure 3 A process diagram of the data center dynamic environment fault root cause tracing method according to an embodiment of the application is schematically shown.

[0021] Figure 4 A workflow diagram of the orchestration intelligent agent according to an embodiment of the application is schematically shown.

[0022] Figure 5 A collaboration process between the multi-type intelligent agents based on the consensus mechanism according to an embodiment of the application is schematically shown.

[0023] Figure 6 A structural diagram of the data center dynamic environment fault root cause tracing system according to an embodiment of the application is schematically shown.

[0024] Figure 7 A block diagram of an electronic device suitable for implementing the data center dynamic environment fault root cause tracing method according to an embodiment of the application is schematically shown. DETAILED DESCRIPTION

[0025] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. It is to be understood, however, that such descriptions are merely exemplary of the application and are intended to provide an overview or framework for understanding the nature and character of the application as it is claimed. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the application. It will be apparent, however, that one or more embodiments can be practiced without

[0026] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the term "includes" and tautological equivalents thereof, means that the named feature, step, operation, and / or component is present, but not excluding the presence or addition of one or more other features, steps, operations, or components.

[0027] All terms used herein including technical and scientific terms have the same meanings as commonly understood by one of ordinary skill in the art unless otherwise defined herein. It should be noted that the terms used herein are merely specific embodiments proposed by the inventor for the purpose of description of the present application and are not intended to limit the present application. Therefore, it is intended that the present application cover modifications and variations of this application provided they include the essential features of the present application.

[0028] In the case of using expressions similar to "at least one of A, B, and C", it is generally to be interpreted as including one or more of the items enumerated in the list (e.g., "a system having at least one of A, B, and C" should be interpreted to include a system having A alone, a system having B alone, a system having C alone, a system having both A and B together, a system having both A and C together, a system having both B and C together, and also a system having all of the A, B, and C together, etc.).

[0029] The existing data center dynamic environment monitoring system has the following problems: (1) alarm island, each subsystem in the existing data center dynamic environment monitoring system independently alarms, lacks cross-domain correlation analysis capability, and is easy to cause "alarm storm"; (2) low efficiency of root cause tracing (or positioning), the existing data center dynamic environment monitoring system relies on manual troubleshooting, which is time-consuming and has high business continuity risk; (3) lack of knowledge deposition, the expert experience stored in the existing data center dynamic environment monitoring system is usually stored in the form of unstructured documents, and there is no way to be reused by the data center dynamic environment monitoring system; (4) single operation and maintenance intelligent agent (AIops, Artificial intelligence for IT operations), the existing data center dynamic environment monitoring system.

[0030] In order to at least solve one of the above problems, it is necessary to provide a data center dynamic environment root cause tracing method and system supporting cross-domain cooperation, having explainability, supporting continuous learning, and being able to realize rapid root cause tracing.

[0031] Figure 1 The diagram illustrates an application scenario of the data center environmental fault root cause tracing method according to an embodiment of the present invention.

[0032] like Figure 1 As shown, application scenario 100 according to this embodiment may include a data center environmental monitoring system operation and maintenance scenario. Network 104 is used as a medium to provide a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0033] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0034] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0035] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0036] It should be noted that the data center environmental fault root cause tracing method provided in this embodiment of the invention can generally be executed by server 105. Correspondingly, the data center environmental fault root cause tracing system provided in this embodiment of the invention can generally be set up in server 105. The data center environmental fault root cause tracing method provided in this embodiment of the invention can also be executed by a server or server cluster that is different from server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the data center environmental fault root cause tracing system provided in this embodiment of the invention can also be set up in a server or server cluster that is different from server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0037] It should be understood that Figure 1 the number of terminal devices, networks and servers in the above-mentioned scenarios is only illustrative. Any number of terminal devices, networks and servers can be provided according to implementation needs.

[0038] The following will describe the data center dynamic environment fault root cause tracing method according to the Figure 1 scenario described above, by Figures 2-5 detailed description of the data center dynamic environment fault root cause tracing method of the disclosed embodiment.

[0039] Figure 2 The flowchart of the data center dynamic environment fault root cause tracing method according to the embodiment of the application is schematically shown.

[0040] As Figure 2 shown, the data center dynamic environment fault root cause tracing method of this embodiment includes operations S210-S240.

[0041] In operation S210, dynamic environment data streams of the data center are collected in real time, and the configuration management data stream and the dynamic environment data stream stored in the time series data lake are used to construct an original data stream with a multi-modal unified view.

[0042] The dynamic environment data stream includes data streams of different modalities and dimensions of multiple subsystems of the data center dynamic environment monitoring system; the configuration management data stream is a data stream associated with the dynamic environment data stream. The dynamic environment data stream collected in real time is stored in the time series data lake, facilitating subsequent root cause tracing or data analysis.

[0043] In operation S220, based on the original data stream, the abnormal event summary generation operation, the directed acyclic graph diagnosis workflow generation operation, the fault root cause confirmation operation and the root cause pre-play verification operation are performed by calling multi-type agents to obtain a fault repair plan scheme.

[0044] In the process of operation S220, a large model (such as a pre-trained large language model) is used for root cause tracing and fault repair plan scheme acquisition; the directed acyclic graph diagnosis workflow can clearly show the root cause tracing path and the diagnosis path, realizing step-by-step positioning of the software and hardware fault root cause.

[0045] In operation S230, the fault root cause confirmation operation and the root cause pre-play verification operation are optimized by using the domain knowledge graph and the large language model fine-tuning mechanism.

[0046] Using the domain knowledge graph and the large language model fine-tuning mechanism can improve the efficiency and accuracy of root cause tracing.

[0047] In operation S240, a digital twin pre-performance operation is performed on the fault repair plan scheme, and a control script of the fault repair plan scheme is executed based on the result of the digital twin pre-performance operation.

[0048] Digital twin pre-performance refers to a process of simulating, testing and verifying an actual running scene in a digital space by constructing a virtual model of a physical entity. In the field of data center dynamic environment monitoring system operation and maintenance technology, it refers to using digital twin technology to virtually test the fault repair plan scheme to ensure operation safety and effectiveness.

[0049] Through the digital twin pre-performance operation, the repair plan scheme can be pre-performed to prevent safety risks caused by directly applying the repair plan scheme. The control script meeting the relevant conditions is stored to realize the closed loop of dynamic environment root cause tracing.

[0050] The data center dynamic environment fault root cause tracing method provided by the application integrates dynamic environment real-time data flow and configuration management data flow to construct a unified data view covering multiple dimensions and multiple modes of data center dynamic environment, solves the data island problem in the traditional data center dynamic environment operation process, and provides a comprehensive and consistent data basis for root cause analysis and tracing. At the same time, the data center dynamic environment fault root cause tracing method provided by the application adopts a multi-type intelligent agent division of labor and cooperation mode, greatly improves the efficiency of root cause tracing and positioning in a closed loop processing mode. In addition, the data center dynamic environment fault root cause tracing method provided by the application combines a large language model fine-tuning mechanism to perform semantic enhancement and logical optimization on the root cause analysis and tracing process, and pre-performs the obtained fault repair plan scheme in a digital twin environment to ensure the effectiveness of the control script and avoid safety risks caused by directly applying the fault repair plan scheme to the data center dynamic environment.

[0051] The specific embodiments will be described below with reference to the accompanying drawings. Figure 3 The data center dynamic environment fault root cause tracing method provided by the application will be described in further detail.

[0052] Figure 3 A data center dynamic environment fault root cause tracing process diagram according to an embodiment of the application is schematically shown.

[0053] As Figure 3As shown, the real-time data of the dynamic environment is detected for anomaly by the perception agent, an event summary is formed, a large language model (LLM) is called by the arrangement agent to analyze the event summary to form a fault hypothesis list, the field expert agent is instantiated by using the fault hypothesis list to generate an agent DAG (Directed Acyclic Graph) workflow (i.e., a directed acyclic graph diagnostic workflow), the agent is arranged to perform a Bayesian root cause aggregation operation by parallel diagnosis and using confidence to obtain a fault repair plan, the obtained fault repair plan is preformed by digital twinning, if the digital twinning preformation is passed, an instruction (or control script) for controlling the fault repair plan is issued, the LLM is fine-tuned by using the issued control script, the fine-tuned LLM is put into an experience replay pool to form a closed loop of data center dynamic environment root cause tracing.

[0054] According to the embodiment of the application, the above real-time acquisition of the dynamic environment data stream of the data center, and the construction of the original data stream with a multi-modal unified view by using the configuration management data stream and the dynamic environment data stream stored in the time series data lake include: real-time acquisition of the dynamic environment data stream of the data center through multiple types of communication protocols, wherein the dynamic environment data stream includes power data stream, heating data stream, security information stream and software and hardware information stream, and the multiple types of communication include Internet of Things communication protocol, industrial automation communication protocol, battery management protocol and network management protocol; writing the dynamic environment data stream into the time series data lake, and synchronously acquiring the configuration management data stream associated with the dynamic environment data stream in the process of real-time acquisition of the dynamic environment data stream, wherein the configuration management data stream includes topology information, configuration information and historical alarm information of the management database; based on a unified metadata specification, performing real-time association operation, data fusion operation and visual mapping operation on the dynamic environment data stream and the configuration management data stream to obtain the original data stream with a multi-modal unified view.

[0055] The acquisition process of the original data stream with a multi-modal unified view provided by the embodiment of the application will be further described in detail through the specific implementation.

[0056] Real-time collection of dynamic environment data (for example, temperature, humidity, equipment power, water leakage, etc.) through IoT (Internet of Things), Modbus (an open communication protocol in the field of industrial automation), SNMP (Simple Network Management Protocol), BMS (Battery Management System), etc. multiple types of communication protocols, forming a dynamic environment data stream and writing the collected dynamic environment data stream into a time series data lake; at the same time, synchronously accessing a configuration management database (CMDB, Configuration Management Database) to obtain related configuration management data streams, such as database topology information, configuration information and historical alarm information; based on the dynamic environment data stream and the configuration management data stream, an original data stream of a multi-modal unified view is constructed.

[0057] The above embodiments or specific embodiments construct a unified data view of the dynamic environment monitoring system of the data center through multi-protocol fusion and multi-modal data association, support the collection of power, heating, security, etc. dynamic environment data stream of heterogeneous protocols such as Internet of Things and industrial automation, and simultaneously associate topological configuration, historical alarm, etc. management data, realize the deep binding of physical equipment and operation and maintenance information; adopt time series data lake to store dynamic environment data stream, support millisecond-level real-time collection and long-term historical data backtracking. Combined with the dynamic update of the configuration management data stream, the fault root cause can be quickly located; in addition, the multi-modal unified view provides high-quality data input for intelligent operation and maintenance analysis.

[0058] According to the embodiments of the present application, the above based on the original data stream, by calling multiple types of intelligent agents to execute abnormal event summary generation operation, directed acyclic graph diagnosis workflow generation operation, fault root cause confirmation operation and root cause pre-play verification operation, the fault repair plan scheme includes: using a perception intelligent agent to execute a data stream compression operation on the original data stream, and based on an unsupervised anomaly detection mechanism, executing an abnormal event detection operation on the compressed original data stream to generate an abnormal event summary, wherein the abnormal event summary includes time information, event object, event feature information and event abnormality degree.

[0059] The above embodiments relate to the abnormal event summary generation operation of the perception intelligent agent (SA, Sense Agent), and the following will make a further detailed description of the above abnormal event summary generation operation through a specific embodiment.

[0060] The perception agent compresses the original data stream based on unsupervised anomaly event detection (VAE (Variational Auto-Encoder) + STL (Seasonal-Trend decomposition procedure based on Loess) decomposition); and outputs an abnormal event summary with high readability and high interpretability, for example, E (abnormal event summary) = [time information, event object, event feature information (or event phenomenon, event representation), event anomaly degree].

[0061] The perception agent is used for compressing the original data stream and performing unsupervised anomaly detection, to automatically generate an abnormal event summary containing key information such as time, object, feature and anomaly degree, so that the workload of manually screening data is reduced, and the timeliness of anomaly identification is improved.

[0062] According to the embodiment of the present application, the above-mentioned based on the original data stream, by calling multiple types of agents to perform abnormal event summary generation operation, directed acyclic graph diagnosis workflow generation operation, fault root cause confirmation operation and root cause pre-play verification operation, the fault repair plan scheme further comprises: based on the thinking chain technology, using the arrangement agent to analyze the abnormal event summary based on the large language model, to obtain a candidate fault hypothesis set, wherein the large language model performs real-time parameter fine-tuning through a low-rank adaptive fine-tuning mechanism; using the candidate fault hypothesis set, hypothesis activity and field dependence to dynamically instantiate a multi-modal expert agent, and using the multi-modal expert agent to dynamically generate a directed acyclic graph diagnosis workflow.

[0063] The specific embodiments will be described below with reference to the accompanying drawings. Figure 4 The arrangement agent generates a directed acyclic graph diagnosis workflow is further described in detail.

[0064] Figure 4 The workflow diagram of the arrangement agent according to the embodiment of the present application is schematically shown.

[0065] The arrangement agent (OA, Orchestrate Agent) uses the thinking chain (CoT, Chain of Thought) prompt technology to analyze the abnormal event summary, to generate a candidate fault hypothesis set H, for example, H = {h1…h n}h i represents the i th candidate fault hypothesis; dynamically instantiates a field expert agent according to the hypothesis activity (a) and the field dependence (b); generates a DAG diagnosis workflow W = <T, E> (T: diagnosis / simulation / forensic task node; E: data node dependent edge).

[0066] AsFigure 4 As shown, the orchestration agent receives an event summary (i.e., an abnormal event summary), parses it using an LLM to obtain a candidate fault hypothesis set (i.e., hypotheses h1, h2, h3 in Figure 4 ), and then processes each fault hypothesis to obtain a corresponding task name and confidence; the orchestration agent performs a Bayesian aggregation operation based on the above task name and confidence to confirm the root cause and generate an interpretable DAR workflow script (i.e., a DAR diagnosis workflow) based on the determined root cause.

[0067] The above embodiment uses the thinking chain technology of a large language model (LLM) to deeply analyze an abnormal event summary and generate a candidate fault hypothesis set. Through a low-rank adaptive (LoRA) fine-tuning mechanism, the model can adapt to dynamic environment field terms in real time; according to the activity (e.g., frequency of occurrence) and field dependence (e.g., power subsystem correlation) of the fault hypothesis, multiple modal expert agents (e.g., power expert agents, etc.) are dynamically instantiated. These expert agents cooperate to generate a DAG diagnosis workflow, enabling cross-domain root cause positioning; in addition, the above embodiment can continuously optimize the diagnosis path through a real-time fine-tuning mechanism.

[0068] According to an embodiment of the present application, the above fault repair plan scheme is obtained based on the original data stream by calling multiple types of agents to perform abnormal event summary generation, directed acyclic graph diagnosis workflow generation, fault root cause confirmation, and root cause pre-play verification operations, which further includes: based on the domain knowledge graph, using a multi-modal expert agent to perform extension processing on the local knowledge graph, and based on the extended local knowledge graph, performing simulation operation on the multi-modal system of the data center, wherein the multi-modal expert agent corresponds to the multi-modal system one by one; during the execution of the simulation operation, based on the directed acyclic graph diagnosis workflow, using the multi-modal expert agent to diagnose the fault root cause set of the data center and the confidence of the fault root cause set in parallel; based on the multi-agent consensus mechanism, using the orchestration agent and the multi-modal expert agent to interactively process the fault root cause set and the confidence of the fault root cause set to obtain the fault target root cause.

[0069] According to the embodiment of the present application, the above-mentioned multi-agent consensus mechanism is used to interactively process the fault root cause set and the confidence of the fault root cause set by the scheduling agent and the multi-modal expert agent, and obtain the fault target root cause, including: based on the confidence of the fault root cause set, using the multi-modal expert agent to vote on the fault root cause set based on the local consensus mechanism, to obtain the fault initial root cause set; using the fault initial root cause set stored in the experience replay pool to optimize the low-rank adaptive fine-tuning mechanism and the extended local knowledge graph; using the optimized low-rank adaptive fine-tuning mechanism to update the parameters of the large language model, and using the optimized local knowledge graph to optimize the simulation operation of the multi-modal expert agent; using the large language model after parameter updating to optimize the scheduling agent, and using the optimized scheduling agent to perform Bayesian aggregation operation on the fault initial root cause set to obtain the fault target root cause.

[0070] The above-mentioned embodiment relates to the consensus collaboration process between the multi-modal expert agent (or domain expert agent) EA (Expert Agent) and the scheduling agent. The following will be described in detail through specific embodiments and in combination with the accompanying drawings Figure 5 The above-mentioned multi-modal expert agent and consensus mechanism provided by the present application are further described in detail.

[0071] According to the subsystems of the data center dynamic monitoring system, i.e. the power subsystem, the heating and ventilation subsystem, the security subsystem, and the IT subsystem, four types of domain expert agents are encapsulated to form multi-modal expert agents. In the encapsulation process, the local knowledge graph (such as device topology information, fault tree information, and fault handling rule information) is used, and a lightweight simulator (such as a power flow simulator, a CFD (Computational Fluid Dynamics) thermal field simulator, and a network connectivity simulator) is constructed in each domain expert agent, and a confidence calculation engine (such as Conf(h i ), the i-th confidence calculation engine) is set. The multi-modal expert agents achieve state consensus through the Raft-Lite (a local consensus protocol) protocol, ensuring that the network and physical partitions of different subsystems can still confirm the root cause through voting.

[0072] Figure 5 The collaboration process between the multi-type agents based on the consensus mechanism according to the embodiment of the present application is schematically shown.

[0073] As shown in FIG. 1, the collaboration process between the multi-type agents based on the consensus mechanism according to the embodiment of the present application is schematically shown. Figure 5 As shown in FIG. 1, the collaboration process between the multi-type agents based on the consensus mechanism according to the embodiment of the present application is schematically shown. Figure 5The local knowledge graph is extended by using the automatically expanded knowledge graph (EA1, EA2, and EA3 shown in FIG. 1) ; then, a root cause is selected through voting based on a Raft-Lite consistency voting module, and a consensus is reached; then, the selected root cause is put into an experience replay pool; the experience replay pool optimizes a LoRA (Low-Rank Adaptation) incremental fine-tuning mechanism, updates LLM parameters by using the optimized LoRA incremental fine-tuning mechanism, optimizes the scheduling agent by using the LLM with updated parameters, and forms a process of continuous learning of the scheduling agent.

[0074] According to the embodiment of the present application, the above-mentioned fault repair plan scheme is obtained based on the original data stream, the abnormal event summary generation operation, the directed acyclic graph diagnosis workflow generation operation, the fault root cause confirmation operation, and the root cause pre-play verification operation are executed by calling multiple types of intelligent agents, and the fault repair plan scheme further includes: generating an initial fault repair plan scheme based on the fault target root cause obtained by the fault root cause confirmation operation, and executing the root cause pre-play verification operation on the fault repair plan scheme to obtain the fault repair plan scheme.

[0075] According to the embodiment of the present application, the above-mentioned digital twin pre-play operation is performed on the fault repair plan scheme, and the control script of the fault repair plan scheme is executed based on the result of the digital twin pre-play operation, and the control script is executed and stored, which includes: using a spatio-temporal graph neural network to perform dynamic global state coding on the fault repair plan scheme to obtain a unified state vector; using a recurrent memory and segment recurrent neural network to perform a digital twin pre-play operation on the unified state vector to predict a future state, to obtain evaluation information of the fault repair plan scheme; and in the case that the difference between the evaluation information of the fault repair plan scheme and the prediction index is within a safety threshold, the control script is executed and stored.

[0076] The verification and pre-play process of the fault repair plan scheme of the above-mentioned embodiment will be further described in detail through a specific implementation manner.

[0077] The confirmed root cause is used to generate a fault repair plan P, and the digital twin pre-play technology is used to verify P to determine the effectiveness of P. Under the premise of the symbol-related effectiveness requirement of P, control instructions (i.e., control steps) are issued to the real data center dynamic environment monitoring system through Ansible (a Python-based automation operation and maintenance tool) / Modbus; at the same time, the execution result is written into a four-tuple memory pool (i.e., an experience replay pool), that is, [risk, action, feedback, timestamp], and the LLM is periodically fine-tuned to realize knowledge evolution.

[0078] According to a second aspect of the present invention, a data center environmental fault root cause tracing system is provided, comprising: a data acquisition module for real-time acquisition of environmental data streams from the data center, and constructing an original data stream with a multimodal unified view using configuration management data streams and environmental data streams stored in a time-series data lake; a multi-type intelligent agent module for obtaining a fault repair plan based on the original data stream by invoking multi-type intelligent agents to perform anomaly event summary generation, directed acyclic graph diagnostic workflow generation, fault root cause confirmation, and root cause pre-simulation verification operations; a domain knowledge base module for optimizing the fault root cause confirmation and root cause pre-simulation verification operations using a domain knowledge graph and a large language model fine-tuning mechanism; and an environmental digital twin module for performing digital twin pre-simulation operations on the fault repair plan, and performing distribution and storage operations on the control script of the fault repair plan based on the results of the digital twin pre-simulation operations.

[0079] According to an embodiment of the present invention, the aforementioned neighborhood knowledge base module includes a dynamic environment fault tree unit, a handling rule unit, and an experience replay pool; wherein, the dynamic environment digital twin module includes a state encoding unit, a state evolution unit, and a closed-loop verification unit; wherein, the multi-type intelligent agent module includes a perception intelligent agent unit, an orchestration intelligent agent unit, a multimodal expert intelligent agent unit, and an action and verification unit.

[0080] The following describes specific implementation methods in conjunction with appendices. Figure 6 The present invention provides a further detailed description of the data center environmental fault root cause tracing system.

[0081] Figure 6 The schematic diagram illustrates the structure of a data center environmental fault root cause tracing system according to an embodiment of the present invention.

[0082] like Figure 6 As shown, the data center environmental fault root cause tracing system includes a data acquisition module, a multi-type intelligent agent module, a domain knowledge base module, and an environmental digital twin module. The domain knowledge base module includes an environmental fault tree unit, a handling rule unit, and an experience replay pool. The environmental fault tree unit formally defines the typical fault modes of the unit using SysML (a standardized modeling language for systems engineering); the handling rule unit uses executable script templates described by YAML (a serialized data format); and the experience replay pool stores four-tuple memories and utilizes these memories for incremental LLM training.

[0083] In the dynamic environment digital twin module, the global state of the dynamic environment is encoded by a spatiotemporal graph neural network (ST-GNN) through a state encoding module, outputting a unified state vector S. tThe state evolution unit predicts future T (T is a positive integer greater than 1) step state evolution based on a Transformer-XL, quickly evaluates the effect of the treatment action; the closed-loop verification unit executes the repair plan P in the digital twin environment, compares the predicted indicators with the safety threshold, and obtains the difference value; if the difference value < epsilon, it is marked as “issuable”, that is, the control script is issued for real data center dynamic ring to perform root cause tracing.

[0084] The data center dynamic ring fault root cause tracing method and system provided by the application realize end-to-end fast root cause tracing and positioning, realize power-heating-security-IT cross-domain collaborative analysis, realize zero-code automatic generation of an interpretable directed acyclic graph diagnostic workflow, and support online incremental learning of root cause knowledge. The application captures abnormal events in real time through a perception intelligent agent, arranges intelligent agent calling a large language model to automatically disassemble tasks and scheduling domain expert intelligent agents, and generates a DAG-form diagnostic workflow. The domain expert intelligent agents diagnose on a local knowledge graph in parallel, and the arrangement intelligent agent aggregates confidence based on Bayesian inference. After confirming the root cause, a repair plan is generated and verified in a digital twin environment, and finally closed-loop execution is performed.

[0085] Figure 7 A block diagram of an electronic device suitable for implementing the data center dynamic ring fault root cause tracing method according to an embodiment of the application is schematically shown.

[0086] As shown in Figure 7 The electronic device 700 according to an embodiment of the application includes a processor 701, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 702 or loaded into a random access memory (RAM) 703 from a storage section 708. The processor 701 can include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a related chipset, and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), and the like. The processor 701 can also include an on-board memory for cache use. The processor 701 can include a single processing unit or a plurality of processing units for performing different actions of the method processes according to embodiments of the application.

[0087] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other via the bus 704. The processor 701 performs various operations of the method flow according to the embodiments of the present application by executing the programs in the ROM 702 and / or the RAM 703. It should be noted that the programs can also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 can also perform various operations of the method flow according to the embodiments of the present application by executing the programs stored in the one or more memories.

[0088] According to the embodiments of the present application, the electronic device 700 can further include an input / output (I / O) interface 705, which is also connected to the bus 704. The electronic device 700 can further include one or more of the following components connected to the input / output (I / O) interface 705: an input part 706 including a keyboard, a mouse, etc.; an output part 707 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage part 708 including a hard disk, etc.; and a communication part 709 including a network interface card such as a LAN card, a modem, etc. The communication part 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as necessary. A removable recording medium 711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 710 as necessary, so that a computer program read therefrom is installed in the storage part 708 as necessary.

[0089] The present application also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, when the one or more programs are executed, the method according to the embodiments of the present application is implemented.

[0090] According to embodiments of the present application, the computer readable storage medium can be a non-transitory computer readable storage medium, such as, for example, without limitation, a portable computer diskette, a hard disk, random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present application, a computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. For example, according to embodiments of the present application, the computer readable storage medium can include the ROM 702 and / or the RAM 703 described above, and / or one or more other memory devices that are not part of the ROM 702 and the RAM 703.

[0091] The flow diagrams and the block diagrams in the drawings are illustrations of possible architectures, functions, and operations for systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may

[0092] Those skilled in the art will appreciate that features recited in the various embodiments of the present application can be combined and / or integrated in various ways, even if such combinations or integrations are not expressly noted in the present application. In particular, features recited in the various embodiments of the present application can be combined and / or integrated in ways that are not expressly noted in the present application, without departing from the spirit and teachings of the present application. All such combinations and / or integrations are within the scope of the present application.

[0093] The embodiments of the present application have been described above. However, these embodiments are merely meant to be illustrative, and are not meant to limit the scope of the present application. Although the various embodiments have been described separately above, this does not mean that the measures in the various embodiments cannot be advantageously combined. Various alternatives and modifications can be made to the embodiments of the present application by those skilled in the art without departing from the scope of the present application, and all such alternatives and modifications are intended to fall within the scope of the present application.

Claims

1. A method for tracing the root causes of environmental failures in a data center, characterized in that, The method includes: Real-time acquisition of environmental data streams from the data center, and construction of raw data streams with a multimodal unified view using configuration management data streams and environmental data streams stored in the time-series data lake; Based on the original data stream, a fault repair plan is obtained by calling multiple types of intelligent agents to perform anomaly event summary generation, directed acyclic graph diagnostic workflow generation, fault root cause confirmation, and root cause pre-simulation verification. The root cause identification operation and the root cause pre-simulation verification operation are optimized by using domain knowledge graphs and large language model fine-tuning mechanisms. A digital twin simulation operation is performed on the fault repair plan, and based on the results of the digital twin simulation operation, the control script of the fault repair plan is distributed and stored.

2. The method according to claim 1, characterized in that, Real-time acquisition of environmental data streams from the data center, and construction of a raw data stream with a multimodal unified view using configuration management data streams and environmental data streams stored in a time-series data lake, including: The environmental data streams of the data center are collected in real time through multiple types of communication protocols. The environmental data streams include power data streams, HVAC data streams, security information streams, and hardware and software information streams. The multiple types of communication protocols include Internet of Things communication protocols, industrial automation communication protocols, battery management protocols, and network management protocols. The environmental data stream is written into the time-series data lake, and the configuration management data stream associated with the environmental data stream is acquired synchronously during the real-time acquisition of the environmental data stream. The configuration management data stream includes the topology information, configuration information and historical alarm information of the management database. Based on the unified metadata specification, the environmental data stream and the configuration management data stream are subjected to real-time association, data fusion, and visualization mapping operations to obtain the original data stream with a multimodal unified view.

3. The method according to claim 1, characterized in that, Based on the original data stream, a fault repair plan is obtained by invoking multiple types of intelligent agents to perform operations such as anomaly event summary generation, directed acyclic graph diagnostic workflow generation, fault root cause confirmation, and root cause pre-simulation verification. The original data stream is compressed using a perceptual agent, and an anomaly detection operation is performed on the compressed original data stream based on an unsupervised anomaly detection mechanism to generate an anomaly event summary. The anomaly event summary includes time information, event object, event feature information, and event anomaly degree.

4. The method according to claim 3, characterized in that, Also includes: Based on the thinking chain technology, the orchestration agent is used to parse the abnormal event summary based on a large language model to obtain a set of candidate fault hypotheses. The large language model is fine-tuned in real time through a low-rank adaptive fine-tuning mechanism. A multimodal expert agent is dynamically instantiated using the candidate fault hypothesis set, hypothesis activity, and domain dependency, and a diagnostic workflow for a directed acyclic graph is dynamically generated using the multimodal expert agent.

5. The method according to claim 3, characterized in that, Also includes: Based on the domain knowledge graph, a multimodal expert agent is used to extend the local knowledge graph, and a simulation operation is performed on the multimodal system of the data center based on the extended local knowledge graph, wherein the multimodal expert agent corresponds one-to-one with the multimodal system; During the simulation operation, based on the diagnostic workflow of the directed acyclic graph, the multimodal expert agent is used to diagnose the root cause set of the data center and the confidence level of the root cause set in parallel. Based on a multi-agent consensus mechanism, the orchestration agent and the multimodal expert agent interactively process the fault root cause set and the confidence level of the fault root cause set to obtain the target root cause of the fault.

6. The method according to claim 5, characterized in that, Based on a multi-agent consensus mechanism, the orchestration agent and the multimodal expert agent interactively process the fault root cause set and its confidence level to obtain the target root causes of the fault, including: Based on the confidence level of the fault root cause set, the multimodal expert agent performs a vote on the fault root cause set using a local consensus mechanism to obtain the initial fault root cause set. The low-rank adaptive fine-tuning mechanism and the extended local knowledge graph are optimized using the initial root cause set of faults stored in the experience replay pool; The parameters of the large language model are updated using an optimized low-rank adaptive fine-tuning mechanism, and the simulation operation of the multimodal expert agent is optimized using an optimized local knowledge graph. The orchestration agent is optimized using the updated large language model, and the optimized orchestration agent is used to perform Bayesian aggregation on the initial root cause set of the fault to obtain the target root cause of the fault.

7. The method according to claim 3, characterized in that, Also includes: Based on the root cause of the fault obtained by the fault root cause confirmation operation, an initial fault repair plan is generated, and a root cause pre-verification operation is performed on the fault repair plan to obtain the fault repair plan.

8. The method according to claim 1, characterized in that, Performing a digital twin simulation of the fault repair plan, and based on the results of the digital twin simulation, performing a distribution and storage operation on the control script of the fault repair plan, includes: The fault repair plan is subjected to dynamic and environmental global state encoding using a spatiotemporal graph neural network to obtain a unified state vector. By using recurrent memory and fragment recurrent neural networks to perform a digital twin pre-simulation operation on the unified state vector to predict future states, the evaluation information of the fault repair plan is obtained. If the difference between the evaluation information and the predicted indicators of the fault repair plan is within a safe threshold, the control script is issued and stored.

9. A data center environmental fault root cause tracing system, characterized in that, The systematic method includes: The data acquisition module is used to collect environmental data streams from the data center in real time, and to construct a raw data stream with a multimodal unified view using configuration management data streams and environmental data streams stored in the time-series data lake. The multi-type intelligent agent module is used to obtain a fault repair plan based on the original data stream by calling multi-type intelligent agents to perform anomaly event summary generation, directed acyclic graph diagnostic workflow generation, fault root cause confirmation, and root cause pre-simulation verification. The domain knowledge base module is used to optimize the fault root cause confirmation operation and the root cause pre-verification operation by utilizing the domain knowledge graph and the large language model fine-tuning mechanism. The dynamic environment digital twin module is used to perform a digital twin pre-drill operation on the fault repair plan, and to perform a distribution operation and a storage operation on the control script of the fault repair plan based on the result of the digital twin pre-drill operation.

10. The system according to claim 9, characterized in that, The neighborhood knowledge base module includes an environmental fault tree unit, a handling rule unit, and an experience replay pool. The dynamic environment digital twin module includes a state encoding unit, a state evolution unit, and a closed-loop verification unit. The multi-type intelligent agent module includes a perception intelligent agent unit, an orchestration intelligent agent unit, a multimodal expert intelligent agent unit, and an action and verification unit.