Intelligent operation and maintenance system and method based on multi-model management

The intelligent operation and maintenance system with multi-model management solves the problems of complex model management, insufficient decision-making accuracy, and inefficient task collaboration in the operation and maintenance system. It realizes the intelligent upgrade of the entire operation and maintenance process, improves the utilization rate of hardware resources and the accuracy of fault diagnosis, and reduces manual intervention and operation and maintenance costs.

CN121937112APending Publication Date: 2026-04-28AVIC AIRBORNE SYST GENERIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AVIC AIRBORNE SYST GENERIC TECH CO LTD
Filing Date
2026-02-12
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

The existing operation and maintenance system suffers from problems such as chaotic model management, insufficient decision-making accuracy, weak task collaboration capabilities, and inefficient knowledge sharing, resulting in low hardware resource utilization, unreliable fault diagnosis results, difficulties in cross-module linkage, and high risks to knowledge sharing security and compliance.

Method used

By constructing an intelligent operation and maintenance system based on multi-model management, and adopting unified model access, dynamic resource scheduling, multi-agent collaboration, and RAG technology, the system achieves intelligent upgrades to the entire operation and maintenance process. This includes the integration of monitoring systems, work order systems, operation and maintenance rule processing modules, operation and maintenance model libraries, and operation and maintenance knowledge bases, supporting unified processing and efficient collaboration of multi-source data.

Benefits of technology

It achieves precise matching between operation and maintenance tasks and model capabilities, improves hardware resource utilization, enhances the accuracy and efficiency of fault diagnosis, reduces manual intervention, forms a closed loop of intelligent operation and maintenance throughout the entire lifecycle, and reduces operation and maintenance costs and misjudgment rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937112A_ABST
    Figure CN121937112A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent operation and maintenance system and method based on multi-model management in the technical field of operation and maintenance. The system comprises a monitoring system which collects the operation state and alarm information of each business system of an office in real time, and displays the operation state and alarm information in a centralized manner through a large visual monitoring screen; the work order system is used for recording, tracking and managing operation and maintenance service requests and forming a structural chemical engineering single flow and a complete operation and maintenance operation log; the operation and maintenance rule processing module is linked with a monitoring system and a work order system respectively and captures operation and maintenance events and structured task information in a full-link manner; the operation and maintenance model library is used for storing multiple types of operation and maintenance models and providing containerized packaging, standardized registration and multi-mode scheduling capabilities; and the operation and maintenance knowledge base is constructed based on an RAG architecture, integrates multi-source knowledge resources, and provides a high-correlation reference basis for model reasoning. According to the system, by constructing a unified multi-model management architecture, accurate matching of operation and maintenance tasks and model capabilities is realized, and the utilization rate of hardware resources is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent operation and maintenance technology for information systems, and in particular to an intelligent operation and maintenance system and method based on multi-model management. Background Technology

[0002] With the rapid development of artificial intelligence technology, large-scale models have gradually penetrated the field of intelligent operations and maintenance (O&M) due to their superior performance and wide applicability. However, single large-scale models have significant limitations when dealing with diverse and complex tasks in O&M scenarios: general-purpose large-scale models are less accurate than specialized models in specific tasks such as log anomaly detection, while specialized models struggle to meet the collaborative analysis needs of multimodal data and cannot cover the needs of the entire O&M process. In high-concurrency scenarios, they are prone to problems such as uneven resource allocation and response delays.

[0003] Currently, traditional AIOps relies on big data analytics and basic machine learning, using single or combined basic machine learning models as its core capabilities. It can only process structured data within preset rules, requiring manual preprocessing of unstructured data. Its scenario adaptability is limited, human-computer interaction is passive, and fault management is limited to post-event diagnosis and manual intervention. In contrast, intelligent operations and maintenance based on large models, leveraging the comprehensive intelligence of these models, can understand all types of data end-to-end, requiring no manual preprocessing. It can dynamically adapt to various scenarios, supports proactive conversational human-computer interaction, and constructs a closed-loop fault management system covering the entire lifecycle of "pre-event prediction - in-event self-healing - post-event review," offering significant advantages. However, current operations and maintenance systems still face the following core challenges: 1. Chaotic model management: The interfaces of different types of large models are not uniform, requiring repeated development of adaptation code, and there is a lack of dynamic load balancing and resource optimization mechanisms, resulting in low utilization of hardware resources; 2. Insufficient accuracy in operation and maintenance decisions: Large models are prone to "illusions," and relying solely on the model's own knowledge is insufficient to handle enterprise operation and maintenance data, resulting in unreliable fault diagnosis results and poor traceability; 3. Weak multi-task collaboration capability: Operation and maintenance tasks require cross-module linkage, and the existing system lacks a flexible task scheduling and intelligent agent collaboration mechanism, making it impossible to achieve full-process automation; 4. Inefficient knowledge management and sharing: Enterprise operation and maintenance knowledge is scattered across multiple sources such as documents and databases. The proportion of unstructured data is high, the retrieval efficiency is low, and there are security and compliance risks in cross-departmental and cross-organizational knowledge sharing. Summary of the Invention

[0004] This application provides an intelligent operation and maintenance system and method based on multi-model management. By using core technologies such as unified model access, dynamic resource scheduling, multi-agent collaboration and RAG, it solves the problems of complex model management, insufficient decision accuracy, inefficient task collaboration and difficulty in knowledge sharing in existing operation and maintenance systems, and realizes intelligent upgrade of the entire operation and maintenance process.

[0005] This application provides an intelligent operation and maintenance system based on multi-model management, including: The monitoring system collects the operational status and alarm information of various office business systems in real time and displays them centrally through a large visual monitoring screen. The work order system is used to record, track, and manage operation and maintenance service requests, forming a structured work order flow and a complete operation and maintenance log. The operation and maintenance rules processing module is linked with the monitoring system and the work order system to capture operation and maintenance events and structured task information across the entire chain. The operation and maintenance model library stores multiple types of operation and maintenance models and provides containerized encapsulation, standardized registration, and multi-mode scheduling capabilities. The operations and maintenance knowledge base is built on the RAG architecture and integrates multi-source knowledge resources to provide highly relevant references for model reasoning.

[0006] The beneficial effects of the above embodiments are as follows: by constructing a unified multi-model management architecture, the problems of complex model management and inconsistent interfaces in the existing operation and maintenance system are solved, the operation and maintenance tasks and model capabilities are accurately matched, and the utilization rate of hardware resources is improved.

[0007] Based on the above embodiments, this application can be further improved as follows: In one embodiment of this application, the operation and maintenance rule processing module establishes an intelligent matching mechanism for "event-data-model", including: The received raw events are parsed and features are extracted to form a structured task description; Model selection is based on the dimensions of task type, data characteristics, and historical model performance.

[0008] The beneficial effects of the above embodiments are: to achieve full-link coverage of operation and maintenance event triggering, to accurately match operation and maintenance tasks with the optimal model, to avoid repeated development of adaptation code, and to improve the accuracy and efficiency of model selection.

[0009] In one embodiment of this application, the task type dimension divides operation and maintenance events into major categories such as fault diagnosis, performance optimization, resource scheduling, and security alarms. Fault diagnosis events are preferentially matched with fault tree models and Bayesian inference models, while performance optimization events focus on regression models and reinforcement learning models in machine learning.

[0010] The beneficial effects of the above embodiments are: matching professional models to different types of operation and maintenance events, improving the accuracy of operation and maintenance decisions, and overcoming the limitations of a single model in handling diverse operation and maintenance tasks.

[0011] In one embodiment of this application, the data characteristic dimension is adapted to traditional statistical models, natural language processing models, and computer vision models for structured data, semi-structured data, and unstructured data, respectively. At the same time, the corresponding model architecture is selected according to the data volume and data integrity of batch data or real-time streaming data.

[0012] The beneficial effects of the above embodiments are: to achieve end-to-end understanding of all types of data, to eliminate the need for manual preprocessing of unstructured data, to improve the efficiency and adaptability of data processing, and to cover more operation and maintenance scenarios.

[0013] In one embodiment of this application, the historical performance dimension of the model is achieved by establishing a model performance evaluation system, recording indicators such as the processing accuracy, response time, resource consumption, and misjudgment rate of each model in similar events, and dynamically updating the model priority using a weighted scoring mechanism.

[0014] The beneficial effects of the above embodiments are: to realize dynamic optimization of the model, ensure the rationality of model selection, improve the overall performance of the model, and reduce the misjudgment rate of operation and maintenance decisions.

[0015] In one embodiment of this application, the operation and maintenance model library uses containerization technology and microservice architecture to containerize and encapsulate operation and maintenance-related models, packaging the model's operating system, dependency libraries, configuration files, and the model itself into an independent container, ensuring consistent deployment of the model in different environments.

[0016] The beneficial effects of the above embodiments are: to achieve efficient deployment, start-up, shutdown, and expansion of the model, improve the reusability of the model, facilitate independent upgrades and iterations of the model, and reduce the complexity of model management.

[0017] In one embodiment of this application, the operation and maintenance model library encapsulates model containers with strong dependencies and sequential collaboration in the operation and maintenance process into the same Pod, and achieves millisecond-level communication through the internal network of the Pod; models without direct collaboration relationships are encapsulated into independent Pods to ensure the flexibility of resource scheduling.

[0018] The beneficial effects of the above embodiments are: improving the communication efficiency between models, avoiding latency loss in cross-Pod network transmission, achieving task load balancing and parallel execution, and improving the overall processing throughput.

[0019] In one embodiment of this application, the specific process of the operation and maintenance knowledge base is divided into four key steps: S1, Based on the information output by the anomaly detection model, a structured fault summary is generated through a natural language processing model; S2, the retrieval system connects the internal knowledge base, database and external authoritative sources. It uses a vector database to semantically encode the query text and multi-source information, converting them into high-dimensional vectors, and quickly retrieves relevant information based on similarity algorithms. S3, the vector model scores the relevance of the retrieved candidate information, and selects highly relevant content by combining the timeliness, credibility and matching degree of the information. S4, the generated model combines the fault summary with the filtered relevant information to perform logical reasoning and comprehensive analysis, locate the root cause of the fault, and generate targeted solutions.

[0020] The beneficial effects of the above embodiments are: breaking through the knowledge boundaries of a single model, improving the accuracy of root cause localization and the feasibility of treatment plans, reducing manual investigation time and the risk of misoperation, and solving the problems of large models being prone to "illusions" and insufficient decision-making accuracy.

[0021] This application also provides an intelligent operation and maintenance method based on multi-model management, including the following steps: S10: Collection of operation and maintenance monitoring data and user events from the work order system; S20: Analysis and processing of operation and maintenance data and user events, and identification of key requests; S30: Trigger model calls and load balancing task allocation based on critical requests; S40: Model inference combined with the RAG framework generates operation and maintenance suggestions and reports; S50: Inter-model communication and collaboration form a closed loop of operation and maintenance workflow and main process.

[0022] The beneficial effects of the above embodiments are: to realize the intelligent upgrade of the entire operation and maintenance process, to build a closed loop of full life cycle fault management of "pre-event prediction - in-event self-healing - post-event review", and to improve operation and maintenance efficiency and decision-making accuracy.

[0023] In one embodiment of this application, the load balancing system in step S30 adopts a hybrid strategy that combines static pre-allocation with dynamic real-time adjustment. In the static stage, based on the historical resource demand profile of the model, a basic resource quota is preset for it, and similar tasks are initially directed to a specific model cluster. In the dynamic stage, based on real-time monitoring data and task priorities, task migration, resource preemption, or dynamic scaling up / down of instances are flexibly implemented.

[0024] The beneficial effects of the above embodiments are: ensuring efficient and flexible resource utilization, avoiding overload of a single node, guaranteeing processing efficiency and stability, and reducing computing costs.

[0025] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: (1) Significantly improved event processing efficiency: Through the dynamic model selection mechanism triggered by operation and maintenance events, the "task-data-model" is accurately matched, avoiding the problem of rigid adaptation of a single model. The initial processing response speed is significantly improved compared with traditional methods. The combination of multi-model parallel scheduling and containerized deployment improves the efficiency of model start-up, shutdown and expansion, shortens the iteration cycle from weekly to daily, and doubles the task throughput in high-concurrency scenarios, effectively solving the problem of operation and maintenance event congestion.

[0026] (2) Significantly improved fault analysis and decision-making accuracy: The RAG-based accurate fault analysis scheme enhances the model's reasoning ability through multi-source knowledge retrieval and vector sorting, improving the accuracy of fault root cause location and the matching degree of handling schemes; the multi-model collaborative collaboration and decision fusion mechanism breaks down data silos, improves the overall decision accuracy, reduces the misjudgment rate, and reduces the cost of manual error correction.

[0027] (3) Intelligent operation and maintenance and reduced labor costs: The entire process from event triggering, model selection, collaborative analysis to solution generation is intelligent, which greatly reduces the manual intervention links; accurate root cause location and disposal solution output shorten the fault investigation time and reduce the dependence on senior operation and maintenance personnel. At the same time, through iterative optimization of model historical performance, a self-evolving intelligent operation and maintenance closed loop is formed, which continuously improves the level of operation and maintenance automation and intelligence. Attached Figure Description

[0028] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0029] Figure 1 This is a general flowchart of the intelligent operation and maintenance method provided in the embodiments of the present invention; Figure 2 This is a schematic diagram of the framework flow of the intelligent operation and maintenance system provided in the embodiments of the present invention; Figure 3 This is a diagram showing the architecture of the intelligent operation and maintenance system provided in this embodiment of the invention. Detailed Implementation

[0030] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.

[0031] This invention addresses the problems of complex model management, insufficient decision-making accuracy, inefficient task collaboration, and difficulties in knowledge sharing in operation and maintenance systems. It provides an intelligent operation and maintenance system and method based on multi-model management, which realizes intelligent upgrade of the entire operation and maintenance process through core technologies such as unified model access, dynamic resource scheduling, multi-agent collaboration, and RAG.

[0032] like Figure 1 As shown, an intelligent operation and maintenance method based on multi-model management includes the following processes: S10: Operation and maintenance monitoring data and work order system user event collection.

[0033] As a fundamental data input link for intelligent operation and maintenance, the operation and maintenance monitoring system deployed across the entire chain captures multi-dimensional data such as server hardware resources, network transmission status, and application operation indicators in real time. At the same time, it seamlessly connects with the work order system to collect event information such as fault reports, request inquiries, and problem feedback submitted by users, covering key content such as fault phenomena, scope of impact, and urgency, thus constructing an input data source and providing complete support for subsequent analysis.

[0034] S20: Analysis and processing of operation and maintenance data and user events, and identification of key requests.

[0035] First, the collected multi-source data is preprocessed by cleaning the data to remove noise, fill in missing values, and unify the format to ensure data quality. Then, a natural language processing model is used to parse user event text and extract core requests. Statistical analysis and anomaly detection models are used to identify abnormal fluctuations and threshold exceedances in the monitoring data. Finally, by combining business rules and historical data, core needs and potential faults are accurately located from massive amounts of information to form clear key requests, providing a basis for model invocation.

[0036] S30: Trigger model calls and load balancing task allocation based on critical requests.

[0037] Based on the type of critical request, the optimal model is matched through an intelligent model selection mechanism. At the same time, relying on the load balancing strategy, the resource utilization rate and task queue length of each model node are monitored in real time. Combined with the task priority, the tasks are dynamically allocated to model nodes with low load and high adaptability, supporting multi-node parallel processing, avoiding overload of a single node, and ensuring processing efficiency and stability.

[0038] S40: Model reasoning combined with the RAG framework generates operation and maintenance suggestions and reports.

[0039] The model nodes assigned to tasks initiate a specialized inference process, conduct analysis and calculations on key requests, and output preliminary results and impact assessments. Simultaneously, the RAG framework is integrated, and based on the fault summary generated by anomaly detection, relevant information is retrieved from the operation and maintenance knowledge base, historical fault database, and external authoritative technical resources. Highly relevant content is sorted and filtered using a vector model to supplement the inference basis. Finally, the model calculation results and retrieved information are combined to generate clear and targeted operation and maintenance handling suggestions and detailed reports, including core content such as root causes, scope of impact, risk warnings, and confidence scores, providing clear guidance for operation and maintenance execution.

[0040] S50: Inter-model communication and collaboration form a closed loop of operation and maintenance workflow and main process.

[0041] For complex operation and maintenance scenarios where a single model cannot complete the entire process, standardized communication protocols are used to achieve parameter transfer, intermediate result exchange, and decision fusion between models. According to the preset process arrangement logic, each link is connected to form a complete operation and maintenance workflow of "data collection → analysis and identification → model inference → solution generation → execution feedback", ensuring seamless connection between each step. After the process is completed, the processing results are fed back to the model evaluation system to update the model's historical performance data, providing a basis for subsequent model selection and optimization, forming a continuously iterative intelligent operation and maintenance closed loop.

[0042] The load balancing system of the described method employs a hybrid strategy combining static pre-allocation and dynamic real-time adjustment. In the static phase, based on the historical resource demand profile of the model, a basic resource quota is preset, and similar tasks are initially directed to specific model clusters. In the dynamic phase, based on real-time monitoring data and task priorities, task migration, resource preemption, or dynamic scaling up / down of instances are flexibly implemented. For example, to ensure high-priority tasks, the system can temporarily reclaim resources from low-priority tasks or quickly scale up new instances; when the overall load decreases, idle instances are automatically reduced to save costs. This mechanism ensures efficient and flexible resource utilization.

[0043] To achieve efficient multi-model collaboration, the system has established a standardized inter-model communication protocol, clearly defining data exchange formats, interface specifications, and verification rules. Communication content mainly includes parameter passing, intermediate result exchange, and decision fusion. In practice, asynchronous, decoupled streaming processing can be achieved through message queues, or large volumes of intermediate data can be transferred using shared storage. The system natively supports pipelined serial and parallel collaboration. Furthermore, a collaborative feedback mechanism has been established, allowing subsequent models to feed back their processing conclusions to preceding models for online parameter optimization and adjustment, thereby achieving co-evolution through collaboration.

[0044] like Figure 2As shown, an intelligent operation and maintenance system based on multi-model management is presented. The core of the system is the operation and maintenance rule processing module, which is bidirectionally integrated: on the one hand, it links in real time with monitoring systems such as performance monitoring and log monitoring to capture immediate events such as operational anomalies, error logs, and service timeouts; on the other hand, it seamlessly integrates with the work order system to obtain structured knowledge such as historical fault records and manual handling trajectories. This achieves full-link coverage and unified processing of real-time operation and maintenance event streams and historical databases.

[0045] Upon receiving the raw event, this module immediately performs parsing and feature extraction, transforming it into a structured task description that clearly depicts the event type, the services or components involved, and the format of the data to be processed.

[0046] Its core model selection logic is a multi-dimensional decision-making system, mainly including: Task type dimension: Categorize events and map them to the corresponding domain's advantageous model.

[0047] Data characteristic dimension: Assign an appropriate model family based on the structure of the data to be analyzed.

[0048] Historical Model Performance Dimension: Through a continuously maintained model performance evaluation system, the accuracy, response speed, resource consumption, and false positive rate of each model in similar historical tasks are recorded. A dynamic weighted scoring mechanism is adopted to update the priority of the model queue in real time, ensuring that the best performer is selected.

[0049] The RAG-based precise fault analysis mechanism is deeply integrated into the system workflow. Its core lies in leveraging external knowledge to enhance the model's intrinsic reasoning capabilities, primarily applied in the root cause localization stage to compensate for the limitations of untimely updates to the model's internal knowledge. The specific process is as follows: First, the NLP model generates a fault summary; then, the retrieval system connects internal and external knowledge sources, performing semantic encoding and similarity matching through a vector database to quickly recall highly relevant information; finally, the generation model integrates all information for reasoning, outputting precise root cause conclusions and corresponding solutions.

[0050] The synergy between the operations and maintenance model and automation tools forms a closed-loop "decision-execution" process. The system transforms the root causes and solutions identified through analysis into standardized execution instructions that can be directly parsed by machines. These instructions include clearly defined operation targets, operation types, specific parameters, execution priorities, and rollback conditions. Operations personnel can evaluate the execution results through the management platform, and this feedback directly influences the optimization direction of the relevant models and the iteration of handling logic.

[0051] The accumulation and reuse of operational knowledge is key to the continuous intelligence of the system. Effective solutions, once completed, undergo a standardized process to be stored in the operational knowledge base. Knowledge screening and verification: The system automatically filters out low-confidence and duplicate content, and combines this with manual review to confirm the feasibility and universality of the solution.

[0052] Knowledge structuring: The approved content is broken down into standardized fields such as fault type, core characteristics, root cause, handling steps, and verification methods, and vector features are labeled.

[0053] Knowledge base update: Synchronize structured knowledge to the knowledge base directory and update the search index.

[0054] Knowledge reuse and association: When similar failures occur in the future, the RAG framework can quickly match the accumulated knowledge, providing direct and accurate references for the model, thereby continuously enriching the knowledge system and continuously evolving the operational decision-making capabilities.

[0055] like Figure 3 The system specifically comprises: Heterogeneous Computing Platform: This platform provides diversified computing power for operational event analysis and handling, integrating heterogeneous resources such as CPU clusters, GPU computing pools, and edge computing nodes to form a flexible and scalable computing support system. It features dynamic resource scheduling capabilities, intelligently allocating suitable computing resources based on the complexity and real-time requirements of operational tasks. The platform supports elastic scaling of computing resources, automatically expanding or reducing computing nodes by monitoring task queue length and resource utilization, ensuring sufficient computing power during peak operational event periods and rationally releasing resources during off-peak periods to reduce computing costs.

[0056] Operations and Maintenance Model Library: This library encompasses diverse models adapted to various operations and maintenance scenarios, and provides full lifecycle model management capabilities. Model types cover fault tree models for fault diagnosis, regression prediction models for performance optimization, and natural language processing models for data processing, allowing for flexible matching based on the type of operations and maintenance event and data characteristics. Simultaneously, the model library incorporates a containerized encapsulation and scheduling module, using Docker and Kubernetes technologies to encapsulate all models and their runtime environments into standardized containers, ensuring consistency and compatibility across environments. Through a visual orchestration engine, users can customize the model execution order and trigger conditions according to the operations and maintenance process, enabling serial / parallel scheduling of multiple models. Furthermore, it integrates with load balancing mechanisms to evenly distribute tasks across different model nodes, preventing overload of a single model and ensuring efficient and stable model invocation.

[0057] Operations and Maintenance Knowledge Base: Based on the RAG architecture, this knowledge base provides accurate and comprehensive data support for operations and maintenance model analysis, and is key to improving analysis efficiency and accuracy. The knowledge base integrates multi-source knowledge resources, including internally accumulated historical fault handling documents, operations and maintenance manuals, configuration specifications, and work order records; external authoritative technical forum posts, vendor troubleshooting guides, industry best practices; and real-time updated system operation data and monitoring logs. The knowledge base uses a vector database to store knowledge content, transforming unstructured and semi-structured data into high-dimensional vectors, and combining similarity algorithms for rapid retrieval. When the model performs fault analysis, the knowledge base accurately extracts relevant knowledge based on the fault summary, sorts and filters it, and feeds it back to the model as a basis for reasoning. Simultaneously, the knowledge base has dynamic update capabilities, automatically receiving effective analysis results and handling solutions accumulated in the operations and maintenance process, standardizing them, and incorporating them into the database to continuously enrich the knowledge coverage, forming a closed loop of "knowledge reuse - update iteration".

[0058] Multi-agent development and orchestration components: Provide end-to-end support from agent development to orchestration and scheduling, enabling multiple agents to collaborate on complex operations and maintenance tasks. The development toolchain includes a visual development interface, pre-built templates, and an API library, lowering the barrier to entry for agent development and allowing developers to quickly build operations and maintenance agents with capabilities such as data collection, analysis and reasoning, and command execution. The orchestration function, through a drag-and-drop workflow design, allows users to define the collaboration logic of multiple agents. Simultaneously, the component supports dynamic scheduling of agents, flexibly adjusting the execution order and resource allocation of agents based on the priority of operations and maintenance events and real-time load, ensuring efficient collaboration among agents in complex operations and maintenance scenarios to form a closed-loop processing capability.

[0059] Automated operation and maintenance tools: Provides operation and maintenance intelligence agents with a set of operation tools that can be directly called to realize the automated execution of operation and maintenance analysis and handling plans, reducing manual intervention.

[0060] Operation and maintenance notification and display components: including message terminals, operation and maintenance dashboards, operation systems, etc., are responsible for the visual presentation of operation and maintenance information and multi-channel notifications, ensuring that operation and maintenance personnel can keep abreast of system status and event progress in real time.

[0061] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. An intelligent operation and maintenance system based on multi-model management, characterized in that, include: The monitoring system collects real-time operational status and alarm information from various office business systems. The work order system is used to record, track, and manage operation and maintenance service requests, forming a structured work order flow and a complete operation and maintenance log. The operation and maintenance rules processing module is linked with the monitoring system and the work order system to capture operation and maintenance events and structured task information across the entire chain. The operation and maintenance model library stores multiple types of operation and maintenance models and provides containerized encapsulation, standardized registration, and multi-mode scheduling capabilities. The operations and maintenance knowledge base is built on the RAG architecture and integrates multi-source knowledge resources to provide highly relevant references for model reasoning.

2. The intelligent operation and maintenance system according to claim 1, characterized in that: The operation and maintenance rule processing module establishes an intelligent matching mechanism for "event-data-model", including: The received raw events are parsed and features are extracted to form a structured task description; Model selection is based on task type, data characteristics, and historical model performance. The task type dimension categorizes operational events into fault diagnosis, performance optimization, resource scheduling, and security alerts. Fault diagnosis events are matched with fault tree models and Bayesian inference models, while performance optimization events are matched with regression models and reinforcement learning models in machine learning. The data characteristic dimension adapts traditional statistical models, natural language processing models, and computer vision models to structured data, semi-structured data, and unstructured data, respectively. The appropriate model architecture is selected based on the data volume and integrity of batch or real-time streaming data. The model historical performance dimension establishes a model performance evaluation system to record the processing accuracy, response time, resource consumption, and misjudgment rate of each model in similar events, and dynamically updates model priorities using a weighted scoring mechanism.

3. The intelligent operation and maintenance system according to claim 1, characterized in that: The operation and maintenance model library uses containerization technology and microservice architecture to containerize and encapsulate operation and maintenance-related models. It packages the model's operating system, dependency libraries, configuration files, and the model itself into an independent container, ensuring consistent deployment of the model in different environments.

4. The intelligent operation and maintenance system according to claim 3, characterized in that: The operation and maintenance model library encapsulates model containers with strong dependencies and sequential collaboration in the operation and maintenance process into the same Pod, and achieves millisecond-level communication through the internal network of the Pod; models without direct collaboration relationships are encapsulated into independent Pods to ensure the flexibility of resource scheduling.

5. The intelligent operation and maintenance system according to claim 1, characterized in that: The specific process of the operation and maintenance knowledge base consists of four steps: S1, Based on the information output by the anomaly detection model, a structured fault summary is generated through a natural language processing model; S2, the retrieval system connects the internal knowledge base, database and external authoritative sources. It uses a vector database to semantically encode the query text and multi-source information, converting them into high-dimensional vectors, and quickly retrieves relevant information based on similarity algorithms. S3, the vector model scores the relevance of the retrieved candidate information, and selects highly relevant content by combining the timeliness, credibility and matching degree of the information. S4, the generated model combines the fault summary with the filtered relevant information to perform logical reasoning and comprehensive analysis, locate the root cause of the fault, and generate targeted solutions.

6. An intelligent operation and maintenance method based on multi-model management, characterized in that: The intelligent operation and maintenance system according to any one of claims 1-5 includes the following steps: S10: Collection of operation and maintenance monitoring data and user events from the work order system; S20: Analysis and processing of operation and maintenance data and user events, and identification of key requests; S30: Trigger model calls and load balancing task allocation based on critical requests; S40: Model inference combined with the RAG framework generates operation and maintenance suggestions and reports; S50: Inter-model communication and collaboration form a closed loop of operation and maintenance workflow and main process.

7. The intelligent operation and maintenance method according to claim 6, characterized in that: The load balancing system in step S30 adopts a hybrid strategy that combines static pre-allocation with dynamic real-time adjustment. In the static stage, based on the historical resource demand profile of the model, a basic resource quota is preset for it, and similar tasks are initially directed to the corresponding model cluster. In the dynamic stage, based on real-time monitoring data and task priority, task migration, resource preemption, or dynamic scaling up / down of instances are flexibly implemented.