Operation and maintenance method, device and equipment of present network system, medium and program product
By using root cause localization and a large language model for operation and maintenance decisions in the existing network system, combined with topology dependency graphs and historical knowledge bases, comprehensive operation and maintenance across systems was achieved. This solved the problems of inaccurate fault location and ineffective operation and maintenance in the existing network system, and improved the accuracy of fault location and the effectiveness of operation and maintenance.
Patent Information
- Application Number
- CN202511384698.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-25
AI Technical Summary
Cross-system operation and maintenance of existing network systems is difficult to achieve integrated operation and maintenance, resulting in inaccurate fault location and ineffective operation and maintenance.
By employing a root cause localization language model and an operation and maintenance decision-making language model, combined with a target topology dependency graph and a historical operation and maintenance knowledge base, anomaly clustering and causal analysis are performed on the live network system to generate an operation and maintenance decision sequence, and comprehensive operation and maintenance is carried out through the target operation and maintenance task sequence.
It improves the accuracy of fault location and the effectiveness of operation and maintenance in the existing network system. Through the closed-loop application of semantic understanding and historical operation and maintenance processes, it enhances the flexibility and feasibility of the system.
Smart Images

Figure CN120875851A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, equipment, medium, and program product for the operation and maintenance of a live network system. Background Technology
[0002] The existing network system is characterized by its cross-system, cross-regional, and long-link characteristics, which brings great challenges to its operation and maintenance. The existing network system includes subsystems such as traffic splitting, traffic parsing, call detail record (CDR) entry, data analysis, and data early warning. These subsystems are highly coupled.
[0003] Currently, cross-system operation and maintenance adopts a general integration platform. Based on the inherent data acquisition process, it acquires data from each subsystem, then uses corresponding fault detection rules to identify faults in each subsystem, and finally performs operation and maintenance on the identified faults.
[0004] The existing cross-system operation and maintenance methods are essentially still independent operation and maintenance of individual subsystems, and have not achieved comprehensive operation and maintenance of the live network system. This makes it difficult to guarantee the accuracy of fault location and the effectiveness of system operation and maintenance. Summary of the Invention
[0005] This invention provides a method, device, equipment, medium, and program product for the operation and maintenance of a live network system, which realizes comprehensive operation and maintenance of the live network system and improves the accuracy of fault location and the effectiveness of system operation and maintenance.
[0006] According to one aspect of the present invention, a method for operating and maintaining a live network system is provided, the method comprising:
[0007] Obtain target network system data, and based on the target network system data, generate target topology dependency graph and target indicator triplet for the target network system;
[0008] The historical operation and maintenance knowledge base is acquired, and the root cause localization big language model is used. Based on the target topology dependency graph, the target indicator triplet and the historical operation and maintenance knowledge base, anomaly clustering and causal analysis are performed on the target live network system to obtain candidate root cause fault information of the target live network system.
[0009] The system acquires historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization targets. Then, using a large language model for operation and maintenance decisions, it makes operation and maintenance decisions on the target network system based on the candidate root cause fault information of the target network system, the historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization targets, to obtain a target executable decision sequence.
[0010] Based on the target executable decision sequence, a target operation and maintenance task sequence for the target live network system is generated, and the target live network system is operated and maintained according to the target operation and maintenance task sequence.
[0011] According to another aspect of the present invention, a network system operation and maintenance device is provided, the device comprising:
[0012] The target live network system data acquisition module is used to acquire target live network system data and generate target topology dependency graph and target indicator triplet of the target live network system based on the target live network system data;
[0013] The fault root cause localization module is used to acquire the historical operation and maintenance knowledge base, and use the root cause localization big language model to perform anomaly clustering and causal analysis on the target topology dependency graph, the target indicator triplet and the historical operation and maintenance knowledge base to obtain candidate root cause fault information of the target live network system.
[0014] The operation and maintenance decision output module is used to acquire historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives. It adopts the operation and maintenance decision big language model to make operation and maintenance decisions for the target network system based on the candidate root cause fault information of the target network system, the historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives, and outputs the target executable decision sequence.
[0015] The target live network system operation and maintenance module is used to generate a target operation and maintenance task sequence for the target live network system based on the target executable decision sequence, and to perform operation and maintenance on the target live network system according to the target operation and maintenance task sequence.
[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the live network system operation and maintenance method according to any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the network system operation and maintenance method described in any embodiment of the present invention.
[0021] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the live network system operation and maintenance method described in any embodiment of the present invention.
[0022] The technical solution of this invention employs a root cause localization language model. Based on the target topology dependency graph, target indicator triples, and historical operation and maintenance knowledge base of the target network system, it performs anomaly clustering and causal analysis on the target network system. It leverages the semantic understanding capability of the root cause localization language model itself to improve the accuracy of root cause localization of faults in the target network system. The target topology dependency graph considers the dependencies between components of each subsystem in the target network system, i.e., the coupling between subsystems. The target indicator triples achieve semantic space unification of the target network system data from three dimensions: "indicator-component-business," and unified orchestration of multi-source data across domains and systems, thereby improving the recognition accuracy of the root cause localization language model. Furthermore, the introduction of a historical operation and maintenance knowledge base enhances the retrieval of root cause faults in the target network system, further improving the accuracy of root cause localization. By employing an operation and maintenance decision-making language model… This model, based on candidate root cause fault information, historical fault operation and maintenance processes, historical effectiveness of historical fault operation and maintenance processes, target operation and maintenance constraints, and reference optimization objectives of the target network system, makes operation and maintenance decisions for the target network system, resulting in a target executable decision sequence. It leverages the semantic understanding capabilities of the root cause localization language model itself to improve the accuracy of operation and maintenance decisions for the target network system. By considering the historical effectiveness of historical fault operation and maintenance processes, it achieves closed-loop application of historical fault operation and maintenance processes. Furthermore, it introduces target operation and maintenance constraints and reference optimization objectives to improve the flexibility of the target network system, thereby enhancing the feasibility of the target executable decision sequence. By transforming the target executable decision sequence into a target operation and maintenance task sequence for the target network system, and performing operation and maintenance according to the target operation and maintenance task sequence, it achieves comprehensive operation and maintenance of the target network system, improving the accuracy of fault localization and the effectiveness of system operation and maintenance.
[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart of a live network system operation and maintenance method provided according to Embodiment 1 of the present invention;
[0026] Figure 2 This is a schematic diagram of the automatic retrieval configuration interface of the automatic acquisition engine provided in Embodiment 1 of the present invention.
[0027] Figure 3 This is a schematic diagram of the automatic login configuration interface of the automatic data collection engine provided in Embodiment 1 of the present invention.
[0028] Figure 4 This is a schematic diagram of the interface for data collection using the manual survey method of the automatic data collection engine provided in Embodiment 1 of the present invention;
[0029] Figure 5 This is an interface display diagram of candidate root cause fault information provided in Embodiment 1 of the present invention;
[0030] Figure 6 This is an interface display diagram of the fault handling tracking content included in the actual execution data provided according to Embodiment 1 of the present invention;
[0031] Figure 7 This is a flowchart of a live network system operation and maintenance method provided according to Embodiment 2 of the present invention;
[0032] Figure 8 This is a schematic diagram of the structure of a network system operation and maintenance device according to Embodiment 3 of the present invention;
[0033] Figure 9 This is a schematic diagram of the structure of an electronic device that implements the network system operation and maintenance method of this invention. Detailed Implementation
[0034] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0035] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data are interchangeable where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0036] Example 1
[0037] Figure 1 This is a flowchart illustrating a live network system operation and maintenance method according to Embodiment 1 of the present invention. This embodiment of the invention is applicable to the comprehensive operation and maintenance of live network systems. The method is executed by a live network system operation and maintenance device, which is implemented in hardware and / or software and can be configured in an electronic device that carries live network system operation and maintenance functions.
[0038] See Figure 1 The live network system operation and maintenance methods shown include:
[0039] S101. Obtain the target network system data, and generate the target topology dependency graph and target indicator triplet of the target network system based on the target network system data.
[0040] The target network system is a carrier-grade service system in operation. The subsystems within the target network system are highly interconnected, including subsystems such as traffic splitting, traffic parsing, call detail record (CDR) entry, data analysis, and data early warning. The target network system exhibits a strongly dependent topology: "Traffic Splitting Subsystem → Traffic Parsing Subsystem → CDR Entry Subsystem → Data Analysis Subsystem → Data Early Warning Subsystem." The network system is characterized by its cross-system, cross-regional, and long-link characteristics. Specifically, during the overall operation and maintenance of the target network system, minor fluctuations in upstream subsystems can be amplified or masked in downstream subsystems, leading to false alarms or delayed fault reports, significantly complicating network system operation and maintenance. Specifically, the traffic splitting subsystem is used to bypass and mirror production traffic for subsequent parsing and analysis of the mirrored traffic. The traffic parsing subsystem is used for protocol restoration, indicator extraction, and content parsing of the mirrored traffic. The CDR entry subsystem is used for standardizing and storing call detail records (CDRs) or session details. The data analysis subsystem is used to analyze and predict metrics, logs, and alarm data. The data early warning subsystem is used to issue early warnings based on metrics, logs, and alarm data. Target live network system data refers to the operational data of the target live network system. Target live network system data is used to characterize the actual operational status of the target live network system. For example, target live network system data includes operational metric data and events of the target live network system.
[0041] The target topology dependency graph is used to reflect the dependencies between components of each subsystem in the target live network system. For example, G(E,V) represents the target topology dependency graph. Here, V represents the components of each subsystem in the target live network system. For example, the components of each subsystem in the target live network system include hardware devices and / or software modules. E represents the dependencies between the components of each subsystem in the target live network system. Specifically, dependencies include data transmission relationships and data control relationships. Optionally, the dependencies between components of each subsystem in the target live network system have corresponding dependency weights. For example, dependency weights include traffic share, call frequency, and historical correlation.
[0042] The target indicator triplet comprises a ternary relationship between the target indicator, target component, and target service. The target indicator includes the data dimensions detected by the target live network system and the corresponding target indicator data. The target component is the data source for the target indicator. For example, target components include target hardware devices and / or target software modules. The target service is the service associated with the target indicator data. This allows for the differentiation of different services for the same indicator. The target indicator triplet is used to uniformly name and map the indicator data, log data, link data, and key service indicators in the target live network system data. This achieves data alignment of multi-source heterogeneous data across subsystems in the target live network system, facilitating the semantic understanding of the target live network system data by the subsequent root cause localization large language model.
[0043] Specifically, target acquisition tasks are registered on the automatic acquisition engine for the target network system, and target acquisition devices are selected and target acquisition strategies are configured based on these tasks. According to the target acquisition strategy, the target acquisition devices are scheduled to execute the corresponding target acquisition tasks to acquire data from the target network system. The target network system data is identified to determine the components of each subsystem within the target network system and the dependencies between these components. Based on the components of each subsystem and the dependencies between them, a target topology dependency graph of the target network system is generated. The target network system data is then used for indicator data identification to determine target indicators, target components, and target services. Based on the target indicators, target components, and target services, target indicator triples are generated. For example, Figure 2 This is a schematic diagram of the automatic retrieval configuration interface for the automatic data acquisition engine. On the automatic retrieval configuration page, the target acquisition strategy for the target acquisition task is configured.
[0044] The target data collection strategy refers to the specific execution strategy for the target data collection task. For example, target data collection strategies include periodic collection, event-triggered collection, and adaptive collection. For instance, adaptive collection reduces the sampling frequency when the load is high. The target data collection task is the task of collecting data from the target live network system. The target data collector is used to execute the target data collection task. For example, target data collectors include web page collectors, software collectors, database collectors, and manual survey collectors.
[0045] Web scrapers collect data through web systems. A web scraper includes a headless browser, browser scraping strategies, and corresponding page structures. Web scrapers support automatic login, CAPTCHA strategies (such as sliders or SMS callback permissions), and data export and parsing (such as in table or chart formats).
[0046] The software data collector collects data through a software client. It interfaces with various subsystems via APIs (Application Programming Interfaces), SDKs (Software Development Kits), or CLIs (Command Line Interfaces). If no public API is available, runtime metrics or events can be collected via an agent or log forwarder. Figure 3 This is a schematic diagram of the automatic login configuration interface for the automatic data acquisition engine. The automatic login configuration interface connects to the APIs of various subsystems for target data acquisition task registration and configuration.
[0047] The database collector collects data from a database. It uses JDBC (Java Database Connectivity), ODBC (Open Database Connectivity), or native drivers to retrieve KPIs (Key Performance Indicators) based on specific metrics. These KPIs include call detail record (CDR) entry success rate, CDR entry latency, CPU usage (Central Processing Unit) usage in the traffic parsing subsystem, memory usage in the traffic parsing subsystem, and queue length in the traffic parsing subsystem.
[0048] The manual survey data collector collects data through manual surveys. It can be configured with questionnaire templates. The manual survey data collector supports time limit settings, data collection retries, and data collection from city or work group levels, and supports automatic structured data import. Figure 4 This is a schematic diagram of the interface for data collection via manual survey using the automated data collection engine. Data is uploaded to the manual survey data collection interface, and after uploading, the entered data is displayed.
[0049] Optionally, after acquiring the target live system data, data cleaning and alignment are performed on the target live system data. For example, data cleaning and alignment methods include missing value alignment, unit conversion, field standardization, alias mapping, time alignment, and time zone unification. Alias mapping is used to align the names of indicator data with different names for the same indicator. Thus, by cleaning and aligning the target live system data, time-series alignment and content alignment of multi-source target live system data across subsystems are achieved, improving the accuracy of target live system operation and maintenance.
[0050] S102. Obtain the historical operation and maintenance knowledge base, and use the root cause localization big language model. Based on the target topology dependency graph, target indicator triples and historical operation and maintenance knowledge base, perform anomaly clustering and causal analysis on the target live network system to obtain candidate root cause fault information of the target live network system.
[0051] The root cause localization large language model is used to locate the root causes of a target network system. The input data for the root cause localization large language model includes the target topology dependency graph, target indicator triples, and a historical operation and maintenance knowledge base. The output of the root cause localization large language model is the candidate root cause fault information of the target network system. The root cause localization large language model performs anomaly clustering and causal analysis on the target network system. Specifically, the root cause localization large language model performs robust statistics, change point detection, and seasonal analysis residuals on the target network system to perform anomaly clustering and determine whether a fault exists in the target network system. Change point detection includes using CUSUM (Cumulative Sum Algorithm) or entropy increment algorithms, etc. The root cause localization large language model performs time lag correlation, path consistency scoring, and conflict evidence suppression along the target topology dependency graph to perform causal analysis on the target network system and determine the candidate root cause fault information of the target network system. For example, the root cause localization large language model is a large language model (LLM).
[0052] The historical operations and maintenance (O&M) knowledge base stores historical O&M knowledge of the target network system. This knowledge base is used for retrieval enhancement by the root cause localization language model and the O&M decision-making language model. For example, the historical O&M knowledge base includes historical general O&M knowledge, historical personalized O&M knowledge, and third-party O&M knowledge. Historical general O&M knowledge refers to O&M knowledge about common faults of the target network system. For example, it includes historical general O&M processes and their historical effectiveness. It could also be the target network system's run book. Historical personalized O&M knowledge refers to O&M knowledge about personalized faults of the target network system. For example, it includes the target network system's postmortem and change logs. Third-party O&M knowledge refers to the O&M knowledge of third-party systems regarding the target network system. For example, it includes equipment manuals and equipment provider's equipment O&M knowledge.
[0053] Candidate root cause fault information refers to the root cause information of potential faults in the target live network system predicted by the root cause localization large language model. For example, Figure 5 This is an interface display diagram for candidate root cause fault information. For example... Figure 5As shown, candidate root cause failure information includes candidate root cause failure links (e.g., dashboard monitoring indicators), candidate failure evidence chain summaries (e.g., source tracing indicators), and candidate root cause failure confidence (included in the warning details).
[0054] In an optional embodiment of the present invention, the candidate root cause fault information includes the candidate root cause fault link, the candidate fault impact area, the candidate fault recurrence probability, the candidate fault evidence chain summary, and the candidate root cause fault confidence level.
[0055] A candidate root cause failure is the root cause of a candidate root cause failure in the target live network system. For example, a candidate root cause failure is a component or a dependency relationship between two components in the target live network system. The candidate failure impact surface characterizes the degree of impact of the candidate root cause failure. The candidate failure recurrence frequency characterizes the frequency of occurrence of the candidate root cause failure. The candidate failure evidence chain summary characterizes the corresponding portion of the data in the target live network system for the candidate root cause failure. For example, the candidate failure evidence chain summary includes mutation index data, key log segments, and link anomaly data corresponding to the candidate root cause failure in the target live network system data. The candidate failure evidence chain summary characterizes the original data content corresponding to the candidate root cause failure in the target live network system data. The candidate root cause failure confidence level is the degree of credibility of the candidate root cause failure.
[0056] This solution improves the accuracy of root cause fault location in the existing network system by precisely pinpointing the specific candidate root cause fault links. By introducing candidate fault impact surface, candidate fault recurrence probability, candidate fault evidence chain summary, and candidate root cause fault confidence, it enhances the comprehensiveness of fault information for candidate root cause fault links, taking into account both the typicality and effectiveness of candidate root cause fault information, and further improves the accuracy of operation and maintenance decisions based on candidate root cause fault information.
[0057] Specifically, a pre-generated historical operation and maintenance knowledge base is acquired. Using a root cause localization language model, semantic understanding is performed on the target topology dependency graph, target indicator triples, and historical operation and maintenance knowledge base within a preset time window with the current time as the endpoint. Anomaly clustering and causal analysis are then performed on the target live network system to obtain candidate root cause fault information. For example, the candidate root cause fault information is the Top-N candidate root cause fault information.
[0058] S103. Obtain historical fault operation and maintenance processes, historical effectiveness of historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives. Using the operation and maintenance decision-making big language model, based on the candidate root cause fault information of the target network system, historical fault operation and maintenance processes, historical effectiveness of historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives, make operation and maintenance decisions for the target network system to obtain the target executable decision sequence.
[0059] The Operations and Maintenance Decision Large Language Model (OLM) is used to generate operations and maintenance decisions for the target live network system. The input data for the OLM includes candidate root cause fault information of the target live network system, historical fault operations and maintenance procedures, historical effectiveness of these procedures, target operations and maintenance constraints, and reference optimization objectives. The output of the OLM is a sequence of executable decisions for the target live network system. For example, the OLM is a Large Language Model (LLM).
[0060] The historical fault maintenance process refers to the actual maintenance process for historical faults in the target network system. For example, the historical fault maintenance process includes historical general fault maintenance processes and historical personalized maintenance processes. The historical effectiveness of the historical fault maintenance process refers to the actual execution effect of the historical fault maintenance process. The target maintenance constraint information refers to the constraint information for maintenance decisions made by the target network system. For example, the target maintenance constraint information includes time windows, approval, and risks. The time window characterizes the size of the time window referenced by the target network system in making maintenance decisions. Approval characterizes whether manual approval is required for a single candidate root cause fault in the target network system during maintenance. Risk characterizes whether the target network system has risks during maintenance and the corresponding risk range. The reference optimization goal characterizes the optimization strategy for the target network system. For example, the reference optimization goal includes prioritizing recovery speed, minimizing recovery risk, and / or minimizing business impact. The reference optimization goal is set by the maintenance team of the target network system or pre-set by this device. The target executable decision sequence consists of executable maintenance decisions for each candidate root cause fault. For example, the target executable decision sequence includes the target execution order between each target root cause failure link, the target failure operation and maintenance process of each target root cause failure link, and the target execution password within each target root cause failure link.
[0061] In an optional embodiment of the present invention, the target executable decision sequence includes the target execution order between each target root cause failure link, the target failure operation and maintenance process of each target root cause failure link, the target execution password within each target root cause failure link, the reference execution index within each target root cause failure link, the reference rollback point corresponding to the reference execution index, and the reference execution effectiveness criterion.
[0062] The target root cause failure stage is a candidate root cause failure stage that requires maintenance. For example, the target root cause failure stage includes a single subsystem or a component within a single subsystem in the target live network system. The target execution order is the execution order among the target root cause failure stages. The target failure maintenance process is the maintenance process for a single target root cause failure stage. The target execution password is the execution password required to execute the target failure maintenance process. The reference execution metric is the data dimension for detecting the maintenance status of the target root cause failure stage. The reference rollback point is the target root cause failure stage to which the target executable decision sequence rolls back when the maintenance effect of the target root cause failure is deemed unsatisfactory based on the reference execution metric. The reference execution effectiveness criterion is used to determine the maintenance effectiveness of the target root cause failure. For example, the reference execution effectiveness criterion includes reference execution metric values and reference recovery time distribution, etc. The reference recovery time distribution is used to characterize the changes in maintenance recovery time.
[0063] This solution concretizes the target executable decision sequence into the target execution order between each target root cause failure link, the target failure operation and maintenance process of each target root cause failure link, and the target execution password within each target root cause failure link, thereby realizing the executability of the target live network system. By introducing reference execution indicators, reference rollback points corresponding to the reference execution indicators, and reference execution effectiveness criteria, it facilitates the execution verification of the target executable decision sequence.
[0064] Specifically, the process involves acquiring historical fault maintenance procedures, their historical effectiveness, target maintenance constraints, and reference optimization goals for the target network system. A large language model for maintenance decision-making is then used to semantically understand the candidate root cause fault information, historical fault maintenance procedures, their historical effectiveness, target maintenance constraints, and reference optimization goals of the target network system. This understanding enables maintenance decisions to be made for the target network system, resulting in a sequence of executable decisions.
[0065] S104. Based on the target executable decision sequence, generate the target operation and maintenance task sequence for the target live network system, and perform operation and maintenance on the target live network system according to the target operation and maintenance task sequence.
[0066] The target operation and maintenance task sequence is used to characterize the transformation result of converting the target executable decision sequence into actual operation and maintenance tasks. In comparison, the target executable decision sequence is mainly used to characterize the execution order of the targets and the target fault operation and maintenance process used to operate and maintain each target root cause fault link. The target operation and maintenance task sequence is mainly used to characterize how to execute each target operation and maintenance task to achieve the operation and maintenance of the target executable decision sequence. For example, the target operation and maintenance task sequence includes each target operation and maintenance task and the target operation and maintenance order of each target operation and maintenance task. The target operation and maintenance order is the execution order of each target operation and maintenance task. For example, target operation and maintenance tasks include scaling up, switching, rate limiting, restarting, rollback, or policy issuance tasks. For example, a target operation and maintenance task is a DAG task (Directed Acyclic Graph Task). Optionally, the target operation and maintenance task corresponds to a work order or change order. Optionally, the target operation and maintenance task supports semi-automatic approval.
[0067] Specifically, based on the target execution order among the target root cause failure links contained in the target executable decision sequence, the target failure operation and maintenance process of each target root cause failure link, and the target execution password within each target root cause failure link, the target executable decision sequence is transformed into target operation and maintenance tasks with a target operation and maintenance sequence, generating a target operation and maintenance task sequence for the target live network system. Operation and maintenance are then performed on the target live network system according to the target operation and maintenance task sequence.
[0068] For target network systems that span multiple systems and regions, fault location and system maintenance require consideration of the coupling between subsystems within the target network system and the complexity of each individual subsystem. Therefore, fault location and system maintenance of target network systems are extremely difficult.
[0069] The technical solution of this invention employs a root cause localization language model. Based on the target topology dependency graph, target indicator triples, and historical operation and maintenance knowledge base of the target network system, it performs anomaly clustering and causal analysis on the target network system. It leverages the semantic understanding capability of the root cause localization language model itself to improve the accuracy of root cause localization of faults in the target network system. The target topology dependency graph considers the dependencies between components of each subsystem in the target network system, i.e., the coupling between subsystems. The target indicator triples achieve semantic space unification of the target network system data from three dimensions: "indicator-component-business," and unified orchestration of multi-source data across domains and systems, thereby improving the recognition accuracy of the root cause localization language model. Furthermore, the introduction of a historical operation and maintenance knowledge base enhances the retrieval of root cause faults in the target network system, further improving the accuracy of root cause localization. By employing an operation and maintenance decision-making language model… This model, based on candidate root cause fault information, historical fault operation and maintenance processes, historical effectiveness of historical fault operation and maintenance processes, target operation and maintenance constraints, and reference optimization objectives of the target network system, makes operation and maintenance decisions for the target network system, resulting in a target executable decision sequence. It leverages the semantic understanding capabilities of the root cause localization language model itself to improve the accuracy of operation and maintenance decisions for the target network system. By considering the historical effectiveness of historical fault operation and maintenance processes, it achieves closed-loop application of historical fault operation and maintenance processes. Furthermore, it introduces target operation and maintenance constraints and reference optimization objectives to improve the flexibility of the target network system, thereby enhancing the feasibility of the target executable decision sequence. By transforming the target executable decision sequence into a target operation and maintenance task sequence for the target network system, and performing operation and maintenance according to the target operation and maintenance task sequence, it achieves comprehensive operation and maintenance of the target network system, improving the accuracy of fault localization and the effectiveness of system operation and maintenance.
[0070] In an optional embodiment of the present invention, the method further includes: periodically acquiring candidate root cause fault information, actual executable decision sequences, actual operation and maintenance task sequences, and actual execution data of the target network system; and using a long-term governance big language model to perform governance analysis on the candidate root cause fault information, actual executable decision sequences, actual operation and maintenance task sequences, and actual execution data of the target network system to obtain a target governance roadmap.
[0071] The actual executable decision sequence is the target executable decision sequence actually used during the operation and maintenance of the target live network system. The actual operation and maintenance task sequence is the operation and maintenance task sequence corresponding to the actual executable decision sequence. The actual execution data is the execution data when performing operation and maintenance on the target live network system using the actual operation and maintenance task sequence. For example, Figure 6This is an interface display of the fault handling tracking content included in the actual execution data. For example, fault handling tracking includes task status, fault category, fault cause, solution, recording time, recording personnel, and related similar tasks. Task status includes unprocessed, in progress, and completed. Fault category represents the type of fault actually managed. Fault cause represents the actual reason for the fault. Related similar tasks represent historical similar operation and maintenance tasks.
[0072] The Long-Term Governance Large Language Model (LLM) is used for long-term governance of the target live network system. The input data for the LLM includes candidate root cause failure information, actual executable decision sequences, actual operation and maintenance task sequences, and actual execution data of the target live network system. The output of the LLM is the target governance roadmap for the target live network system. For example, the LLM is a Large Language Model (LLM).
[0073] The target governance roadmap is used to periodically govern the target live network system. For example, the target governance roadmap includes at least two weak points, governance priorities and recommendations for each weak point, and milestone weak points within each weak point. Weak points are those in the target live network system that require long-term governance. Governance priorities characterize the impact of a weak point. For example, governance priorities include the scope of impact and frequency of occurrence. Governance recommendations characterize how to govern weak points. For example, governance recommendations include historical fault maintenance processes and historical results corresponding to the weak point. Milestone weak points characterize the critical weak points within each weak point. These can be understood as weak points in the target governance roadmap that need to be addressed at specific stages.
[0074] Specifically, candidate root cause failure information, actual executable decision sequences, actual operation and maintenance task sequences, and actual execution data of the target live network system are collected periodically. A long-term governance big language model is used to perform semantic understanding on the candidate root cause failure information, actual executable decision sequences, actual operation and maintenance task sequences, and actual execution data of the target live network system, and output the target governance roadmap.
[0075] This solution periodically acquires candidate root cause fault information, actual executable decision sequences, actual operation and maintenance task sequences, and actual execution data of the target network system. It employs a long-term governance big language model to perform governance analysis on these data, enabling periodic structured assessment of the target network system. By obtaining a target governance roadmap, the operation and maintenance of the target network system is transformed from post-fault governance to engineering-based periodic governance.
[0076] Example 2
[0077] Figure 7 This is a flowchart of a live network system operation and maintenance method provided in Embodiment 2 of the present invention. Based on the above embodiments, this embodiment of the present invention specifies "operating and maintaining the target live network system according to the target operation and maintenance task sequence" as "selecting and executing the current target operation and maintenance task according to the target operation and maintenance task sequence; when the current target operation and maintenance task is completed, detecting the current execution effectiveness of the current execution indicators of the current target operation and maintenance task; updating and executing the current target operation and maintenance task according to the current execution effectiveness; returning to the step of detecting the current execution effectiveness of the current target operation and maintenance task when the current target operation and maintenance task is completed, until all target operation and maintenance tasks in the target operation and maintenance task sequence are completed." This achieves quantitative detection of the operation and maintenance effect of the current target operation and maintenance task, forms a closed loop between the operation and maintenance effect of the current operation and maintenance task and the actual operation and maintenance process of the target live network system, realizes adaptive adjustment of the current target operation and maintenance task, and thus improves the overall operation and maintenance effect of the target live network system. It should be noted that parts not described in detail in this embodiment of the present invention can be referred to in the descriptions of other embodiments.
[0078] See Figure 7 The live network system operation and maintenance methods shown include:
[0079] S701. Obtain the target network system data, and generate the target topology dependency graph and target indicator triplet of the target network system based on the target network system data.
[0080] S702. Using a root cause localization language model, obtain the historical operation and maintenance knowledge base, and based on the target topology dependency graph, target indicator triples and the historical operation and maintenance knowledge base, perform anomaly clustering and causal analysis on the target live network system to obtain candidate root cause fault information of the target live network system.
[0081] S703. Using a large language model for operation and maintenance decision-making, obtain historical fault operation and maintenance processes, historical effectiveness of historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives. Based on the candidate root cause fault information of the target network system, historical fault operation and maintenance processes, historical effectiveness of historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives, make operation and maintenance decisions for the target network system to obtain the target executable decision sequence.
[0082] S704. Select and execute the current target maintenance task according to the target maintenance task sequence.
[0083] The current target maintenance task is the target maintenance task being executed at the current moment. For example, the current target maintenance task is the first target maintenance task in the target maintenance task sequence that can be executed.
[0084] Specifically, based on the target operation and maintenance tasks contained in the target operation and maintenance task sequence and the target operation and maintenance order of each target operation and maintenance task, the target operation and maintenance task with the earliest target operation and maintenance order is selected from among the executable target operation and maintenance tasks, and the current target operation and maintenance task is determined and executed.
[0085] S705. When the current target maintenance task is completed, check the current execution effectiveness of the current execution indicators of the current target maintenance task.
[0086] The current execution metric is a data dimension used to detect the operational effectiveness of the current target operational task. The current execution effectiveness characterizes the operational effectiveness of the current target operational task. Reference execution metrics in the target executable decision sequence correspond to the current execution metrics. For example, the current execution effectiveness includes whether the operational effectiveness of the current target operational task reaches the preset operational effectiveness or whether the operational effectiveness of the current target operational task does not reach the preset operational effectiveness.
[0087] Specifically, when the current target maintenance task is completed, the reference execution effectiveness criteria in the target executable decision sequence are used to detect the current execution indicators of the current target maintenance task and determine the current execution effectiveness of the current target maintenance task.
[0088] S706. Based on the current execution results, update and execute the current target operation and maintenance task, and return to the step of detecting the current execution results of the current execution indicators of the current target operation and maintenance task when the current target operation and maintenance task is completed, until all target operation and maintenance tasks in the target operation and maintenance task sequence are completed.
[0089] Specifically, when the operational effect of the current target operation and maintenance task reaches the preset operational effect, the next target operation and maintenance task in the target operation and maintenance task sequence is updated to the current target operation and maintenance task and executed. When the operational effect of the current target operation and maintenance task does not reach the preset operational effect, the current target operation and maintenance task is adjusted, updated to the current target operation and maintenance task, and executed. The process then returns to the step of checking the current execution effectiveness of the current execution metrics of the current target operation and maintenance task upon completion, until all target operation and maintenance tasks in the target operation and maintenance task sequence have been completed.
[0090] In an optional embodiment of the present invention, updating and executing the current target operation and maintenance task based on the current execution results includes: when the current execution indicator value of the current target operation and maintenance task reaches the first reference key indicator value, updating the next target operation and maintenance task in the target operation and maintenance task sequence to the current target operation and maintenance task, and executing it; when the current execution indicator value of the current target operation and maintenance task does not reach the first reference key indicator value, performing a rollback on the target operation and maintenance task sequence according to the reference rollback point contained in the target executable decision sequence, and updating the current target operation and maintenance task.
[0091] The first reference key performance indicator (KPI) is used to measure the operational effectiveness of the current target maintenance task. The KPI is the lower limit of the current execution KPI value when the operational effectiveness of the current target maintenance task reaches the preset operational effectiveness. If the current execution KPI value of the current target maintenance task reaches the first reference KPI value, it can be understood that the operational effectiveness of the current target maintenance task has reached the preset operational effectiveness. If the current execution KPI value of the current target maintenance task does not reach the first reference KPI value, it can be understood that the operational effectiveness of the current target maintenance task has not reached the preset operational effectiveness. The reference rollback point is the candidate root cause failure node to which the rollback occurs when the current execution KPI value does not reach the first reference KPI value.
[0092] Specifically, when the current execution metric value of the current target operation and maintenance task reaches the first reference key metric value, the next target operation and maintenance task in the target operation and maintenance task sequence is updated to the current target operation and maintenance task and executed. When the current execution metric value of the current target operation and maintenance task has not reached the first reference key metric value, the target operation and maintenance task sequence is rolled back according to the reference rollback point contained in the target executable decision sequence, so as to roll back to the target operation and maintenance task corresponding to the reference rollback point, and the current target operation and maintenance task is updated.
[0093] This solution introduces a detection mechanism to check whether the current execution indicator value of the current target maintenance task has reached the first reference key indicator value. This allows for rapid detection of the maintenance effectiveness of the current target maintenance task. When the current execution indicator value of the current target maintenance task reaches the first reference key indicator value, the next target maintenance task in the target maintenance task sequence is updated to the current target maintenance task. When the current execution indicator value of the current target maintenance task has not reached the first reference key indicator value, the target maintenance task sequence is rolled back based on the reference rollback point contained in the target executable decision sequence, and the current target maintenance task is updated. The introduction of the reference rollback point enables rapid rollback of the target maintenance task sequence when the current execution indicator value of the current target maintenance task has not reached the first reference key indicator value, improving the executability of the target network system's maintenance, thereby improving the efficiency and accuracy of maintenance adjustments to the target network system.
[0094] In an optional embodiment of the present invention, an operation and maintenance decision-making big language model is employed to make operation and maintenance decisions on the target network system based on candidate root cause fault information, historical fault operation and maintenance processes, historical effectiveness of historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives, and output a target executable decision sequence. This includes: using the operation and maintenance decision-making big language model, making operation and maintenance decisions on the target network system based on candidate root cause fault information, historical fault operation and maintenance processes, historical effectiveness of historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives, and outputting at least two candidate executable decisions. The process includes: determining the confidence level of each candidate executable decision sequence and each candidate executable decision sequence; selecting the target executable decision sequence from among the candidate executable decision sequences based on the confidence level of each candidate decision; and, after detecting the current execution effectiveness of the current execution indicator of the current target operation and maintenance task, further including: when the current execution indicator value of the current target operation and maintenance task does not reach the second reference key indicator value, selecting the target executable decision sequence from among the candidate executable decisions other than the target executable decision sequence based on the confidence level of each candidate decision, and returning to the execution step of generating the target operation and maintenance task sequence of the target live network system based on the target executable decision sequence.
[0095] Candidate executable decision sequences are selectable executable operation and maintenance decisions when performing operation and maintenance on each candidate root cause failure. Candidate executable decision sequences are used to screen target executable decision sequences. Candidate decision confidence is used to characterize the reliability of candidate executable decision sequences. The second reference key indicator value is used to measure the operation and maintenance effectiveness of the current target operation and maintenance task. The second reference key indicator value is the lower limit of the current execution indicator value when the target executable decision sequence needs to be changed. If the current execution indicator value of the current target operation and maintenance task reaches the second reference key indicator value, it can be understood that the target operation and maintenance task sequence corresponding to the target executable decision sequence should continue to be used to operate and maintain the target live network system. If the current execution indicator value of the current target operation and maintenance task does not reach the second reference key indicator value, it can be understood that the target executable decision sequence needs to be changed.
[0096] Specifically, a large language model for operation and maintenance (O&M) decisions is used to semantically understand the candidate root cause fault information, historical fault O&M processes, historical effectiveness of historical fault O&M processes, target O&M constraints, and reference optimization objectives of the target live network system. This allows for O&M decision-making on the target live network system, resulting in at least two candidate executable decision sequences and the confidence scores of each candidate executable decision sequence. The candidate executable decision sequence with the highest confidence score is selected as the executable decision sequence. After detecting the current execution effectiveness of the current execution metric for the current target O&M task, when the current execution metric value of the current target O&M task reaches the second reference key metric value, the current execution metric value of the current target O&M task is compared with the first reference key metric value, and the process returns to the step of detecting the current execution effectiveness of the current execution metric of the current target O&M task. When it is detected that the current execution indicator value of the current target operation and maintenance task has not reached the second reference key indicator value, among all candidate executable decisions other than the target executable decision sequence, the candidate executable decision sequence with the highest confidence is selected as the target executable decision sequence, and the process of generating the target operation and maintenance task sequence of the target network system based on the target executable decision sequence is returned.
[0097] This solution introduces a comparison process between the current execution indicator value of the current target operation and maintenance task and the second reference key indicator value. When the current execution indicator value of the current target operation and maintenance task does not reach the second reference key indicator value, it is adjusted to the suboptimal executable decision sequence, which takes into account the fault tolerance and executability of the target network system operation and maintenance, and improves the operation and maintenance efficiency of the target network system.
[0098] The technical solution of this invention, during the operation and maintenance of the target network system, detects the current execution effectiveness of the current execution indicators of the current target operation and maintenance task upon completion, thereby achieving quantitative detection of the operation and maintenance effect of the current target operation and maintenance task. Based on the current execution effectiveness of the current execution indicators of the current target operation and maintenance task, the current target operation and maintenance task is updated and executed, forming a closed loop between the operation and maintenance effect of the current operation and maintenance task and the actual operation and maintenance process of the target network system. This achieves adaptive adjustment of the current target operation and maintenance task, thereby improving the overall operation and maintenance effect of the target network system.
[0099] For example, based on the above embodiments, the present invention also provides a preferred embodiment. This network system operation and maintenance method includes the following technical features:
[0100] S1. Register the target acquisition task on the target network system, and select the target collector and configure the target acquisition strategy according to the target acquisition task.
[0101] For example, target data collectors include web page collectors, software collectors, database collectors, and manual survey collectors.
[0102] S2. Schedule the target collector, execute the corresponding target collection task according to the target collection strategy, obtain the target network system data, and preprocess the target network system data.
[0103] S3. Generate a target topology dependency graph and target index triples based on the preprocessed target network system data.
[0104] The target topology dependency graph reflects the dependencies between components of various subsystems in the target live network system. The target metric triple includes the ternary relationship between the target metric, the target component, and the target service.
[0105] S4. Obtain historical operation and maintenance knowledge base.
[0106] The historical operations and maintenance (O&M) knowledge base includes historical general O&M knowledge, historical personalized O&M knowledge, and third-party O&M knowledge. Historical general O&M knowledge includes historical general O&M processes and their historical effectiveness. For example, historical general O&M knowledge is the target network system's run book. Historical personalized O&M knowledge includes the target network system's postmortem and change logs. Third-party O&M knowledge includes equipment manuals and equipment provider's O&M knowledge.
[0107] S5. Using the LLM analysis, based on the target topology dependency graph, target indicator triples and historical operation and maintenance knowledge base, perform anomaly clustering and causal analysis on the target live network system, and output the Top-N candidate root cause fault information of the target live network system.
[0108] The analysis focuses on the LLM (Large Language Model for Root Cause Analysis). Candidate root cause fault information includes the faulty component, its impact area, recurrence probability, a summary of the evidence chain, and its confidence level. The summary of the evidence chain is a readable string of evidence obtained by concatenating mutation index data, key log segments, and link anomaly data corresponding to the candidate root cause fault in the target live network system data.
[0109] S6. Using Decision LLM, based on the Top-N candidate root cause fault information of the target network system, historical general operation and maintenance processes, historical effectiveness of historical general operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives, operation and maintenance decisions are made for the target network system, and the target executable decision sequence is output.
[0110] Decision LLM, or Operations and Maintenance Decision Large Language Model, includes the target executable decision sequence, which comprises the target execution order between each target root cause failure link, the target failure operation and maintenance process for each target root cause failure link, the target execution password within each target root cause failure link, the reference execution indicators within each target root cause failure link, the reference rollback points corresponding to the reference execution indicators, and the reference execution effectiveness criteria.
[0111] S7. Based on the target executable decision sequence, generate the target operation and maintenance task sequence, execute the current target operation and maintenance task, and detect the current execution indicators; when the current execution indicator reaches the corresponding reference execution effectiveness criterion, update and execute the current target operation and maintenance task based on the target operation and maintenance task sequence; when the current execution indicator does not reach the corresponding reference execution effectiveness criterion, trigger the target operation and maintenance task sequence rollback and update the target executable decision sequence.
[0112] S8. When the target operation and maintenance task sequence is completed, update the historical general operation and maintenance process and the historical effectiveness of the historical general operation and maintenance process based on the actual execution data of the target operation and maintenance task sequence.
[0113] S9. Periodically acquire Top-N candidate root cause fault information, target executable decision sequence, target operation and maintenance task sequence, and actual execution data of the target operation and maintenance task sequence of the target live network system. Use Long-Term Governance (LLM) to analyze the Top-N candidate root cause fault information, target executable decision sequence, target operation and maintenance task sequence, and actual execution data of the target live network system, and output the target governance roadmap.
[0114] Among them, Long-Term Governance LLM is the Long-Term Governance Big Language Model. The target governance roadmap includes the Top-N weak fault links, the governance priorities and governance recommendations for each weak fault link, and the milestone weak fault links in each weak fault link.
[0115] This solution integrates web systems, software systems, database systems, and manual survey systems into a single orchestration framework using four types of automatic data collectors. It provides unified methods for automatic login, retrieval, extraction, cleaning, entity alignment, and time alignment, significantly improving the consistency and comparability of data arrival. This achieves unified orchestration of live network data and enhances the consistency of data across live network systems. By inputting the target topology dependency graph, target indicator triples, and historical operation and maintenance knowledge base for long-haul telecommunications links into the root cause localization language model for analysis, the solution generates a target data set along the links from the traffic splitting subsystem to the traffic parsing subsystem, call detail record (CDR) database entry subsystem, data analysis subsystem, and data early warning subsystem. This is achieved by combining topology dependency, time lag correlation, historical similarity, and conflict evidence suppression. By selecting root cause failure confidence and candidate root cause failure evidence chain summaries, the interpretability and hit rate of cross-domain problems are improved. Combining historical failure operation and maintenance processes, historical effectiveness of historical failure operation and maintenance processes, and target operation and maintenance constraints, a large language model for operation and maintenance decision-making is used to generate a target executable decision sequence, which is then transformed into an executable, rollbackable, and verifiable target operation and maintenance task sequence, turning "suggestions" into "implementable" ones. By conducting regular structured analysis of high-frequency failures, capacity bottlenecks, strategy or model drift, a target governance roadmap is generated, and the governance results are fed back into the historical operation and maintenance knowledge base and the long-term governance large language model, enabling the operation and maintenance of the target network system to shift from "firefighting" to "engineering governance," thus achieving long-term analysis and stability governance of the target network system.
[0116] Example 3
[0117] Figure 8 This is a schematic diagram of a network system operation and maintenance device provided in Embodiment 3 of the present invention. This embodiment of the present invention is applicable to the comprehensive operation and maintenance of existing network systems. The device executes network system operation and maintenance methods and is implemented in hardware and / or software. The device can be configured in electronic devices that carry out network system operation and maintenance functions.
[0118] See Figure 8The network system operation and maintenance device shown includes: a target network system data acquisition module 801, a fault root cause localization module 802, an operation and maintenance decision output module 803, and a target network system operation and maintenance module 804. Specifically, the target network system data acquisition module 801 acquires target network system data and generates a target topology dependency graph and target indicator triples based on the data; the fault root cause localization module 802 acquires historical operation and maintenance knowledge base and uses a root cause localization large language model to perform anomaly clustering and causal analysis on the target topology dependency graph, target indicator triples, and historical operation and maintenance knowledge base to obtain candidate root cause fault information of the target network system; the operation and maintenance decision output module 803 acquires historical fault operation and maintenance processes, historical fault operation and maintenance procedures, and other relevant information. The system analyzes the historical effectiveness of maintenance processes, target maintenance constraints, and reference optimization goals. Using a large language model for maintenance decisions, it makes maintenance decisions for the target network system based on candidate root cause fault information, historical fault maintenance processes, historical effectiveness of historical fault maintenance processes, target maintenance constraints, and reference optimization goals, outputting a target executable decision sequence. The target network system maintenance module 804 generates a target maintenance task sequence for the target network system based on the target executable decision sequence and performs maintenance on the target network system according to the target maintenance task sequence.
[0119] The technical solution of this invention employs a root cause localization language model. Based on the target topology dependency graph, target indicator triples, and historical operation and maintenance knowledge base of the target network system, it performs anomaly clustering and causal analysis on the target network system. It leverages the semantic understanding capability of the root cause localization language model itself to improve the accuracy of root cause localization of faults in the target network system. The target topology dependency graph considers the dependencies between components of each subsystem in the target network system, i.e., the coupling between subsystems. The target indicator triples achieve semantic space unification of the target network system data from three dimensions: "indicator-component-business," and unified orchestration of multi-source data across domains and systems, thereby improving the recognition accuracy of the root cause localization language model. Furthermore, the introduction of a historical operation and maintenance knowledge base enhances the retrieval of root cause faults in the target network system, further improving the accuracy of root cause localization. By employing an operation and maintenance decision-making language model… This model, based on candidate root cause fault information, historical fault operation and maintenance processes, historical effectiveness of historical fault operation and maintenance processes, target operation and maintenance constraints, and reference optimization objectives of the target network system, makes operation and maintenance decisions for the target network system, resulting in a target executable decision sequence. It leverages the semantic understanding capabilities of the root cause localization language model itself to improve the accuracy of operation and maintenance decisions for the target network system. By considering the historical effectiveness of historical fault operation and maintenance processes, it achieves closed-loop application of historical fault operation and maintenance processes. Furthermore, it introduces target operation and maintenance constraints and reference optimization objectives to improve the flexibility of the target network system, thereby enhancing the feasibility of the target executable decision sequence. By transforming the target executable decision sequence into a target operation and maintenance task sequence for the target network system, and performing operation and maintenance according to the target operation and maintenance task sequence, it achieves comprehensive operation and maintenance of the target network system, improving the accuracy of fault localization and the effectiveness of system operation and maintenance.
[0120] In an optional embodiment of the present invention, the target network system operation and maintenance module 804 includes: a current target operation and maintenance task filtering unit, used to select and execute a current target operation and maintenance task according to the target operation and maintenance task sequence; a current execution effectiveness detection unit, used to detect the current execution effectiveness of the current execution indicators of the current target operation and maintenance task when the current target operation and maintenance task is completed; and a current target operation and maintenance task updating unit, used to update and execute the current target operation and maintenance task according to the current execution effectiveness, and return to the step of detecting the current execution effectiveness of the current execution indicators of the current target operation and maintenance task when the current target operation and maintenance task is completed, until all target operation and maintenance tasks in the target operation and maintenance task sequence are completed.
[0121] In an optional embodiment of the present invention, the current target operation and maintenance task update unit includes: a first current target operation and maintenance task update unit, configured to update the next target operation and maintenance task in the target operation and maintenance task sequence to the current target operation and maintenance task when the current execution indicator value of the current target operation and maintenance task reaches the first reference key indicator value, and execute it; and a second current target operation and maintenance task update unit, configured to perform a rollback on the target operation and maintenance task sequence and update the current target operation and maintenance task when the current execution indicator value of the current target operation and maintenance task does not reach the first reference key indicator value.
[0122] In an optional embodiment of the present invention, the operation and maintenance decision output module 803 includes: a candidate executable decision sequence output unit, used to use an operation and maintenance decision big language model to make operation and maintenance decisions on the target network system based on candidate root cause fault information, historical fault operation and maintenance processes, historical effectiveness of historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives, and output at least two candidate executable decision sequences and the candidate decision confidence scores of each candidate executable decision sequence; an operation and maintenance decision output unit, used to filter target executable decision sequences from each candidate executable decision sequence according to the confidence scores of each candidate decision; the target network system operation and maintenance module 804 further includes: a target executable decision sequence adjustment unit, used to, after detecting the current effectiveness of the current execution indicator of the current target operation and maintenance task, when the current execution indicator value of the current target operation and maintenance task does not reach the second reference key indicator value, filter target executable decision sequences from each candidate executable decision sequence other than the target executable decision sequence according to the confidence scores of each candidate decision, and return to the step of generating the target operation and maintenance task sequence of the target network system according to the target executable decision sequence.
[0123] In an optional embodiment of the present invention, the candidate root cause fault information includes the candidate root cause fault stage, the candidate fault impact area, the candidate fault recurrence probability, the candidate fault evidence chain summary, and the candidate root cause fault confidence level; the target executable decision sequence includes the target execution order between each target root cause fault stage, the target fault operation and maintenance process of each target root cause fault stage, the target execution password within each target root cause fault stage, the reference execution index within each target root cause fault stage, the reference rollback point corresponding to the reference execution index, and the reference execution effectiveness criterion.
[0124] In an optional embodiment of the present invention, the device further includes: an actual operation and maintenance data acquisition module, used to periodically acquire candidate root cause fault information, actual executable decision sequence, actual operation and maintenance task sequence, and actual execution data of the target live network system; and a long-term fault governance module, used to perform governance analysis on the candidate root cause fault information, actual executable decision sequence, actual operation and maintenance task sequence, and actual execution data of the target live network system using a long-term governance big language model to obtain a target governance roadmap.
[0125] The network system operation and maintenance device provided in the embodiments of the present invention can execute the network system operation and maintenance method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0126] In the technical solutions of this invention, the acquisition, storage, and application of target network system data, historical operation and maintenance knowledge base, historical fault operation and maintenance process, historical effectiveness of historical fault operation and maintenance process, target operation and maintenance constraint information, reference optimization targets, candidate root cause fault information of target network system, actual executable decision sequence, actual operation and maintenance task sequence, and actual execution data all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0127] Example 4
[0128] Figure 9 A schematic diagram of an electronic device 900 for implementing embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device also represents various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0129] like Figure 9As shown, the electronic device 900 includes at least one processor 901 and a memory, such as a read-only memory (ROM) 902 or a random access memory (RAM) 903, communicatively connected to the at least one processor 901. The memory stores computer programs executable by the at least one processor. The processor 901 performs various appropriate actions and processes based on the computer program stored in the ROM 902 or loaded into the RAM 903 from storage unit 908. The RAM 903 may also store various programs and data required for the operation of the electronic device 900. The processor 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0130] Multiple components in electronic device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of displays, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows electronic device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0131] Processor 901 is a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 901 include, but are not limited to, central processing units (CPUs), graphics processing units (GPUs), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processors, controllers, microcontrollers, etc. Processor 901 performs the various methods and processes described above, such as live network system operation and maintenance methods.
[0132] In some embodiments, the network system operation and maintenance method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 908. In some embodiments, part or all of the computer program is loaded into and / or installed on electronic device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by processor 901, one or more steps of the network system operation and maintenance method described above are performed. Alternatively, in other embodiments, processor 901 may be configured to perform the network system operation and maintenance method by any other suitable means (e.g., by means of firmware).
[0133] Various embodiments of the systems and techniques described above herein are implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which is a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.
[0134] Computer programs used to implement the methods of the present invention are written in any combination of one or more programming languages. These computer programs are provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may execute entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0135] In the context of this invention, a computer-readable storage medium is a tangible medium that contains or stores a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. Computer-readable storage media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium is a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0136] To provide interaction with a user, the systems and techniques described herein are implemented on an electronic device, which includes: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices are also used to provide interaction with the user; for example, feedback provided to the user is any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user is received in any form (including sound input, voice input, or tactile input).
[0137] The systems and technologies described herein are implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with a graphical user interface or web browser through which a user interacts with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system are interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0138] A computing system consists of clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. A cloud server, also known as a cloud computing server or cloud host, is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability.
[0139] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0140] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for operating and maintaining a live network system, characterized in that, The method includes: Obtain target network system data, and based on the target network system data, generate target topology dependency graph and target indicator triplet for the target network system; The historical operation and maintenance knowledge base is acquired, and the root cause localization big language model is used. Based on the target topology dependency graph, the target indicator triplet and the historical operation and maintenance knowledge base, anomaly clustering and causal analysis are performed on the target live network system to obtain candidate root cause fault information of the target live network system. The system acquires historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization targets. Then, using a large language model for operation and maintenance decisions, it makes operation and maintenance decisions on the target network system based on the candidate root cause fault information of the target network system, the historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization targets, to obtain a target executable decision sequence. Based on the target executable decision sequence, a target operation and maintenance task sequence for the target live network system is generated, and the target live network system is operated and maintained according to the target operation and maintenance task sequence.
2. The method according to claim 1, characterized in that, The step of performing maintenance on the target existing network system according to the target maintenance task sequence includes: Select and execute the current target operation and maintenance task according to the target operation and maintenance task sequence; When the current target operation and maintenance task is completed, the current execution effectiveness of the current execution indicators of the current target operation and maintenance task is detected; Based on the current execution results, update and execute the current target operation and maintenance task, and return to the step of detecting the current execution results of the current execution indicators of the current target operation and maintenance task when the current target operation and maintenance task is completed, until all target operation and maintenance tasks in the target operation and maintenance task sequence are completed.
3. The method according to claim 2, characterized in that, The step of updating and executing the current target maintenance task based on the current execution results includes: When the current execution indicator value of the current target operation and maintenance task reaches the first reference key indicator value, the next target operation and maintenance task in the target operation and maintenance task sequence is updated to the current target operation and maintenance task and executed. When the current execution indicator value of the current target operation and maintenance task does not reach the first reference key indicator value, the target operation and maintenance task sequence is rolled back according to the reference rollback point contained in the target executable decision sequence, and the current target operation and maintenance task is updated.
4. The method according to claim 2, characterized in that, The method employs a large language model for operation and maintenance decisions. Based on the candidate root cause fault information of the target network system, the historical fault operation and maintenance processes, the historical effectiveness of the historical fault operation and maintenance processes, the target operation and maintenance constraints, and the reference optimization objectives, it makes operation and maintenance decisions for the target network system, resulting in a target executable decision sequence, including: Using a large language model for operation and maintenance decision-making, based on the candidate root cause fault information of the target network system, the historical fault operation and maintenance process, the historical effectiveness of the historical fault operation and maintenance process, the target operation and maintenance constraint information, and the reference optimization objective, operation and maintenance decisions are made for the target network system, and at least two candidate executable decision sequences and the candidate decision confidence of each candidate executable decision sequence are output. Based on the confidence levels of each candidate decision, a target executable decision sequence is selected from each of the candidate executable decision sequences; After detecting the current execution effectiveness of the current execution metrics of the current target operation and maintenance task, the method further includes: When the current execution indicator value of the current target operation and maintenance task does not reach the second reference key indicator value, the target executable decision sequence is selected from the candidate executable decisions other than the target executable decision sequence according to the confidence level of each candidate decision, and the process of generating the target operation and maintenance task sequence of the target network system according to the target executable decision sequence is returned.
5. The method according to claim 1, characterized in that, The candidate root cause fault information includes the candidate root cause fault stage, the impact area of the candidate fault, the recurrence probability of the candidate fault, the summary of the evidence chain of the candidate fault, and the confidence level of the candidate root cause fault; the target executable decision sequence includes the target execution order between each target root cause fault stage, the target fault operation and maintenance process of each target root cause fault stage, the target execution password within each target root cause fault stage, the reference execution indicators within each target root cause fault stage, the reference rollback point corresponding to the reference execution indicators, and the reference execution effectiveness criteria.
6. The method according to claim 1, characterized in that, Also includes: Periodically acquire candidate root cause fault information, actual executable decision sequences, actual operation and maintenance task sequences, and actual execution data of the target live network system; Using a long-term governance big language model, governance analysis is performed on the candidate root cause fault information, the actual executable decision sequence, the actual operation and maintenance task sequence, and the actual execution data of the target network system to obtain the target governance roadmap.
7. A network system operation and maintenance device, characterized in that, The device includes: The target live network system data acquisition module is used to acquire target live network system data and generate target topology dependency graph and target indicator triplet of the target live network system based on the target live network system data; The fault root cause localization module is used to acquire the historical operation and maintenance knowledge base, and use the root cause localization big language model to perform anomaly clustering and causal analysis on the target topology dependency graph, the target indicator triplet and the historical operation and maintenance knowledge base to obtain candidate root cause fault information of the target live network system. The operation and maintenance decision output module is used to acquire historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives. It adopts the operation and maintenance decision big language model to make operation and maintenance decisions for the target network system based on the candidate root cause fault information of the target network system, the historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives, and outputs the target executable decision sequence. The target live network system operation and maintenance module is used to generate a target operation and maintenance task sequence for the target live network system based on the target executable decision sequence, and to perform operation and maintenance on the target live network system according to the target operation and maintenance task sequence.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the live network system operation and maintenance method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the network system operation and maintenance method according to any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the network system operation and maintenance method according to any one of claims 1-6.
Citation Information
Patent Citations
Large model enhanced equipment operation and maintenance multi-modal knowledge graph construction method and system and storage medium
CN120450011A
Operation and maintenance data exception positioning method and computing equipment
CN120596301A
Intelligent operation and maintenance method and device based on knowledge graph and large model and electronic equipment
CN120596306A
Network fault root cause determining method and apparatus, device, and storage medium
WO2023010823A1