Network system operation and maintenance method, device, equipment, medium and program product
By combining the root cause localization language model and the operation and maintenance decision-making language model with the topology dependency graph and historical operation and maintenance knowledge base, the comprehensive problems of operation and maintenance of the live network system are solved, and the accuracy of fault location and the effectiveness of operation and maintenance are achieved.
Patent Information
- Application Number
- CN202511384698.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-09-25
AI Technical Summary
Cross-system operation and maintenance of existing network systems is difficult to achieve integrated operation and maintenance, resulting in inaccurate fault location and ineffective operation and maintenance.
By employing a root cause localization language model and an operation and maintenance decision-making language model, combined with a target topology dependency graph and a historical operation and maintenance knowledge base, anomaly clustering and causal analysis are performed on the live network system to generate an operation and maintenance decision sequence, and comprehensive operation and maintenance is carried out through the target operation and maintenance task sequence.
It improves the accuracy of fault location and the effectiveness of operation and maintenance in the existing network system. Through the closed-loop application of semantic understanding capabilities and historical operation and maintenance processes, it enhances the flexibility and feasibility of operation and maintenance.
Smart Images

Figure CN120875851B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a method and device for operating a live network system, a medium and a program product. BACKGROUND
[0002] The live network system has the characteristics of cross-system, cross-region and long link, which brings great difficulty to the operation and maintenance of the live network system. The live network system includes a traffic splitting subsystem, a traffic analysis subsystem, a call detail record (CDR) warehousing subsystem, a data analysis subsystem and a data early warning subsystem. The subsystems of the live network system have strong coupling.
[0003] At present, the system operation and maintenance across systems adopts a general integrated platform, based on an inherent data acquisition process, to acquire data of each subsystem, then uses corresponding fault detection rules to identify faults for each subsystem, and finally performs operation and maintenance on the identified faults.
[0004] The existing cross-system operation and maintenance method is essentially independent operation and maintenance for a single subsystem, and does not realize comprehensive operation and maintenance of the live network system, making it difficult to ensure the accuracy of live network system fault positioning and the effectiveness of system operation and maintenance. SUMMARY
[0005] The present application provides a method and device for operating a live network system, a medium and a program product, which realizes comprehensive operation and maintenance of the live network system and improves the accuracy of live network system fault positioning and the effectiveness of system operation and maintenance.
[0006] According to an aspect of the present application, a method for operating a live network system is provided, the method comprising:
[0007] acquiring target live network system data, and generating a target topology dependency graph and a target index triple of a target live network system based on the target live network system data;
[0008] acquiring a historical operation and maintenance knowledge base, and using a root cause positioning large language model to perform abnormal clustering and causal analysis on the target live network system based on the target topology dependency graph, the target index triple and the historical operation and maintenance knowledge base, to obtain candidate root cause fault information of the target live network system;
[0009] acquiring a historical fault operation and maintenance process, historical effectiveness of the historical fault operation and maintenance process, target operation and maintenance constraint information and a reference optimization target, and using an operation and maintenance decision large language model to perform operation and maintenance decision for the target live network system based on the candidate root cause fault information of the target live network system, the historical fault operation and maintenance process, the historical effectiveness of the historical fault operation and maintenance process, the target operation and maintenance constraint information and the reference optimization target, to obtain a target executable decision sequence;
[0010] According to the target executable decision sequence, a target operation and maintenance task sequence of the target live network system is generated, and the target live network system is operated and maintained according to the target operation and maintenance task sequence.
[0011] According to another aspect of the present application, a live network system operation and maintenance device is provided, and the device comprises:
[0012] A target live network system data acquisition module is configured to acquire target live network system data, and generate a target topology dependency graph and a target index triple of a target live network system according to the target live network system data;
[0013] A fault root cause positioning module is configured to acquire a historical operation and maintenance knowledge base, and perform abnormal clustering and cause analysis on the target topology dependency graph, the target index triple and the historical operation and maintenance knowledge base by using a root cause positioning large language model, to obtain candidate root cause fault information of the target live network system;
[0014] An operation and maintenance decision output module is configured to acquire a historical fault operation and maintenance process, historical effectiveness of the historical fault operation and maintenance process, target operation and maintenance constraint information and a reference optimization target, and perform operation and maintenance decision on the target live network system based on the candidate root cause fault information of the target live network system, the historical fault operation and maintenance process, the historical effectiveness of the historical fault operation and maintenance process, the target operation and maintenance constraint information and the reference optimization target by using an operation and maintenance decision large language model, to output a target executable decision sequence;
[0015] A target live network system operation and maintenance module is configured to generate a target operation and maintenance task sequence of the target live network system according to the target executable decision sequence, and operate and maintain the target live network system according to the target operation and maintenance task sequence.
[0016] According to another aspect of the present application, an electronic device is provided, and the electronic device comprises:
[0017] At least one processor; and
[0018] A memory in communication connection with the at least one processor; wherein,
[0019] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the live network system operation and maintenance method described in any embodiment of the present application.
[0020] According to another aspect of the present application, a computer readable storage medium is provided, and the computer readable storage medium stores computer instructions for enabling a processor to implement the live network system operation and maintenance method described in any embodiment of the present application when executed.
[0021] According to another aspect of the present application, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the method for operating a live system according to any of the embodiments of the present application.
[0022] The technical scheme of the embodiment of the present application adopts the root cause positioning large language model, performs abnormal clustering and cause analysis on the target live system based on the target topology dependency graph, target index triplets and historical operation and maintenance knowledge base of the target live system, utilizes the semantic understanding ability of the root cause positioning large language model itself, improves the accuracy of root cause positioning of the fault of the target live system, considers the dependency relationship between components of each subsystem in the target live system through the target topology dependency graph, that is, considers the coupling between each subsystem in the target live system, realizes semantic space unification of the target live system data from three dimensions of "index-component-business" through the target index triplets, realizes unified arrangement of multi-source data across domains and systems, thereby improving the recognition accuracy of the root cause positioning large language model; moreover, the historical operation and maintenance knowledge base is introduced to enhance the search of root cause faults in the target live system, further improving the accuracy of root cause positioning of the fault of the target live system; the operation and maintenance decision large language model is adopted to make operation and maintenance decisions for the target live system based on the candidate root cause fault information, historical fault operation and maintenance process, historical effectiveness of the historical fault operation and maintenance process, target operation and maintenance constraint information and reference optimization target of the target live system, obtain a target executable decision sequence, utilize the semantic understanding ability of the root cause positioning large language model itself, improve the accuracy of operation and maintenance decisions for the target live system, consider the historical effectiveness of the historical fault operation and maintenance process through the historical fault operation and maintenance process and the historical effectiveness of the historical fault operation and maintenance process, realize closed-loop application of the historical fault operation and maintenance process, moreover, the target operation and maintenance constraint information and reference optimization target are introduced, improving the flexibility of the target live system, thereby improving the implementability of the target executable decision sequence; the target executable decision sequence is converted into a target operation and maintenance task sequence of the target live system, and the target live system is operated according to the target operation and maintenance task sequence, realizing comprehensive operation and maintenance of the target live system, improving the accuracy of fault positioning of the live system and the effectiveness of system operation and maintenance.
[0023] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0024] In order to make the technical solutions in the embodiments of the present application clearer, the accompanying drawings needed in the embodiment description will be briefly introduced. Obviously, the accompanying drawings in the following description only show some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without any creative effort based on these drawings.
[0025] Figure 1 is a flow chart of a method for operating an existing network system according to the first embodiment of the present application;
[0026] Figure 2 is an interface diagram of an automatic search configuration interface of an automatic collection engine according to the first embodiment of the present application;
[0027] Figure 3 is an interface diagram of an automatic login configuration interface of an automatic collection engine according to the first embodiment of the present application;
[0028] Figure 4 is an interface diagram of a manual investigation mode data collection of an automatic collection engine according to the first embodiment of the present application;
[0029] Figure 5 is an interface display diagram of candidate root cause fault information according to the first embodiment of the present application;
[0030] Figure 6 is an interface display diagram of fault handling tracking content contained in actual execution data according to the first embodiment of the present application;
[0031] Figure 7 is a flow chart of a method for operating an existing network system according to the second embodiment of the present application;
[0032] Figure 8 is a structural schematic diagram of a device for operating an existing network system according to the third embodiment of the present application;
[0033] Figure 9 is a structural schematic diagram of an electronic device for implementing the method for operating an existing network system according to the embodiments of the present application. DETAILED DESCRIPTION
[0034] In order to make the technical solutions in the embodiments of the present application clearer, the accompanying drawings needed in the embodiment description will be briefly introduced. Obviously, the accompanying drawings in the following description only show some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without any creative effort based on these drawings.
[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and in the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used are interchangeable under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0036] Embodiment one
[0037] Figure 1 A flowchart of a method for operating a live network system is provided for the first embodiment of the present application. The embodiment of the present application can be applicable to the case of comprehensive operation and maintenance of a live network system. The method is executed by a live network system operation and maintenance device, which is realized in the form of hardware and / or software. The live network system operation and maintenance device can be configured in an electronic device that carries the live network system operation and maintenance function.
[0038] Referring to Figure 1 The method for operating a live network system includes the following steps.
[0039] In S101, target live network system data is acquired, and a target topology dependency graph and a target index triple of the target live network system are generated according to the target live network system data.
[0040] The target live network system is a telecom-level service system running in a network. The target live network system is interactively coupled between subsystems, for example, the target live network system includes a traffic splitting subsystem, a traffic analysis subsystem, a call detail record (CDR) storage subsystem, a data analysis subsystem, and a data early warning subsystem. The target live network system has a strong dependence topology of “traffic splitting subsystem → traffic analysis subsystem → CDR storage subsystem → data analysis subsystem → data early warning subsystem”. The live network system has the characteristics of cross-system, cross-region, and long link. Specifically, when the target live network system as a whole is maintained, a small fluctuation of an upstream subsystem is amplified or hidden in a downstream subsystem, resulting in a fault misreporting or a fault delay reporting of the target live network system, which brings great difficulty to the live network system maintenance. The traffic splitting subsystem is used for bypass mirroring and splitting of production traffic, for subsequent analysis and analysis of the mirror traffic. The traffic analysis subsystem is used for protocol restoration, index extraction, and content analysis of the mirror traffic. The CDR storage subsystem is used for standardization and storage of CDRs or session details. The data analysis subsystem is used for analysis and prediction of indexes, logs, and alarm data. The data early warning subsystem is used for early warning based on indexes, logs, and alarm data. The target live network system data is the running data of the target live network system. The target live network system data is used to represent the actual running state of the target live network system. For example, the target live network system data includes running index data and events of the target live network system.
[0041] The target topology dependency graph is used to reflect the dependency relationship between the components of the subsystems in the target live network system. For example, the target topology dependency graph is represented by G(E, V). V is the component of the subsystems in the target live network system. For example, the components of the subsystems in the target live network system include hardware devices and / or software modules. E is the dependency relationship between the components of the subsystems in the target live network system. Specifically, the dependency relationship includes a data transmission relationship and a data control relationship. Optionally, the dependency relationship between the components of the subsystems in the target live network system has a corresponding dependency relationship weight. For example, the dependency relationship weight includes a traffic proportion, a call frequency, and a historical correlation degree.
[0042] The target index triad includes a three-way relationship among a target index, a target component, and a target service. The target index includes a data dimension detected by a target live network system and corresponding target index data. The target component is a data source of the target index. For example, the target component includes a target hardware device and / or a target software module, etc. The target service is a service associated with the target index data. Thus, the same index in different services is distinguished. By using the target index triad, the index data, log data, link data, and service key indicators in the target live network system data are uniformly mapped in naming and unit. Thus, the data alignment of multi-source heterogeneous data across subsystems in the target live network system is realized, facilitating the subsequent semantic understanding of the target live network system data by the root cause positioning large language model.
[0043] Specifically, a target collection task is registered on the automatic collection engine for the target live network system, and a target collector is selected and a target collection strategy is configured according to the target collection task. According to the target collection strategy, the target collector is dispatched to execute the corresponding target collection task to obtain target live network system data. The target live network system data is identified to determine the components of each subsystem in the target live network system and the dependency relationship between the components of each subsystem in the target live network system. According to the components of each subsystem in the target live network system and the dependency relationship between the components of each subsystem in the target live network system, a target topology dependency graph of the target live network system is generated. The target live network system data is identified to determine the target index, the target component, and the target service. According to the target index, the target component, and the target service, a target index triad is generated. For example, Figure 2 The interface schematic diagram of the automatic search configuration interface of the automatic collection engine is configured. On the automatic search configuration page, the target collection strategy of the target collection task is configured.
[0044] The target collection strategy is a specific execution strategy of the target collection task. For example, the target collection strategy includes periodic collection, event-triggered collection, and adaptive collection, etc. For example, the adaptive collection reduces the sampling frequency when the load is high. The target collection task is a task of collecting data from the target live network system. The target collector is used to execute the target collection task. For example, the target collector includes a web page collector, a software collector, a database collector, and a manual investigation collector, etc.
[0045] The web page collector collects data through a web page system. The web page collector includes a headless browser, a browser collector strategy, and a corresponding page structure. The web page collector supports automatic login, verification code strategy (such as a sliding block or a short message callback permission), and data (such as a table form or a chart form) export and parsing, etc.
[0046] The software collector collects data through a software client. The software collector interfaces with APIs (Application Programming Interface), SDKs (Software Development Kit), or CLIs (Command Line Interface) of each subsystem. If there is no public API, the running index or event can be collected through an agent or a log forwarder. Figure 3 The interface schematic diagram of the automatic login configuration interface for the automatic collection engine. The automatic login configuration interface is used to register and configure target collection tasks for the APIs of each subsystem.
[0047] The database collector collects data through a database. The database collector uses JDBC (Java Database Connectivity), ODBC (Open Database Connectivity), or a native driver to pull KPIs (Key Performance Indicators) according to the criteria. For example, the KPIs include the success rate of call detail record (CDR) storage, the CDR storage delay, the CPU (Central Processing Unit) of the traffic analysis subsystem, the memory of the traffic analysis subsystem, and the queue length of the traffic analysis subsystem.
[0048] The manual investigation collector collects data through manual investigation. The manual investigation collector can configure a questionnaire template. The manual investigation collector supports time limit setting, collection retry, and data recovery from city or team dimensions, and supports automatic structured storage. Figure 4 The interface schematic diagram of the manual investigation data collection of the automatic collection engine. The data is uploaded in the manual investigation data collection interface, and the filled data is displayed after the data is uploaded.
[0049] Optionally, after obtaining the target live network system data, the target live network system data is cleaned and aligned. For example, the data cleaning and alignment methods include missing value alignment, unit conversion, field standardization, alias mapping, time alignment, and time zone unification. The alias mapping is used to align the index names of the same index data. Thus, through the data cleaning and alignment of the target live network system data, the time sequence alignment and content alignment of the multi-source target live network system data across subsystems in the target live network system data are achieved, and the accuracy of the target live network system operation and maintenance is improved.
[0050] S102, acquire a historical operation and maintenance knowledge base, and adopt a root cause positioning large language model, based on a target topology dependency graph, a target index triple and the historical operation and maintenance knowledge base, perform abnormal clustering and causal analysis on the target live network system to obtain candidate root cause fault information of the target live network system.
[0051] The root cause positioning large language model is used for root cause positioning of the target live network system. The input data of the root cause positioning large language model includes a target topology dependency graph, a target index triple and a historical operation and maintenance knowledge base. The output result of the root cause positioning large language model is candidate root cause fault information of the target live network system. The root cause positioning large language model performs abnormal clustering and causal analysis on the target live network system. Specifically, the root cause positioning large language model performs robust statistics, change point detection and seasonal analysis residual on the target live network system to perform abnormal clustering on the target live network system and determine whether the target live network system has a fault. The change point detection includes using CUSUM (Cumulative Sum Algorithm) or entropy increment algorithm, etc. The root cause positioning large language model performs time lag correlation, path consistency scoring and conflict evidence suppression along the target topology dependency graph to perform causal analysis on the target live network system and determine the candidate root cause fault information of the target live network system. Exemplarily, the root cause positioning large language model is a large language model (LLM).
[0052] The historical operation and maintenance knowledge base is used for storing historical operation and maintenance knowledge of the target live network system. The historical operation and maintenance knowledge base is used for retrieval enhancement of the root cause positioning large language model and the operation and maintenance decision-making large language model. Exemplarily, the historical operation and maintenance knowledge base includes historical general operation and maintenance knowledge, historical individualized operation and maintenance knowledge and third-party operation and maintenance knowledge. The historical general operation and maintenance knowledge is operation and maintenance knowledge of general faults of the target live network system. For example, the historical general operation and maintenance knowledge includes historical general operation and maintenance processes and historical effectiveness of the historical general operation and maintenance processes. For example, the historical general operation and maintenance knowledge is a run book of the target live network system. The historical individualized operation and maintenance knowledge is operation and maintenance knowledge of individualized faults of the target live network system. For example, the historical individualized operation and maintenance knowledge is a postmortem of the target live network system and a change record. The third-party operation and maintenance knowledge is operation and maintenance knowledge of the target live network system by a third-party system. For example, the third-party operation and maintenance knowledge includes a device manual and device operation and maintenance knowledge of a device provider.
[0053] The candidate root cause fault information is root cause information of a possible fault in the target live network system predicted by the root cause positioning large language model. Exemplarily, Figure 5 is an interface display diagram of the candidate root cause fault information. As Figure 5As shown, the candidate root cause failure information includes a candidate root cause failure link (such as a board monitoring indicator), a candidate failure evidence chain digest (such as a traceability indicator), and a candidate root cause failure confidence (included in the early warning details), etc.
[0054] In an optional embodiment of the present application, the candidate root cause failure information includes a candidate root cause failure link, a candidate failure impact scope, a candidate failure recurrence probability, a candidate failure evidence chain digest, and a candidate root cause failure confidence.
[0055] The candidate root cause failure link is a root cause link of the candidate root cause failure in the target live network system. For example, the candidate root cause failure link is a certain component in the target live network system or a dependency relationship between two components. The candidate failure impact scope is used to represent the impact degree of the candidate root cause failure. The candidate failure recurrence frequency is used to represent the occurrence frequency of the candidate root cause failure. The candidate failure evidence chain digest is used to represent the corresponding part of the content in the target live network system data of the candidate root cause failure. For example, the candidate failure evidence chain digest includes mutation indicator data, log key fragments, and link anomaly data corresponding to the candidate root cause failure in the target live network system data. The candidate failure evidence chain digest is used to represent the original data content corresponding to the candidate root cause failure in the target live network system data. The candidate root cause failure confidence is the credibility of the candidate root cause failure.
[0056] The present scheme improves the accuracy of root cause failure positioning of the live network system by accurately positioning the candidate root cause failure link of the live network system. By introducing the candidate failure impact scope, the candidate failure recurrence probability, the candidate failure evidence chain digest, and the candidate root cause failure confidence, the comprehensiveness of the failure information of the candidate root cause failure link is improved, the typicality and effectiveness of the candidate root cause failure information are taken into account, and the accuracy of the operation and maintenance decision based on the candidate root cause failure information is further improved.
[0057] Specifically, a pre-generated historical operation and maintenance knowledge base is obtained. A root cause positioning large language model is used to perform semantic understanding on the target topology dependency graph, the target indicator triple, and the historical operation and maintenance knowledge base in a preset time window with the current time as the terminal time, to perform anomaly clustering and causal analysis on the target live network system, and to obtain candidate root cause failure information of the target live network system. For example, the candidate root cause failure information is Top-N candidate root cause failure information.
[0058] S103, obtain historical failure operation and maintenance processes, historical effects of the historical failure operation and maintenance processes, target operation and maintenance constraint information, and reference optimization targets, and use an operation and maintenance decision large language model to make an operation and maintenance decision for the target live network system based on the candidate root cause failure information of the target live network system, the historical failure operation and maintenance processes, the historical effects of the historical failure operation and maintenance processes, the target operation and maintenance constraint information, and the reference optimization targets, to obtain a target executable decision sequence.
[0059] The operation and maintenance decision large language model is used to generate an operation and maintenance decision for a target live network system. The input data of the operation and maintenance decision large language model includes candidate root cause fault information of the target live network system, a historical fault operation and maintenance process, historical effectiveness of the historical fault operation and maintenance process, target operation and maintenance constraint information, and a reference optimization target. The output result of the operation and maintenance decision large language model is a target executable decision sequence of the target live network system. Exemplarily, the operation and maintenance decision large language model is a large language model (LLM).
[0060] The historical fault operation and maintenance process is an actual operation and maintenance process of a historical fault of the target live network system. Exemplarily, the historical fault operation and maintenance process includes a historical general fault operation and maintenance process and a historical individualized operation and maintenance process. The historical effectiveness of the historical fault operation and maintenance process is an actual execution effect of the historical fault operation and maintenance process. The target operation and maintenance constraint information is constraint information for the target live network system to make an operation and maintenance decision. Exemplarily, the target operation and maintenance constraint information includes a time window, an approval, and a risk. The time window is used to represent a time window size referred to by the target live network system to make an operation and maintenance decision. The approval is used to represent whether a single candidate root cause fault in the target live network system needs to be manually approved when performing operation and maintenance. The risk is used to represent whether the target live network system has a risk when performing operation and maintenance and a corresponding risk range. The reference optimization target is used to represent an optimization strategy of the target live network system. Exemplarily, the reference optimization target includes a recovery speed priority, a minimum recovery risk, and / or a minimum business impact. The reference optimization target is set by an operation and maintenance party of the target live network system or pre-set by the device. The target executable decision sequence is an executable operation and maintenance decision of each candidate root cause fault. Exemplarily, the target executable decision sequence includes a target execution order between target root cause fault links, a target fault operation and maintenance process of each target root cause fault link, and a target execution password in each target root cause fault link.
[0061] In an optional embodiment of the present application, the target executable decision sequence includes a target execution order between target root cause fault links, a target fault operation and maintenance process of each target root cause fault link, a target execution password in each target root cause fault link, a reference execution index in each target root cause fault link, a reference rollback point corresponding to the reference execution index, and a reference execution effectiveness criterion.
[0062] The target root cause failure link is a candidate root cause failure link that needs to be maintained. Illustratively, the target root cause failure link includes a single subsystem in the target live network system or a component in the single subsystem. The target execution order is an execution order between the target root cause failure links. The target fault maintenance process is a fault maintenance process of a single target root cause failure link. The target execution password is an execution password required when executing the target fault maintenance process. The reference execution indicator is a data dimension for detecting the maintenance of the target root cause failure link. The reference rollback point is a target root cause failure link to which the target executable decision sequence can be rolled back when it is determined based on the reference execution indicator that the maintenance of the target root cause failure is not effective. The reference execution effectiveness criterion is used to judge the effectiveness of the maintenance of the target root cause failure. Illustratively, the reference execution effectiveness criterion includes a reference execution indicator value and a reference recovery time distribution. The reference recovery time distribution is used to represent the change of the maintenance recovery time.
[0063] The present scheme realizes the executability of the target live network system by specifying the target executable decision sequence as the target execution order between the target root cause failure links, the target fault maintenance process of each target root cause failure link, and the target execution password in each target root cause failure link. By introducing the reference execution indicator, the reference rollback point corresponding to the reference execution indicator, and the reference execution effectiveness criterion, the execution verification of the target executable decision sequence is facilitated.
[0064] Specifically, the historical fault maintenance process, the historical effectiveness of the historical fault maintenance process, the target operation constraint information of the target live network system, and the reference optimization target are obtained. A maintenance decision large language model is used to perform semantic understanding on the candidate root cause failure information of the target live network system, the historical fault maintenance process, the historical effectiveness of the historical fault maintenance process, the target operation constraint information, and the reference optimization target, and to make a maintenance decision for the target live network system to obtain a target executable decision sequence.
[0065] S104, according to the target executable decision sequence, a target operation task sequence of the target live network system is generated, and the target live network system is maintained according to the target operation task sequence.
[0066] The target operation task sequence is used to represent the conversion result of converting the target executable decision sequence into actual operation tasks. In comparison, the target executable decision sequence is mainly used to represent how to execute each target root cause fault link according to a target execution order and using a target fault operation process. The target operation task sequence is mainly used to represent how to execute each target operation task to implement the operation of the target executable decision sequence. For example, the target operation task sequence includes each target operation task and a target operation order of each target operation task. The target operation order is the execution order of each target operation task. For example, the target operation task includes a capacity expansion task, a switching task, a flow limiting task, a restart task, a rollback task, or a policy issuing task, etc. For example, the target operation task is a DAG task (Directed Acyclic Graph Task). Optionally, the target operation task corresponds to a work order or a change order. Optionally, the target operation task supports semi-automatic approval.
[0067] Specifically, according to the target execution order between each target root cause fault link included in the target executable decision sequence, the target fault operation process of each target root cause fault link, and the target execution password in each target root cause fault link, the target executable decision sequence is converted into each target operation task with a target operation order, and a target operation task sequence of the target live network system is generated. According to the target operation task sequence, the target live network system is operated.
[0068] For the target live network system across systems and regions, when fault locating and system operation are performed, the coupling between each subsystem in the target live network system and the complexity of a single subsystem need to be considered. Therefore, the fault locating and system operation of the target live network system are extremely difficult.
[0069] The technical scheme of the embodiment of the application adopts a root cause positioning large language model, performs abnormal clustering and cause analysis on a target live network system based on a target topology dependency graph, target index triplets and a historical operation and maintenance knowledge base of the target live network system, utilizes the semantic understanding capability of the root cause positioning large language model itself, improves the accuracy of root cause positioning of faults of the target live network system, considers the dependency relationship between components of each subsystem in the target live network system through the target topology dependency graph, that is, considers the coupling between each subsystem in the target live network system, realizes semantic space unification of target live network system data from three dimensions of "index-component-business" through the target index triplets, and unifies multi-source data across domains and systems, thereby improving the recognition accuracy of the root cause positioning large language model; moreover, the historical operation and maintenance knowledge base is introduced to enhance the search of root cause faults in the target live network system, and the accuracy of root cause positioning of faults of the target live network system is further improved; the operation and maintenance decision large language model is adopted to perform operation and maintenance decision of the target live network system based on candidate root cause fault information, historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information and reference optimization targets, and a target executable decision sequence is obtained, the semantic understanding capability of the root cause positioning large language model itself is utilized, the accuracy of operation and maintenance decision of the target live network system is improved, the historical effectiveness of the historical fault operation and maintenance processes is considered through the historical fault operation and maintenance processes and the historical effectiveness of the historical fault operation and maintenance processes, closed loop application of the historical fault operation and maintenance processes is realized, and the target operation and maintenance constraint information and the reference optimization targets are introduced, the flexibility of the target live network system is improved, and the implementability of the target executable decision sequence is improved; the target executable decision sequence is converted into a target operation and maintenance task sequence of the target live network system, and the target live network system is operated and maintained according to the target operation and maintenance task sequence, comprehensive operation and maintenance of the target live network system is realized, and the accuracy of fault positioning of the live network system and the effectiveness of system operation and maintenance are improved.
[0070] In an optional embodiment of the application, the method further comprises: periodically acquiring candidate root cause fault information, an actual executable decision sequence, an actual operation and maintenance task sequence and actual execution data of the target live network system; and adopting a long-term governance large language model to perform governance analysis on the candidate root cause fault information, the actual executable decision sequence, the actual operation and maintenance task sequence and the actual execution data of the target live network system, and obtaining a target governance roadmap.
[0071] The actual executable decision sequence is a target executable decision sequence actually adopted when the target live network system is operated and maintained. The actual operation and maintenance task sequence is an operation and maintenance task sequence corresponding to the actual executable decision sequence. The actual execution data is execution data when the target live network system is operated and maintained by using the actual operation and maintenance task sequence. For example, Figure 6The interface display figure of the fault handling trace content contained in the actual execution data is displayed. For example, the fault handling trace includes task status, fault category, fault reason, solution, recording time, recording personnel, and associated similar tasks. The task status includes unprocessed, processing, and completed. The fault category is used to represent the category of the actual fault to be handled. The fault reason is used to represent the actual cause of the fault. The associated similar tasks are used to represent the historical actual operation tasks of the same type.
[0072] The long-term governance large language model is used to perform long-term governance on the target live system. The input data of the long-term governance large language model includes candidate root cause fault information of the target live system, an actual executable decision sequence, an actual operation task sequence, and actual execution data. The output result of the long-term governance large language model is a target governance roadmap of the target live system. For example, the long-term governance large language model is a large language model (LLM).
[0073] The target governance roadmap is used to periodically perform governance on the target live system. For example, the target governance roadmap includes at least two weak fault links, a governance priority and a governance suggestion of each weak fault link, and a milestone weak fault link in each weak fault link. The weak fault link is a fault link in the target live system that needs long-term governance. The governance priority is used to represent the fault impact degree of the weak fault link. For example, the governance priority includes fault impact range and fault occurrence frequency. The governance suggestion is used to represent how to govern the weak fault link. For example, the governance suggestion includes a historical fault operation process and historical effectiveness corresponding to the weak fault link. The milestone weak fault link is used to represent a key fault link in each weak fault link. It can be understood that the milestone weak fault link is a weak fault link that needs to be periodically governed in the target governance roadmap.
[0074] Specifically, the candidate root cause fault information, the actual executable decision sequence, the actual operation task sequence, and the actual execution data of the target live system are periodically collected. The long-term governance large language model is used to perform semantic understanding on the candidate root cause fault information, the actual executable decision sequence, the actual operation task sequence, and the actual execution data of the target live system, and outputs the target governance roadmap.
[0075] The scheme periodically acquires the candidate root cause fault information, the actual executable decision sequence, the actual operation task sequence, and the actual execution data of the target live system, and uses the long-term governance large language model to perform governance analysis on the candidate root cause fault information, the actual executable decision sequence, the actual operation task sequence, and the actual execution data of the target live system. The periodic structured research and judgment of the target live system are realized, and the target live system operation is converted from fault handling to engineering periodic governance through the target governance roadmap.
[0076] Embodiment Two
[0077] Figure 7 A flowchart of a method for operating an existing network system according to Embodiment Two of the present application is provided. Based on the above-mentioned embodiments, the present application further includes the following steps: according to the target operation task sequence, selecting and executing the current target operation task; when the current target operation task is completed, detecting the current execution effect of the current execution index of the current target operation task; according to the current execution effect, updating and executing the current target operation task, returning to the step of detecting the current execution effect of the current execution index of the current target operation task when the current target operation task is completed, until each target operation task in the target operation task sequence is completed. The quantitative detection of the operation effect of the current target operation task is achieved, the operation effect of the current operation task and the actual operation process of the target existing network system form a closed loop, the adaptive adjustment of the current target operation task is achieved, and the comprehensive operation effect of the target existing network system operation is improved. It should be noted that the parts not described in detail in the embodiments of the present application can be referred to the descriptions of other embodiments.
[0078] Referring to Figure 7 The method for operating an existing network system shown in the figure comprises the following steps:
[0079] S701, obtaining target existing network system data, and generating a target topology dependency graph and a target index triple of the target existing network system according to the target existing network system data.
[0080] S702, obtaining a historical operation knowledge base by using a root cause positioning large language model, and performing abnormal clustering and cause analysis on the target existing network system based on the target topology dependency graph, the target index triple and the historical operation knowledge base, to obtain candidate root cause fault information of the target existing network system.
[0081] S703, obtaining a historical fault operation process, a historical effect of the historical fault operation process, target operation constraint information and a reference optimization target by using an operation decision large language model, and performing operation decision on the target existing network system based on the candidate root cause fault information of the target existing network system, the historical fault operation process, the historical effect of the historical fault operation process, the target operation constraint information and the reference optimization target, to obtain a target executable decision sequence.
[0082] S704, according to the target operation task sequence, selecting and executing the current target operation task.
[0083] The current target operation and maintenance task is a target operation and maintenance task executed at the current time. For example, the current target operation and maintenance task is a target operation and maintenance task with the earliest target operation and maintenance sequence among the executable target operation and maintenance tasks in the target operation and maintenance task sequence.
[0084] Specifically, according to the target operation and maintenance tasks included in the target operation and maintenance task sequence and the target operation and maintenance sequences of the target operation and maintenance tasks, a target operation and maintenance task with the earliest target operation and maintenance sequence is selected from the executable target operation and maintenance tasks, and the current target operation and maintenance task is determined and executed.
[0085] S705, when the current target operation and maintenance task is executed, the current execution effect of the current execution index of the current target operation and maintenance task is detected.
[0086] The current execution index is a data dimension for detecting the operation and maintenance effect of the current target operation and maintenance task. The current execution effect is used to represent the operation and maintenance effect of the current target operation and maintenance task. The reference execution index in the target executable decision sequence corresponds to the current execution index. For example, the current execution effect includes that the operation and maintenance effect of the current target operation and maintenance task reaches a preset operation and maintenance effect or the operation and maintenance effect of the current target operation and maintenance task does not reach the preset operation and maintenance effect.
[0087] Specifically, when the current target operation and maintenance task is executed, the reference execution effect criterion in the target executable decision sequence is used to detect the current execution index of the current target operation and maintenance task, and the current execution effect of the current target operation and maintenance task is determined.
[0088] S706, according to the current execution effect, the current target operation and maintenance task is updated and executed, and the step of detecting the current execution effect of the current execution index of the current target operation and maintenance task when the current target operation and maintenance task is executed is returned until all the target operation and maintenance tasks in the target operation and maintenance task sequence are executed.
[0089] Specifically, when the operation and maintenance effect of the current target operation and maintenance task reaches the preset operation and maintenance effect, the next target operation and maintenance task of the current target operation and maintenance task in the target operation and maintenance task sequence is updated as the current target operation and maintenance task and executed. When the operation and maintenance effect of the current target operation and maintenance task does not reach the preset operation and maintenance effect, the current target operation and maintenance task is adjusted and updated as the current target operation and maintenance task and executed. The step of detecting the current execution effect of the current execution index of the current target operation and maintenance task when the current target operation and maintenance task is executed is returned until all the target operation and maintenance tasks in the target operation and maintenance task sequence are executed.
[0090] In an optional embodiment of the present application, the current target operation and maintenance task is updated and executed according to the current execution effect, including: when the current execution index value of the current target operation and maintenance task reaches the first reference key index value, the next target operation and maintenance task of the current target operation and maintenance task in the target operation and maintenance task sequence is updated as the current target operation and maintenance task and is executed; when the current execution index value of the current target operation and maintenance task does not reach the first reference key index value, the target operation and maintenance task sequence is executed to return according to the reference return point contained in the target executable decision sequence, and the current target operation and maintenance task is updated.
[0091] The first reference key index value is used to measure the operation and maintenance effect of the current target operation and maintenance task. The first reference key index value is the lower limit value of the current execution index value when the operation and maintenance effect of the current target operation and maintenance task reaches the preset operation and maintenance effect. It can be understood that the operation and maintenance effect of the current target operation and maintenance task reaches the preset operation and maintenance effect when the current execution index value of the current target operation and maintenance task reaches the first reference key index value. It can be understood that the operation and maintenance effect of the current target operation and maintenance task does not reach the preset operation and maintenance effect when the current execution index value of the current target operation and maintenance task does not reach the first reference key index value. The reference return point is a candidate root cause fault node returned to when the current execution index value does not reach the first reference key index value.
[0092] Specifically, when the current execution index value of the current target operation and maintenance task reaches the first reference key index value, the next target operation and maintenance task of the current target operation and maintenance task in the target operation and maintenance task sequence is updated as the current target operation and maintenance task and is executed. When the current execution index value of the current target operation and maintenance task does not reach the first reference key index value, the target operation and maintenance task sequence is executed to return according to the reference return point contained in the target executable decision sequence, so as to return to the target operation and maintenance task corresponding to the reference return point and update the current target operation and maintenance task.
[0093] The present scheme introduces detection of whether the current execution index value of the current target operation and maintenance task reaches the first reference key index value, which can quickly detect the operation and maintenance effect of the current target operation and maintenance task. By updating the next target operation and maintenance task of the current target operation and maintenance task in the target operation and maintenance task sequence as the current target operation and maintenance task when the current execution index value of the current target operation and maintenance task reaches the first reference key index value, and executing the next target operation and maintenance task, and by executing the target operation and maintenance task sequence to return according to the reference return point contained in the target executable decision sequence when the current execution index value of the current target operation and maintenance task does not reach the first reference key index value, and updating the current target operation and maintenance task to the target operation and maintenance task corresponding to the reference return point, the present scheme introduces the reference return point and realizes quick return of the target operation and maintenance task sequence when the current execution index value of the current target operation and maintenance task does not reach the first reference key index value, improves the operation and maintenance executability of the target existing network system, and thus improves the operation and maintenance adjustment efficiency and the operation and maintenance accuracy of the target existing network system.
[0094] In an optional embodiment of the present application, a large language model for operation and maintenance decision is used to make operation and maintenance decisions for the target live network system based on candidate root cause fault information of the target live network system, historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization targets, and output a target executable decision sequence, including: using a large language model for operation and maintenance decision, based on candidate root cause fault information of the target live network system, historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization targets, making operation and maintenance decisions for the target live network system, and outputting at least two candidate executable decision sequences and candidate decision confidence of each candidate executable decision sequence; according to the candidate decision confidence, screening the target executable decision sequence from the candidate executable decision sequence; after detecting the current execution effectiveness of the current execution indicator of the current target operation and maintenance task, further comprising: when the current execution indicator value of the current target operation and maintenance task does not reach the second reference key indicator value, screening the target executable decision sequence from each candidate executable decision sequence according to the candidate decision confidence, and returning to the step of generating the target operation and maintenance task sequence of the target live network system according to the target executable decision sequence.
[0095] The candidate executable decision sequence is an executable operation and maintenance decision that can be selected when operation and maintenance is performed for each candidate root cause fault. The candidate executable decision sequence is used to screen the target executable decision sequence. The candidate decision confidence is used to represent the credibility of the candidate executable decision sequence. The second reference key indicator value is used to measure the operation and maintenance effect of the current target operation and maintenance task. The second reference key indicator value is the lower limit value of the current execution indicator value when the target executable decision sequence needs to be replaced. When the current execution indicator value of the current target operation and maintenance task reaches the second reference key indicator value, it can be understood that the target live network system is operated and maintained by continuing to use the target operation and maintenance task sequence corresponding to the target executable decision sequence. When the current execution indicator value of the current target operation and maintenance task does not reach the second reference key indicator value, it can be understood that the target executable decision sequence needs to be replaced.
[0096] Specifically, the operation and maintenance decision large language model is used for semantic understanding of candidate root cause fault information, historical fault operation and maintenance process, historical effect of the historical fault operation and maintenance process, target operation and maintenance constraint information and reference optimization target of the target live network system, operation and maintenance decision is made for the target live network system, at least two candidate executable decision sequences and candidate decision confidence of each candidate executable decision sequence are obtained. The candidate executable decision sequence with the highest candidate decision confidence is screened as the executable decision sequence. After detecting the current execution effect of the current execution index of the current target operation and maintenance task, when it is detected that the current execution index value of the current target operation and maintenance task reaches the second reference key index value, the current execution index value of the current target operation and maintenance task is compared with the first reference key index value, and the step of detecting the current execution effect of the current execution index of the current target operation and maintenance task is returned. When it is detected that the current execution index value of the current target operation and maintenance task does not reach the second reference key index value, in each candidate executable decision sequence except the target executable decision sequence, the candidate executable decision sequence with the highest candidate decision confidence is screened as the target executable decision sequence, and the step of generating the target operation and maintenance task sequence of the target live network system according to the target executable decision sequence is returned.
[0097] The comparison process of the current execution index value of the current target operation and maintenance task and the second reference key index value is introduced, when the current execution index value of the current target operation and maintenance task does not reach the second reference key index value, the suboptimal executable decision sequence is adjusted, the fault tolerance and executability of the target live network system operation are considered, and the operation and maintenance efficiency of the target live network system is improved.
[0098] In the process of operation and maintenance of the target live network system, the quantitative detection of the operation and maintenance effect of the current target operation and maintenance task is realized by detecting the current execution effect of the current execution index of the current target operation and maintenance task when the current target operation and maintenance task is executed, the current target operation and maintenance task is updated and executed based on the current execution effect of the current execution index of the current target operation and maintenance task, the operation and maintenance effect of the current operation and maintenance task and the actual operation and maintenance process of the target live network system form a closed loop, the adaptive adjustment of the current target operation and maintenance task is realized, and the comprehensive operation and maintenance effect of the target live network system operation is improved.
[0099] Exemplarily, on the basis of the above-mentioned embodiment, the present application also provides a preferred embodiment. The live network system operation method comprises the following technical features:
[0100] S1, a target acquisition task is registered for a target live network system, and a target acquisition device is selected and a target acquisition strategy is configured according to the target acquisition task.
[0101] Exemplarily, the target collector includes a webpage collector, a software collector, a database collector, and a manual investigation collector, etc.
[0102] S2, a target collector is scheduled to perform a corresponding target collection task according to a target collection strategy, to obtain target live system data, and to pre-process the target live system data.
[0103] S3, a target topology dependency graph and a target index triple are generated according to the pre-processed target live system data.
[0104] The target topology dependency graph is used to reflect the dependency relationship between components of each subsystem in the target live system. The target index triple includes a three-way relationship between a target index, a target component, and a target service.
[0105] S4, a historical operation and maintenance knowledge base is obtained.
[0106] The historical operation and maintenance knowledge base includes historical general operation and maintenance knowledge, historical individualized operation and maintenance knowledge, and third-party operation and maintenance knowledge. The historical general operation and maintenance knowledge includes historical general operation and maintenance processes and historical effectiveness of the historical general operation and maintenance processes. Exemplarily, the historical general operation and maintenance knowledge is a run book of the target live system. The historical individualized operation and maintenance knowledge is an operation and maintenance analysis report (Postmortem) and a change record of the target live system. The third-party operation and maintenance knowledge includes a device manual and device operation and maintenance knowledge provided by a device provider.
[0107] S5, a research and judgment LLM is used to perform abnormal clustering and cause-effect analysis on the target live system according to the target topology dependency graph, the target index triple, and the historical operation and maintenance knowledge base, to output Top-N candidate root cause failure information of the target live system.
[0108] The research and judgment LLM is a root cause positioning large language model. The candidate root cause failure information includes a candidate root cause failure link, a candidate failure impact area, a candidate failure recurrence probability, a candidate failure evidence chain summary, and a candidate root cause failure confidence. The candidate failure evidence chain summary is a readable evidence string obtained by splicing corresponding mutation index data, log key fragments, and link anomaly data of the candidate root cause failure in the target live system data.
[0109] S6, a decision LLM is used to perform operation and maintenance decision-making on the target live system based on the Top-N candidate root cause failure information of the target live system, the historical general operation and maintenance processes, the historical effectiveness of the historical general operation and maintenance processes, target operation and maintenance constraint information, and a reference optimization target, to output a target executable decision sequence.
[0110] The decision LLM is an operation and maintenance decision large language model. The target executable decision sequence includes a target execution order between target root cause failure links, a target failure operation and maintenance process of each target root cause failure link, a target execution password in each target root cause failure link, a reference execution index in each target root cause failure link, a reference rollback point corresponding to the reference execution index, and a reference execution effectiveness criterion.
[0111] S7, generating a target operation and maintenance task sequence according to the target executable decision sequence, executing a current target operation and maintenance task, and detecting a current execution index; when the current execution index reaches the corresponding reference execution effectiveness criterion, updating and executing the current target operation and maintenance task based on the target operation and maintenance task sequence; when the current execution index does not reach the corresponding reference execution effectiveness criterion, triggering the target operation and maintenance task sequence rollback, and updating the target executable decision sequence.
[0112] S8, when the target operation and maintenance task sequence is executed, updating the historical general operation and maintenance process and the historical effectiveness of the historical general operation and maintenance process based on actual execution data of the target operation and maintenance task sequence.
[0113] S9, periodically acquiring Top-N candidate root cause failure information of a target live network system, a target executable decision sequence, a target operation and maintenance task sequence, and actual execution data of the target operation and maintenance task sequence, using a long-term governance LLM to analyze the Top-N candidate root cause failure information of the target live network system, the target executable decision sequence, the target operation and maintenance task sequence, and the actual execution data of the target operation and maintenance task sequence, and outputting a target governance roadmap.
[0114] The long-term governance LLM is a long-term governance large language model. The target governance roadmap includes Top-N weak failure links, a governance priority and a governance suggestion of each weak failure link, and a milestone weak failure link in each weak failure link.
[0115] The present scheme integrates the webpage system, software system, database system and manual investigation system into a single arrangement framework through four types of automatic collectors, provides a unified method of automatic login, retrieval, extraction, cleaning, entity alignment and time alignment, significantly improves the consistency and comparability of data arrival, realizes the unified arrangement of the existing network data, and improves the consistency of the existing network system data; by inputting the target topology dependency graph, target index triple and historical operation and maintenance knowledge base facing the long link of the telecommunications into the root cause positioning large language model for research and judgment, on the link of the flow splitting subsystem→flow analysis subsystem→call entry subsystem→data analysis subsystem→data early warning subsystem, combined with the topology dependency, time lag correlation, historical similarity and conflict evidence suppression, the root cause positioning large language model generates candidate root cause fault information with candidate root cause fault confidence and candidate fault evidence chain abstract, improves the explanation and hit rate of cross-domain problems; combined with the historical fault operation and maintenance process, historical effectiveness of the historical fault operation and maintenance process and target operation and maintenance constraint information, the operation and maintenance decision large language model is used to generate output target executable decision sequence, and is converted into executable, reversible and verifiable target operation and maintenance task sequence, and the "suggestion" is converted into "implementable"; by periodically structuring the research and judgment of high-frequency faults, capacity bottlenecks, strategy or model drift, the target governance roadmap is generated, and the governance effectiveness is backfilled to the historical operation and maintenance knowledge base and the long-term governance large language model, so that the operation and maintenance of the target existing network system is changed from "fire fighting" to "engineering governance", and the long-term research and judgment and stability governance of the target existing network system are realized.
[0116] Embodiment three
[0117] Figure 8 A structural schematic diagram of an existing network system operation and maintenance device provided by the third embodiment of the present application. The third embodiment of the present application can be applied to the case of comprehensive operation and maintenance of the existing network system, the device executes the existing network system operation and maintenance method, the device is realized in the form of hardware and / or software, and the device can be configured in an electronic device carrying the existing network system operation and maintenance function.
[0118] Reference Figure 8The illustrated live network system operation and maintenance device comprises: a target live network system data acquisition module 801, a fault root cause positioning module 802, an operation and maintenance decision output module 803, and a target live network system operation and maintenance module 804. Among them, the target live network system data acquisition module 801 is configured to acquire target live network system data, and generate a target topology dependency graph and a target index triple of the target live network system according to the target live network system data; the fault root cause positioning module 802 is configured to acquire a historical operation and maintenance knowledge base, and perform abnormal clustering and causal analysis on the target topology dependency graph, the target index triple and the historical operation and maintenance knowledge base by using a root cause positioning large language model, to obtain candidate root cause fault information of the target live network system; the operation and maintenance decision output module 803 is configured to acquire a historical fault operation and maintenance process, historical effectiveness of the historical fault operation and maintenance process, target operation and maintenance constraint information and a reference optimization target, and perform operation and maintenance decision on the target live network system based on the candidate root cause fault information of the target live network system, the historical fault operation and maintenance process, the historical effectiveness of the historical fault operation and maintenance process, the target operation and maintenance constraint information and the reference optimization target by using an operation and maintenance decision large language model, and output a target executable decision sequence; and the target live network system operation and maintenance module 804 is configured to generate a target operation and maintenance task sequence of the target live network system according to the target executable decision sequence, and perform operation and maintenance on the target live network system according to the target operation and maintenance task sequence.
[0119] The technical scheme of the embodiment of the application adopts the root cause positioning large language model, performs abnormal clustering and cause analysis on the target live network system based on the target topology dependency graph, the target index triple of the target live network system and the historical operation and maintenance knowledge base, utilizes the semantic understanding ability of the root cause positioning large language model itself, improves the accuracy of root cause positioning of the fault of the target live network system, considers the dependency relationship between the components of each subsystem in the target live network system through the target topology dependency graph, that is, considers the coupling between the subsystems in the target live network system, realizes semantic space unification of the target live network system data from three dimensions of "index-component-business" through the target index triple, realizes unified arrangement of multi-source data across domains and systems, and thus improves the recognition accuracy of the root cause positioning large language model; moreover, the historical operation and maintenance knowledge base is introduced to enhance the search of the root cause fault in the target live network system, and the accuracy of root cause positioning of the fault of the target live network system is further improved; the operation and maintenance decision large language model is adopted to perform operation and maintenance decision on the target live network system based on the candidate root cause fault information of the target live network system, the historical fault operation and maintenance process, the historical effect of the historical fault operation and maintenance process, the target operation and maintenance constraint information and the reference optimization target, obtain the target executable decision sequence, utilize the semantic understanding ability of the root cause positioning large language model itself, improve the accuracy of operation and maintenance decision of the target live network system, consider the historical effect of the historical fault operation and maintenance process through the historical fault operation and maintenance process and the historical effect of the historical fault operation and maintenance process, realize closed loop application of the historical fault operation and maintenance process, moreover, the target operation and maintenance constraint information and the reference optimization target are introduced, and the flexibility of the target live network system is improved, and thus the implementability of the target executable decision sequence is improved; the target executable decision sequence is converted into the target operation and maintenance task sequence of the target live network system, and the target live network system is operated and maintained according to the target operation and maintenance task sequence, comprehensive operation and maintenance of the target live network system is realized, and the accuracy of fault positioning of the live network system and the effectiveness of system operation and maintenance are improved.
[0120] In an optional embodiment of the application, the target live network system operation and maintenance module 804 comprises: a current target operation and maintenance task screening unit, configured to select and execute a current target operation and maintenance task according to the target operation and maintenance task sequence; a current execution effect detection unit, configured to detect the current execution effect of the current execution index of the current target operation and maintenance task when the current target operation and maintenance task is executed; and a current target operation and maintenance task updating unit, configured to update and execute the current target operation and maintenance task according to the current execution effect, and return to the step of detecting the current execution effect of the current execution index of the current target operation and maintenance task when the current target operation and maintenance task is executed, until each target operation and maintenance task in the target operation and maintenance task sequence is executed.
[0121] In an optional embodiment of the present application, the current target operation and maintenance task updating unit comprises: a first current target operation and maintenance task updating unit configured to update the next target operation and maintenance task of the current target operation and maintenance task in the target operation and maintenance task sequence to the current target operation and maintenance task and perform the update when the current execution index value of the current target operation and maintenance task reaches the first reference key index value; and a second current target operation and maintenance task updating unit configured to perform rollback on the target operation and maintenance task sequence according to the reference rollback point included in the target executable decision sequence and update the current target operation and maintenance task when the current execution index value of the current target operation and maintenance task does not reach the first reference key index value.
[0122] In an optional embodiment of the present application, the operation and maintenance decision output module 803 comprises: a candidate executable decision sequence output unit configured to perform operation and maintenance decision on the target network system based on the candidate root cause fault information of the target network system, the historical fault operation and maintenance process, the historical effectiveness of the historical fault operation and maintenance process, the target operation and maintenance constraint information and the reference optimization target by using the operation and maintenance decision large language model, and output at least two candidate executable decision sequences and candidate decision confidence of each candidate executable decision sequence; and an operation and maintenance decision output unit configured to filter the target executable decision sequence from the candidate executable decision sequences according to the candidate decision confidence. The target network system operation and maintenance module 804 further comprises: a target executable decision sequence adjustment unit configured to filter the target executable decision sequence from the candidate executable decisions other than the target executable decision sequence according to the candidate decision confidence when the current execution index value of the current target operation and maintenance task does not reach the second reference key index value after detecting the current execution effectiveness of the current execution index of the current target operation and maintenance task, and return to the step of generating the target operation and maintenance task sequence of the target network system according to the target executable decision sequence.
[0123] In an optional embodiment of the present application, the candidate root cause fault information comprises a candidate root cause fault link, a candidate fault influence area, a candidate fault recurrence probability, a candidate fault evidence chain abstract and a candidate root cause fault confidence; and the target executable decision sequence comprises a target execution order between each target root cause fault link, a target fault operation and maintenance process of each target root cause fault link, a target execution password in each target root cause fault link, a reference execution index in each target root cause fault link, a reference rollback point corresponding to the reference execution index and a reference execution effectiveness criterion.
[0124] In an optional embodiment of the present application, the device further comprises: an actual operation data acquisition module, configured to periodically acquire candidate root cause fault information, an actual executable decision sequence, an actual operation task sequence and actual execution data of the target live network system; and a long-term fault management module, configured to use a long-term management large language model to perform management analysis on the candidate root cause fault information, the actual executable decision sequence, the actual operation task sequence and the actual execution data of the target live network system, to obtain a target management roadmap.
[0125] The live network system operation device provided in the embodiments of the present application can execute the live network system operation method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0126] In the technical solution of the embodiments of the present application, the acquisition, storage and application of the target live network system data, the historical operation knowledge base, the historical fault operation process, the historical effectiveness of the historical fault operation process, the target operation constraint information, the reference optimization target, the candidate root cause fault information of the target live network system, the actual executable decision sequence, the actual operation task sequence and the actual execution data, etc., all conform to the relevant legal regulations and do not violate public order and good customs.
[0127] Embodiment Four
[0128] Figure 9 A structural schematic diagram of an electronic device 900 used to implement an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device also represents various forms of mobile devices, such as personal digital assistants, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.
[0129] As Figure 9As shown, the electronic device 900 includes at least one processor 901, and a memory, such as a read-only memory (ROM) 902, a random access memory (RAM) 903, etc., connected to the at least one processor 901 in communication. The memory stores computer programs executable by the at least one processor, and the processor 901 performs various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 902 or loaded from the storage unit 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The processor 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0130] A plurality of components in the electronic device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc., an output unit 907, such as various types of displays, a speaker, etc., a storage unit 908, such as a magnetic disk, an optical disk, etc., and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0131] The processor 901 is various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 901 performs various methods and processes described above, such as the in-network system operation method.
[0132] In some embodiments, the in-network system operation method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 908. In some embodiments, part or all of the computer program is loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the processor 901, one or more steps of the in-network system operation method described above are performed. Alternatively, in other embodiments, the processor 901 is configured to perform the in-network system operation method by any other appropriate means, such as by means of firmware.
[0133] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0134] Computer programs used to implement the processes of the application are written in any combination of one or more programming languages. These computer programs provide instructions that cause a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to perform the processes specified in the flow diagrams and / or block diagrams. The computer programs are completely executed on a machine, partially executed on a machine, partially executed on a machine as a standalone software package and partially executed on a remote machine, or completely executed on a remote machine or server.
[0135] In the context of the present application, a computer-readable storage medium is a tangible medium that contains or stores a computer program for use by or in connection with an instruction execution system, apparatus, or device. Computer-readable storage media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium is a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0136] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0137] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0138] The computing system includes a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server is one of client-server relationship, where the server is a cloud server, also known as cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS (Virtual Private Server, virtual private server) service.
[0139] It should be understood that the steps shown in the various forms above can be reordered, added to, or deleted from. For example, the steps described in the present application can be executed in parallel, in sequence, or in different orders, as long as the desired results of the technical solutions of the present application can be achieved, and the present application is not limited herein.
[0140] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for operating and maintaining a live network system, characterized in that, The method includes: Obtain target network system data, and based on the target network system data, generate target topology dependency graph and target indicator triplet for the target network system; The historical operation and maintenance knowledge base is acquired, and the root cause localization big language model is used. Based on the target topology dependency graph, the target indicator triplet and the historical operation and maintenance knowledge base, anomaly clustering and causal analysis are performed on the target live network system to obtain candidate root cause fault information of the target live network system. The system acquires historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives. Then, using a large language model for operation and maintenance decisions, it makes operation and maintenance decisions on the target network system based on the candidate root cause fault information of the target network system, the historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives, thereby obtaining a target executable decision sequence. Based on the target executable decision sequence, a target operation and maintenance task sequence for the target live network system is generated, and the target live network system is operated and maintained according to the target operation and maintenance task sequence.
2. The method according to claim 1, characterized in that, The step of performing maintenance on the target existing network system according to the target maintenance task sequence includes: Select and execute the current target operation and maintenance task according to the target operation and maintenance task sequence; When the current target operation and maintenance task is completed, the current execution effectiveness of the current execution indicators of the current target operation and maintenance task is detected; Based on the current execution results, update and execute the current target operation and maintenance task, and return to the step of detecting the current execution results of the current execution indicators of the current target operation and maintenance task when the current target operation and maintenance task is completed, until all target operation and maintenance tasks in the target operation and maintenance task sequence are completed.
3. The method according to claim 2, characterized in that, The step of updating and executing the current target maintenance task based on the current execution results includes: When the current execution indicator value of the current target operation and maintenance task reaches the first reference key indicator value, the next target operation and maintenance task in the target operation and maintenance task sequence is updated to the current target operation and maintenance task and executed. When the current execution indicator value of the current target operation and maintenance task does not reach the first reference key indicator value, the target operation and maintenance task sequence is rolled back according to the reference rollback point contained in the target executable decision sequence, and the current target operation and maintenance task is updated.
4. The method according to claim 2, characterized in that, The method employs a large language model for operation and maintenance decisions. Based on the candidate root cause fault information of the target network system, the historical fault operation and maintenance processes, the historical effectiveness of the historical fault operation and maintenance processes, the target operation and maintenance constraints, and the reference optimization objectives, it makes operation and maintenance decisions for the target network system, resulting in a target executable decision sequence, including: Using a large language model for operation and maintenance decision-making, based on the candidate root cause fault information of the target network system, the historical fault operation and maintenance process, the historical effectiveness of the historical fault operation and maintenance process, the target operation and maintenance constraint information, and the reference optimization objective, operation and maintenance decisions are made for the target network system, and at least two candidate executable decision sequences and the candidate decision confidence of each candidate executable decision sequence are output. Based on the confidence levels of each candidate decision, a target executable decision sequence is selected from each of the candidate executable decision sequences; After detecting the current execution effectiveness of the current execution metrics of the current target operation and maintenance task, the method further includes: When the current execution indicator value of the current target operation and maintenance task does not reach the second reference key indicator value, the target executable decision sequence is selected from the candidate executable decisions other than the target executable decision sequence according to the confidence level of each candidate decision, and the process of generating the target operation and maintenance task sequence of the target network system according to the target executable decision sequence is returned.
5. The method according to claim 1, characterized in that, The candidate root cause fault information includes the candidate root cause fault stage, the impact area of the candidate fault, the recurrence probability of the candidate fault, the summary of the evidence chain of the candidate fault, and the confidence level of the candidate root cause fault; the target executable decision sequence includes the target execution order between each target root cause fault stage, the target fault operation and maintenance process of each target root cause fault stage, the target execution password within each target root cause fault stage, the reference execution indicators within each target root cause fault stage, the reference rollback point corresponding to the reference execution indicators, and the reference execution effectiveness criteria.
6. The method according to claim 1, characterized in that, Also includes: Periodically acquire candidate root cause fault information, actual executable decision sequences, actual operation and maintenance task sequences, and actual execution data of the target live network system; Using a long-term governance big language model, governance analysis is performed on the candidate root cause fault information, the actual executable decision sequence, the actual operation and maintenance task sequence, and the actual execution data of the target network system to obtain the target governance roadmap.
7. A network system operation and maintenance device, characterized in that, The device includes: The target live network system data acquisition module is used to acquire target live network system data and generate target topology dependency graph and target indicator triplet of the target live network system based on the target live network system data; The fault root cause localization module is used to acquire the historical operation and maintenance knowledge base, and use the root cause localization big language model to perform anomaly clustering and causal analysis on the target topology dependency graph, the target indicator triplet and the historical operation and maintenance knowledge base to obtain candidate root cause fault information of the target live network system. The operation and maintenance decision output module is used to acquire historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives. It adopts the operation and maintenance decision big language model to make operation and maintenance decisions for the target network system based on the candidate root cause fault information of the target network system, the historical fault operation and maintenance processes, historical effectiveness of the historical fault operation and maintenance processes, target operation and maintenance constraint information, and reference optimization objectives, and outputs the target executable decision sequence. The target live network system operation and maintenance module is used to generate a target operation and maintenance task sequence for the target live network system based on the target executable decision sequence, and to perform operation and maintenance on the target live network system according to the target operation and maintenance task sequence.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the live network system operation and maintenance method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the network system operation and maintenance method according to any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the network system operation and maintenance method according to any one of claims 1-6.
Citation Information
Patent Citations
Large model enhanced equipment operation and maintenance multi-modal knowledge graph construction method and system and storage medium
CN120450011A
Operation and maintenance data exception positioning method and computing equipment
CN120596301A