Transmission network fault intelligent pipeline closed-loop processing method and device

By building an intelligent pipeline closed-loop processing method for transmission network faults, automated and intelligent fault processing is achieved, solving the problem of low fault analysis and processing efficiency in existing technologies, improving network maintenance efficiency and reducing operating costs.

CN116208467BActive Publication Date: 2025-09-09WUHAN OPTICAL NETWORK INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310215216.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-07
Publication Date
2025-09-09
Estimated Expiration
2043-03-07

AI Technical Summary

Technical Problem

Existing technologies lack automated and intelligent closed-loop processing methods for transmission network fault handling, resulting in low fault analysis and processing efficiency, reliance on manual experience, and increased network operation and maintenance costs.

Method used

Construct an intelligent pipeline closed-loop processing method for transmission network faults. By performing a two-level classification of fault types, using knowledge analysis methods to generate typical fault analysis and processing processes, building a simulated network environment, and establishing fault intelligent pipeline nodes, including nodes for alarm generation, location, analysis, processing, and elimination, to achieve automated and intelligent fault processing.

Benefits of technology

It improves the automation level of fault analysis and processing, improves network maintenance efficiency, reduces dependence on the experience of fault maintenance personnel, and reduces network operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116208467B_ABST
    Figure CN116208467B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for intelligent pipeline closed-loop processing of transmission network faults. The method comprises the following steps: classifying fault types and performing secondary classification on various fault scenarios; employing a knowledge analysis method to generate general nodes and processes for typical fault analysis and processing flows based on fault handling cases and user help text; generating a simulated network environment based on the topology, configuration, and operating status of the transmission network managed by the network management system; constructing an intelligent pipeline for transmission network faults; constructing an intelligent pipeline node operating status monitor and scheduler responsible for pipeline execution status monitoring and exception handling scheduling; constructing a pipeline node operating status monitor and scheduler, wherein the status monitor is responsible for node operating status monitoring, the scheduler is responsible for exception handling, and provides manual scheduling and adjustment functions for the fault handling flow. The present invention also provides a corresponding intelligent pipeline closed-loop processing device for transmission network faults.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent operation and maintenance technology, and more specifically, relates to a method and device for intelligent pipeline closed-loop processing of transmission network faults. Background Art

[0002] With the development of intelligent transmission networks, the Telecommunications Management Forum (TMF) has proposed the concept of autonomous networks and a series of standard recommendations. Autonomous networks place higher demands on intelligent fault handling in transmission networks, requiring automation and intelligence. Therefore, a closed-loop approach to transmission network fault handling, from alarm generation to fault resolution, is required. Traditional manual, semi-manual, multi-step, coordinated fault handling should be developed into an automated, intelligent, closed-loop approach. This will improve fault analysis and handling efficiency, reduce reliance on the experience of maintenance personnel, and ultimately reduce network operation and maintenance costs. Summary of the Invention

[0003] In response to the above-mentioned defects or improvement needs of the existing technology, the present invention provides a method and device for intelligent pipeline closed-loop processing of transmission network faults to achieve automated and intelligent closed-loop processing of faults, thereby improving the efficiency of fault analysis and processing, reducing the strong dependence on the experience of fault maintenance personnel, and thus reducing network operation and maintenance costs.

[0004] To achieve the above objectives, according to one aspect of the present invention, a method for intelligent pipeline closed-loop processing of transmission network faults is provided, the method comprising the following steps:

[0005] Classify fault types and categorize various fault scenarios into two levels. Use knowledge analysis methods to generate common nodes and processes for typical fault analysis and processing flows based on fault handling cases and user help text. Generate a simulated network environment based on the topology, configuration, and operating status of the transmission network managed by the network management system.

[0006] Build an intelligent pipeline for transmission network faults, including alarm generation nodes, alarm reporting nodes, alarm reduction nodes, root alarm location nodes, fault analysis and identification nodes, fault handling solution nodes, fault handling execution nodes, and fault elimination nodes. Build an intelligent pipeline node operation status monitor and scheduler for fault handling, responsible for pipeline execution status monitoring and exception handling scheduling.

[0007] Construct a pipeline node operation status monitor and scheduler. The status monitor is responsible for monitoring the node operation status, and the scheduler is responsible for exception handling and providing manual arrangement and adjustment functions for the fault handling process.

[0008] In one embodiment of the present invention, the fault types are divided into two levels, and various fault scenarios are classified, specifically including:

[0009] Fault scenarios are classified using a two-level approach. The first level is based on the role of the faulty object in the network, and is divided into service, equipment, line, environment, and network management categories.

[0010] The second-level scenario is divided based on the specific impact and root cause of the fault under the first-level scenario;

[0011] Secondary service scenarios include optical layer service interruption, electrical layer service interruption, tunnel layer service interruption, pseudo wire layer service interruption, and client layer service interruption; optical layer service performance degradation, electrical layer service performance degradation, tunnel layer service performance degradation, pseudo wire layer service performance degradation, and client layer service performance degradation; and protection group failure.

[0012] Secondary equipment scenarios include single disk failure, master / slave disk switchover failure, power disk failure, service disk signal loss, lightning protection module failure, device power outage, and module aging.

[0013] Line-related secondary scenarios include line interruption, abnormal line optical power, excessive line loss, line relay, and pigtail.

[0014] Environmental secondary scenarios are divided into temperature anomalies, voltage anomalies, and humidity anomalies;

[0015] Secondary network management scenarios include: network element outage, single disk outage, DCN network anomalies, and network management service anomalies.

[0016] The fault type is determined by the combined values ​​of the first and second levels of the fault scenario.

[0017] In one embodiment of the present invention, the knowledge analysis method is used to generate the general nodes and processes of a typical fault analysis and processing flow through fault handling cases and fault handling user help text, specifically including:

[0018] The titles of troubleshooting cases and user help texts should use the fault type format. The format of each description should be as follows: "Number + Action + Specific Object + Result Judgment Branch + Next Step Number of Branch". For pure operation statements, only "Number + Action + Specific Object" is required.

[0019] Each type of action and object + result judgment can generate a common node for the fault analysis and processing process. The common process nodes are divided into two major categories: fault troubleshooting and fault recovery. Each major category is further divided into network general category, OTN network category, and packet network category. These common process nodes are deduplicated and stored in the common process node component library.

[0020] Each common node in the process is marked as either automated or manual. Automated nodes can be executed online through automated programs. Software code must be developed to implement this functionality, and a parameterized call interface must be provided to invoke the operation. Manual nodes currently require offline manual operation, with the results entered into the system.

[0021] Through the analysis of fault handling cases and fault handling user help texts, a fault troubleshooting process table indexed by root alarms and derivative alarm codes and a fault recovery process table indexed by fault scenarios are generated and stored in the process general node component library.

[0022] In one embodiment of the present invention, generating a simulated network environment based on the topology, configuration, and operating status of the transmission network managed by the network management system specifically includes:

[0023] Based on the transmission network topology scope of fault management, the network simulation service is started, and the configuration and operating status of the current network nodes are synchronized to generate a simulated network environment that can be operated through the management and control system; fault recovery node operations during troubleshooting are all performed in the simulated network environment during the troubleshooting.

[0024] In one embodiment of the present invention,

[0025] The alarm generating node is responsible for collecting alarm information on the network element device node, deduplicating the collected information, and transmitting the collected information to the alarm reporting node; the node is deployed on the network element device;

[0026] The alarm reporting node reports the acquired alarm information to the management and control system through the reporting protocol agreed upon with the management and control system, stores the alarm information in the original alarm information database, and transmits the alarm information to the alarm reduction node. The node consists of two parts: a server and a client. The server is deployed on the network element device and is responsible for protocol assembly and sending of alarm information. The client is deployed on the management and control system and is responsible for receiving alarm information and protocol unsealing.

[0027] The alarm reduction node is responsible for deduplicating the received alarm information and removing the oscillation alarm according to the reduction strategy, and transmitting the processed alarm information to the root alarm location node;

[0028] The root alarm location node is responsible for analyzing a group of alarms to determine the root-derived relationships based on network topology information, service path information, static root-derived relationships of alarms, alarm occurrence time, and acquired alarm information, determining the root alarm and derived alarms, and passing this group of root-derived relationships to the fault analysis and identification node;

[0029] The fault analysis and identification node searches for the corresponding troubleshooting process in the process general node component library based on a set of root alarms and derived alarms, and calls the corresponding process general node component to instantiate and execute it according to the process; thereby, the root cause of the fault is identified, the fault scenario is determined, and the fault scenario is passed to the fault handling solution node;

[0030] The fault handling solution node searches for the corresponding fault recovery process in the process general node component library according to the fault scenario, and calls the corresponding fault recovery process general node component according to the process to instantiate it, generates a fault handling solution, and provides it to the fault handling execution node;

[0031] The fault handling execution node: This node executes the fault handling solution in the simulated network environment and evaluates the execution results. After the fault is eliminated after the execution in the simulated network environment, the solution can be executed in the real physical network environment. After the execution is completed, the fault elimination node is notified.

[0032] The fault elimination node: after receiving the notification of completion of a certain fault processing sent by the fault execution node, confirms that the fault has been eliminated and stores the fault data in the historical fault database.

[0033] In one embodiment of the present invention, the specific execution methods of the fault analysis and identification, fault handling solution, fault handling execution, and fault elimination are as follows:

[0034] (3.1) The fault analysis identifies one or more root alarms determined based on the root alarm location, combines the analysis and troubleshooting nodes and processes in the fault analysis and processing general node component library, and determines the type of fault;

[0035] (3.2) The fault handling solution uses the fault scenario type as an index, finds the corresponding fault recovery process from the fault recovery process table of the process general node component library, determines the instantiation parameters of each node according to the execution order of the process general nodes recorded in the fault recovery process, and generates a fault handling solution;

[0036] (3.3) The fault handling execution is carried out in accordance with the fault handling solution generated in (3.2) above, in a simulated network environment, displaying and evaluating the execution results of each node, thereby evaluating whether the entire fault handling solution is effective; if effective, proceed to (3.4); if not, proceed to (3.5); the fault handling process is manually arranged and adjusted through the pipeline node operation status monitor and scheduler;

[0037] (3.4) The fault elimination is to execute the fault handling solution in (3.3) in the simulated network environment in the real network environment to eliminate the fault;

[0038] (3.5) The fault handling process is adjusted through the pipeline node operation status monitor and scheduler to form an adjusted fault handling plan, and then transferred to (3.3) for execution.

[0039] In one embodiment of the present invention, the fault analysis and identification specifically includes:

[0040] Based on one or more root alarms identified by root alarm positioning, a "troubleshooting table alarm code index" is generated, and the corresponding troubleshooting process is found from the troubleshooting process table of the process general node component library. According to the execution order of the process general nodes recorded in the troubleshooting process, the instantiation parameters of each node are determined, and the call execution is performed, and finally the root cause of the fault is found, thereby determining the type of fault.

[0041] The construction of the fault intelligent pipeline node operation status monitor and scheduler specifically includes:

[0042] The status monitor is responsible for monitoring the status of all nodes, recording and displaying the current node and status of the process execution; when the management and control system runs abnormally or a user shuts down the system and logs back in, the scheduler is responsible for continuing to run the process under the current node.

[0043] In one embodiment of the present invention, the status monitor is also responsible for monitoring the abnormalities of each process and node execution, and quickly restarting the nodes with service abnormalities. The scheduler also provides manual orchestration services to optimize the execution process of the node.

[0044] According to another aspect of the present invention, there is also provided an intelligent pipeline closed-loop processing device for transmission network faults, comprising at least one processor and a memory, wherein the at least one processor and the memory are connected via a data bus, and the memory stores instructions that can be executed by the at least one processor. After being executed by the processor, the instructions are used to complete the intelligent pipeline closed-loop processing method for transmission network faults.

[0045] In general, the above technical solutions conceived by the present invention have the following beneficial effects compared with the prior art:

[0046] (1) The present invention constructs an intelligent pipeline for transmission network faults, defines pipeline nodes, and implements end-to-end closed-loop fault processing;

[0047] (2) The present invention provides key node processing methods such as fault analysis and identification, fault handling solutions, fault handling execution, and fault elimination;

[0048] (3) The present invention realizes an intelligent pipeline for transmission network faults, improves the automation level of fault recovery, and thus improves network maintenance efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 1 is a flow chart of a method for intelligent pipeline closed-loop processing of transmission network faults according to an embodiment of the present invention;

[0050] Figure 2 This is a diagram showing the operating principle of a transmission network fault intelligent pipeline closed-loop processing node in an embodiment of the present invention. DETAILED DESCRIPTION

[0051] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0052] In order to solve the problems existing in the prior art, the present invention defines each node of the fault closed-loop processing from alarm generation, alarm reporting, alarm volume reduction, root alarm location, fault analysis and identification, fault handling plan, fault handling execution, to fault elimination. By introducing automation and intelligent technology, an intelligent pipeline closed-loop processing solution for transmission network faults is formed.

[0053] like Figure 1 As shown, the present invention provides a method for intelligent pipeline closed-loop processing of transmission network faults, comprising the following steps:

[0054] (1) Classify fault types and classify various fault scenarios into two levels; use knowledge analysis methods to generate common nodes and processes of typical fault analysis and processing processes through fault handling cases and fault handling user help texts; generate a simulated network environment based on the transmission network topology, configuration, and operating status managed by the network management system. Specifically, it includes:

[0055] (1.1) Fault type classification: various fault scenarios are classified into two levels.

[0056] Fault scenarios are classified using a two-level approach. The first level is based on the role of the faulty object in the network, and is divided into service, equipment, line, environment, and network management categories.

[0057] The second-level scenario is divided based on the specific impact and root cause of the fault under the first-level scenario.

[0058] Secondary service scenarios include optical layer service interruption, electrical layer service interruption, tunnel layer service interruption, pseudo wire layer service interruption, and customer layer service interruption; optical layer service performance degradation, electrical layer service performance degradation, tunnel layer service performance degradation, pseudo wire layer service performance degradation, and customer layer service performance degradation; and protection group failure.

[0059] Secondary equipment scenarios include single disk failure, master / slave disk switchover failure, power disk failure, service disk signal loss, lightning protection module failure, equipment power outage, and module aging.

[0060] Line-related secondary scenarios include line interruption, abnormal line optical power, excessive line loss, line relay, and fiber pigtail.

[0061] Secondary environmental scenarios include abnormal temperature, abnormal voltage, and abnormal humidity.

[0062] Secondary network management scenarios include network element disconnection, single disk disconnection, DCN (Data Communication Network) network anomalies, and network management service anomalies.

[0063] Fault scenario classification is based on root alarms and derivative alarms, combined with experience in fault cause analysis and troubleshooting. In addition, the occurrence of new network faults and the accumulation of experience in manual fault cause analysis and troubleshooting can expand the classification of secondary scenarios through the experience of fault points.

[0064] The fault type is determined by the combined values ​​of the first and second levels of the fault scenario.

[0065] (1.2) Using knowledge analysis methods, generate common nodes and processes of typical fault analysis and processing procedures through fault handling cases and fault handling user help texts.

[0066] The titles of troubleshooting cases and user help texts use the fault type format. The format requirements for each description are as follows:

[0067] "Number + action + specific object + result judgment branch + branch next step number", among which for pure operation statements, there is only "number + action + specific object".

[0068] Each type of action and object, combined with result judgment, generates a common node for the fault analysis and processing process. These common process nodes are categorized into two main categories: troubleshooting and recovery. Each category is further divided into general network, OTN (Optical Transport Network), and packet network categories. These common process nodes are deduplicated and stored in a common process node component library.

[0069] In addition, each common process node is identified as either automated or manual. Automated nodes can be executed online through program automation. These nodes require the development of corresponding software code to implement this functionality and provide a parameterized call interface for invoking the operation, such as checking an alarm node. Manually operated nodes, such as checking a pigtail node, currently require manual offline operation and the results must be entered into the system.

[0070] For example, optical port P1 on an XGE line tray on a packet network element A reports R_LOS and LINK_LOS alarms, and a large number of tunnel service switching alarms occur simultaneously. The network management system performs root cause analysis on the alarms and finds that R_LOS is the root alarm, LINK_LOS is a level 1 derivative alarm, and the tunnel service switching alarm is a level 2 derivative alarm. Maintenance personnel resolve the fault by finding an optical cable break, which is restored after fiber splicing. The case is summarized as a "Line Type - Line Disconnection" fault type. The following table shows an example of the common nodes in the process for generating this case:

[0071] Table 1 Line Class - Line Interruption Fault Type

[0072]

[0073] By adopting the knowledge analysis method and analyzing the fault handling cases and fault handling user help texts, a fault troubleshooting process table indexed by root alarms and derived alarm codes and a fault recovery process table indexed by fault scenarios are generated. Both are stored in the process general node component library.

[0074] Each automated operation node will develop an interface function to complete the execution operation. For example, the definition of "check alarm" is:

[0075] bool checkAlarm(int objectID, int alarmType)

[0076] {

[0077] Execute the alarm found for the object and determine whether the alarm exists;

[0078] }

[0079] As shown in the example in the table above, a 1-9 processing flow table will be generated with the R_LOS alarm code and the first-level derived alarm code in the root-derived relationship tree as the index.

[0080] Note: There may be multiple root alarms and multiple first-level derived alarms. To increase the efficiency of generation and query, the index is implemented using an 8-byte integer, and a maximum of four alarm codes are used as indexes, as follows:

[0081] Table 2 Troubleshooting Flowchart Alarm Code Index

[0082]

[0083] As shown in the example in the table above, a fault recovery process table indexed by the fault scenario is generated. For example, if the index is the two field values ​​corresponding to the line type-line interruption fault type, the fault recovery process only includes the "Process Optical Cable" node and the "Process End" node.

[0084] (1.3) Generate a simulated network environment based on the transmission network topology, configuration, and operating status managed by the network management system.

[0085] Based on the transmission network topology scope of fault management, the network simulation service is started, and the configuration and operating status (including single disk status, current alarms, current performance, etc.) of the current network nodes (including single disks) are synchronized (manually or scheduled), generating a simulated network environment that can be operated through the management and control system.

[0086] Fault recovery node operations during troubleshooting are all performed in a simulated network environment. (This requires that the troubleshooting process for troubleshooting must be performed on a simulated network. Only after the fault handling has been confirmed to have effectively eliminated the fault in the simulated network can it be performed on the physical network.)

[0087] (2) Construct an intelligent pipeline for transmission network faults to achieve closed-loop fault processing; the intelligent pipeline includes an alarm generation node, an alarm reporting node, an alarm reduction node, a root alarm location node, a fault analysis and identification node, a fault handling solution node, a fault handling execution node, and a fault elimination node; construct an intelligent pipeline node operation status monitor and scheduler to be responsible for pipeline execution status monitoring and exception handling scheduling. Specifically, Figure 2 As shown, including:

[0088] The intelligent fault pipeline includes eight nodes: alarm generation, alarm reporting, alarm volume reduction, root alarm location, fault analysis and identification, fault handling plan, fault handling execution, and fault elimination.

[0089] The alarm generating node is responsible for collecting alarm information on the network element device node, performing deduplication processing on the collected information, and transmitting the collected information to the alarm reporting node. This node is deployed on the network element device.

[0090] The alarm reporting node reports the acquired alarm information to the control system through the reporting protocol agreed upon with the control system, stores the alarm information in the original alarm information database, and transmits the alarm information to the alarm reduction node. This node consists of two parts: a server and a client. The server is deployed on the network element device and is mainly responsible for protocol assembly and transmission of alarm information. The client is deployed on the control system and is mainly responsible for receiving alarm information and protocol decryption.

[0091] The alarm reduction node is responsible for deduplicating the received alarm information, removing oscillation alarms, etc. according to the reduction strategy, and transmitting the processed alarm information to the root alarm positioning node.

[0092] The root alarm location node is responsible for analyzing the root-derivative relationship of a group of alarms based on network topology information, service path information, static root-derivative relationship of alarms (which can be generated by human experience and AI training), alarm occurrence time and acquired alarm information, determining the root alarm and derived alarms, and passing this group of root-derivative relationships to the fault analysis and identification node.

[0093] The fault analysis and identification node searches for the corresponding fault troubleshooting process in the process general node component library in step (1) (1.2) based on a set of root alarms and derived alarms, and calls the corresponding process general node component to instantiate and execute it according to the process. This will identify the root cause of the fault and determine the fault scenario. The fault scenario is then passed to the fault handling solution node.

[0094] The fault handling solution node searches for the corresponding fault recovery process in the process general node component library according to the fault scenario, and calls the corresponding fault recovery process general node component to instantiate it according to the process, generates a fault handling solution, and provides it to the fault handling execution node.

[0095] The fault handling execution node executes the fault handling solution in a simulated network environment (the simulated network is used to troubleshoot and confirm effective fault resolution, ensuring that the physical network is not affected before an effective solution is determined. The simulated network should be as similar as possible to the local network in the physical network to be operated. The simulation software for network elements and single disks should be consistent with the software versions on the physical network elements, and the network element and single disk configuration information, status, and current alarms should be synchronized). This node then evaluates the execution results. Only after the fault is resolved in the simulated network environment can the solution be executed in the real physical network environment. Upon completion, the fault resolution node is notified.

[0096] The fault elimination node: after receiving the notification of completion of a certain fault processing sent by the fault execution node, confirms that the fault has been eliminated and stores the fault data in the historical fault database.

[0097] (3) Construct a pipeline node operation status monitor and scheduler. The status monitor is responsible for monitoring the node operation status, and the scheduler is responsible for exception handling and providing manual arrangement and adjustment functions for the fault handling process.

[0098] The status monitor is responsible for monitoring the status of all nodes in step (2) above, recording and displaying the current node and status of the process execution. When the management and control system runs abnormally or a user shuts down the system and logs back in, the scheduler is responsible for continuing to run the process at the current node.

[0099] The status monitor is also responsible for monitoring the execution of each process and node for abnormalities, and quickly restarting nodes with abnormal services.

[0100] The scheduler also provides manual orchestration services, allowing users to manually orchestrate the execution sub-processes of the fault analysis and identification nodes and the fault resolution nodes. Certain common node components can be omitted or added to these sub-processes. Users can choose whether to store these orchestrated fault analysis and resolution processes in the troubleshooting and recovery process tables.

[0101] The pipeline node operation is responsible for monitoring the node operation status and handling exceptions, and also provides manual arrangement and adjustment functions for the fault handling process.

[0102] Alarm generation and reporting can be achieved using currently available, mature technologies. Alarm volume reduction is achieved by removing duplicate alarms and suppressing oscillating alarms. Root alarm location uses a manually or intelligently determined alarm root-derived relationship table to determine the alarm derivation tree, thereby identifying one or more root alarms. These technologies are already widely used in the industry and will not be detailed here. This technology primarily addresses the less mature stages of fault analysis and identification, fault resolution, fault resolution execution, and fault elimination.

[0103] (3.1) Fault analysis identifies one or more root alarms based on root alarm location, and combines the analysis and troubleshooting nodes and processes in the fault analysis and processing general node component library to determine the fault type.

[0104] Specifically, based on one or more root alarms determined by the root alarm location, generate the "Troubleshooting Table Alarm Code Index" in Table 2, find the corresponding troubleshooting process from the troubleshooting process table of the process general node component library, and determine the instantiation parameters of each node according to the execution order of the process general nodes recorded in the troubleshooting process. Call and execute, and finally find out the root cause of the fault, thereby determining the type of fault.

[0105] For example, in the "Check Alarm" process node, if the process node is "Check whether the R_LOS alarm on this end has disappeared. If it has, go to 2; if not, go to 3," parameter 1, the network object ID, is the identifier of the local network element port; parameter 2, the alarm code, is instantiated as the R_LOS alarm code. The corresponding execution function of the "Check Alarm" process node is then called, and the next node of the process execution is determined based on the return value:

[0106] If(checkAlarm(local NE port ID, R_LOS alarm code))

[0107] {

[0108] Turn 2;

[0109] }

[0110] Else

[0111] {

[0112] Turn 3;

[0113] }

[0114] (3.2) The fault handling plan uses the fault scenario type as an index to find the corresponding fault recovery process from the fault recovery process table of the process general node component library. According to the execution order of the process general nodes recorded in the fault recovery process, the instantiation parameters of each node are determined to generate a fault handling plan.

[0115] (3.3) Fault handling is executed in a simulated network environment according to the fault handling plan generated in (3.2). The execution results of each node are displayed and evaluated to determine the effectiveness of the entire fault handling plan. If effective, proceed to (3.4); otherwise, proceed to (3.5). Manual adjustments to the fault handling process are made through the pipeline node operation status monitor and scheduler.

[0116] (3.4) Fault elimination is to execute the fault handling solution in (3.3) in the simulated network environment in the real network environment to eliminate the fault.

[0117] (3.5) The fault handling process is adjusted through the pipeline node operation status monitor and scheduler to form an adjusted fault handling plan, and then transferred to (3.3) for execution.

[0118] Furthermore, the present invention also provides an intelligent pipeline closed-loop processing device for transmission network faults, comprising at least one processor and a memory, wherein the at least one processor and the memory are connected via a data bus, and the memory stores instructions that can be executed by the at least one processor. After being executed by the processor, the instructions are used to complete the intelligent pipeline closed-loop processing method for transmission network faults.

[0119] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A transmission network fault intelligent pipeline closed-loop processing method, characterized in that: The method comprises the following steps: Classify fault types and classify various fault scenarios into two levels; use knowledge analysis methods to generate common nodes and processes for typical fault analysis and processing flows through fault handling cases and fault handling user help texts; Generate a simulated network environment based on the transmission network topology, configuration, and operating status managed by the network management system; Build an intelligent pipeline for transmission network faults, including alarm generation nodes, alarm reporting nodes, alarm reduction nodes, root alarm location nodes, fault analysis and identification nodes, fault resolution nodes, fault processing execution nodes, and fault elimination nodes. Build an intelligent pipeline node status monitor and scheduler for fault status monitoring and exception handling scheduling. The status monitor is responsible for monitoring the node operation status, and the scheduler is responsible for exception handling and providing manual arrangement and adjustment functions for the fault handling process; The fault handling solution node searches for the corresponding fault recovery process in the process general node component library according to the fault scenario, and calls the corresponding fault recovery process general node component according to the process to instantiate it, generates a fault handling solution, and provides it to the fault handling execution node; The fault handling execution node: This node executes the fault handling plan in the simulated network environment and evaluates the execution results; after the fault is eliminated after execution in the simulated network environment, it can be executed in the real physical network environment, and after execution is completed, the fault elimination node is notified.

2. The transmission network fault intelligent pipeline closed-loop processing method according to claim 1, characterized in that: The fault types are divided into two levels, and various fault scenarios are classified into the following categories: Fault scenarios are classified using a two-level approach. The first level is based on the role of the faulty object in the network, and is divided into service, equipment, line, environment, and network management categories. The second-level scenario is divided based on the specific impact and root cause of the fault under the first-level scenario; Secondary service scenarios include optical layer service interruption, electrical layer service interruption, tunnel layer service interruption, pseudo wire layer service interruption, and client layer service interruption; optical layer service performance degradation, electrical layer service performance degradation, tunnel layer service performance degradation, pseudo wire layer service performance degradation, and client layer service performance degradation; and protection group failure. Secondary equipment scenarios include single disk failure, master / slave disk switchover failure, power disk failure, service disk signal loss, lightning protection module failure, device power outage, and module aging. Line-related secondary scenarios include line interruption, abnormal line optical power, excessive line loss, line relay, and pigtail. Environmental secondary scenarios are divided into temperature anomalies, voltage anomalies, and humidity anomalies; Secondary network management scenarios include network element outage, single disk outage, DCN network anomalies, and network management service anomalies. The fault type is determined by the combined values ​​of the first and second levels of the fault scenario.

3. The transmission network fault intelligent pipeline closed-loop processing method according to claim 1 or 2, characterized in that: The knowledge analysis method is used to generate the general nodes and processes of a typical fault analysis and processing flow through fault handling cases and fault handling user help texts, specifically including: The titles of troubleshooting cases and user help texts should use the fault type format. The format of each description should be as follows: "number + action + specific object + result judgment branch + next step number of the branch." For pure operation statements, only "number + action + specific object" is required. Each type of action and object + result judgment can generate a common node for the fault analysis and processing process. The common process nodes are divided into two major categories: fault troubleshooting and fault recovery. Each major category is further divided into network general category, OTN network category, and packet network category. These common process nodes are deduplicated and stored in the common process node component library. Each common node in the process is marked as either automated or manual. Automated nodes can be executed online through automated programs. These nodes require corresponding software code to implement the functionality and provide a parameter-based calling interface to invoke the operation. Manual nodes currently require offline manual operation and the results must be entered into the system. Through the analysis of fault handling cases and fault handling user help texts, a fault troubleshooting process table indexed by root alarm codes and derived alarm codes and a fault recovery process table indexed by fault scenarios are generated and stored in the process general node component library.

4. The transmission network fault intelligent pipeline closed-loop processing method according to claim 1 or 2, characterized in that: Generating a simulated network environment based on the topology, configuration, and operating status of the transmission network managed by the network management system specifically includes: Based on the transmission network topology scope of fault management, the network simulation service is started, and the configuration and operating status of the current network nodes are synchronized to generate a simulated network environment that can be operated through the management and control system; fault recovery node operations during troubleshooting are all performed in the simulated network environment during the troubleshooting.

5. The transmission network fault intelligent pipeline closed-loop processing method according to claim 1 or 2, characterized in that: The alarm generating node is responsible for collecting alarm information on the network element device node, deduplicating the collected information, and transmitting the collected information to the alarm reporting node; the node is deployed on the network element device; The alarm reporting node reports the acquired alarm information to the management and control system through the reporting protocol agreed upon with the management and control system, stores the alarm information in the original alarm information database, and transmits the alarm information to the alarm reduction node. The node consists of two parts: a server and a client. The server is deployed on the network element device and is responsible for protocol assembly and sending of alarm information. The client is deployed on the management and control system and is responsible for receiving alarm information and protocol unsealing. The alarm reduction node is responsible for deduplicating the received alarm information and removing the oscillation alarm according to the reduction strategy, and transmitting the processed alarm information to the root alarm location node; The root alarm location node is responsible for analyzing a group of alarms to determine the root-derived relationships based on network topology information, service path information, static root-derived relationships of alarms, alarm occurrence time, and acquired alarm information, determining the root alarm and derived alarms, and passing this group of root-derived relationships to the fault analysis and identification node; The fault analysis and identification node searches for a corresponding troubleshooting process in the process general node component library based on a set of root alarms and derived alarms, and calls the corresponding process general node component to instantiate and execute it according to the process; This helps identify the root cause of the fault and determine the fault scenario; And pass the fault scenario to the fault handling solution node; The fault elimination node: after receiving the notification of completion of a certain fault processing sent by the fault execution node, confirms that the fault has been eliminated and stores the fault data in the historical fault database.

6. The transmission network fault intelligent pipeline closed-loop processing method according to claim 5, characterized in that: The specific execution methods for fault analysis and identification, fault handling plan, fault handling execution, and fault elimination are as follows: The fault analysis in step 3.1 identifies one or more root alarms determined based on the root alarm location, combines the analysis and troubleshooting nodes and processes in the fault analysis and processing general node component library, and determines the fault type; The fault handling solution described in step 3.2 uses the fault scenario type as an index, finds the corresponding fault recovery process from the fault recovery process table of the process general node component library, determines the instantiation parameters of each node according to the execution order of the process general nodes recorded in the fault recovery process, and generates a fault handling solution; Step 3.3: The fault handling execution is to execute the fault handling solution generated in step 3.2 above in the simulated network environment, display and evaluate the execution results of each node, and thus evaluate whether the entire fault handling solution is effective; if it is effective, go to step 3.4; if not, go to step 3.5; Manually orchestrate and adjust the fault handling process through the pipeline node operation status monitor and scheduler; The fault elimination described in step 3.4 is to execute the fault handling solution in the simulated network environment in step 3.3 in the real network environment to eliminate the fault; In step 3.5, the fault handling process is adjusted through the pipeline node operation status monitor and scheduler to form an adjusted fault handling plan, and then the process goes to step 3.3 for execution.

7. The transmission network fault intelligent pipeline closed-loop processing method according to claim 6, characterized in that: The fault analysis and identification specifically includes: Based on one or more root alarms identified by root alarm location, a "troubleshooting table alarm code index" is generated. The corresponding troubleshooting process is found from the troubleshooting process table in the process general node component library. According to the execution order of the process general nodes recorded in the troubleshooting process, the instantiation parameters of each node are determined, and the call execution is finally carried out. The root cause of the fault is finally found, thereby determining the fault type.

8. The transmission network fault intelligent pipeline closed-loop processing method according to claim 1 or 2, characterized in that: The construction of the fault intelligent pipeline node operation status monitor and scheduler specifically includes: The status monitor is responsible for monitoring the running status of all nodes, recording and displaying the current node and status of process execution; when the management and control system runs abnormally or a user shuts down the system and logs back in, the scheduler is responsible for continuing to run the process under the current node.

9. The transmission network fault intelligent pipeline closed-loop processing method according to claim 8, characterized in that: The status monitor is also responsible for monitoring the abnormalities of each process and node execution, and quickly restarting the nodes with abnormal services. The scheduler also provides manual orchestration services to optimize the execution process of the node.

10. A transmission network fault intelligent pipeline closed-loop processing device, characterized by: The method comprises at least one processor and a memory, wherein the at least one processor and the memory are connected via a data bus, and the memory stores instructions that can be executed by the at least one processor, and after being executed by the processor, the instructions are used to complete the intelligent pipeline closed-loop processing method for transmission network faults according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Centralized alarm monitoring system and method of power system terminal communication access network

    CN107196804A

  • A method and system for determining fault of heterogeneous system based on machine learning

    CN111209131A