A method and apparatus for handling faults

By identifying and reporting root alarms within network elements, and combining the linkage mechanism between network elements and the control system, the problem of slow processing of non-atomic function root alarms in traditional fault handling processes has been solved, achieving fast and efficient fault handling.

CN119676056BActive Publication Date: 2026-03-31FIBERHOME TELECOMMUNICATION TECHNOLOGIES CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional fault handling processes are inefficient at quickly processing root alarms caused by non-atomic functions and lack a linkage mechanism between network elements and the control system.

Method used

By performing alarm correlation analysis on network elements, root alarms can be directly identified within the network element and the fault diagnosis process can be initiated. At the same time, the fault is reported to the management and control system. The management and control system performs fault identification and diagnosis based on network connection, providing a linkage mechanism between network elements and the management and control system.

Benefits of technology

It shortens the time for fault identification and diagnosis, improves the efficiency of fault handling, and does not affect the traditional fault identification process, thus achieving rapid fault handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119676056B_ABST
    Figure CN119676056B_ABST
Patent Text Reader

Abstract

The application relates to a fault processing method and device, and relates to the technical field of network management and control.The method comprises the following steps: after a root alarm in a network element is reported to a management and control system, the fault causing the root alarm of the non-atomic function is analyzed and recognized in the network element, and a network element fault diagnosis process is started; meanwhile, the fault is reported to the management and control system.The management and control system carries out a fault recognition process based on network connection after receiving the root alarm in the network element, and according to the network element fault diagnosis state, the fault causing the root alarm of the non-atomic function is removed, so that the processing efficiency of the fault is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network management and control technology, specifically to a method and apparatus for fault handling. Background Technology

[0002] Traditional fault handling procedures involve the control system analyzing alarms obtained from devices to identify root alarms related to network connections, and then diagnosing and handling the faults associated with these root alarms. Alarms generated by non-atomic functions are generally root alarms and can cause widespread business disruptions.

[0003] Currently, traditional fault handling procedures are also used to handle faults that cause root alarms in non-atomic functions, making it difficult to quickly resolve such faults. Furthermore, traditional fault handling methods do not provide a fault handling mechanism that links network elements and the control system, resulting in low fault handling efficiency. Summary of the Invention

[0004] This application provides a fault handling method and apparatus that can quickly handle faults that cause non-atomic functions to generate root alarms. It also provides a fault handling mechanism that links network elements and the control system, thereby improving the efficiency of fault handling.

[0005] In a first aspect, embodiments of this application provide a fault handling method, the method comprising:

[0006] The network element performs alarm correlation analysis to obtain the root alarm within the network element and reports it to the control system. When the root alarm within the network element is generated by a non-atomic function, the network element analyzes the fault that caused the root alarm and initiates the network element fault diagnosis process, while simultaneously reporting the fault to the control system.

[0007] After receiving a root alarm within a network element, the control system performs a fault identification process based on network connection and clears the faults that cause non-atomic functions to generate root alarms according to the fault diagnosis status of the network element.

[0008] In conjunction with the first aspect, in one implementation, the network element performs alarm correlation analysis to obtain the root alarm within the network element, including:

[0009] Correlation analysis is performed on alarms within network elements to obtain root alarms and corresponding derived alarms within the network element. Isolated alarms are also identified as root alarms within the network element.

[0010] In conjunction with the first aspect, in one implementation, reporting root alarms within a network element to the control system includes:

[0011] Report the basic information, alarm source classification, and diagnostic status of root alarms within the network element to the control system;

[0012] The alarm source classification includes atomic function generation and non-atomic function generation;

[0013] The diagnostic status includes not started, diagnostic in progress, and completed.

[0014] In conjunction with the first aspect, in one implementation, the control system, upon receiving a root alarm within a network element, performs a network connection-based fault identification process, including:

[0015] The control system performs alarm correlation analysis based on the root alarms reported by each network element to obtain the root alarms and corresponding derived alarms of the network connection. Further analysis reveals the faults that cause the root alarms of the network connection.

[0016] In conjunction with the first aspect, in one implementation, if the control system receives a network element fault diagnosis result before completing the network connection-based fault identification process, it includes:

[0017] If the fault that caused the non-atomic function to generate a root alarm has been cleared, the corresponding fault in the control system will be entered into the historical faults.

[0018] If the fault causing the non-atomic function to generate a root alarm is not cleared, the control system updates the information of the fault causing the non-atomic function to generate a root alarm based on the received network element fault diagnosis results, and performs manual diagnosis on the fault until the fault is cleared.

[0019] In conjunction with the first aspect, in one implementation, if the control system receives a network element fault diagnosis result after completing the network connection-based fault identification process, it includes:

[0020] The control system analyzes each fault that causes a root alarm in the network connection and determines whether it is the fault that causes a root alarm in the non-atomic function. If it is, the system updates the information of the fault that causes a root alarm in the non-atomic function based on the information of the fault that causes a root alarm in the network connection. If not, the control system handles the fault that causes a root alarm in the network connection in the traditional way.

[0021] After receiving the fault diagnosis results of the network element, the control system executes the process that it would have executed when it received the fault diagnosis results of the network element before completing the fault identification process based on network connection.

[0022] In conjunction with the first aspect, in one implementation, updating the information about the fault causing the non-atomic function to generate a root alarm based on the information about the fault causing the network connection to generate a root alarm includes:

[0023] Update the fault-related alarm information in the information of the fault that causes the non-atomic function to generate a root alarm to the fault-related alarm information in the information of the fault that causes the network connection to generate a root alarm.

[0024] In conjunction with the first aspect, in one implementation, if the fault causing the non-atomic function to generate a root alarm is not cleared, the control system updates the information of the fault causing the non-atomic function to generate a root alarm based on the received network element fault diagnosis results. The updated information includes:

[0025] Fault diagnosis status and fault diagnosis results.

[0026] In conjunction with the first aspect, in one implementation, the corresponding fault in the control system is entered into the historical fault list, including:

[0027] Based on the received network element fault diagnosis results and corresponding alarm clearing messages, the control system updates the corresponding fault diagnosis status, diagnosis result information and status information of all associated alarms in the control system, fills in the fault recovery time information, and transfers the fault from the current fault to the historical fault.

[0028] Secondly, embodiments of this application provide a fault handling apparatus based on any of the above methods, the apparatus comprising:

[0029] The network element layer analysis module is used to perform correlation analysis on alarms within the network element, obtain the root alarm within the network element and report it to the control system; it is also used to analyze the fault that caused the root alarm when the root alarm within the network element is generated by a non-atomic function and execute the network element fault diagnosis process.

[0030] The network layer analysis module is used to execute a fault identification process based on network connectivity.

[0031] The update module is used to mark the fault causing the non-atomic function to generate a root alarm as a historical fault when the fault causing the non-atomic function to generate a root alarm is successfully cleared; it is also used to update the information of the fault causing the non-atomic function to generate a root alarm based on the information of the fault causing the network connection to generate a root alarm when the identified fault causing the network connection to generate a root alarm is the fault causing the non-atomic function to generate a root alarm.

[0032] The processing module is used to clear the fault that causes the non-atomic function to generate a root alarm based on the sequential completion of the network element fault diagnosis process and the network connection-based fault identification process.

[0033] The beneficial effects of the technical solutions provided in this application include at least the following:

[0034] This method reports root alarms within network elements to the management and control system, then directly analyzes the faults causing root alarms in non-atomic functions within the network element and initiates the network element fault diagnosis process, while simultaneously reporting the faults to the management and control system. Compared to traditional fault identification and diagnosis processes, this method eliminates the need for the management and control system to analyze the root and derivative relationships between alarms reported by each network element before identifying the faults causing network connection root alarms (including root alarms generated by non-atomic functions) and finally diagnosing and processing the faults. This significantly shortens the identification and diagnosis time for faults causing root alarms in non-atomic functions. Furthermore, the rapid fault handling mechanism provided by this method does not affect the traditional fault identification and diagnosis process and works in conjunction with it, improving fault handling efficiency. Attached Figure Description

[0035] Figure 1 This is a flowchart illustrating the fault handling method of the first embodiment of this application;

[0036] Figure 2 This is a flowchart illustrating the fault handling method according to the second embodiment of this application;

[0037] Figure 3 This is a flowchart illustrating the fault handling method according to the third embodiment of this application;

[0038] Figure 4 This is a schematic diagram of the fault handling device according to an embodiment of this application. Detailed Implementation

[0039] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0040] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0041] Firstly, please refer to Figure 1 , Figure 1 This is a flowchart illustrating the fault handling method according to the first embodiment of this application. The fault handling method provided in this embodiment includes the following steps:

[0042] Step S1: The network element performs alarm correlation analysis to obtain the root alarm within the network element and reports it to the control system.

[0043] Step S2: When a root alarm within a network element is generated by a non-atomic function, the network element analysis identifies the fault that caused the root alarm and initiates the network element fault diagnosis process. At the same time, the fault is reported to the control system.

[0044] Step S3: The control system initiates a fault identification process based on network connectivity.

[0045] Step S4: The control system determines the completion status of the network element fault diagnosis process and the network connection-based fault identification process. If the network element fault diagnosis process is completed first, proceed to step S5. Otherwise, proceed to step S8.

[0046] Step S5: The control system determines whether the fault that caused the non-atomic function to generate the root alarm has been successfully cleared based on the received network element fault diagnosis results and the corresponding alarm clearing message. If yes, proceed to step S6; otherwise, proceed to step S7.

[0047] Step S6: The corresponding fault in the control system is entered into the historical fault list, and the process ends.

[0048] Step S7: The control system updates the information of the fault that caused the non-atomic function to generate a root alarm based on the received network element fault diagnosis results, and performs manual diagnosis on the fault until the fault is cleared, after which the process ends.

[0049] Step S8: The control system judges each fault that causes a root alarm in the network connection, and determines whether it is a fault that causes a root alarm in a non-atomic function. If yes, proceed to step S9; otherwise, proceed to step S10.

[0050] Step S9: Update the information on the faults that cause non-atomic functions to generate root alarms based on the information on the faults that cause network connection to generate root alarms, and then proceed to step S5.

[0051] Step S10: The control system handles the fault that causes the network connection to generate a root alarm in the traditional way, and the process ends.

[0052] It should be noted that, in the above steps, atomic functions refer to functions within the device related to signal transmission, including adaptation, termination, and connection. Non-atomic functions refer to objects within the device other than atomic functions, such as devices and circuit boards.

[0053] This method reports root alarms within network elements to the management and control system, then directly analyzes the faults causing root alarms in non-atomic functions within the network element and initiates the network element fault diagnosis process, while simultaneously reporting the faults to the management and control system. Compared to traditional fault identification and diagnosis processes, this method eliminates the need for the management and control system to analyze the root and derivative relationships between alarms reported by each network element before identifying the faults causing root alarms in network connections (including root alarms generated by non-atomic functions) and finally diagnosing and processing the faults. This significantly shortens the identification and diagnosis time for faults causing root alarms in non-atomic functions. Furthermore, the rapid fault handling mechanism provided by this method does not affect the traditional fault identification and diagnosis process and works in conjunction with it, improving fault handling efficiency.

[0054] In some embodiments, in step S1 above, the network element performs alarm correlation analysis to obtain the root alarm within the network element, including the following steps:

[0055] Correlation analysis is performed on alarms within network elements to obtain root alarms and corresponding derived alarms within the network element. Isolated alarms are also identified as root alarms within the network element.

[0056] In some embodiments, reporting root alarms within network elements to the management system in step S1 above includes the following steps:

[0057] The basic information, alarm source classification, and diagnostic status of root alarms within network elements are reported to the control system.

[0058] Alarm sources are classified into atomic function generation and non-atomic function generation.

[0059] Diagnostic status includes not enabled, diagnostic in progress, and completed.

[0060] It should be noted that the diagnostic status is valid when the alarm source of the root alarm is classified as non-atomic function generation.

[0061] In some embodiments, in steps S2-S9 above, the fault information includes basic fault information, fault handling suggestions, fault-related alarm information, fault diagnosis status, and fault diagnosis results.

[0062] The basic fault information mentioned above includes fault serial number, fault name, fault status, fault occurrence time, and fault recovery time.

[0063] The alarm information associated with the above-mentioned faults includes root alarm information and its derived alarm information.

[0064] The above fault diagnosis status includes not started, diagnosis in progress, and completed.

[0065] It should be noted that when the fault diagnosis status is not enabled or in progress, the fault diagnosis result will be empty.

[0066] In some embodiments, in step S6 above, the corresponding fault in the control system is entered into historical faults, including the following steps:

[0067] The fault diagnosis status in the information of the fault that causes the non-atomic function to generate the root alarm is updated to "completed". The network element fault diagnosis result is added to the network connection-based fault diagnosis result. All alarm statuses associated with the fault are refreshed (cleared). The alarm recovery time is refreshed. After filling in the fault recovery time, the fault is transferred from the current fault to the historical fault.

[0068] In some embodiments, in step S7 above, the control system updates the information of the fault that caused the non-atomic function to generate a root alarm based on the received network element fault diagnosis results, including the following steps:

[0069] Update the fault diagnosis status to "completed" and supplement the network element fault diagnosis results into the network connection-based fault diagnosis results.

[0070] In some embodiments, in step S9 above, updating the information about faults that cause non-atomic functions to generate root alarms based on the information about faults that cause network connectivity to generate root alarms includes the following steps:

[0071] The information associated with the root alarm and derived alarms in the information of the fault that caused the non-atomic function to generate a root alarm will be updated to the information associated with the root alarm and derived alarms in the information of the fault that caused the network connection to generate a root alarm.

[0072] In a more specific embodiment, please refer to Figure 2 , Figure 2 This is a flowchart illustrating the fault handling method of the second embodiment of this application. Assuming that the network element fault diagnosis process in step S4 is completed before the network connection-based fault identification process, the specific implementation steps of this embodiment are as follows:

[0073] Step A01: Each network element monitors the alarms generated by its own network element.

[0074] Step A02: For any network element, the network element performs correlation analysis on the alarms generated within the network element to obtain the root alarm within the network element, and determines whether the root alarm is an alarm generated by a non-atomic function.

[0075] Specifically, in this embodiment, there is a networking relationship between a first network element and a second network element. The first network element generates "overheating" alarms and "low transmit optical power" alarms, while the second network element generates a "low receive optical power" alarm. The first and second network elements respectively perform correlation analysis on the alarms generated within their respective network elements. The analysis by the first network element concludes that the "overheating" alarm is the root alarm of the "low transmit optical power" alarm, that is, the "overheating" alarm is the root alarm of the first network element. The "low receive optical power" alarm in the second network element is an isolated alarm, and the "low receive optical power" alarm is taken as the root alarm of the second network element.

[0076] Subsequently, it was determined that the "overheating" root alarm corresponding to the first network element was an alarm generated by a non-atomic function.

[0077] Step A03: Each network element reports the root alarm within the network element to the control system.

[0078] Specifically, the first network element reports an "overheating" alarm to the control system, and the reported information includes, but is not limited to:

[0079] Alarm name: Overheat alarm.

[0080] Alarm clearance status: Not cleared.

[0081] Alarm location information: First network element.

[0082] Alarm occurrence time: yy-mm-dd hh:mm:ss.

[0083] Alarm clearing time: None.

[0084] Alarm source classification: non-atomic function.

[0085] Diagnostic status: Under diagnosis.

[0086] The second network element reports a "low received optical power" alarm to the control system, and the reported information includes, but is not limited to:

[0087] Alarm name: Low received optical power alarm.

[0088] Alarm clearance status: Not cleared.

[0089] Alarm location information: Second network element - x single disk - y port.

[0090] Alarm occurrence time: yy-mm-dd hh:mm:ss.

[0091] Alarm clearing time: None.

[0092] Alarm source classification: Atomic function.

[0093] Step A04: When the root alarm in a network element is an alarm generated by a non-atomic function, each network element analyzes the fault that caused the root alarm and starts the network element fault diagnosis process, while reporting the fault to the control system.

[0094] Specifically, the first network element analysis determines that the fault causing the "overheating" alarm is a "hardware fault". The network element initiates the diagnostic process for the "hardware fault" and reports the "hardware fault" to the control system.

[0095] Information regarding "hardware failure" includes, but is not limited to:

[0096] Fault serial number: 1222.

[0097] Fault name: Hardware failure.

[0098] Fault status: Not recovered.

[0099] Fault occurrence time: yy-mm-dd hh:mm:ss.

[0100] Fault recovery time: None.

[0101] Fault-related alarms: "Overheating" alarm and its corresponding information, "Low Transmit Optical Power" derivative alarm of the first network element corresponding to "Overheating" alarm and its corresponding information.

[0102] Fault diagnosis status: Diagnosing.

[0103] Fault diagnosis result: empty.

[0104] Step A05: The control system initiates a network connection-based fault identification process to analyze and identify the faults that cause the network connection to generate root alarms.

[0105] Specifically, the control system analyzes the root and derivative relationships between the "overheating" alarm reported by the first network element and the "low received optical power" alarm reported by the second network element, and concludes that the "overheating" alarm of the first network element is the root alarm of the "low received optical power" alarm of the second network element. That is, the root alarm based on network connection is the "overheating" alarm, and the analysis shows that the fault that caused the root alarm is "hardware fault".

[0106] Step A06: When the above network element fault diagnosis process is completed before the network connection-based fault identification process, the first network element reports the network element fault diagnosis result to the control system after the network element fault diagnosis process is completed.

[0107] Specifically, in this embodiment, the diagnosis result of the "hardware fault" of the first network element is a fan fault, which requires the replacement of the fan plate.

[0108] Step A07: Based on the received network element fault diagnosis results and the status of associated alarms, the control system determines whether the fault of the first network element and the associated alarms have been successfully cleared. If yes, proceed to step A08; otherwise, proceed to step A09.

[0109] Step A08: The corresponding fault in the control system is entered into the historical fault list.

[0110] Specifically, update the fault diagnosis status (completed), fault diagnosis result, fault recovery time, and the clearing time of each alarm in the fault-related alarms in the fault information corresponding to the "hardware fault", and transfer the fault from the current fault to the historical fault.

[0111] Step A09: The control system updates the fault information of the fault based on the received network element fault diagnosis results, and initiates manual diagnosis of the fault in the control system.

[0112] Specifically, the fault diagnosis status and fault diagnosis results in the fault information corresponding to "hardware failure" are updated. At the same time, the control system initiates a manual diagnosis operation for the fault. In this embodiment, this means that the fan plate is replaced manually at the station.

[0113] In a more specific embodiment, please refer to Figure 3 , Figure 3 This is a flowchart illustrating the fault handling method according to the third embodiment of this application. Assuming that the network element fault diagnosis process in step S4 is completed after the network connection-based fault identification process, the specific implementation steps of this embodiment are as follows:

[0114] Steps B01-B05 are the same as steps A01-A05 in the second embodiment above, so they will not be described again here.

[0115] Step B06: When the fault identification process based on network connection is completed before the fault diagnosis process of network element, the control system judges each fault that causes the network connection to generate a root alarm and determines whether it is a fault that causes a non-atomic function to generate a root alarm. If yes, proceed to step B07; otherwise, proceed to step B11.

[0116] Step B07: Update the information on faults that cause root alarms in non-atomic functions based on the information on faults that cause root alarms in network connectivity.

[0117] Specifically, a derivative alarm message of "received optical power too low" for the second network element is added to the alarms associated with "hardware failure".

[0118] Step B08: After the network element fault diagnosis process is completed, determine whether all faults that caused root alarms in non-atomic functions have been successfully cleared. If yes, proceed to step B09; otherwise, proceed to step B10.

[0119] Step B09: The corresponding fault in the control system is entered into the historical fault list.

[0120] Specifically, update the fault diagnosis status (completed), fault diagnosis result, fault recovery time, and clearing time of each alarm in the fault-related alarm information corresponding to the "hardware fault".

[0121] Step B10: The management system updates the fault diagnosis status (completed) and fault diagnosis results in the fault information that caused the non-atomic function to generate a root alarm, and initiates manual diagnosis for the fault in the management system.

[0122] Step B11: The control system handles faults that cause root alarms in network connectivity in a traditional manner.

[0123] Secondly, please refer to Figure 4 , Figure 4 This is a schematic diagram of the fault handling apparatus according to an embodiment of this application. The fault handling apparatus provided in this embodiment includes the following modules:

[0124] The network element layer analysis module is used to perform correlation analysis on alarms within the network element, obtain the root alarm within the network element, and report it to the control system; it is also used to analyze the fault that caused the root alarm when the root alarm within the network element is generated by a non-atomic function and execute the network element fault diagnosis process.

[0125] The network layer analysis module is used to perform fault identification processes based on network connections.

[0126] The update module is used to mark the fault that caused the root alarm of the non-atomic function as a historical fault when the fault that caused the root alarm of the non-atomic function is successfully cleared; it is also used to update the information of the fault that caused the root alarm of the non-atomic function based on the information of the fault that caused the root alarm of the network connection when the identified fault that caused the root alarm of the network connection is not the fault that caused the root alarm of the non-atomic function.

[0127] The processing module is used to clear faults that cause root alarms in non-atomic functions, based on the order of completion of the network element fault diagnosis process and the network connection-based fault identification process.

[0128] It should be noted that the sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0129] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus. The terms "first," "second," and "third," etc., are used to distinguish different objects, etc., and do not indicate a sequence, nor do they limit "first," "second," and "third" to different types.

[0130] In the description of the embodiments of this application, terms such as "exemplary," "for example," or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary," "for example," or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a concrete manner.

[0131] In the description of the embodiments of this application, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.

[0132] In some processes described in the embodiments of this application, multiple operations or steps are included in a specific order. However, it should be understood that these operations or steps may not be executed in the order they appear in the embodiments of this application, or they may be executed in parallel. The sequence number of the operation is only used to distinguish different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed sequentially or in parallel, and these operations or steps may be combined.

[0133] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device to execute the methods described in the various embodiments of this application.

[0134] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method of fault handling, characterized by, The method comprises: The network element performs alarm correlation analysis to obtain a root alarm in the network element and report the root alarm to the management and control system, when the root alarm in the network element is caused by non-atomic function, the network element analyzes a fault causing the root alarm and starts a network element fault diagnosis process, and the fault is reported to the management and control system; The management and control system performs a network connection-based fault identification process after receiving the root alarm in the network element, and clears the fault causing the root alarm of non-atomic function according to the network element fault diagnosis state; If the management and control system receives the network element fault diagnosis result before completing the network connection-based fault identification process, the network element fault diagnosis result comprises: If the fault causing the root alarm of non-atomic function has been cleared, the corresponding fault in the management and control system enters a historical fault; If the fault causing the root alarm of non-atomic function has not been cleared, the management and control system updates information of the fault causing the root alarm of non-atomic function based on the received network element fault diagnosis result, and performs manual diagnosis on the fault until the fault is cleared; If the management and control system receives the network element fault diagnosis result after completing the network connection-based fault identification process, the network element fault diagnosis result comprises: The management and control system judges each fault causing the root alarm of network connection obtained by analysis, if the fault is the fault causing the root alarm of non-atomic function, updates information of the fault causing the root alarm of non-atomic function based on information of the fault causing the root alarm of network connection; After the management and control system receives the network element fault diagnosis result, the management and control system performs the process performed when the management and control system receives the network element fault diagnosis result before completing the network connection-based fault identification process.

2. The method of fault handling of claim 1, wherein: The network element performs alarm correlation analysis to obtain a root alarm in the network element, comprising: Correlation analysis is performed on alarms in the network element to obtain root alarms in the network element and corresponding derivative alarms, and isolated alarms are also identified as root alarms in the network element.

3. The method of fault handling of claim 1, wherein: The root alarm in the network element is reported to the management and control system, comprising: Basic information, alarm source classification and diagnosis state of the root alarm in the network element are reported to the management and control system; The alarm source classification comprises atomic function generation and non-atomic function generation; The diagnosis state comprises not started, in diagnosis and completed.

4. The method of fault handling of claim 1, wherein: The management and control system performs a network connection-based fault identification process after receiving the root alarm in the network element, comprising: The management and control system performs network connection-based alarm correlation analysis based on the root alarms reported by each network element to obtain root alarms of network connection and corresponding derivative alarms, and further analyzes faults causing the root alarms of network connection.

5. The method of fault handling as described in claim 1, wherein: The information of the fault causing the root alarm of non-atomic function is updated based on information of the fault causing the root alarm of network connection, comprising: Alarm information associated with the fault in the information of the fault causing the root alarm of non-atomic function is updated to alarm information associated with the fault in the information of the fault causing the root alarm of network connection.

6. The method of fault handling as described in claim 1, wherein: If the fault causing the root alarm of non-atomic function has not been cleared, the management and control system updates information of the fault causing the root alarm of non-atomic function based on the received network element fault diagnosis result, and the updated information comprises: Fault diagnosis state and fault diagnosis result.

7. The method of fault handling of claim 1, wherein: The corresponding fault in the management and control system enters a history fault, including: The management and control system updates the diagnosis state, diagnosis result information and state information of all alarms associated with the corresponding fault in the management and control system based on the received network element fault diagnosis result and corresponding alarm clearing message, fills in the fault recovery time information, and converts the fault from the current fault to the history fault.

8. An apparatus for failure handling based on the method of any of claims 1-7, characterized by: The device comprises: A network element layer analysis module, configured to perform correlation analysis on alarms in a network element, obtain a root alarm in the network element and report to the management and control system, and further configured to, when the root alarm in the network element is caused by a non-atomic function, analyze a fault causing the root alarm and execute a network element fault diagnosis process; A network layer analysis module, configured to execute a network connection-based fault identification process; An update module, configured to, when the fault causing the root alarm of the non-atomic function is successfully cleared, mark the fault causing the root alarm of the non-atomic function as a history fault, and further configured to, when the fault causing the root alarm of the network connection is identified and the fault causing the root alarm of the non-atomic function is included, update information of the fault causing the root alarm of the non-atomic function based on information of the fault causing the root alarm of the network connection; A processing module, configured to clear the fault causing the root alarm of the non-atomic function according to completion of the network element fault diagnosis process and the network connection-based fault identification process.

Citation Information

Patent Citations

  • Method for discriminating related alert

    CN101075902A

  • Operation maintenance device, network element equipment and method thereof for processing reported alarms

    CN101656976A

  • Data correlation analysis method and system

    CN115037592A