Network equipment monitoring device, network equipment monitoring method, and program
The network equipment monitoring device addresses the issue of repeated alarms in optical transmission systems by filtering and analyzing alarm patterns, thereby reducing the time to identify causes of abnormalities and alleviating the burden on maintainers.
Patent Information
- Application Number
- JP2024520111
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-05-10
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2042-05-10
AI Technical Summary
In optical transmission systems, repeated fluctuations in optical signal intensity lead to frequent abnormal detection and recovery detection alarms, overwhelming maintainers with a high volume of alerts and increasing the time required to identify the cause of abnormalities.
A network equipment monitoring device that includes an alarm database, an alarm pruning unit, a repeated alarm extraction unit, a recovery alarm pattern database, and an alarm pattern collation unit. This device periodically prunes alarms, extracts repeated alarm groups, collates them with past cases, and notifies the cause of similar past alarms when similarity exceeds a threshold.
The solution significantly reduces the time needed to identify the cause of abnormalities by filtering out repetitive alarms and providing immediate notification of the likely cause based on past patterns, thus streamlining maintenance processes.
Smart Images

Figure 0007694822000004 
Figure 0007694822000005 
Figure 0007694822000006
Abstract
Description
Technical Field
[0001] The present invention relates to network equipment monitoring technology for optical transmission systems, and particularly relates to a network equipment monitoring device, a network equipment monitoring method, and a program.
Background Art
[0002] Conventionally, technologies for realizing system monitoring, alert notification, performance visualization, etc. are known (see Non-Patent Document 1 and Non-Patent Document 2). The monitoring device described in Non-Patent Document 1 collects data from monitoring targets such as network equipment, determines threshold values, executes alert notification actions, and stores various data in a database.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in an optical transmission system that processes optical signals in one go, fluctuations occur in signals with analog characteristics such as optical signal intensity. Therefore, in the case of a failure mode that fluctuates near the threshold due to this fluctuation, it is conceivable that abnormal detection and abnormal recovery detection will occur repeatedly. In this state, the network equipment of the optical transmission system repeatedly outputs an abnormal detection alarm notifying of the occurrence of an abnormality and a recovery detection alarm notifying of the recovery of the abnormality. Therefore, the monitoring device continuously notifies the maintainer of a large number of alarms per unit time. This may increase the amount of work for the maintainer to confirm the alarms, and ultimately increase the time required to identify the cause of the abnormality in the network equipment.
[0005] Therefore, an object of the present invention is to solve the above problems and reduce the time required to identify the cause of an abnormality when a large number of repeated alarms are output from network equipment.
Means for Solving the Problems
[0006] The network equipment monitoring device according to the present invention includes an alarm database that accumulates, as alarm information, an abnormal detection alarm notifying of the occurrence of an abnormality and a recovery detection alarm notifying of the recovery of an abnormality output from the network equipment of the optical transmission system; an alarm pruning unit that periodically prunes the alarm information accumulated in the alarm database and transmits it to a predetermined notification destination; a repeated alarm extraction unit that extracts, from the alarm database, an alarm group in which an abnormal detection alarm and a recovery detection alarm are repeatedly generated a predetermined number of times within a predetermined time for the same alarm content of the same network equipment as a repeated alarm; a generation recovery alarm pattern database that accumulates information on repeated alarms of past cases together with the causes of the repeated alarms; and an alarm pattern collation unit that collates the information on the extracted repeated alarms with the information on the repeated alarms of the past cases, and when it is determined that the similarity is higher than a predetermined threshold, notifies the cause of the repeated alarm of the past case to the predetermined notification destination.
Effects of the Invention
[0007] According to the present invention, it is possible to reduce the time required to identify the cause of an abnormality when a large number of repeated alarms are output from network equipment.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 3D
Figure 3E
Figure 3F
Figure 4A
Figure 4B
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 10
Embodiments for Carrying Out the Invention
[0009] Hereinafter, the network equipment monitoring device according to the present embodiment will be described in detail with reference to the drawings. [Outline of System Configuration] As shown in FIG. 1, the optical transmission system 1 includes a network equipment monitoring device 10, network equipment 20, and a host device 30. The network equipment monitoring device 10 is configured by, for example, NE-OpS (Network element operation system), notifies the host device 30 of the alarms acquired from the network equipment 20, and notifies the host device 30 of the causes of abnormalities in past cases as necessary. The host device 30 is a device that monitors the network equipment monitoring device 10 and presents alarms and the like to the maintainer 2.
[0010] The network equipment 20 includes, for example, a transponder (TPND) 21, a wavelength selective switch (WSS) 22, and an optical amplifier (AMP) 23. The number of network equipment 20 is arbitrary. When distinguishing between the two network equipment shown in FIG. 1, they are denoted as NE1 and NE2, and when not distinguishing, they are denoted as network equipment 20.
[0011] For example, when an electrical signal is input from an external communication device to the transponder 21 of the network facility NE1, this electrical signal is converted into an optical signal by the transponder 21, multiplexed by the wavelength selective switch 22, amplified by the optical amplification unit 23, and then transmitted externally. This optical signal is amplified, for example, by the optical amplification unit 23 of the network facility NE2, then demultiplexed by the wavelength selective switch 22, received by the transponder 21, and transmitted to a communication device (not shown). Note that this optical transmission system 1 is for two-way communication.
[0012] [Network Facility Monitoring Device 10] The network facility monitoring device 10 includes an alarm database 11, an alarm pruning unit 12, a GUI (Graphical User Interface) unit 13, a repeated alarm extraction unit 14, an occurrence / recovery alarm pattern database 15, and an alarm pattern matching unit 16. These units may be housed in the same housing called the network facility monitoring device, or may be housed in different housings. In the following description and drawings, the database is denoted as DB.
[0013] The alarm DB 11 accumulates, as alarm information, an abnormality detection alarm notifying the occurrence of an abnormality output from the network facility 20 of the optical transmission system 1 and a recovery detection alarm notifying the recovery of the abnormality. The alarm DB 11 accumulates the raw data of the alarms. The alarm pruning unit 12 periodically prunes the alarm information accumulated in the alarm DB 11 and transmits it to a predetermined notification destination such as the upper device 30. The GUI unit 13 presents predetermined information such as alarms to the user. The GUI unit 13 receives, for example, a user's instruction from an input device such as a mouse and outputs information corresponding to the user's instruction to an output device such as a display. Note that the alarm DB 11, the alarm pruning unit 12, and the GUI unit 13 have conventionally known configurations, and thus further description thereof is omitted.
[0014] The repeated alarm extraction unit 14 extracts, from the alarm DB 11, a group of alarms in which abnormal detection alarms and recovery detection alarms have occurred repeatedly a predetermined number of times within a predetermined time for the same alarm content of the same network facility as repeated alarms. The occurrence and recovery alarm pattern DB 15 accumulates information on repeated alarms of past cases together with the causes of the repeated alarms. The alarm pattern matching unit 16 matches the information on the repeated alarms extracted by the repeated alarm extraction unit 14 with the information on the repeated alarms of past cases, and when it is determined that the similarity is higher than a predetermined threshold, notifies the cause of the repeated alarm of the past case to a predetermined notification destination. Note that the cause of an alarm refers not to the content notified by the alarm but to the real reason for the occurrence of the alarm, that is, the cause that has caused an abnormality in the network facility. Hereinafter, the occurrence and recovery alarm pattern DB 15, the repeated alarm extraction unit 14, and the alarm pattern matching unit 16 will be described in detail.
[0015] [Details of the Occurrence and Recovery Alarm Pattern DB 15] The information on the repeated alarms of past cases includes basic information, the cause of the abnormality for the repeated alarms of past cases, and statistical information. The basic information includes alarm content information, identification information of the network facility where the alarm has occurred, and information for specifying the location where the alarm has occurred. These pieces of information are each recorded in association with the time information of the abnormal detection alarm or the recovery detection alarm.
[0016] (Basic Information) FIG. 2 is a diagram showing an example of past occurrence / recovery repeated alarms that are basic information registered in the occurrence / recovery alarm pattern DB 15. The basic information 210 shown in FIG. 2 includes information regarding items 211 to 215. Item 211 indicates the date and time when the alarm occurred. The date and time is represented by, for example, year / month / day and hour / minute / second. Item 212 indicates the type of the alarm. In this example, "occurrence" represents an abnormality detection alarm, and "recovery" represents a recovery detection alarm. In item 211, the time t0 indicates the date and time when the first abnormality detection alarm in this repeated alarm occurred. Also, the time t1 indicates the date and time when the first recovery detection alarm in this repeated alarm occurred. Item 213 indicates the alarm content. LOS represents Loss of Signal. Item 214 indicates the name of the device that output the alarm. OXC represents optical cross-connect. Item 215 indicates the location where the alarm occurred. 10G path CH03 represents 10 Giga bits per second path channel 03. Note that the values of each item of the basic information 210 are examples.
[0017] (Cause of abnormality) The causes of abnormalities in the repeated alarms of past cases are, for example, a failure of the wavelength selective switch (WSS) 22, a failure of the optical amplifier section (AMP) 23, etc. The repeated patterns of alarms notifying abnormalities and alarms notifying recovery vary depending on the location of the device failure, and it is difficult to identify the cause of the abnormality. Even if the alarm content indicates "LOS", it is difficult to identify whether the cause of the abnormality is a failure of the wavelength selective switch (WSS) or a failure of the optical amplifier section (AMP).
[0018] For example, when there is fluctuation in the output of the wavelength selective switch (WSS), in a predetermined network facility 20, alarms of LOS, recovery, LOS, recovery,... are repeated. Also, when there is fluctuation in the output of the optical amplification section (AMP), in a predetermined network facility 20, an event occurs where alarms such as LOS, recovery, LOS, recovery,... are repeated, and alarms such as LOS, recovery, LOS, recovery,... are also repeated in another network facility 20. That is, repeated alarms may spread to other devices. The spread is, for example, an event where repeated alarms regarding abnormalities occur from multiple locations for one cause such as a failure of the optical amplification section.
[0019] (Statistical information) The statistical information registered in the occurrence recovery alarm pattern DB 15 is information calculated using time information as characteristic information of the repetition pattern for repeated alarms of past cases. Here, the repetition pattern will be described with reference to FIG. 3A. FIG. 3A is a graph showing the time change of repeated alarms. The horizontal axis is the time axis and shows the time information of the repeated alarms. The vertical axis shows the binarized alarm type. The abnormality detection alarm is represented by 1, and the recovery detection alarm is represented by 0. FIGS. 3B to 3F are diagrams showing other repetition patterns. Details of these repetition patterns will be described later.
[0020] Hereinafter, the period from the occurrence time of the abnormality detection alarm to the occurrence time of the recovery detection alarm will be referred to as the abnormality continuation period. The period from the occurrence time of the recovery detection alarm to the occurrence time of the next abnormality detection alarm will be referred to as the recovery continuation period. The period obtained by combining the abnormality continuation period and the subsequent recovery continuation period will be referred to as the cycle. Specifically, in FIG. 3A, the period T1 from time t0 to time t1 is the first abnormality continuation period. The period T2 from time t1 to time t2 is the first recovery continuation period. The period T3 from time t2 to time t3 is the second abnormality continuation period. The period T4 from time t3 to time t4 is the second recovery continuation period. The period T5 from time t4 to time t5 is the third abnormality continuation period. The period T6 from time t5 to time t6 is the third recovery continuation period. The period T7 from time t6 to time t7 is the fourth abnormality continuation period. Also, the sum of period T1 and period T2 is the first cycle. The sum of period T3 and period T4 is the second cycle. The sum of period T4 and period T5 is the third cycle. Note that the cycle is not necessarily constant.
[0021] (Feature information) The feature information of the repeating pattern is, for example, the average value of the alarm occurrence intervals, the standard deviation of the alarm occurrence intervals, the ratio of the abnormal duration in the repeating alarms, etc. The average value of the alarm occurrence intervals is the average value of the intervals between the output times of the anomaly detection alarms and the output times of the recovery detection alarms. In the case of the repeating pattern up to time t7 shown in FIG. 3A, the average value of the alarm occurrence intervals is the average value of the abnormal duration periods T1, T3, T5, T7 and the recovery duration periods T2, T4, T6. When using the time information from time t0 to time tn, the average value of the alarm occurrence intervals is calculated by the following formula (1).
[0022] [Number]
[0023] The standard deviation of the alarm occurrence intervals indicates the variation of the alarm occurrence intervals. Using the average value of the alarm occurrence intervals shown in formula (1), the standard deviation σ of the alarm occurrence intervals is represented by the following formula (2).
[0024] [Number]
[0025] The ratio of the abnormal duration in the repeating alarms is the abnormal duration with respect to the cycle of the repeating alarms. More specifically, the ratio r of the abnormal duration in the repeating alarms is calculated by the following formula (3).
[0026] [Number]
[0027] In Equation (3), E1 is the average value of the abnormal duration. E2 is the average value of the recovery duration. In the case of the repeating pattern up to time t7 shown in FIG. 3A, E1 = (T1 + T3 + T5 + T7) / 4, and E2 = (T2 + T4 + T6) / 3.
[0028] (Range of similarity coefficient) The information registered in the occurrence / recovery alarm pattern DB 15 may further include a range of similarity coefficients preset corresponding to the allowable range from the similarity reference value (statistical information). The range of similarity coefficients is a range including the case where the statistical information of the repeating alarm to be compared and the statistical information of the past case are the same (100%). The range of similarity coefficients is defined by a lower limit value and an upper limit value. The calculation result obtained by multiplying the value of the statistical information of the past case by the lower limit value of the similarity coefficient becomes a threshold value (lower limit value) for determining high similarity. The calculation result obtained by multiplying the value of the statistical information of the past case by the upper limit value of the similarity coefficient becomes a threshold value (upper limit value) for determining high similarity. Note that the alarm pattern collation unit 16 calculates statistical information as feature information of the repeating pattern using the time information of the extracted repeating alarm, and collates it with the statistical information of the repeating alarm of the past case to determine similarity.
[0029] (Propagating event) The information registered in the occurrence / recovery alarm pattern DB 15 may further include a propagating event. That is, the information of the repeating alarm of the past case may further include information indicating whether there was propagation of the repeating alarm from the network facility 20 where the repeating alarm of the past case occurred to another network facility 20. When there is propagation of the repeating alarm, it includes the identification information of the other network facility 20 where the propagation of the repeating alarm occurred, and the feature information of the repeating pattern of the repeating alarm that occurred as the propagating event.
[0030] [Details of the repeating alarm extraction unit 14] The repeated alarm extraction unit 14 periodically searches for repeated alarms from the alarm DB 11. Here, as extraction conditions for repeated alarms, the repetition time and the number of repetitions of the alarm are defined in advance. It is preferable that the repetition time of the alarm is short, for example, it may be 60 seconds. The same alarm content output from the same network facility may be, for example, "LOS in 10G path CH03". The number of times of any alarm may be, for example, 5 times, 7 times, or 10 times. When the repeated alarm extraction unit 14 finds an alarm group that satisfies the extraction conditions by searching, it extracts an alarm group recorded more than a predetermined number of times (for example, 5 times) as one repeated alarm and passes it to the alarm pattern matching unit 16.
[0031] Figure 4A is a diagram showing the data of the alarm group recorded in the alarm DB 11. The alarm group data 310 shown in Figure 4A includes information regarding items 311 to 315. Items 311 to 315 are the same as items 211 to 215 of the basic information 210 for the repeated alarms of past cases registered in advance. Note that the values of each item of the alarm group data 310 are examples.
[0032] Here, it is assumed that the alarm group data 310 is such that abnormal detection alarms and recovery detection alarms have occurred repeatedly a predetermined number of times (for example, 10 times) within a predetermined time (for example, 60 seconds) for the same alarm content (LOS) of the same device (OXC-01). The repeated alarm extraction unit 14 summarizes the alarm group data 310 into the information 320 of one repeated alarm as shown in Figure 4B, and passes this information 320 of the repeated alarm to the alarm pattern matching unit 16.
[0033] As shown in FIG. 4B, the repeated alarm information 320 includes information on items 321 to 326. Item 321 indicates the date and time when the repeated alarm occurred. This date and time is represented, for example, as the date and time when the first abnormality detection alarm of the repeated alarm occurred. Item 322 indicates the type of alarm. In this example, "repeated occurrence and recovery" indicates the repetition of an abnormality detection alarm and a recovery detection alarm. Item 323 indicates the alarm content. Item 324 indicates the name of the device that output the alarm. Item 325 indicates the location where the alarm occurred. Item 326 indicates the time information of each alarm in the repeated alarm. This time information is information on the date and time when each of the alarms occurred a predetermined number of times (for example, five times). Note that the values of each item of the repeated alarm information 320 are merely examples.
[0034] When the repeat alarm extraction unit 14 finds an alarm group that satisfies the extraction condition by the search, the repeat alarm extraction unit 14 may record the information of the extracted repeat alarm in an identifiable manner. In this case, when a repeat alarm related to the same alarm content output from the same network equipment is extracted every time by the periodic search, the repeat alarm extraction unit 14 determines that the abnormality of the network equipment continues. On the other hand, when the repeat alarm that was extracted every time cannot be extracted by a search performed within a predetermined period, the repeat alarm extraction unit 14 can determine that the network equipment has been restored. As described later, when the repeat alarm extraction unit 14 determines that the network equipment that outputted the repeat alarm has been restored after the repeat alarm harvesting has been stopped, the repeat alarm extraction unit 14 can also instruct the alarm harvesting unit 12 to cancel the stop state of the repeat alarm harvesting.
[0035] [Details of the alarm pattern matching unit 16] When the alarm pattern matching unit 16 acquires the repetitive alarm information 320 shown in FIG. 4B from the repetitive alarm extraction unit 14, it starts the similarity determination process. The alarm pattern matching unit 16 narrows down candidates to be matched with the repetitive alarm information 320 from among the past cases previously stored in the occurrence / recovery alarm pattern DB 15. The alarm pattern matching unit 16 extracts, as record data of candidates to be matched, past cases in which the alarm content is described, based on the information of the alarm content 323 described in the repetitive alarm information 320.
[0036] FIG. 5 is a diagram showing record data of candidates narrowed down from the occurrence / recovery alarm pattern DB 15. The table 250 shown in FIG. 5 has columns of serial number, [1], [2], [3], and [4]. Specific data is recorded for each of the candidates of entry No. 1, entry No. 2, and entry No. 3. Item [1] includes basic information and statistical information about the repetitive alarm of the past case. The basic information about the repetitive alarm includes information on the alarm content and information on the occurrence location.
[0037] The statistical information in item [1] has items [1-1], [1-2], and [1-3]. Item [1-1] indicates the average value of the alarm occurrence interval. The average value of the alarm occurrence interval is calculated by the above-described formula (1). Item [1-2] indicates the standard deviation of the alarm occurrence interval. The standard deviation of the alarm occurrence interval is calculated by the above-described formula (2). Item [1-3] indicates the ratio of the abnormal duration in the repetitive alarm. The ratio of the abnormal duration in the repetitive alarm is calculated by the above-described formula (3).
[0038] Item [2] indicates the range of the similarity coefficient used for the similarity determination. The range of the similarity coefficient for the candidate of entry No. 1 is "0.9 to 1.1". The ranges of the similarity coefficients for the candidates of entry No. 2 and No. 3 are "0.8 to 1.2", respectively.
[0039] Item [3] indicates whether there has been repeated spread of the alarm to other devices as a cascading event. For the candidate of Entry No. 1, there has been no repeated spread of the alarm to other devices. For the candidate of Entry No. 2, there has actually been repeated spread of the alarm to the device of Entry No. 3. For the candidate of Entry No. 3, there has actually been repeated spread of the alarm to the device of Entry No. 2.
[0040] Item [4] indicates the cause of the abnormality regarding the repeated alarms of past cases. The cause of the abnormality for the candidate of Entry No. 1 is a failure of the wavelength selection switch (WSS). The causes of the abnormalities for the candidates of Entry No. 2 and No. 3 are failures of the optical amplification section (AMP), respectively.
[0041] Next, with reference to FIGS. 3B to 3F, the relationship between the statistical information used by the alarm pattern matching unit 16 to determine similarity and the repetition pattern of the repeated alarm will be described. The alarm pattern matching unit 16 performs pattern matching of the repeated alarm using the criteria from the following three viewpoints as an example for the repetition pattern.
[0042] [First viewpoint] The repeated alarm having the repetition pattern shown in FIG. 3B is not similar to the repeated alarm having the repetition pattern shown in FIG. 3C. The period (the first cycle) from time t0 to time t2 shown in FIG. 3B is shorter than the period (the first cycle) from time t0 to time t2 shown in FIG. 3C. That is, it can be seen that the repeated alarm having the repetition pattern shown in FIG. 3B has a shorter period and lower similarity compared to the repeated alarm having the repetition pattern shown in FIG. 3C. Therefore, in the first viewpoint, the alarm pattern matching unit 16 determines similarity based on the average value of the alarm occurrence intervals shown in FIG. 5 for item [1-1].
[0043] [Second viewpoint] The repetitive alarm having the repetitive pattern shown in FIG. 3D is not similar to the repetitive alarm having the repetitive pattern shown in FIG. 3E. The period from time t0 to time t2 (the first cycle) shown in FIG. 3D is not much different from the period from time t2 to time t4 (the second cycle). The period from time t0 to time t2 (the first cycle) shown in FIG. 3E is longer than the period from time t2 to time t4 (the second cycle). That is, it can be seen that the repetitive alarm having the repetitive pattern shown in FIG. 3E has a greater variation in period and lower similarity compared to the repetitive alarm having the repetitive pattern shown in FIG. 3D. Therefore, from the second perspective, the alarm pattern matching unit 16 determines similarity based on the standard deviation of the alarm occurrence intervals (variation in time intervals), item [1-2] shown in FIG. 5.
[0044] [Third perspective] The repetitive alarm having the repetitive pattern shown in FIG. 3F is characterized in that the abnormal continuation period is longer than the recovery continuation period. Focusing on the period from time t0 to time t2 (the first cycle) shown in FIG. 3F, the abnormal continuation period from time t0 to time t1 is approximately twice as long as the recovery continuation period from time t1 to time t2. Also, this repetitive alarm has a similar tendency in the second cycle and the third cycle as well. That is, it can be seen that the repetitive alarm having the repetitive pattern shown in FIG. 3F has a large ratio of the abnormal continuation period to the period. Therefore, from the third perspective, the alarm pattern matching unit 16 determines similarity based on the ratio of the abnormal continuation time in the repetitive alarm, item [1-3] shown in FIG. 5.
[0045] In the similarity determination process, the alarm pattern matching unit 16 can determine similarity with less computational effort compared to the prior art, without imposing a processing load on the CPU, and can quickly detect whether it matches a typical failure case. Note that it is also possible to utilize AI technologies such as machine learning for similarity determination. The alarm pattern matching unit 16 has functions such as executing a notification process to the upper device 30, etc. along with the similarity determination process. These functions will be described later in conjunction with the processing operation of the alarm pattern matching unit 16.
[0046] [Hardware Configuration] The network facility monitoring device 10 according to the embodiment is realized by a computer 900 configured as shown in FIG. 6, for example. FIG. 6 is a hardware configuration diagram showing an example of a computer 900 that realizes the functions of the network facility monitoring device 10 according to the present embodiment. The computer 900 includes a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, a RAM (Random Access Memory) 903, an HDD (Hard Disk Drive) 904, an input / output I / F (Interface) 905, a communication I / F 906, and a media I / F 907.
[0047] The CPU 901 operates based on a program stored in the ROM 902 or the HDD 904. The ROM 902 stores a boot program executed by the CPU 901 when the computer 900 is started up, a program related to the hardware of the computer 900, and the like.
[0048] The CPU 901 controls an input device 910 such as a mouse and a keyboard, and an output device 911 such as a display and a printer via the input / output I / F 905. The CPU 901 acquires data from the input device 910 via the input / output I / F 905, and outputs the generated data to the output device 911. Note that, as the processor, a GPU (Graphics Processing Unit) or the like may be used together with the CPU 901.
[0049] The HDD 904 stores a program executed by the CPU 901 and data used by the program. The communication I / F 906 receives data from another device via the communication network 920 and outputs it to the CPU 901, and transmits the data generated by the CPU 901 to another device via the communication network 920.
[0050] The media I / F 907 reads the program or data stored in the recording medium 912 and outputs it to the CPU 901 via the RAM 903. The CPU 901 loads the program related to the target process from the recording medium 912 onto the RAM 903 via the media I / F 907 and executes the loaded program. The recording medium 912 is an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto Optical disk), a magnetic recording medium, or a semiconductor memory, etc.
[0051] For example, when the computer 900 functions as the network facility monitoring device 10 according to the above embodiment, the CPU 901 realizes the functions of the network facility monitoring device 10 by executing the program loaded onto the RAM 903. Also, the data in the RAM 903 is stored in the HDD 904. The CPU 901 reads and executes the program related to the target process from the recording medium 912. In addition, the CPU 901 may read the program related to the target process from another device via the communication network 920.
[0052] [Operation of Network Facility Monitoring Device] Next, the flow of the monitoring process by the network facility monitoring device 10 will be described with reference to FIG. 7 (refer to FIG. 1 as appropriate). In the network facility monitoring device 10, the alarm harvesting unit 12 starts and executes a harvesting process and a transmission process (step S10). Here, for example, abnormal detection alarms and recovery detection alarms are repeated for the same alarm content of the network facility NE1, and it is assumed that each alarm is stored in the alarm DB11. When the alarm harvesting unit 12 performs the alarm harvesting process, the corresponding alarm is deleted from the alarm DB11. When the periodic harvesting process is not performed, the repeated alarm extraction unit 14 determines that, for example, the alarm group from time t0 to time t6 satisfies the condition. In this case, the repeated alarm extraction unit 14 extracts a repeated alarm obtained by grouping the alarm group from time t0 to time t6 (step S20). Then, the repeated alarm extraction unit 14 passes the extracted repeated alarm to the alarm pattern matching unit 16 (step S30).
[0053] The alarm pattern matching unit 16 executes a detection process for a similar pattern by comparing the received repeated alarm with the repeated alarms of past cases recorded in the occurrence / recovery alarm pattern DB15 (step S40). Details of this process will be described later. When it is determined that the similarity is high, the alarm pattern matching unit 16 notifies the upper device 30 of the cause of the repeated alarm of the past case (step S50). Further, the alarm pattern matching unit 16 instructs the alarm harvesting unit 12 to stop the harvesting process (step S60). The content of the harvesting stop instruction is, for example, to stop the harvesting for the abnormal detection alarm and the recovery detection alarm having entries of "device name: OXC-01", "occurrence location: 10G path CH03", and alarm content "LOS". Thereby, the alarm harvesting unit 12 stops the alarm harvesting process and the transmission process (step S70). Note that when the alarm pattern matching unit 16 determines that the similarity between the received repeated alarm and the repeated alarm of the past case is low, steps S50 and S60 are not executed.
[0054] Even after the harvesting stop state is reached, the output of anomaly detection alarms and recovery detection alarms from the network facility NE1 continues for a while, but for example, the alarms stop during a predetermined period after time t31. At this time, in the alarm DB11, alarms from the network facility NE1 are not accumulated over a predetermined period. As a result, the repeated alarm extraction unit 14 detects the recovery of the network facility NE1 (step S80). Then, the repeated alarm extraction unit 14 instructs the alarm harvesting unit 12 to cancel the harvesting stop state (step S90). The content of the harvesting stop cancellation instruction is, for example, to cancel the harvesting stop for anomaly detection alarms and recovery detection alarms having entries of "Device name: OXC-01", "Occurrence location: 10G path CH03", and alarm content "LOS". As a result, the alarm harvesting unit 12 restarts and executes the alarm harvesting process and the transmission process (step S100).
[0055] Note that the alarm pattern matching unit 16 may notify the repeated alarm extraction unit 14 that it has instructed the alarm harvesting unit 12 to stop the harvesting process. This can surely prevent a malfunction in which the repeated alarm extraction unit 14 gives an instruction to cancel when the alarm harvesting unit 12 is not in the harvesting stop state.
[0056] Next, the flow of processing by the alarm pattern matching unit 16 will be described with reference to FIG. 8 (appropriately refer to FIG. 7). Note that the alarm pattern matching unit 16 does not necessarily execute all the processes described in FIG. 8 every time. First, when the alarm pattern matching unit 16 acquires the repeated alarm information 320 (FIG. 4B) from the repeated alarm extraction unit 14, it narrows down the candidates for past cases (step S41). At this time, the alarm pattern matching unit 16 creates a table 250 (FIG. 5) that stores the record data of each candidate for past cases to be matched with the repeated alarm information 320. Then, the alarm pattern matching unit 16 executes a similarity determination process (step S42) with each candidate to be matched with the repeated alarm information 320.
[0057] In this similarity determination process (step S42), as shown in FIG. 9A, the alarm pattern matching unit 16 calculates the characteristic information of the repeated alarm from the time information 326 of the repeated alarm information 320 (FIG. 4B) (step S421). Here, the alarm pattern matching unit 16 calculates each statistical information using the above-described formulas (1) to (3) as the characteristic information of the problematic repeated alarm.
[0058] In addition, the alarm pattern matching unit 16 calculates the range of the characteristic information of the past case for each candidate of the past case (step S422). Here, the alarm pattern matching unit 16 calculates the range of the characteristic information of the past case using the statistical information in item [1] and the range of the similarity coefficient in item [2] in the table 250 (FIG. 5). Specifically, in the case of entry No. 1 shown in FIG. 5, the range of the similarity coefficient is set to "0.9 to 1.1", and the average value of the alarm occurrence interval is "1.0 sec". In this case, the range of the characteristic information of the past case (the range of the threshold value) is "0.9 to 1.1 sec". At this time, if the average value of the occurrence interval calculated from the problematic repeated alarm is within the range of "0.9 to 1.1 sec", it is determined that the similarity with the past case is high.
[0059] Subsequently, the alarm pattern matching unit 16 detects a past case with a high similarity to the problematic repeated alarm (step S423). Here, the alarm pattern matching unit 16 sequentially performs a similarity determination process between the repeated alarm information 320 shown in FIG. 4B and each candidate (No, 1, No. 2, No. 3 shown in FIG. 5) to be matched. Then, when it can be determined that the similarity is high from all the above-described first to third viewpoints, the alarm pattern matching unit 16 determines that the repeated alarm extracted from the alarm DB 11 has a high similarity to the past case. If there is no candidate that can be determined to have a high similarity from all viewpoints, the alarm pattern matching unit 16 ends the process.
[0060] On the one hand, if there is a candidate that can be determined to be highly similar in all aspects, the alarm pattern matching unit 16 determines the presence or absence of spread (step S43: Fig. 8). That is, the alarm pattern matching unit 16 determines whether a spread event is registered in the past case with high similarity. For example, if the past case with high similarity is Entry No. 1 shown in Fig. 5, since no spread event is registered, the alarm pattern matching unit 16 notifies the upper device 30 of the cause of the repeated alarm "WSS failure" in the past case of Entry No. 1 (step S50).
[0061] Also, for example, if the past case with high similarity is Entry No. 2 shown in Fig. 5, since Entry No. 3 is registered as a spread case, the alarm pattern matching unit 16 executes standby mode processing (step S44: Fig. 8). A predetermined standby period (for example, 120 seconds) is set for the standby mode processing (step S44). During this standby period, as shown in Fig. 9B, the alarm pattern matching unit 16 determines whether a new repeated alarm is received from the repeated alarm extraction unit 14 (step S441). If no new repeated alarm is received (step S441: No), the alarm pattern matching unit 16 determines whether the standby period has ended (step S442). If the standby period has ended (step S442: Yes), the alarm pattern matching unit 16 ends the processing.
[0062] On the other hand, if the waiting period has not ended in step S442 (step S442: No), the alarm pattern matching unit 16 returns to step S441. Then, during the waiting period, if a new repeated alarm is received repeatedly (step S441: Yes), the alarm pattern matching unit 16 executes a similarity determination process (step S443) and ends the standby mode process. The similarity determination process in step S443 is the same as the similarity determination process in step S42 (FIG. 9A). In this step S443, the alarm pattern matching unit 16 similarly determines the similarity using all of the above-described first to third viewpoints. However, the candidate for the past case is the information of the repeated alarm accumulated as the ripple event of the past case (for example, the information of entry No. 3). Also, the problematic repeated alarm is a new repeated alarm that occurred during the waiting period.
[0063] In step S443, if it is determined that the similarity between the new repeated alarm and the repeated alarm of the past ripple case is low, the alarm pattern matching unit 16 ends the process. On the other hand, in step S443, if it is determined that the similarity between the new repeated alarm and the repeated alarm of the past ripple case is high, in step S50, the alarm pattern matching unit 16 notifies the upper device 30 of the cause of the abnormality of the past case of entry No. 2, "AMP failure".
[0064] Also, after the alarm pattern matching unit 16 notifies the upper device 30 of the cause of the abnormality of the past case, the alarm pattern matching unit 16 instructs the alarm pruning unit 12 to stop each of the process of pruning and the transmission process of the repeated alarm determined to have a high similarity with the past case (step S60). After notifying the cause of the repeated alarm, it is not necessary for the alarm pruning unit 12 to notify the upper device 30 of the repeated alarm. By stopping the process of the alarm pruning unit 12 from notifying the upper device 30 of the repeated alarm, the system load can be reduced.
[0065] Next, the advantageous points of the operation of the network facility monitoring device according to the present embodiment will be described with reference to FIGS. 10 and 7. FIG. 10 is a schematic diagram showing the operation in which a conventional network facility monitoring device notifies an abnormality detection alarm and a recovery detection alarm. As shown in FIG. 10, a conventional network facility monitoring device 110 accumulates alarm data notified from network facilities 20 in an alarm DB 11. In the network facility monitoring device 110, an alarm harvesting unit 12 periodically harvests alarms from the alarm DB 11 and transmits them to a host device 30.
[0066] As shown in FIG. 10, in a conventional network facility monitoring device 110, for example, when an abnormality detection alarm and a recovery detection alarm repeatedly occur for the same alarm content of a network facility NE1, a large number of alarms are notified from the network facility monitoring device 110 to the host device 30. As a result, since the host device 30 displays a large number of alarms, the confirmation work by the maintenance personnel is increased. Therefore, it takes time for the maintenance personnel to identify the cause of the abnormality of the network facility NE1. Also, when the abnormality detection alarm and the recovery detection alarm repeatedly occur, the load of the process in which the alarm harvesting unit 12 harvests a large number of alarms from the alarm DB 11 increases. As a result, there may be problems such as a delay in alarms from the alarm harvesting unit 12 to the host device 30 or a leakage of alarms to be notified from the alarm harvesting unit 12 to the host device 30.
[0067] On the other hand, in the network facility monitoring device 10 shown in FIG. 7, the occurrence recovery alarm pattern DB15 stores the repetition pattern of repeated alarms based on past cases and the cause of the abnormality. Also, when abnormal detection alarms and recovery detection alarms are repeatedly generated, the repeated alarm extraction unit 14 extracts the repeated alarms and passes them to the alarm pattern matching unit 16. Then, the alarm pattern matching unit 16 determines the similarity with the repetition pattern in past failure cases, and notifies the maintainer of the cause of the abnormality for the past cases determined to have a high similarity. Furthermore, after detecting a past case with a high similarity, the network facility monitoring device 10 stops the pruning process of the repeated alarms and stops notifying the upper device 30 of the alarms, thereby reducing the system load. Therefore, according to the network facility monitoring device 10, it is easy for the maintainer to identify the failure location for a failure occurring due to the main signal in the optical transmission system 1, and a large amount of alarm output can be suppressed.
[0068] [Effect] As described above, the network facility monitoring device includes an alarm database 11 that accumulates, as alarm information, an abnormal detection alarm that notifies the occurrence of an abnormality output from the network facility 20 of the optical transmission system 1 and a recovery detection alarm that notifies the recovery of the abnormality; an alarm pruning unit 12 that periodically prunes the alarm information accumulated in the alarm database 11 and transmits it to a predetermined notification destination (upper device 30); a repeated alarm extraction unit 14 that extracts, as repeated alarms, an alarm group in which abnormal detection alarms and recovery detection alarms are repeatedly generated within a predetermined time for the same alarm content of the same network facility from the alarm database 11; a generation recovery alarm pattern database 15 that stores the information of the repeated alarms of past cases together with the cause of the repeated alarms; and an alarm pattern matching unit 16 that compares the information of the extracted repeated alarms with the information of the repeated alarms of past cases, and when it is determined that the similarity is higher than a predetermined threshold value, notifies the cause of the repeated alarms of past cases to a predetermined notification destination (upper device 30).
[0069] By doing so, in the network facility monitoring device, when there are accumulated past cases with a similarity higher than a predetermined threshold for repetitive alarms that notify the occurrence and recovery of anomalies, the alarm pattern matching unit 16 can notify the cause of the repetitive alarms of the past cases. Therefore, the maintainer of the network facility can quickly identify the cause of the repetitive alarms.
[0070] In the network facility monitoring device, the information on the repetitive alarms of past cases includes basic information containing the alarm content information, the identification information of the network facility 20 where the alarm occurred, and the information for specifying the location where the alarm occurred, each recorded in association with the time information of the anomaly detection alarm or the recovery detection alarm, statistical information calculated using the time information as the characteristic information of the repetitive pattern, and the alarm pattern matching unit 16 calculates the statistical information as the characteristic information of the repetitive pattern using the time information of the extracted repetitive alarm, and determines the similarity by comparing it with the statistical information on the repetitive alarms of past cases.
[0071] By doing so, in the network facility monitoring device, the occurrence / recovery alarm pattern database 15 can accumulate detailed information on the anomaly detection alarm or the recovery detection alarm itself and the statistical information calculated using the time information. Also, since the alarm pattern matching unit 16 determines the similarity between the extracted repetitive alarm and the repetitive alarms of past cases by comparing the statistical information, the processing load of the calculation can be reduced.
[0072] In a network equipment monitoring device, the information on repeated alarms of past cases includes information indicating whether there has been a spread of repeated alarms from the network equipment 20 where the repeated alarms of past cases occurred to other network equipment 20, identification information of other network equipment 20 where the spread of repeated alarms has occurred, and characteristic information of the repetition pattern of the repeated alarms that occurred as the spreading event. When information indicating that there has been a spread is stored for a repeated alarm of a past case determined to have a similarity higher than a predetermined threshold with the extracted repeated alarm, the alarm pattern matching unit 16 waits for a predetermined period. When a new repeated alarm occurs within the predetermined period, the information of the new repeated alarm is compared with the information of the repeated alarm stored as the spreading event of the past case to determine the similarity.
[0073] By doing so, in the network equipment monitoring device, when information indicating that there has been a spread is stored for a repeated alarm of a past case determined to have a similarity higher than a predetermined threshold with the extracted repeated alarm, the alarm pattern matching unit 16 can determine whether the extracted repeated alarm spreads to other network equipment. Therefore, when repeated alarms notifying abnormalities from a plurality of network equipment regarding one cause are output, the maintainer of the network equipment can quickly identify the cause of the repeated alarms.
[0074] In a network equipment monitoring device, when it is determined that the similarity between the information of the extracted repeated alarm and the information of the repeated alarm of the past case is higher than a predetermined threshold, the alarm pattern matching unit 16 instructs the alarm pruning unit 12 to stop the pruning process and the transmission process of the repeated alarm determined to have a high similarity, respectively.
[0075] By doing so, it is possible to prevent a situation where, at a predetermined notification destination (upper device 30), the confirmation work is increased by a large number of alarm notifications, and it takes time to identify the cause of repeated alarms. Further, in the alarm culling unit 12, it is possible to prevent the occurrence of alarm delays, alarm omissions, etc. due to an increase in the load of the alarm culling process.
[0076] In the network facility monitoring device, the repeated alarm extraction unit 14 is characterized in that, after the culling of repeated alarms has stopped, if the repeated alarm cannot be extracted from the alarm database 11 within a predetermined period, it instructs the alarm culling unit 12 to cancel the stopped state of the culling of the repeated alarm.
[0077] By doing so, in the network facility monitoring device, when the repeated alarm extraction unit 14 detects that the network facility that has been outputting repeated alarms has been restored, it can promptly collate the next repeated alarm with the repeated alarms of past cases.
[0078] The network facility monitoring method is a network facility monitoring method of the network facility monitoring device 10. The network facility monitoring device 10 includes an alarm database 11 that accumulates an abnormality detection alarm notifying the occurrence of an abnormality output from the network facility 20 of the optical transmission system 1 and a recovery detection alarm notifying the recovery of the abnormality as alarm information, and an occurrence and recovery alarm pattern database 15 that accumulates information on repeated alarms of past cases together with the causes of the repeated alarms. The method includes the steps of periodically culling the alarm information accumulated in the alarm database 11 and transmitting it to a predetermined notification destination 30; extracting, from the alarm database 11, an alarm group in which an abnormality detection alarm and a recovery detection alarm have occurred repeatedly a predetermined number of times within a predetermined time for the same alarm content of the same network facility as a repeated alarm; and collating the information of the extracted repeated alarm with the information of the repeated alarm of the past case, and if it is determined that the similarity is higher than a predetermined threshold, notifying the cause of the repeated alarm of the past case to a predetermined notification destination (upper device 30).
[0079] By doing so, in the network facility monitoring method, when there are accumulated past cases with a similarity higher than a predetermined threshold for repeated alarms that notify the occurrence and recovery of anomalies, the network facility monitoring device 10 can notify the cause of the repeated alarms of the past cases. Therefore, the maintainer of the network facility can quickly identify the cause of the repeated alarms.
[0080] Note that the present invention is not limited to the embodiments described above, and many modifications are possible by those with ordinary knowledge in the art within the technical idea of the present invention. For example, although the alarm pattern matching unit is configured to notify the cause of past cases of repeated alarms to the upper device 30, it is not limited thereto, and it may be configured to notify the GUI unit 13. Also, the alarm pattern matching unit may be configured to notify the cause of past cases of repeated alarms to the upper device 30 and the GUI unit 13. Further, the alarm pattern matching unit may notify the cause only to the upper device 30 or only to the GUI unit 13 according to the cause of past cases of repeated alarms.
[0081] In the above embodiment, as the characteristic information of the repeated pattern, the average value of the alarm occurrence interval, the standard deviation of the alarm occurrence interval, and the ratio of the abnormal duration in the repeated alarm are exemplified, but it is not limited thereto. The characteristic information may be statistical information that can be calculated using the time information of the repeated alarm. This network facility monitoring method can easily add statistical information obtained by other applicable calculation methods. Also, although the alarm pattern matching unit determines that a past case has a high similarity to the problematic repeated alarm when the similarity is high in all three viewpoints regarding the characteristic information, it is not limited to this. The number of viewpoints regarding the characteristic information necessary for determining the similarity is predetermined. The alarm content is not limited to LOS, and may be other failure states such as loss of frame (LOF).
Explanation of Reference Numerals
[0082] 1 Optical transmission system 2 Maintainer 10 Network equipment monitoring device 11 Alarm database 12 Alarm harvesting section 13 GUI section (destination) 14 Repeated alarm extraction section 15 Occurrence / recovery alarm pattern database 16 Alarm pattern matching section 20 Network equipment 21 Transponder 22 Wavelength selection switch 23 Optical amplification section 30 Higher-level device (destination)
Claims
1. An alarm database that accumulates, as alarm information, an anomaly detection alarm that notifies of the occurrence of an anomaly output from the network facilities of an optical transmission system and a recovery detection alarm that notifies of the recovery of the anomaly, An alarm pruning unit that periodically prunes the alarm information accumulated in the alarm database and transmits it to a predetermined notification destination, A repeated alarm extraction unit that extracts, from the alarm database, an alarm group in which an anomaly detection alarm and a recovery detection alarm repeatedly occur a predetermined number of times within a predetermined time for the same alarm content of the same network facility, as a repeated alarm, An occurrence / recovery alarm pattern database that accumulates information on repeated alarms of past cases together with the causes of the repeated alarms, An alarm pattern matching unit that collates the information of the extracted repeated alarm with the information of the repeated alarm of the past case, and when it is determined that the similarity is higher than a predetermined threshold, notifies the cause of the repeated alarm of the past case to the predetermined notification destination. A network facility monitoring device characterized by comprising:
2. The information on the repeated alarms of the past cases Includes basic information in which alarm content information, identification information of the network facility where the alarm occurred, and information for specifying the location where the alarm occurred are each recorded in association with the time information of the anomaly detection alarm or the recovery detection alarm, And statistical information calculated using the time information as characteristic information of the repeated pattern, The alarm pattern matching unit calculates statistical information as characteristic information of the repeated pattern using the time information of the extracted repeated alarm, and determines the similarity by collating with the statistical information on the repeated alarms of the past cases. The network facility monitoring device according to claim 1, characterized in that.
3. The information on the repeated alarms of the past cases Information indicating whether there was a spread of the repeated alarm from the network facility where the repeated alarm of the past case occurred to other network facilities, The information includes identification information of other network equipment to which the repeated alarm has been propagated, and characteristic information of the repeat pattern of the repeated alarm that has occurred as a propagating event, The network equipment monitoring device according to claim 2, characterized in that when information indicating that a repeat warning of a past case whose similarity to the extracted repeat warning is determined to be higher than a predetermined threshold has been accumulated, the alarm pattern matching unit waits for a predetermined period of time, and when a new repeat warning occurs during the predetermined period, it compares information of the new repeat warning with information of the repeat warning accumulated as a spillover event of the past case to determine similarity.
4. The network equipment monitoring device according to claim 1, characterized in that when the alarm pattern matching unit determines that the similarity between the extracted repetitive alarm information and the repetitive alarm information of the past case is higher than a predetermined threshold, it instructs the alarm pruning unit to stop both the process of pruning and the process of transmitting the repetitive alarm that is determined to have a high similarity.
5. The network equipment monitoring device according to claim 4, characterized in that if the repeat alarm cannot be extracted from the alarm database within a predetermined period after the pruning of the repeat alarm is stopped, the repeat alarm extraction unit instructs the alarm pruning unit to cancel the stop state of pruning of the repeat alarm.
6. A network equipment monitoring method for a network equipment monitoring device, comprising: The network equipment monitoring device includes: The optical transmission system includes an alarm database that stores, as alarm information, an abnormality detection alarm that notifies an abnormality occurrence and a recovery detection alarm that notifies a recovery from the abnormality, which are output from network equipment of the optical transmission system, and an occurrence and recovery alarm pattern database that stores information on repeated alarms of past cases together with the causes of the repeated alarms, a step of periodically collecting the alarm information stored in the alarm database and transmitting the information to a predetermined notification destination; A step of extracting, as repeated alarms, an alarm group in which abnormal detection alarms and recovery detection alarms repeatedly occur a predetermined number of times within a predetermined time for the same alarm content of the same network facility from the alarm database; A step of collating the information of the extracted repeated alarms with the information of the repeated alarms of the past cases, and when it is determined that the similarity is higher than a predetermined threshold, notifying the cause of the repeated alarms of the past cases to the predetermined notification destination; A network facility monitoring method characterized by executing the above.
7. A program for causing a computer to function as the network facility monitoring device according to any one of claims 1 to 5.
Citation Information
Patent Citations
Power distribution communication network fault analysis method and system
CN111431754A
Fault supporting device
JP1996221295A
Control method of computer system, and computer system
JP2008009842A
Transmitting apparatus, and unit mounted on the same
JP2010147804A
Alarm aggregation device, and alarm aggregation method
JP2012213112A