Failure monitoring device and failure monitoring method
The fault monitoring device enhances standby switch unit failure detection by generating and modifying monitoring packets to cover all buffer bits, addressing the accuracy issues in existing methods and ensuring comprehensive fault detection without increased computational load.
Patent Information
- Application Number
- PCT/JP2024/029425
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-02-26
AI Technical Summary
Existing fault monitoring methods in standby switch units of packet transmission devices are unable to accurately detect hardware failures outside the range of bits stored by monitoring packets due to reduced data flow rates, leading to undetected failures.
A fault monitoring device and method that generates and modifies monitoring packets to ensure they cover all bits in the standby switch unit's buffer, using techniques like altering packet length or inverting data, enabling accurate parity checks.
High-accuracy detection of hardware failures in standby switch units without increasing computational load, ensuring reliable fault detection across all buffer bits.
Smart Images

Figure JP2024029425_26022026_PF_FP_ABST
Abstract
Description
Fault monitoring device and fault monitoring method
[0001] The present disclosure relates to a fault monitoring device and a fault monitoring method.
[0002] In order to maintain communication in the event of a hardware failure (hereinafter simply referred to as a "failure"), a packet transmission device is provided with a redundant common unit including an active switch unit (active SW unit) and a standby switch unit (standby SW unit). That is, under normal circumstances, the active SW unit is operated to output packets input from an input I / F to a downstream output I / F. Furthermore, if a failure occurs in the active SW unit, the system switches to the standby SW unit, which is a non-active system, and outputs packets.
[0003] The active and standby SW units each have a buffer unit that stores packets input from the input I / F and then transmits the packets from the output I / F to downstream devices. Also, a parity check is performed in the buffer unit to detect any faults that occur in the buffer unit.
[0004] The active SW performs a parity check on the entire signal to be communicated. This allows it to detect faults in all bits in the buffer section installed in the active SW. On the other hand, in order to reduce the monitoring load, the standby SW outputs a fixed pattern monitoring packet to the buffer section and performs a parity check.
[0005] Furthermore, Patent Documents 1 and 2 disclose that a parity bit is added to a main signal in order to detect a fault occurring in a communication line, and when a certain number of errors are detected, it is determined that a fault has occurred.
[0006] JP-A No. 63-269640 JP-A No. 60-072460
[0007] In the above-described packet transmission device, the active SW unit transmits packets during operation, allowing parity checks to be performed on each bit in the buffer unit and detecting failures in each bit. However, the standby SW unit reduces the data flow rate to reduce the computational load required for failure monitoring. As a result, the data length of the monitoring packet is limited, and it may not be possible to perform parity checks on each bit in the buffer unit. For example, if a failure occurs outside the range of bits in the buffer unit where the monitoring packet is stored, it cannot be detected. Furthermore, if the data in the monitoring packet matches data caused by a hardware failure, the failure cannot be detected by parity checks.
[0008] Furthermore, even with the methods disclosed in the above-mentioned Patent Documents 1 and 2, if a failure occurs outside the range of bits in which the monitoring packets used for parity checking are stored, and if the data in the monitoring packet matches the data caused by a hardware failure, the failure cannot be detected.
[0009] The present disclosure has been made in consideration of the above circumstances, and its purpose is to provide a fault monitoring device and a fault monitoring method that are capable of detecting faults in a buffer section installed in a standby switch section with high accuracy.
[0010] A fault monitoring device according to one aspect of the present disclosure is a fault monitoring device that detects faults that occur in a standby switch unit installed in a packet transmission device, and includes a generation unit that generates a monitoring packet, a modification unit that modifies the monitoring packet, and a transmission unit that transmits the monitoring packet generated by the generation unit and the monitoring packet modified by the modification unit to the standby switch unit.
[0011] A fault monitoring method according to one aspect of the present disclosure is a fault monitoring method in which a fault monitoring device detects a fault occurring in a standby switch unit of a packet transmission device, generates a monitoring packet, modifies the monitoring packet, and transmits the generated monitoring packet and the modified monitoring packet to the standby switch unit.
[0012] According to the present disclosure, it is possible to detect a failure in a buffer unit mounted on a standby switch unit with high accuracy.
[0013] FIG. 1 is a block diagram showing the configuration of a fault monitoring device according to an embodiment and a transmission device to be monitored for faults. FIG. 2A is an explanatory diagram showing data accumulated in a buffer unit mounted in an active SW unit according to the first embodiment. FIG. 2B is an explanatory diagram showing data accumulated in a buffer unit mounted in a standby SW unit according to the first embodiment. FIG. 3 is a flowchart showing a processing procedure performed by the fault monitoring device according to the first embodiment. FIG. 4 is an explanatory diagram showing how a hardware fault in a buffer unit mounted in a standby SW unit is detected by the fault monitoring device according to the first embodiment. FIG. 5A is an explanatory diagram showing data accumulated in a buffer unit mounted in an active SW unit according to the second embodiment. FIG. 5B is an explanatory diagram showing data accumulated in a buffer unit mounted in a standby SW unit according to the second embodiment. FIG. 6 is a flowchart showing a processing procedure performed by the fault monitoring device according to the second embodiment. FIG. 7 is an explanatory diagram showing how a hardware fault in a buffer unit mounted in a standby SW unit is detected by the fault monitoring device according to the second embodiment. FIG. 8 is a block diagram showing the hardware configuration of this embodiment.
[0014] [Description of First Embodiment] Hereinafter, an embodiment will be described with reference to the drawings. Fig. 1 is a block diagram showing the configuration of a fault monitoring device 1 according to the first embodiment and a transmission device 100 (packet transmission device) to be monitored for faults. As shown in Fig. 1, the transmission device 100 has a redundant configuration including two switch units. Specifically, the transmission device 100 includes two switch units: an active SW unit 2 (active switch unit) and a standby SW unit 3 (standby switch unit). Furthermore, the transmission device 100 includes an input I / F 4 and output I / Fs 5 and 6. "I / F" refers to an interface.
[0015] The operational SW unit 2 normally reads data input from the input I / F 4 and outputs the data to the output I / Fs 5 and 6. The operational SW unit 2 includes a buffer unit 21. The operational SW unit 2 accumulates packets input from the input I / F 4 in the buffer unit 21 and performs a parity check to determine whether or not a failure (hardware failure) has occurred in each bit of the buffer unit 21.
[0016] Since packets are constantly transmitted to the active SW unit 2, data is accumulated in each bit of the buffer unit 21 as shown in FIG. 2A. The shaded areas in FIG. 2A indicate bits in which data is accumulated. Therefore, by performing a parity check on the data accumulated in the buffer unit 21, it is possible to detect whether a failure has occurred in each bit of the buffer unit 21. For example, if a failure has occurred in the bit indicated by symbol q1 in FIG. 2A, this failure can be detected. For the sake of simplicity, FIG. 2A shows an example in which the number of bits in the buffer unit 21 is 18. The number of bits in the buffer unit 21 is not limited to 18.
[0017] The standby SW unit 3 operates during non-operational periods when, for example, a failure occurs in the active SW unit 2 and the active SW unit 2 is unable to operate. That is, when a failure occurs in the active SW unit 2, the transmission device 100 switches to the standby SW unit 3. When the active SW unit 2 is unavailable, the standby SW unit 3 reads data input from the input I / F 4 and outputs the data to the output I / Fs 5 and 6. The standby SW unit 3 includes a buffer unit 31.
[0018] The standby SW unit 3 does not transmit packets when the active SW unit 2 is operating. Therefore, the standby SW unit 3 acquires monitoring packets output from the fault monitoring device 1 at any time interval (for example, at regular time intervals) and stores them in the buffer unit 31. By performing a parity check on the data stored in the buffer unit 31 by the monitoring packets, it is possible to detect whether or not a failure has occurred in each bit in the buffer unit 31.
[0019] 1 includes a generating unit 11, a changing unit 12, and a transmitting unit 13. The fault monitoring device 1 is connected to the standby SW unit 3, and outputs monitoring packets for fault monitoring to the standby SW unit 3 at any time intervals, as described above. The fault monitoring device 1 detects a fault that occurs in the standby SW unit 3 mounted on the transmission device 100 (packet transmission device) as described below.
[0020] The generation unit 11 generates a monitoring packet for detecting a failure occurring in each bit of the buffer unit 31 mounted in the standby SW unit 3. The generation unit 11 limits the packet length to reduce the calculation load during monitoring. Therefore, as shown in Fig. 2B , when a monitoring packet is transmitted to the standby SW unit 3, data is not stored in all bits of the buffer unit 31 mounted in the standby SW unit 3, and the data of the monitoring packet is stored in, for example, the bit indicated by symbol q2 in Fig. 2B .
[0021] The modification unit 12 modifies the monitoring packet generated by the generation unit 11. The modification unit 12 executes a process of modifying the data length of the monitoring packet generated by the generation unit 11. That is, if the monitoring packet generated by the generation unit 11 is output as is to the standby SW unit 3, data is accumulated in the bit indicated by symbol q2 in FIG. 2B in the buffer unit 31. For this reason, if a failure occurs in the bit indicated by symbol q3, for example, this failure cannot be detected even by a parity check. For this reason, the modification unit 12 modifies the data length of the monitoring packet.
[0022] Specifically, when a monitoring packet is sent to the standby SW unit 3, once every several times (for example, 10 times), the processing for changing the number of bits of the monitoring packet to N times is executed. By increasing the number of bits of the monitoring packet to N times, the data of the monitoring packet is accumulated in all bits of the buffer unit 31 shown in FIG. 2B. That is, data is accumulated in the bit indicated by symbol q3. The changing unit 12 changes the packet length of the monitoring packet to be longer at any point when the transmitting unit 13 (details will be described later) periodically transmits the monitoring packet.
[0023] The transmitter 13 periodically transmits a monitoring packet to the standby SW unit 3 at an arbitrary time interval (e.g., at a fixed time interval). The transmitter 13 transmits the monitoring packet generated by the generator 11 and the monitoring packet modified by the modifying unit 12 (a monitoring packet with N times the number of bits) to the standby SW unit 3. Specifically, the transmitter 13 transmits the monitoring packet generated by the generator 11 to the standby SW unit 3 at an arbitrary interval, and transmits the monitoring packet modified by the modifying unit 12 to the standby SW unit 3, for example, once every M monitoring packet transmissions. That is, the monitoring packet transmitted from the transmitter 13 accumulates data in the bit indicated by symbol q1 of the buffer unit 31 shown in FIG. 2B. Furthermore, data accumulates in all bits of the buffer unit 31 shown in FIG. 2B once every M transmissions.
[0024] 3 is a flowchart showing the procedure for fault determination in the fault monitoring device 1 and the standby SW unit 3 according to the first embodiment. The procedure for fault determination will be described below with reference to the flowchart shown in Fig. 3. First, in step S11 of Fig. 3, the generation unit 11 of the fault monitoring device 1 generates a monitoring packet for fault detection in the buffer unit 31 mounted in the standby SW unit 3.
[0025] In step S12, the transmitter 13 determines whether the number of times the monitoring packet has been transmitted has reached a certain number (for example, M times). If the number of times has reached M (S12; YES), the process proceeds to step S13. If not (S12; NO), the process proceeds to step S14. Initially, the number of times the monitoring packet has been transmitted has not reached M, so the process proceeds to step S14.
[0026] In step S14, the transmitter 13 transmits the monitoring packet to the standby SW unit 3. Monitoring packet data is stored in each bit of the buffer unit 31 of the standby SW unit 3. As described above, the packet length of the monitoring packet is limited to reduce the calculation load. For this reason, as shown in FIG. 2B , packet data is not stored in all bits of the buffer unit 31, and data is stored only in, for example, the bit indicated by symbol q2.
[0027] In step S15, the standby SW unit 3 performs a parity check on the monitoring packet to determine whether or not a failure (hardware failure) has been detected in the buffer unit 31. If a failure has been detected (S15; YES), the process proceeds to step S16; otherwise (S15; NO), the process returns to step S12.
[0028] As described above, since packet data is not stored in all bits of the buffer unit 31, if a failure occurs in, for example, the bit indicated by symbol q3, this failure cannot be detected by parity check. Therefore, the determination in step S15 is NO.
[0029] If the number of times the monitoring packet has been transmitted reaches M, the determination in step S12 is YES, and the process proceeds to step S13.
[0030] In step S13, the change unit 12 changes the packet length of the monitoring packet by a factor of N. As a result, the packet length of the monitoring packet becomes longer, and the data of the monitoring packet is stored in all bits of the buffer unit 31 of the standby SW unit 3, as shown in Fig. 4. Therefore, even if a failure occurs in the bit indicated by symbol q3, for example, this failure can be detected by a parity check.
[0031] If a hardware failure is detected in the buffer unit 31 (S15; YES), the standby SW unit 3 notifies the user of the occurrence of the failure in step S16. As a result, the user can recognize that a failure has occurred in the buffer unit 31.
[0032] As such, the fault monitoring device 1 of this embodiment is a fault monitoring device 1 that detects faults that occur in the standby SW unit 3 mounted on the transmission device 100 (packet transmission device), and is equipped with a generation unit 11 that generates a monitoring packet, a modification unit 12 that modifies the monitoring packet, and a transmission unit 13 that transmits the monitoring packet generated by the generation unit 11 and the monitoring packet modified by the modification unit 12 to the standby SW unit 3.
[0033] In this embodiment, the monitoring packet generated by the generation unit 11 is modified by the modification unit 12 and output to the standby SW unit 3, making it possible to detect failures in the buffer unit 31 installed in the standby SW unit 3 with high accuracy.
[0034] In this embodiment, the data length of the monitoring packet is set to be long at a rate of once every several times, so even if the packet length of the monitoring packet is short, it is possible to detect with high accuracy a failure occurring in the buffer unit 31 mounted in the standby SW unit 3. In other words, it is possible to detect a failure occurring in the buffer unit 31 without increasing the calculation load due to the monitoring packet.
[0035] [Description of Second Embodiment] Next, a second embodiment will be described. The device configuration is the same as that shown in Fig. 1, and therefore a description of the configuration will be omitted. In the second embodiment, an inverted monitoring packet is generated by inverting the data of the monitoring packet to be output to the standby SW unit 3, and a parity check is performed using both the monitoring packet and the inverted monitoring packet to detect a failure occurring in the buffer unit 31 mounted in the standby SW unit 3. A detailed description will be given below.
[0036] 5A is an explanatory diagram showing data stored in the buffer unit 21 of the active SW unit 2 shown in FIG. 1. For example, assume that a failure (hardware failure) occurs in the bit indicated by symbol q11 in the buffer unit 21. Since packets are constantly being transmitted to the active SW unit 2, a parity check can be performed to detect the failure that has occurred in the bit indicated by symbol q11.
[0037] FIG. 5B is an explanatory diagram showing data stored in the buffer unit 31 of the standby SW unit 3. A monitoring packet with fixed data is input to the standby SW unit 3 and stored in the buffer unit 31. Therefore, if the data of the monitoring packet matches the data generated by a bit failure, this failure cannot be detected by a parity check. For example, if the data of the monitoring packet is "010111" and a failure occurs in which the bit indicated by symbol q12 in FIG. 5B is fixed to "1," the data of this bit matches the data of the monitoring packet, making it impossible to detect whether or not a failure has occurred. In the second embodiment, failures are detected by using an inverted monitoring packet, which is an inverted version of the monitoring packet, as described above.
[0038] The fault monitoring device 1 according to the second embodiment differs from that of the first embodiment in the function of the change unit 12 shown in Fig. 1. The change unit 12 will be described below.
[0039] The change unit 12 executes a process of inverting each piece of data included in the monitoring packet generated by the generation unit 11. For example, if the data of the monitoring packet is "010111", the data of the inverted monitoring packet becomes "101000".
[0040] The transmitter 13 transmits both the monitoring packet generated by the generator 11 and the inverted monitoring packet generated by the changer 12 to the buffer 31 of the standby SW unit 3. That is, the changer 12 generates an inverted monitoring packet by inverting the data of the monitoring packet, and the transmitter 13 transmits the monitoring packet and the inverted monitoring packet to the standby SW unit 3.
[0041] 6 is a flowchart showing the procedure for fault determination in the fault monitoring device 1 and the standby SW unit 3 according to the second embodiment. The procedure for fault determination according to the second embodiment will be described below with reference to the flowchart shown in FIG. 6. First, in step S31 of FIG. 6, the generation unit 11 of the fault monitoring device 1 generates a monitoring packet for detecting a fault in the buffer unit 31 mounted in the standby SW unit 3. The data of the monitoring packet is, for example, "010111".
[0042] In step S32 , the transmitting unit 13 transmits a monitoring packet to the standby SW unit 3 .
[0043] In step S33, the standby SW unit 3 performs a parity check on the data of the monitoring packet stored in the buffer unit 31 to determine whether a failure has been detected. If a failure has been detected (S33; YES), the process proceeds to step S36; if not (S33; NO), the process proceeds to step S34.
[0044] In step S34, the change unit 12 generates an inverted monitoring packet by inverting the data of the monitoring packet. For example, the change unit 12 generates an inverted monitoring packet having data "101000" obtained by inverting the data "010111" of the monitoring packet. The transmission unit 13 transmits the inverted monitoring packet to the standby SW unit 3.
[0045] In step S35, the standby SW unit 3 performs a parity check on the data of the monitoring packet stored in the buffer unit 31 to determine whether a failure has been detected. If a failure has been detected (S35; YES), the process proceeds to step S36; if not (S35; NO), the process returns to step S32.
[0046] In step S36, the standby SW unit 3 notifies the user of the occurrence of the failure, so that the user can recognize that a failure has occurred in the buffer unit 31.
[0047] For example, if a fault occurs in the bit indicated by q12 in the buffer unit 31 shown in FIG. 5B, where the data is fixed to "1," using a monitoring packet containing the data "010111" would prevent the fault from being detected. On the other hand, using an inverted monitoring packet would result in the data "101000" being stored in each bit of the buffer unit 31. However, in reality, the data "101100" is stored, as shown in FIG. 7. That is, the bit indicated by q13 is "1" instead of "0." Therefore, performing a parity check can detect a fault in the buffer unit 31.
[0048] In this way, in the failure detection device according to the second embodiment, a parity check is performed on the buffer unit 31 mounted in the standby SW unit 3 using the monitoring packet generated by the generation unit 11, and further, a parity check is performed on the buffer unit 31 using an inverted monitoring packet obtained by inverting the data of the monitoring packet. Therefore, if a failure occurs in the buffer unit 31, it is possible to reliably detect this failure.
[0049] The fault monitoring device 1 of the present embodiment described above can be, for example, a general-purpose computer system including a CPU (Central Processing Unit, processor) 901, a memory 902, a storage 903 (HDD: Hard Disk Drive, SSD: Solid State Drive), a communication device 904, an input device 905, and an output device 906, as shown in Fig. 8. The memory 902 and the storage 903 are storage devices. In this computer system, the CPU 901 executes a predetermined program loaded on the memory 902, thereby realizing each function of the fault monitoring device 1.
[0050] The fault monitoring device 1 may be implemented in one computer or in multiple computers, or may be a virtual machine implemented in a computer.
[0051] The program for the fault monitoring device 1 can be stored in a computer-readable recording medium such as a HDD, SSD, USB (Universal Serial Bus) memory, CD (Compact Disc), or DVD (Digital Versatile Disc), or can be distributed via a network. The computer-readable recording medium is, for example, a non-transitory recording medium.
[0052] The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the present disclosure.
[0053] REFERENCE SIGNS LIST 1 Fault monitoring device 2 Operational system SW unit (operational system switch unit) 3 Standby system SW unit (standby system switch unit) 11 Generation unit 12 Change unit 13 Transmission unit 21, 31 Buffer unit 100 Transmission device (packet transmission device)
Claims
1. A fault monitoring device that detects faults that occur in a standby switch unit installed in a packet transmission device, comprising: a generation unit that generates a monitoring packet; a modification unit that modifies the monitoring packet; and a transmission unit that transmits the monitoring packet generated by the generation unit and the monitoring packet modified by the modification unit to the standby switch unit.
2. The fault monitoring device of claim 1, wherein the transmitting unit periodically transmits the monitoring packet to the standby switch unit after an arbitrary time has elapsed, and the modifying unit modifies the packet length of the monitoring packet to be longer at any point when the transmitting unit periodically transmits the monitoring packet.
3. The fault monitoring device according to claim 1, wherein the change unit generates an inverted monitoring packet by inverting the data of the monitoring packet, and the transmission unit transmits the monitoring packet and the inverted monitoring packet to the standby system switch unit.
4. A fault monitoring method for detecting a fault occurring in a standby switch unit of a packet transmission device by a fault monitoring device, the fault monitoring method comprising: generating a monitoring packet; modifying the monitoring packet; and transmitting the generated monitoring packet and the modified monitoring packet to the standby switch unit.
Citation Information
Patent Citations
Switching device for LAN
JP1997181771A
Packet relay device and fault diagnosis method
JP2011166514A