Monitoring device, monitoring program, and monitoring method

JP2026126552APending Publication Date: 2026-08-05OKI ELECTRIC INDUSTRY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
OKI ELECTRIC INDUSTRY CO LTD
Filing Date
2025-01-24
Publication Date
2026-08-05

AI Technical Summary

Benefits of technology

【0011】 本発明によれば、IPネットワーク上の様々な装置を監視する監視装置が同一ネットワーク上の監視対象サーバの障害を検知したときに、障害発生した監視対象サーバへのアクセスを制限できる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026126552000001_ABST
    Figure 2026126552000001_ABST
Patent Text Reader

Abstract

This system allows monitoring devices that monitor various devices on a network to restrict access to a monitored server on the same network when they detect a failure in that server. [Solution] The present invention is a monitoring device that monitors the status of multiple redundant monitored devices via a switch device, and is characterized by comprising: a monitoring unit that monitors whether or not a failure has occurred in each monitored device; and a port blocking unit that, when a failure is detected by the monitoring unit, blocks the communication port of the switch device to which the failed monitored device is connected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a monitoring device, a monitoring program, and a monitoring method, and can be applied, for example, to a monitoring device of a redundant system that prepares a plurality of servers having the same function and switches to another server when a failure occurs in one server.

Background Art

[0002] Patent Document 1 discloses a method in which, when a failure occurs in an IP encoder in a video service, an NMS (Network Management System) controls a video signal path by controlling the IP encoder or a switch and a router.

[0003] Patent Document 2 discloses a method for switching between communication boards having a communication function, in which, when a CPU congestion state occurs in another board, a connection port of the other board connected from its own board to a switch board is stopped.

Prior Art Documents

Patent Documents

[0004] [[ID=,25]]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] By the way, in a redundant system, when a certain server detects its own failure (for example, a disk failure, etc.), there is a method of promoting connection to another server such as an SBY system by shutting down itself.

[0006] In this situation, the failed server may access its disk (hereinafter referred to as "disk access") while shutting down, which can cause the shutdown process to remain incomplete. This can delay the switchover to another server, which has a significant impact on systems where the accuracy of the process and the possibility of interruptions are unacceptable.

[0007] Therefore, there is a need for monitoring devices, monitoring programs, and monitoring methods that can restrict access to a monitored server that has failed when a monitoring device that monitors various devices on a network detects a failure in that server on the same network. [Means for solving the problem]

[0008] To solve these problems, the first invention provides a monitoring device that monitors the status of a plurality of redundant monitored devices via a switch device, characterized by comprising: (1) a monitoring unit that monitors whether or not a failure has occurred in each monitored device; and (2) a port blocking unit that, when a failure is detected by the monitoring unit, blocks the communication port of the switch device to which the failed monitored device is connected.

[0009] The second aspect of the present invention relates to a monitoring program for a monitoring device that monitors the status of multiple redundant monitored devices via a switch device, characterized in that the computer functions as (1) a monitoring unit that monitors whether or not a failure has occurred in each monitored device, and (2) a port blocking unit that, when a failure is detected by the monitoring unit, blocks the communication port of the switch device to which the failed monitored device is connected.

[0010] The third aspect of the present invention relates to a monitoring method for monitoring the status of multiple redundant monitored devices via a switch device, characterized in that (1) a monitoring unit monitors whether or not a failure has occurred in each monitored device, and (2) when a failure is detected by the monitoring unit, a port blocking unit blocks the communication port of the switch device to which the failed monitored device is connected. [Effects of the Invention]

[0011] According to the present invention, when a monitoring device that monitors various devices on an IP network detects a failure in a monitored server on the same network, it can restrict access to the failed monitored server. [Brief explanation of the drawing]

[0012] [Figure 1] This is an overall configuration diagram showing the overall configuration of the communication system according to the embodiment. [Figure 2] This is a configuration diagram showing an example of the configuration of alarm information recorded in the alarm journal DB according to the embodiment. [Figure 3] This is a flowchart showing the operation of port blocking processing by a monitoring device in a communication system according to an embodiment. [Figure 4] This is a configuration diagram showing the host port correspondence information configuration according to the embodiment. [Modes for carrying out the invention]

[0013] (A) Main embodiment The main embodiments of the monitoring device, monitoring program, and monitoring method according to the present invention will be described in detail below with reference to the drawings.

[0014] (A-1) Configuration of the embodiment Figure 1 is an overall configuration diagram showing the overall configuration of the communication system according to the embodiment.

[0015] In FIG. 1, the communication system 1 includes a monitoring device 100, a switch device 200 such as a Layer 2 switch (L2 switch), a first monitored server 300, a second monitored server 400, and a server connection device 500.

[0016] The communication system 1 is a communication system in which the server connection device 500 can be connected to monitored servers (the first monitored server 300 and the second monitored server 400).

[0017] On the side of the monitored servers, a redundancy configuration is constituted by the first monitored server 300 and the second monitored server 400 having the same function. That is, when a failure occurs in one of the monitored servers, a redundancy system is adopted in which the processing (services) performed on one of the monitored servers can be taken over by switching to the other monitored server.

[0018] In this embodiment, the case where the side of the monitored servers has a two-system redundancy configuration is illustrated, but the number of redundancies is not particularly limited. Also, the configuration method of the redundancy is illustrated by the case of active / active (ACT / ACT), but is not limited thereto. The services provided by the side of the monitored servers are not particularly limited, but for example, those that need to avoid the risk of interruption of business / operations such as telephone services (including IP telephone services), billing management, and financial transactions can be applied.

[0019] The switch device 200 has a plurality of physical communication ports (hereinafter also referred to as "switch ports"), and is connected to and communicates with the first monitored server 300 and the second monitored server 400. For example, it is assumed that the first monitored server 300 is connected to a first switch port 210 as a first communication port, and the second monitored server 400 is connected to a second switch port 220 as a second communication port.

[0020] The monitoring device 100 is connected to the first monitored server 300 and the second monitored server 400 via the switch device 200, and monitors whether the first monitored server 300 and the second monitored server 400 are operating normally.

[0021] Although the hardware configuration of the monitoring device 100 is omitted, a personal computer equipped with a CPU, ROM, RAM, EEPROM, etc., a general-purpose server, etc. can be applied to the monitoring device 100. The monitoring device 100 can realize monitoring processing, redundancy processing, etc. by reading and executing a processing program (for example, a monitoring program, a redundancy program, etc.) stored in the ROM by the CPU.

[0022] The monitoring device 100 monitors the operating states of the first monitored server 300 and the second monitored server 400.

[0023] Here, various monitoring methods can be applied. In this embodiment, polling monitoring and trap monitoring can be applied. For example, SNMP (Simple Network Management Protocol) Trap can be used. For example, when a failure occurs in either the first monitored server 300 or the second monitored server 400, the monitored server in which the failure has occurred transmits a failure notification (hereinafter also referred to as "SNMP trap") to the monitoring device 100 and performs a shutdown process. The monitoring device 100 grasps that something abnormal has occurred in the monitored server that is the source of the SNMP trap by receiving the SNMP trap as a failure notification.

[0024] Furthermore, if the monitoring device 100 matches the pre-set conditions, it will block the switch port of the switch device 200 connected to the failed monitored server. The process of blocking the switch port will be explained in detail in the operation section. By blocking the switch port of the switch device 200 connected to the failed monitored server in this way, access to the monitored server can be blocked, and the shutdown process of the monitored server can be completed.

[0025] In Figure 1, the monitoring device 100 includes a control unit 110, an alarm journal database (DB) 120, and a communication unit 130 that sends and receives information to and from a communication network.

[0026] The control unit 110 is responsible for the functions of the monitoring device 100. The control unit 110 includes a trap monitoring unit 111 and a port blocking unit 112.

[0027] The Trap monitoring unit 111 monitors the operating status of the first monitored server 300 and the second monitored server 400 by polling trap monitoring. The Trap monitoring unit 111 functions, for example, as an SNMP manager using the SNMP trap method. When the Trap monitoring unit 111 receives an SNMP trap as a fault notification, it records the received SNMP trap in the alarm journal DB 120.

[0028] For example, an SNMP trap includes at least a "TrapID" that identifies the type of failure information and a "Trap message content" that indicates the cause of the failure. The Trap monitoring unit 111 records alarm information, including the "TrapID" and "Trap message content" contained in the SNMP trap, in the alarm journal DB 120. An example of the alarm information configuration will be described later.

[0029] The port blocking unit 112 refers to alarm information recorded in the alarm journal DB 120 and, when predetermined conditions set in advance are met, blocks the switch port to which the monitored server that sent the SNMP trap (i.e., the monitored server that experienced a failure) is connected. For example, the port blocking unit 112 notifies the switch device 200 of a command to block the switch port and causes it to perform the blocking process.

[0030] Furthermore, the port blocking unit 112 has server-port correspondence information 112a that associates the identification information of the monitored servers (first monitored server 300, second monitored server 400) with the switch port number to which the monitored server is connected. Here, the switch port number is a number assigned to each physical communication port among the multiple physical communication ports of the switch device 200 (note that the switch port number may also be information such as alphanumeric characters assigned to each physical communication port).

[0031] Since it is possible to know in advance which switch port of the switch device 200 the monitored device will be connected to, the server port correspondence information 112a can be configured in advance. Note that if the connected switch port is changed, the server port correspondence information 112a will be changed each time.

[0032] The alarm journal DB120 records alarm information for each SNMP trap it receives.

[0033] Figure 2 is a diagram showing an example of the configuration of alarm information recorded in the alarm journal DB120 according to the embodiment. As illustrated in Figure 2, the alarm information includes items such as "No." which is a serial number, "TrapID" which indicates identification information (identification number) that uniquely identifies the SNMP trap type (i.e., event), "Trap Occurrence Time" which indicates the time of the failure, such as year, month, and day (yyyy / mm / dd) and time (hh:mm:ss), "Trap Message Content" which indicates the cause of the failure, and "Trap Sending Server" which indicates the source of the SNMP trap. Note that other items may also be set in addition to these items.

[0034] The "Trap message content" contains information indicating the cause of the failure. The cause of the failure is information indicating various failures that may occur in the monitored server, and can be hardware failures, software failures, etc.

[0035] For example, hardware failures could include disk failures (recording media) where data stored on the monitored server's disk is inaccessible or files and folders cannot be opened, or hardware failures such as the server failing to start or restart, or running slowly. Similarly, software failures could include application failures such as files and folders not opening or messages not being displayed. The "Trap message content" contains information that identifies the cause of the failure as described above. Alternatively, identification information for identifying the cause of the failure may be set up in advance and included in the "Trap message content."

[0036] The "Trap Sending Server" field contains identification information that identifies the monitored server that sent the SNMP trap. For example, this could be the server name (hostname) or the address information (MAC address, IP address, etc.) of the monitored server. In any case, as long as it can identify the monitored server that sent the SNMP trap, it can be various pieces of information.

[0037] The first monitored server 300 and the second monitored server 400 are servers monitored by the monitoring device 100, are connectable to the server connection device 500, and accept access from the server connection device 500 and exchange data.

[0038] The hardware configuration of the monitored servers (first monitored server 300, second monitored server 400) is omitted, but the monitoring device 100 can be a personal computer or general-purpose server equipped with a CPU, ROM, RAM, EEPROM, etc. The monitoring device 100 can perform monitoring processing, redundancy processing, etc. by having the CPU read and execute processing programs (e.g., monitoring program, redundancy program, etc.) stored in ROM.

[0039] The first monitored server 300 and the second monitored server 400 each have a control unit 310 and a control unit 410, a disk device 320 and a disk device 420, and a communication unit 330 and a communication unit 430.

[0040] Here, the first monitored server 300 and the second monitored server 400 are servers with the same functionality. The internal configuration of the first monitored server 300 will be described below, and the second monitored server 400 will be omitted, but the second monitored server 400 has the same components as the first monitored server 300.

[0041] The disk device 320 is a recording medium that records and stores data necessary for the first monitored server 300 to execute processing. For example, if the first monitored server 300 is a call control server such as SIP (Session Initiation Protocol), the disk device 320 can record subscriber information, call information, call detail record (CDR), etc. Note that the data recorded and stored in the disk device 320 is not limited and varies depending on the function of the monitored server.

[0042] The control unit 310 is responsible for the functions of the first monitored server 300. The control unit 310 is equipped with an SNMP trap monitoring function and has a first disk monitoring unit 311, a first trap sending unit 312, and a first automatic shutdown unit 313. Here, various causes can be applied to the failure factor, but disk failure is given as an example.

[0043] The first disk monitoring unit 311 monitors the status of the disk drive 320. For example, if the first disk monitoring unit 311 detects that the disk drive 320 has failed, it notifies the first trap sending unit 312 of this fact.

[0044] The first trap sending unit 312, when a fault is detected, creates an SNMP trap corresponding to the detected fault cause and sends it to the monitoring device 100. For example, it functions as an SNMP agent using the SNMP trap method.

[0045] The first automatic shutdown unit 313 sends an SNMP trap as a fault notification to the monitoring device 100, then terminates processing on the first monitored server 300 and puts it into a power-off state (i.e., shuts it down).

[0046] The server connection device 500 connects to either the first monitored server 300 or the second monitored server 400 to exchange data, and can be a personal computer, server, or the like. The server connection device 500 includes a communication unit 520 that sends and receives communication signals to and from a communication network, and a connection bypass unit 510 that bypasses the connection path when a failure occurs in the monitored server (the first monitored server 300 or the second monitored server 400).

[0047] (A-2) Operation of the embodiment Next, the operation of the redundancy processing of the monitoring device 100 in the communication system 1 according to the embodiment will be explained with reference to the drawings.

[0048] Here, we will illustrate a case where the cause of the failure is a disk failure. For example, we will illustrate a case where the first monitored server 300 and the second monitored server 400 employ an ACT / ACT configuration, and the disk device 320 of the first monitored server 300 fails, resulting in a disk failure.

[0049] First, in the first monitored server 300, the first disk monitoring unit 311 monitors the status of the disk device 320. When it detects a failure in the disk device 320, the first trap sending unit 312 sends an SNMP trap indicating the disk failure to the monitoring device 100. At this time, after sending the SNMP trap, the first automatic shutdown unit 313 in the first monitored server 300 performs the shutdown process.

[0050] Here, when the first monitored server 300 shuts down, it is basically inaccessible. However, for some reason, the disk device 320 of the first monitored server 300 may be accessed while the shutdown process is in progress, and the shutdown process may not be completed. In that case, access to the first monitored server 300 becomes possible, which may affect the system, such as causing delays in server switching or preventing the switch from being performed correctly. Therefore, to prevent access to the failed first monitored server 300 even if the shutdown process remains incomplete, the following measures are taken.

[0051] Figure 3 is a flowchart showing the operation of port blocking processing by the monitoring device 100 in the communication system 1 according to this embodiment.

[0052] When the monitoring device 100 receives an SNMP trap from the first monitored server 300, the Trap monitoring unit 111 records the received SNMP trap in the alarm journal DB 120.

[0053] [Step S101] In the monitoring device 100, the port blocking unit 112 is activated continuously or intermittently. When activated, the port blocking unit 112 performs a conflict check (step S101).

[0054] In other words, it checks whether the port blocking function by the port blocking unit 112 is already running multiple times. If there is a conflict (step S101 / conflict found), the port blocking unit 112 terminates its process; if there is no conflict (step S101 / no conflict), it proceeds to step S102.

[0055] For example, if the port blocking function is activated when using an SNMP trap monitoring tool, the port blocking function will be executed twice. To avoid this duplication, the port blocking unit 112 performs a conflict check. However, the conflict check process is not mandatory and may be omitted depending on how the monitoring method is implemented.

[0056] [Step S102] The port blocking unit 112 checks the alarm journal DB 120 to confirm the operating status of the first monitored server 300 and the second monitored server 400 (step S102).

[0057] In this example, if, for example, an SNMP trap related to a disk failure is recorded, the port blocking unit 112 performs port blocking processing. On the other hand, if no SNMP trap related to a disk failure is recorded, the port blocking unit 112 terminates the port blocking function processing.

[0058] The time when the last alarm journal DB120 was checked is recorded, and the port blocking unit 112 checks whether or not SNMP traps sent by the first monitored server 300 and the second monitored server 400 exist since that time.

[0059] For example, when checking the alarm journal DB120 at one-minute intervals, the port blocking unit 112 checks whether the SNMP trap's TrapID and Trap message content have been recorded within the past minute using the identification information of the first monitored server 300 or the second monitored server 400.

[0060] Here, it is assumed that an SNMP Trap sent by the first monitored server 300 or the second monitored server 400 is recorded in the alarm journal DB 120. If no such trap is recorded, the process may be terminated.

[0061] [Step S103] The port blocking unit 112 checks whether the SNMP trap sent by the first monitored server 300 or the second monitored server 400 is the Trap ID of the target fault (step S103).

[0062] For example, if the target of port blocking is a disk failure, the port blocking unit 112 checks whether the TrapID indicates a disk failure.

[0063] Then, if the port blocking unit 112 does not have a target TrapID (step S103 / no target Trap), it terminates processing, and if it does have a target TrapID (step S103 / target TrapID exists), it proceeds to step S104.

[0064] [Step S104] For SNMP traps sent by the first monitored server 300 and the second monitored server 400, if the trap ID matches the target trap ID, the port blocking unit 112 checks whether the trap message content is the target cause of the failure (step S104).

[0065] Then, if the port blocking unit 112 is not the target cause of failure (step S104 / not the target), it terminates processing, and if it is the target cause of failure (step S104 / is the target), it proceeds to step S105.

[0066] The TrapID is identification information that identifies the event name of the trap, and the Trap message content describes the cause of the failure. Therefore, based on the TrapID in step S103 and the Trap message content in step S104, the port blocking unit 112 can determine that there is a disk failure.

[0067] [Step S105] Based on the TrapID and Trap message content, if the failure that is the target of the port blocking is a disk failure, the port blocking unit 112 refers to the host port correspondence information 112a to identify the switch port (step S105).

[0068] Figure 4 is a configuration diagram showing the configuration of host port correspondence information 112a according to the embodiment. As illustrated in Figure 4, the host port correspondence information 112a includes at least "server identification information" that uniquely identifies the server to be monitored and "communication port number" that identifies the communication port of the switch device 200, and associates the "server identification information" with the "communication port number".

[0069] The "server identification information" is the same information recorded in the "Trap sending server" item of the alarm information in Figure 2. For example, it can be the server name (hostname) of the monitored server, the address information of the monitored server (MAC address, IP address, etc.), etc. This makes it possible to quickly identify the switch port of the switch device 200 to which the first monitored server 300 that has experienced a failure is connected.

[0070] [Step S106] Next, the port blocking unit 112 instructs the switch device 200 to block the switch port to which the first monitored server 300 that has experienced a failure is connected (step S106).

[0071] For example, the port blocking unit 112 sends a port blocking command to the switch device 200 that includes a port number that identifies the first switch port (communication port) 210 connected to the first monitored server 300 identified in step S105, thereby blocking the first switch port 210.

[0072] Port blocking refers to preventing data from passing through the switch port (communication port) to which a failed monitored server is connected. Port blocking means closing the communication port and shutting it down.

[0073] Furthermore, port blocking includes both logical and physical blocking. As an example of logical blocking, if the switch device 200 uses an input / output port connection mapping table to establish connections between switch ports, the connection mapping table can be modified to prevent connection to the first switch port 210. Note that port blocking can be broadly applied to any method that prevents data from being sent or received through the switch port in question.

[0074] [Step S107] The port blocking unit 112 checks the port blocking result of the switch port instructed to the switch device 200 (step S107). If the port blocking result is OK (step S107 / OK), the port blocking process is terminated. On the other hand, if the port blocking result is NG (step S107 / OFF), the port blocking unit 112 proceeds to process S108.

[0075] [Steps S108, S109] If the port closure result for the switch port is NG, the port closure unit 112 retries shutting down the switch port until it reaches a predetermined number of retries (switch S108 / YES) (step S108).

[0076] If the switch port shutdown fails a predetermined number of times (step S108 / NO), the port blocking unit 112 receives an SNMP trap indicating port blocking failure sent from the switch device 200 to the monitoring device 100 and terminates processing (step S109).

[0077] If the shutdown of the switch port is successful, the port blocking unit 112 terminates the port blocking function.

[0078] In this way, the first switch port 210 to which the first monitored server 300 that has failed can be shut down, thereby restricting disk access, and disk access can also be restricted even if the first monitored server 300 remains shut down for some reason.

[0079] The server connection device 500 will no longer be able to access the first monitored server 300, but the connection bypass processing by the connection bypass unit 510 will enable it to access the other second monitored server 400.

[0080] (A-3) Effects of the Embodiment As described above, according to this embodiment, by shutting down the switch port of the switch device to which the failed monitored server is connected, access from the server connection device can be restricted even if the shutdown process on the failed monitored server remains incomplete.

[0081] (B) Other embodiments Although various modified embodiments have been mentioned in the embodiments described above, the present invention can also be applied to the following modified embodiments.

[0082] (B-1) In the above embodiment, the example given was a case where the disk device in the monitored server failed, but the cause of failure is not limited to disk failure.

[0083] (B-2) In the above-described embodiment, the application of SNMP traps as a fault notification method was illustrated as an example, but the method is not limited to any method that enables the exchange of fault monitoring information between the monitoring device and the monitored server and the monitoring device, and that can transmit the cause of the fault, and various methods can be applied. [Explanation of Symbols]

[0084] 1... Communication systems, 100...Monitoring device, 110...Control unit, 111...Trap monitoring unit, 112...Port blocking unit, 112a...Server port correspondence information, 120...Alarm journal DB, 130...Communication unit, 200... Switching device, 210... First switch port, 220... Second switch port, 300...First monitored server, 310...Control unit, 311...First disk monitoring unit, 312...First trap sending unit, 313...First automatic shutdown unit, 320...Disk device, 330...Communication unit, 400...Second monitored server, 410...Control unit, 411...Second disk monitoring unit, 412...Second trap sending unit, 413...Second automatic shutdown unit, 420...Disk device, 430...Communication unit, 500...Server connection device, 510...Connection bypass unit, 520...Communication unit.

Claims

1. In a monitoring device that monitors the status of multiple redundant monitored devices via a switch device, A monitoring unit that monitors whether or not a failure has occurred in each of the aforementioned monitored devices, When the monitoring unit detects a fault, the port blocking unit blocks the communication port of the switch device to which the faulty monitored device is connected. A monitoring device characterized by being equipped with the following features.

2. The system includes server port correspondence information which associates the device identification information of each of the monitored devices with the communication port number of the switch device to which the monitored device is connected, or information linked to the communication port number of the switch device. The port blocking unit refers to the server port correspondence information and instructs the switch device to block the communication port corresponding to the device identification information of the monitored device that has experienced a failure. The monitoring device according to feature 1.

3. The system includes a storage unit that stores fault notifications obtained from one of the multiple monitored devices that has experienced a fault. When the port monitoring unit has stored the fault notification containing the fault cause for port closure in the storage unit, it reads the identification information of the monitored device that sent the fault notification and instructs the switch device to close the corresponding communication port. The monitoring device according to feature 2.

4. In a monitoring program for a monitoring device that monitors the status of multiple redundant monitored devices via a switch device, Computers, A monitoring unit that monitors whether or not a failure has occurred in each of the aforementioned monitored devices, When the monitoring unit detects a fault, the port blocking unit blocks the communication port of the switch device to which the faulty monitored device is connected. A monitoring program characterized by its ability to function in this way.

5. In a monitoring method that monitors the status of multiple redundant monitored devices via a switch device, The monitoring unit monitors whether or not a failure has occurred in each of the aforementioned monitored devices. When the monitoring unit detects a fault, the port blocking unit blocks the communication port of the switch device to which the faulty monitored device is connected. A monitoring method characterized by the following features.