Cluster system, monitoring system, monitoring method, and program
The cluster system coordinates recovery operations across multiple systems using shared determination criteria to address overlapping issues, ensuring appropriate recovery for shared server devices.
Patent Information
- Application Number
- JP2021080395
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-05-11
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-05-11
AI Technical Summary
When a server device is shared and managed by multiple cluster systems, overlapping or conflicting recovery operations occur upon failure, preventing appropriate recovery operations from being executed.
A cluster system with a management unit, monitoring unit, determination unit, and control unit that coordinate recovery operations across multiple cluster systems using shared determination criteria to ensure a single appropriate recovery operation is executed.
Ensures appropriate and coordinated recovery operations for shared server devices across multiple cluster systems, preventing duplicate efforts and maintaining system integrity.
Smart Images

Figure 0007707640000001 
Figure 0007707640000002 
Figure 0007707640000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a cluster system, a monitoring system, a monitoring method, and a program.
Background Art
[0002] When a company or the like constructs an in-house network, a cluster system may be used to ensure scalability and availability. The cluster system manages server devices and the like within the cluster system using a predetermined policy or specific parameters or the like. Also, a server device for which availability is not ensured in the cluster system is excluded from management by the cluster system, and the policy applied to the cluster system is not applied. Thus, a server device that is excluded from management by the cluster system has a recovery operation during a failure executed by a procedure different from that in the case where a failure occurs in a server device or the like within the cluster system.
[0003] Patent Document 1 discloses a configuration in which a plurality of computers connected via a network perform distributed processing. When determining the output order of data, the computers disclosed in Patent Document 1 perform semi-ordered distribution to ensure the consistency of data output from each computer and continue processing even when a failure occurs in some of the computers.
[0004] Also, Patent Document 2 discloses a configuration of a system having two computers that distribute a plurality of functions and a common auxiliary storage device. Patent Document 1 discloses a backup operation mode in which when a failure occurs in one computer, the other computer takes over the functions being executed in the computer in which the failure occurred and operates.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
[0006] Here, when a plurality of cluster systems are included in an in-house network or the like, there are cases where a server device that is not subject to management by the cluster systems is shared and managed by the plurality of cluster systems. In this case, when a failure occurs in the server device, each cluster system executes a recovery operation for the server device, so there is a problem that the recovery operations overlap or conflict, and an appropriate recovery operation cannot be performed. Here, the computer disclosed in Patent Document 2 performs a function takeover according to a predetermined procedure when a failure occurs, so a plurality of recovery operations are not executed for the computer in which the failure has occurred. Therefore, even if the recovery operation at the time of failure disclosed in Patent Document 2 is executed, the problem that an appropriate recovery operation cannot be performed when a failure occurs in the server device shared and further managed by the plurality of cluster systems cannot be solved.
[0007] One of the objects of the present disclosure is to provide a cluster system, a monitoring system, a monitoring method, and a program capable of executing an appropriate recovery operation for a server device when a failure occurs in the server device shared by a plurality of cluster systems. [Means for Solving the Problems]
[0008] The cluster system according to the first aspect of the present disclosure includes a management unit that manages the execution state indicating the monitoring state of server devices in a plurality of cluster systems and the execution state of a first cluster system that executes a recovery operation on a server device when the server device is in an abnormal state, a monitoring unit that monitors whether the server device is in a normal state or an abnormal state, reflects the monitoring result in the monitoring state, and reflects the monitoring result of the server device received from another cluster system in the monitoring state, a determination unit that determines the first cluster system that executes a recovery operation on the server device according to the same determination criteria used by the other cluster system that manages the monitoring state when the monitoring result in at least one of the plurality of cluster systems indicates an abnormal state, and reflects the determination result in the execution state, and a control unit that determines whether to execute a recovery operation on the server device according to the managed execution state.
[0009] The monitoring system according to the second aspect of the present disclosure is a monitoring system including a plurality of cluster systems and server devices managed by the plurality of cluster systems. Each of the cluster systems manages the execution state indicating the monitoring state of the server devices in the plurality of cluster systems and the execution state of a first cluster system that executes a recovery operation on a server device when the server device is in an abnormal state, monitors whether the server device is in a normal state or an abnormal state, reflects the monitoring result in the monitoring state, and reflects the monitoring result of the server device received from another cluster system in the monitoring state. When the monitoring result in at least one of the plurality of cluster systems indicates an abnormal state, the first cluster system that executes a recovery operation on the server device is determined according to the same determination criteria used by the other cluster system that manages the monitoring state, the determination result is reflected in the execution state, and it is determined whether to execute a recovery operation on the server device according to the managed execution state.
[0010] The monitoring method according to the third aspect of the present disclosure manages the monitoring state of server devices in a plurality of cluster systems and the execution state indicating the first cluster system that performs a recovery operation on the server device when the server device is in an abnormal state, monitors whether the server device is in a normal state or an abnormal state, reflects the monitoring result in the monitoring state, and reflects the monitoring result of the server device received from other cluster systems in the monitoring state. When the monitoring result in at least one of the plurality of cluster systems indicates an abnormal state, the first cluster system that performs a recovery operation on the server device according to the same determination criteria used by the other cluster system that manages the monitoring state is determined, the determination result is reflected in the execution state, and it is determined whether to perform a recovery operation on the server device according to the managed execution state.
[0011] The program according to the fourth aspect of the present disclosure manages the monitoring state of server devices in a plurality of cluster systems and the execution state indicating the first cluster system that performs a recovery operation on the server device when the server device is in an abnormal state, monitors whether the server device is in a normal state or an abnormal state, reflects the monitoring result in the monitoring state, and reflects the monitoring result of the server device received from other cluster systems in the monitoring state. When the monitoring result in at least one of the plurality of cluster systems indicates an abnormal state, the first cluster system that performs a recovery operation on the server device according to the same determination criteria used by the other cluster system that manages the monitoring state is determined, the determination result is reflected in the execution state, and causes a computer to determine whether to perform a recovery operation on the server device according to the managed execution state.
Advantages of the Invention
[0012] According to the present disclosure, there can be provided a cluster system, a monitoring system, a monitoring method, and a program that can execute an appropriate recovery operation for a server device when a failure occurs in the server device shared by a plurality of cluster systems.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Modes for Carrying Out the Invention
[0014] (Embodiment 1) Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. A configuration example of the cluster system 10 according to Embodiment 1 will be described with reference to FIG. 1. The cluster system 10 is a system that realizes flexible scalability or high availability by the coordinated operation of one or more computer devices. The cluster system 10 may be a system in which a plurality of computer devices perform distributed processing. Alternatively, the cluster system 10 may be a system having one computer device performing an active operation and a computer device for backing up the computer device performing the active operation. The components of the cluster system 10 described below may be functions or the like that are distributed and executed in a plurality of computer devices, or may be functions or the like that are executed in one computer device performing an active operation.
[0015] A computer device is a device that operates by a processor executing a program stored in a memory. The computer device may be, for example, a server device.
[0016] The cluster system 10, which is a computer device or a set of computer devices, has a management unit 11, a monitoring unit 12, a determination unit 13, and a control unit 14. Components of the cluster system 10 such as the management unit 11, the monitoring unit 12, the determination unit 13, and the control unit 14 may be software or modules in which processing is executed by a processor executing a program stored in a memory. Alternatively, the components of the cluster system 10 may be hardware such as a circuit or a chip.
[0017] The management unit 11 manages the execution state indicating the monitoring state of server devices in a plurality of cluster systems and the execution state of a first cluster system that executes a recovery operation on a server device when the server device is in an abnormal state. Each of the plurality of cluster systems included in the plurality of cluster systems may realize scalability or availability using a policy or system configuration different from other cluster systems. A server device is a computer device that is not subject to management in order to ensure scalability or availability in each cluster system. The server device may be, for example, a DNS (Domain Name System) server device. The server device is managed by each cluster system. In other words, when a failure occurs in the server device, each cluster system detects the failure of the server device, and further, the recovery operation of the server device is executed by each cluster system.
[0018] The monitoring state indicates the monitoring results in each cluster system, and for example, indicates whether the server device is in a normal state or an abnormal state. The abnormal state may be, for example, a state in which a failure or malfunction has occurred in the server device. The recovery operation may be, for example, restarting some functions, services, or applications that the server device has, or restarting the server device itself. The execution state indicates, for example, which cluster system executes the recovery operation for the server device in which a failure has occurred.
[0019] The management unit 11 may manage, for example, the monitoring state and the execution state for each cluster system. Specifically, the management unit 11 may manage flag information indicating the monitoring state and the execution state for each cluster system using a database.
[0020] The monitoring unit 12 monitors whether the server device is in a normal state or an abnormal state, reflects the monitoring results in the monitoring state, and reflects the monitoring results of the server device received from other cluster systems in the monitoring state.
[0021] The monitoring unit 12 may determine whether the server device is in a normal state or an abnormal state according to, for example, whether it can send a message to the server device and receive a response message. Or, when the server device is a DNS server device, the monitoring unit 12 may send a virtual host name to the DNS server device and determine whether the server device is normal or in an abnormal state according to whether it can receive address information for the virtual host name.
[0022] The monitoring unit 12 reflects the monitoring result in the monitoring state of the server device in the cluster system 10 managed by the management unit 11. Furthermore, the monitoring unit 12 receives the monitoring result of the server device from another cluster system different from the cluster system 10. That is, like the monitoring unit 12, the other cluster system also monitors the server device. When receiving the monitoring result, the monitoring unit 12 reflects it in the monitoring state of the server device in another cluster system managed by the management unit 11.
[0023] When the monitoring result in at least one of the plurality of cluster systems indicates an abnormal state, the determination unit 13 determines a cluster system that executes a recovery operation for the server device. The determination unit 13 determines a cluster system that executes a recovery operation for the server device in an abnormal state according to the same determination criteria used by other cluster systems that manage the monitoring state. When determining a cluster system that executes a recovery operation, the determination unit 13 reflects the determination result in the execution state managed by the management unit 11.
[0024] Each cluster system may monitor the server device using a different method. Therefore, there are a cluster system that can detect an abnormal state of the server device and a cluster system that cannot detect an abnormal state of the server device.
[0025] The determination criterion is a criterion that can uniquely determine the cluster system that executes the recovery operation. For example, in the determination criterion, the priority order of each cluster system is determined, and the determination unit 13 may determine the cluster system with the highest priority order as the cluster system that executes the recovery operation. A plurality of cluster systems have the same determination criterion. That is, a plurality of cluster systems share the same determination criterion.
[0026] The control unit 14 determines whether to execute a recovery operation on the server device according to the execution state. When it is shown in the execution state that the cluster system 10 executes the recovery operation, the control unit 14 executes the recovery operation on the server device. Also, when it is shown in the execution state that another cluster system executes the recovery operation, the control unit 14 does not execute the recovery operation on the server device.
[0027] As described above, the cluster system 10 manages the monitoring state of the server device in all cluster systems including the cluster system 10. Thereby, even when the cluster system 10 cannot detect an abnormal state of the server device, the cluster system 10 can grasp that an abnormal state of the server device has been detected in another cluster system.
[0028] Furthermore, the cluster system 10 determines the cluster system that executes the recovery operation for the server device in which an abnormal state has been detected, using the same determination criterion as the determination criterion that the other cluster systems have. Thereby, a plurality of cluster systems including the cluster system 10 can uniquely determine the cluster system that executes the recovery operation. As a result, it is possible to avoid the recovery operation for the server device in an abnormal state from being executed repeatedly by a plurality of cluster systems. That is, each cluster system can appropriately determine the cluster system that executes the recovery operation for the server device in an abnormal state.
[0029] (Embodiment 2) Next, a configuration example of the monitoring system according to Embodiment 2 will be described with reference to FIG. 2. The monitoring system of FIG. 2 includes a cluster system 10, a cluster system 20, a cluster system 30, and a shared server device 40. The cluster system 10, the cluster system 20, the cluster system 30, and the shared server device 40 may be included in, for example, one in-house system or the like.
[0030] The cluster system 10, the cluster system 20, the cluster system 30, and the shared server device 40 are connected via a network. The network may be, for example, an IP network. The cluster system 20 and the cluster system 30 have the same configuration as the cluster system 10. The shared server device 40 is a server device that is not subject to management for ensuring scalability or availability in the cluster system 10, the cluster system 20, and the cluster system 30. The shared server device 40 is managed by the cluster system 10, the cluster system 20, and the cluster system 30. The shared server device 40 may be, for example, a DNS server device.
[0031] For example, the cluster system 10 may obtain address information for identifying the cluster system 20 or 30 from the shared server device 40 operating as a DNS server device in order to access the cluster system 20 or 30. Accessing the cluster system 20 may mean accessing any computer device managed within the cluster system 20. Alternatively, accessing the cluster system 20 may mean accessing a computer device having a function of communicating with other cluster systems in the cluster system 20.
[0032] Next, the monitoring maps managed by the cluster system 10, the cluster system 20, and the cluster system 30 will be described with reference to FIG. 3. In the following, the monitoring map mainly managed by the cluster system 10 will be described, but the monitoring maps managed by the cluster system 20 and the cluster system 30 also have the same configuration as the monitoring map managed by the cluster system 10.
[0033] The cluster system 10 manages the monitoring map in the management unit 11. The monitoring map associates each cluster system with a monitoring state, an execution state, and an execution order. The numerical value set in the column of the cluster system indicates the identification information of the cluster system, indicating that the cluster system 10, the cluster system 20, and the cluster system 30 shown in FIG. 2 are managed in the monitoring map.
[0034] The numerical value set in the column of the execution order indicates the order in which the recovery operation is executed. The cluster system with 1 set is the cluster system that executes the recovery operation with the highest priority, and the cluster system with 3 set is the cluster system with the lowest priority.
[0035] The numerical value set in the monitoring state will be described with reference to FIG. 4. The numerical value set in the monitoring state may be regarded as flag information. FIG. 4 shows that there are parameters for normal, suspended, and abnormal as the monitoring states. Also, FIG. 4 shows that the flag indicating normal as the monitoring state is 0, the flag indicating suspended is 1, and the flag indicating abnormal is 2. Normal indicates that the shared server device 40 is not in an abnormal state, that is, no failure or malfunction has occurred in the shared server device 40. Suspended indicates that the monitoring of the shared server device 40 has been temporarily stopped. Abnormal indicates that the shared server device 40 is not normal, that is, a failure or malfunction has occurred in the shared server device 40.
[0036] Next, the numerical values set to the execution state will be described with reference to FIG. 5. The numerical values set to the execution state may be paraphrased as flag information. FIG. 5 shows that there are parameters in the states of not yet implemented, ready for execution, in execution, and completed as the execution state. Further, FIG. 5 shows that the flag indicating the state of not yet implemented is 0, the flag indicating ready for execution is 1, the flag indicating in execution is 2, and the flag indicating completed is 3 as the execution state. Not yet implemented indicates that the recovery operation for the shared server device 40 in the abnormal state is not executed. Ready for execution indicates that the preparation for executing the recovery operation for the shared server device 40 in the abnormal state is in progress. In execution indicates that the recovery operation for the shared server device 40 in the abnormal state is being executed. Completed indicates that the recovery operation for the shared server device 40 in the abnormal state has been completed.
[0037] Subsequently, with reference to FIGS. 6 and 7, the flow of the execution process of the recovery operation when only the cluster system 10 detects an abnormality in the shared server device 40 will be described. Further, with reference to FIG. 8, the transition of the values set in the monitoring map will be described. FIG. 8 shows that the execution order of the cluster system 10 is 1, the execution order of the cluster system 20 is 2, and the execution order of the cluster system 30 is 3. Further, FIG. 8 shows the steps in which the monitoring map is updated in FIGS. 6 and 7 in association with the flag information of the monitoring map.
[0038] First, the cluster system 10 detects that the shared server device 40 is in an abnormal state (S11). For example, when the cluster system 10 cannot obtain the address information corresponding to the virtual host name from the shared server device 40, it determines that the shared server device 40 is in an abnormal state.
[0039] Next, the cluster system 10 sends a message indicating that it has detected an abnormal state of the shared server device 40 to the cluster system 20 and the cluster system 30 (S12).
[0040] Next, the cluster systems 10, 20, and 30 update the monitoring status in the monitoring map (S13). For example, the cluster system 10 updates the monitoring map when it transmits a message indicating that an abnormal state has been detected. Also, the cluster systems 20 and 30 update the monitoring map when they receive a message indicating that an abnormal state has been detected. In FIG. 6, it is shown that the timings at which the cluster systems 10, 20, and 30 update the monitoring map are the same, but the monitoring map does not have to be updated at exactly the same timing. Similarly, in the following description, even if it is shown that the timings of the processes executed in the cluster systems 10, 20, and 30 are the same, they do not have to be exactly the same timing.
[0041] Specifically, the cluster systems 10, 20, and 30 set the monitoring status of the cluster system 10 to 2 as shown in the column of step S12 in the monitoring map of FIG. 8.
[0042] Also, in FIG. 6, although the cluster system 10 updates the monitoring map after transmitting the message, it may update the monitoring map when an abnormal state is detected in step S11 and before transmitting the message in step S12.
[0043] Next, the cluster systems 10, 20, and 30 execute a monitoring process for the shared server device 40 (S14). In FIG. 6, for the purpose of explaining an example in which only the cluster system 10 detects an abnormal state of the shared server device 40, it is assumed that the cluster systems 20 and 30 do not detect an abnormal state in step S14.
[0044] Next, the cluster system 20 sends a message including the monitoring result to the cluster system 10 and the cluster system 30 (S15). Further, the cluster system 30 sends a message indicating the monitoring result to the cluster system 10 and the cluster system 20 (S16). The cluster system 20 and the cluster system 30 send a message indicating that the shared server device 40 is normal. Also, FIG. 6 shows an example in which after the cluster system 20 sends a message in step S15, the cluster system 30 sends a message in step S16, but the order of steps S15 and S16 may be reversed. Alternatively, steps S15 and S16 may be executed at substantially the same timing.
[0045] Next, the cluster system 10, the cluster system 20, and the cluster system 20 update the monitoring state in the monitoring map (S17). The cluster system 10 reflects the monitoring results received from the cluster system 20 and the cluster system 30 in the monitoring state of the monitoring map. The cluster system 20 reflects the monitoring result in step S14 and the monitoring result received from the cluster system 30 in the monitoring state of the monitoring map. The cluster system 30 reflects the monitoring result in step S14 and the monitoring result received from the cluster system 20 in the monitoring state of the monitoring map.
[0046] Specifically, the cluster system 10, the cluster system 20, and the cluster system 30 have a monitoring map in the same state as the monitoring state in step S12, as shown in the column of step S17 in the monitoring map of FIG. 8.
[0047] Next, the cluster systems 10, 20, and 30 determine a cluster system that executes a recovery operation and update the execution status of the monitoring map (S18). The cluster systems 10, 20, and 30 determine a cluster system that executes a recovery operation from among the cluster systems that have detected an abnormal state. When a plurality of cluster systems have detected an abnormal state of the shared server device 40, the cluster systems 10, 20, and 30 determine a cluster system that executes a recovery operation according to the execution order. In FIG. 6, only the cluster system 10 has detected an abnormal state of the shared server device 40. Therefore, the cluster systems 10, 20, and 30 set the execution state of the cluster system 10 to 1 and update the execution status of the monitoring map, assuming that the cluster system 10 is preparing to execute a recovery operation.
[0048] Specifically, as shown in the column of step S18 of the monitoring map in FIG. 8, the cluster systems 10, 20, and 30 set the execution state of the cluster system 10 to 1. That is, the cluster systems 10, 20, and 30 assume that the cluster system 10 is preparing to execute a recovery operation.
[0049] Next, since the cluster system 20 does not execute the recovery operation, it sends a message indicating that the monitoring of the shared server device 40 is temporarily stopped to the cluster system 10 and the cluster system 30 (S19). Also, the cluster system 30 sends a message indicating that the monitoring of the shared server device 40 is temporarily stopped to the cluster system 10 and the cluster system 20 (S20). Steps S19 and S20 may be executed in the reverse order or may be performed at substantially the same timing. When the recovery operation is executed, the shared server device 40 may be restarted. In this case, if a cluster system that does not execute the recovery operation has been monitoring the shared server device 40, it may recognize that an abnormal state has occurred in the shared server device 40 and detect the abnormal state of the shared server device 40. Therefore, the cluster system that does not execute the recovery operation can avoid detecting an abnormal state regarding the shared server device 40 during the recovery operation by temporarily stopping the monitoring.
[0050] Next, the cluster system 10, the cluster system 20, and the cluster system 30 update the monitoring states of the cluster system 20 and the cluster system 30 in the monitoring map (S21). Specifically, the cluster system 10, the cluster system 20, and the cluster system 30 set the monitoring states of the cluster system 20 and the cluster system 30 to 1 as shown in the column of step S21 in the monitoring map of FIG. 8.
[0051] Next, the cluster system 10 sends a message indicating that the recovery operation is to start to the cluster system 20 and the cluster system 30 (S22).
[0052] Next, the cluster system 10, the cluster system 20, and the cluster system 30 update the execution state of the cluster system 10 in the monitoring map during execution (S23). Specifically, the cluster system 10, the cluster system 20, and the cluster system 30 set the execution state of the cluster system 10 to 2 as shown in the column of step S23 in the monitoring map of FIG. 8. Also, the cluster system 10 may set the execution state of the cluster system 10 to 2 before sending a message indicating that the recovery operation is started in step S22.
[0053] Next, the cluster system 10 executes a recovery operation on the shared server device 40 (S24). For example, the cluster system 10 may restart some applications that the shared server device 40 has, or may restart the shared server device 40. Next, the cluster system 10 completes the recovery operation on the shared server device 40 (S25).
[0054] Next, the cluster system 10 sends a message indicating that the recovery operation on the shared server device 40 has been completed to the cluster system 20 and the cluster system 30 (S26).
[0055] Next, the cluster system 10, the cluster system 20, and the cluster system 30 update the execution state of the cluster system 10 in the monitoring map to executed (S27). Specifically, the cluster system 10, the cluster system 20, and the cluster system 30 set the execution state of the cluster system 10 to 3 as shown in the column of step S27 in the monitoring map of FIG. 8. Also, the cluster system 10 may set the execution state of the cluster system 10 to 3 before sending a message indicating that the recovery operation has been completed in step S27.
[0056] Next, the cluster systems 10, 20, and 30 update the execution states of the cluster systems 10, 20, and 30 in the monitoring map to executed (S27). Specifically, the cluster systems 10, 20, and 30 set the execution state of the cluster system 10 to 3 as shown in the column of step S27 in the monitoring map of FIG. 8.
[0057] Next, the cluster systems 10, 20, and 30 execute monitoring of the shared server device 40 (S28). When the cluster systems 10, 20, and 30 determine that the shared server device 40 is operating normally, they reset the monitoring state and execution state in the monitoring map (S29). Specifically, the cluster systems 10, 20, and 30 set 0 for the monitoring state and execution state as shown in the column of step S29 in the monitoring map of FIG. 8.
[0058] Subsequently, the flow of the execution process of the recovery operation when the cluster systems 10 and 20 detect an abnormal state of the shared server device 40 will be described. For example, the case where the cluster system 10 first detects an abnormal state of the shared server device 40 and then the cluster system 20 explains the abnormal state of the shared server device 40 will be described.
[0059] The flow of the execution process of the recovery operation when the cluster systems 10 and 20 detect an abnormal state of the shared server device 40 is the same as that in FIGS. 6 and 7. Here, the difference from the case where the cluster system 10 detects an abnormal state will be described regarding the transition of the values set in the monitoring map when the cluster systems 10 and 20 detect an abnormal state of the shared server device 40.
[0060] Regarding the execution process flow of the recovery operation when the cluster system 10 and the cluster system 20 detect an abnormal state of the shared server device 40, steps S1 to S13 in FIG. 6 are the same as the case where only the cluster system 10 detects the abnormal state.
[0061] The cluster system 20 detects an abnormal state of the shared server device 40 in step S14 of FIG. 6. Further, in step S15, the cluster system 20 sends a message indicating that it has detected the abnormal state of the shared server device 40 to the cluster system 10 and the cluster system 30.
[0062] In this case, the cluster system 10, the cluster system 20, and the cluster system 30 set the monitoring states of the cluster system 10 and the cluster system 20 to 2 as shown in the column of step S17 in FIG. 9.
[0063] Next, the cluster system 10, the cluster system 20, and the cluster system 30 determine the cluster system that executes the recovery operation for the shared server device 40 in step S18. At the time of step S17, the cluster systems that detected the abnormal state of the shared server device 40 are the cluster system 10 and the cluster system 20. Also, since the execution order of 1 is set for the cluster system 10, the priority of the execution order is higher than that of the cluster system 20. Therefore, the cluster system 10, the cluster system 20, and the cluster system 30 update the execution state of the monitoring map of the cluster system 10 as the cluster system that executes the recovery operation for the shared server device 40.
[0064] Specifically, the cluster system 10, the cluster system 20, and the cluster system 30 set the execution state of the cluster system 10 to 1 as shown in the column of step S18 in the monitoring map of FIG. 9. That is, the cluster system 10, the cluster system 20, and the cluster system 30 assume that the cluster system 10 is preparing to execute the recovery operation.
[0065] For the steps after step S19, since the processing after step 19 is the same as that when only the cluster system 10 detects an abnormal state, detailed description thereof will be omitted.
[0066] Subsequently, a description will be given of the execution process flow of the recovery operation when the cluster system 10 and the cluster system 20 detect an abnormal state of the shared server device 40 and, furthermore, the shared server device 40 does not transition to the normal state in the recovery operation. In this case, the processing up to step S28 in FIGS. 6 and 7 is the same as the processing when the cluster system 10 and the cluster system 20 detect an abnormal state of the shared server device 40, and thus detailed description thereof will be omitted. Hereinafter, the processing after step S28 will be described with reference to FIGS. 10 and 11.
[0067] FIG. 10 shows the processing after step S28 in FIG. 7. When the cluster system 10 and the cluster system 20 execute monitoring of the shared server device 40 in step S28, they detect an abnormal state of the shared server device 40 (S31). That is, the cluster system 10 executed a recovery operation on the shared server device 40, but the abnormal state of the shared server device 40 has not been recovered.
[0068] Next, the cluster system 10 transmits a message indicating that the shared server device 40 is in an abnormal state to the cluster system 20 and the cluster system 30 (S32). Further, the cluster system 20 also transmits a message indicating that the shared server device 40 is in an abnormal state to the cluster system 10 and the cluster system 30 (S33). Also, the cluster system 30 that has not detected an abnormal state may transmit a monitoring result indicating that no abnormal state has been detected to the cluster system 10 and the cluster system 20.
[0069] Next, cluster systems 10, 20, and 30 update the monitoring status of cluster systems 10 and 20 in the monitoring map (S34). Specifically, cluster systems 10, 20, and 30 update from the state of the monitoring map shown in the column of step S27 in FIG. 9 to the state of the monitoring map shown in the column of step S34 in FIG. 12. Specifically, cluster systems 10, 20, and 30 update the monitoring status of cluster systems 10 and 20 in FIG. 12 to 2.
[0070] Next, cluster systems 10, 20, and 30 determine the cluster system that executes the recovery operation and update the execution status of the monitoring map (S35). In step S31, cluster systems 10 and 20 detect the abnormal state of the shared server device 40. Also, in the execution status in the column of step S34 in FIG. 12, 3 is set for cluster system 10, indicating that the recovery operation in cluster system 10 has been executed. Therefore, in step S35, cluster systems 10, 20, and 30 set cluster system 20, whose execution order is set to 2, as the cluster system that executes the recovery operation.
[0071] Specifically, cluster systems 10, 20, and 30 update the execution status of cluster system 20 in the column of step S35 in FIG. 12 to 1.
[0072] Next, since the cluster system 10 does not execute a recovery operation, it transmits a message indicating that the monitoring of the shared server device 40 is temporarily stopped to the cluster system 20 and the cluster system 30 (S36). Also, the cluster system 30 transmits a message indicating that the monitoring of the shared server device 40 is temporarily stopped to the cluster system 10 and the cluster system 20 (S37). Steps S36 and S37 may be executed in the reverse order or may be performed at substantially the same timing.
[0073] Next, the cluster system 10, the cluster system 20, and the cluster system 30 update the monitoring states of the cluster system 10 and the cluster system 30 in the monitoring map (S38). Specifically, the cluster system 10, the cluster system 20, and the cluster system 30 set the monitoring states of the cluster system 10 and the cluster system 30 to 1 as shown in the column of step S38 in the monitoring map of FIG. 12.
[0074] Next, the cluster system 20 transmits a message indicating that the recovery operation is started to the cluster system 10 and the cluster system 30 (S39).
[0075] Next, the cluster system 10, the cluster system 20, and the cluster system 30 update the execution state of the cluster system 10 in the monitoring map to in-execution (S40). Specifically, the cluster system 10, the cluster system 20, and the cluster system 30 set the execution state of the cluster system 20 to 2 as shown in the column of step S40 in the monitoring map of FIG. 12. Also, the cluster system 20 may set the execution state of the cluster system 20 to 2 before transmitting the message indicating that the recovery operation is started in step S39.
[0076] Next, the cluster system 20 executes a recovery operation on the shared server device 40. Next, the cluster system 20 completes the recovery operation on the shared server device 40 (S42).
[0077] Next, the cluster system 20 transmits a message indicating that the recovery operation for the shared server device 40 has been completed to the cluster system 10 and the cluster system 30 (S43).
[0078] Next, the cluster system 10, the cluster system 20, and the cluster system 30 update the execution state of the cluster system 20 in the monitoring map to "executed" (S44). Specifically, the cluster system 10, the cluster system 20, and the cluster system 30 set the execution state of the cluster system 20 to 3 as shown in the column of step S44 in the monitoring map of FIG. 12. Also, the cluster system 20 may set the execution state of the cluster system 20 to 3 before transmitting the message indicating that the recovery operation has been completed in step S43.
[0079] Next, the cluster system 10, the cluster system 20, and the cluster system 30 execute monitoring of the shared server device 40 (S45). When the cluster system 10, the cluster system 20, and the cluster system 30 determine that the shared server device 40 is operating normally, they reset the monitoring state and the execution state in the monitoring map (S46). Specifically, the cluster system 10, the cluster system 20, and the cluster system 30 set 0 for the monitoring state and the execution state as shown in the column of step S46 in the monitoring map of FIG. 12.
[0080] As described above, the monitoring maps held by the cluster system 10, the cluster system 20, and the cluster system 30 are the same. Also, the order in which the recovery operation is executed is defined in the monitoring map. Therefore, the cluster system 10, the cluster system 20, and the cluster system 30 can uniquely determine the cluster system that executes the recovery operation by using the monitoring map. Thus, the cluster system 10, the cluster system 20, and the cluster system 30 do not execute duplicate recovery operations on the shared server device 40 and can appropriately execute the recovery operation on the shared server device 40.
[0081] Furthermore, a cluster system that does not execute a recovery operation temporarily stops monitoring the shared server device 40. Thereby, a cluster system that does not execute a recovery operation can avoid detecting a server device that is executing a recovery operation as being in an abnormal state.
[0082] Also, in the monitoring system according to the second embodiment, since each cluster system has a monitoring map, a higher-level server device or a server device that becomes a leader is unnecessary. Thereby, it is possible to eliminate a sequence or the like until a leader determined in general distributed processing is executed, and it is possible to reduce the cost for installing a higher-level server device or the like.
[0083] FIG. 13 is a block diagram showing a configuration example of a cluster system 10 that operates as a single computer device. Referring to FIG. 13, the cluster system 10 includes a network interface 1201, a processor 1202, and a memory 1203. The network interface 1201 may be used to communicate with network nodes (e.g., eNB, MME, P-GW,). The network interface 1201 may include, for example, a network interface card (NIC) compliant with the IEEE 802.3 series. Here, eNB represents evolved Node B, MME represents Mobility Management Entity, and P-GW represents Packet Data Network Gateway. IEEE represents Institute of Electrical and Electronics Engineers.
[0084] The processor 1202 reads and executes software (computer program) from the memory 1203 to perform the processing of the cluster system 10 described using the flowchart in the above-described embodiment. The processor 1202 may be, for example, a microprocessor, MPU, or CPU. The processor 1202 may include a plurality of processors.
[0085] The memory 1203 is composed of a combination of a volatile memory and a non-volatile memory. The memory 1203 may include storage arranged separately from the processor 1202. In this case, the processor 1202 may access the memory 1203 via an I / O (Input / Output) interface (not shown).
[0086] In the example of FIG. 13, the memory 1203 is used to store a group of software modules. The processor 1202 can perform the processing of the cluster system 10 described in the above embodiments by reading out and executing these groups of software modules from the memory 1203.
[0087] As described with reference to FIG. 13, each of the processors included in the cluster system 10 in the above embodiments executes one or more programs including a group of instructions for causing a computer to perform the algorithms described with reference to the drawings.
[0088] In the above example, when the program is loaded into a computer, it includes a set of instructions (or software code) for causing the computer to perform one or more functions described in the embodiment. The program may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, the computer-readable medium or tangible storage medium includes random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD), or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray (registered trademark) disc, or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage devices. The program may also be transmitted on a transitory computer-readable medium or a communication medium. By way of example and not limitation, the transitory computer-readable medium or communication medium includes electrical, optical, acoustic, or other forms of propagated signals.
[0089] Note that the present disclosure is not limited to the above embodiments and can be appropriately modified without departing from the spirit.
[0090] Some or all of the above embodiments may be described as follows, but are not limited thereto. (Appendix 1) A management unit that manages the execution state indicating the monitoring state of the server device in a plurality of cluster systems and the execution state of a first cluster system that executes a recovery operation on the server device when the server device is in an abnormal state; A monitoring unit that monitors whether the server device is in a normal state or an abnormal state, reflects the monitoring result in the monitoring state, and reflects the monitoring result of the server device received from another cluster system in the monitoring state; When the monitoring result in at least one of the plurality of cluster systems indicates an abnormal state, a first cluster system that executes a recovery operation on the server device according to the same determination criteria used by the other cluster system that manages the monitoring state is determined, and a determination unit that reflects the determination result in the execution state; A cluster system including a control unit that determines whether to execute a recovery operation on the server device according to the managed execution state. (Appendix 2) The monitoring unit The cluster system according to Appendix 1, wherein when it is indicated in the execution state that the other cluster system executes a recovery operation on the server device, monitoring of the server device is stopped. (Appendix 3) The monitoring unit The cluster system according to Appendix 2, wherein the monitoring state of at least one second cluster system that does not execute a recovery operation on the server device is updated to information indicating that the monitoring of the server device is stopped. (Appendix 4) The determination criteria The cluster system according to any one of Appendices 1 to 3, which determines the priority of the first cluster system that executes the recovery operation. (Appendix 5) The determination unit The cluster system according to any one of Appendices 1 to 4, which determines the first cluster system that executes a recovery operation on the server device from among at least one third cluster system that has detected that the server device is in an abnormal state according to the determination criteria. (Appendix 6) The recovery operation The cluster system according to any one of Appendices 1 to 5, which is a restart of an application provided in the server device or a restart of the server device. (Appendix 7) The monitoring unit When the server device is a DNS server device, the DNS server device determines whether it is in a normal state or an abnormal state according to whether the address resolution of the virtual host name is successful. The cluster system according to any one of Appendices 1 to 6. (Appendix 8) A plurality of cluster systems, A monitoring system including a server device managed by the plurality of cluster systems, Each of the cluster systems, Manages the execution state indicating the first cluster system that monitors the monitoring state of the server device in the plurality of cluster systems and performs a recovery operation on the server device when the server device is in an abnormal state, Monitors whether the server device is in a normal state or an abnormal state, reflects the monitoring result in the monitoring state, and reflects the monitoring result of the server device received from another cluster system in the monitoring state, When the monitoring result in at least one of the plurality of cluster systems indicates an abnormal state, the first cluster system that performs a recovery operation on the server device according to the same determination criterion used by the other cluster system that manages the monitoring state is determined, and the determination result is reflected in the execution state, A monitoring system that determines whether to perform a recovery operation on the server device according to the managed execution state. (Appendix 9) Each of the cluster systems, When it is indicated in the execution state that the other cluster system performs a recovery operation on the server device, the monitoring of the server device is stopped. The monitoring system according to Appendix 8. (Appendix 10) Manages the execution state indicating the first cluster system that monitors the monitoring state of the server device in the plurality of cluster systems and performs a recovery operation on the server device when the server device is in an abnormal state, Monitors whether the server device is in a normal state or an abnormal state, Reflect the monitoring result in the monitoring state, and reflect the monitoring result of the server device received from another cluster system in the monitoring state. When the monitoring result in at least one of the plurality of cluster systems indicates an abnormal state, determine the first cluster system that executes a recovery operation on the server device according to the same determination criteria used by the other cluster system that manages the monitoring state. Reflect the determination result in the execution state. A monitoring method executed in a cluster system that determines whether to execute a recovery operation on the server device according to the managed execution state. (Appendix 11) Manage the monitoring state of the server device in a plurality of cluster systems and the execution state indicating the first cluster system that executes a recovery operation on the server device when the server device is in an abnormal state. Monitor whether the server device is in a normal state or an abnormal state. Reflect the monitoring result in the monitoring state, and reflect the monitoring result of the server device received from another cluster system in the monitoring state. When the monitoring result in at least one of the plurality of cluster systems indicates an abnormal state, determine the first cluster system that executes a recovery operation on the server device according to the same determination criteria used by the other cluster system that manages the monitoring state. Reflect the determination result in the execution state. A program that causes a computer to determine whether to execute a recovery operation on the server device according to the managed execution state.
Explanation of Signs
[0091] 10 Cluster system 11 Management unit 12 Monitoring unit 13 Determination unit 14 Control unit 20 Cluster system 30 Cluster System 40 Shared Server Device
Claims
1. A management unit that manages the execution state indicating the monitoring state of a server device in a plurality of cluster systems and the execution state of a first cluster system that performs a recovery operation on the server device when the server device is in an abnormal state; A monitoring unit that monitors whether the server device is in a normal state or an abnormal state, reflects the monitoring result in the monitoring state, and reflects the monitoring result of the server device received from another cluster system in the monitoring state; When the monitoring result in at least one of the plurality of cluster systems indicates an abnormal state, a determination unit that determines the first cluster system that performs a recovery operation on the server device according to the same determination criteria used by the other cluster system that manages the monitoring state, and reflects the determination result in the execution state; A control unit that determines whether to perform a recovery operation on the server device according to the managed execution state, a cluster system comprising the same.
2. The monitoring unit, When it is shown in the execution state that the other cluster system performs a recovery operation on the server device, stops monitoring the server device, the cluster system according to claim 1.
3. The monitoring unit, Updates the monitoring state of at least one second cluster system that does not perform a recovery operation on the server device with information indicating that the monitoring of the server device has been stopped, the cluster system according to claim 2.
4. The determination criteria are Determine the priority order of the first cluster system that performs the recovery operation, the cluster system according to any one of claims 1 to 3.
5. The determination unit, Among the plurality of cluster systems, determines the first cluster system that performs a recovery operation on the server device according to the determination criteria from among at least one first cluster system that has detected that the server device is in an abnormal state, the cluster system according to any one of claims 1 to 4.
6. The recovery operation is Restarting the application provided in the server device or restarting the server device, the cluster system according to any one of claims 1 to 5.
7. The monitoring unit, When the server device is a DNS server device, the DNS server device determines whether it is in a normal state or an abnormal state according to whether the address resolution of the virtual host name is successful. The cluster system according to any one of claims 1 to 6.
8. A monitoring system including a plurality of cluster systems and a server device managed by the plurality of cluster systems, wherein each of the cluster systems manages an execution state indicating a first cluster system that executes a recovery operation on the server device when the monitoring state of the server device in the plurality of cluster systems and the server device is in an abnormal state, monitors whether the server device is in a normal state or an abnormal state, reflects the monitoring result in the monitoring state, and reflects the monitoring result of the server device received from another cluster system in the monitoring state, when the monitoring result in at least one of the plurality of cluster systems indicates an abnormal state, determines a first cluster system that executes a recovery operation on the server device according to the same determination criterion used by the other cluster system that manages the monitoring state, reflects the determination result in the execution state, and determines whether to execute a recovery operation on the server device according to the managed execution state. The monitoring system.
9. A cluster system, manages an execution state indicating a first cluster system that executes a recovery operation on the server device when the monitoring state of the server device in the plurality of cluster systems and the server device is in an abnormal state, monitors whether the server device is in a normal state or an abnormal state, reflects the monitoring result in the monitoring state, and reflects the monitoring result of the server device received from another cluster system in the monitoring state, when the monitoring result in at least one of the plurality of cluster systems indicates an abnormal state, determines a first cluster system that executes a recovery operation on the server device according to the same determination criterion used by the other cluster system that manages the monitoring state, reflects the determination result in the execution state, and determines whether to execute a recovery operation on the server device according to the managed execution state. The monitoring method executed in the cluster system.
10. Manages the execution state indicating the monitoring state of a server device in a plurality of cluster systems and the recovery operation for the server device when the server device is in an abnormal state, monitors whether the server device is in a normal state or an abnormal state, reflects the monitoring result in the monitoring state, and also reflects the monitoring result of the server device received from another cluster system in the monitoring state, when the monitoring result in at least one of the plurality of cluster systems indicates an abnormal state, determines the first cluster system that executes a recovery operation for the server device according to the same determination criteria used by the other cluster system that manages the monitoring state, reflects the determination result in the execution state, A program that causes a computer to determine whether to execute a recovery operation for the server device according to the managed execution state.
Citation Information
Patent Citations
Backup method for decentralized control system
JP1997244910A
Shared resource failure detection system and method
JP2004287980A
Information processor, monitoring method, and program
JP2007304837A
Mutual monitoring system
JP2012168907A
Operation service provision system, method for recovering operation service, and operation service recovery program
JP2020135287A