A cluster node failure processing method, device, equipment and medium
By introducing a new heartbeat sending mechanism into the CTDB cluster, the misjudgment problem caused by the master node failure is solved, and the cluster is highly available and reliable.
Patent Information
- Application Number
- CN202210879795.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-07-25
AI Technical Summary
In a CTDB cluster, the failure of the main node leads to incorrect judgment, and the wrong judgment of the normal node is a failed node, resulting in waste of resources and cluster operation failure.
By modifying the processing logic inside the CTDB, a new heartbeat sending mechanism is introduced. Each node synchronizes the heartbeat interaction data between nodes when sending heartbeat data. When the main node cannot receive the heartbeat data, it judges the faulty node by the saved interactive data.
Avoid misjudgment caused by master node failure, maintain the normal state of the cluster, quickly remove the faulty nodes, and ensure high availability and reliability of the cluster.
Smart Images

Figure CN115242820B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of distributed storage, and in particular to a cluster node failure processing method, device, equipment and medium. Background Art
[0002] CTDB (Cluster Trivial Database) is a lightweight cluster database implementation and the cluster database component of cluster Samba. It is primarily used to handle Samba's cross-node messages and implement a distributed TDB database across all cluster nodes. CTDB's main functions include: providing a clustered version of the TDB database and automatically rebuilding and recovering the database in the event of a node failure; monitoring the nodes in the cluster and the services running on each node; and managing the public IP address pool used to provide services to clients.
[0003] In the current CTDB high-availability solution, each node uses heartbeats to determine if its peer is abnormal. Specifically, in a CTDB cluster, all nodes, including the master node, periodically send heartbeats to each other. When a node receives a heartbeat from another node, it assumes that the peer and the cluster are both functioning normally and continues to operate. If a node fails to receive a heartbeat from another node, it assumes that the peer is faulty. When a fault occurs, only the master node can handle it. For example, if master node A fails to receive heartbeat data from node B within a preset time, master node A will identify node B as faulty and remove it from the cluster. However, in some cases, master node A's failure to receive heartbeat data from node B may be due to a fault at master node A. In this case, master node A mistakenly identifies node B as faulty and removes it from the cluster, resulting in unnecessary resource waste. The master node will then continue to operate in a faulty state, potentially causing cluster failure and service interruption, with potentially serious consequences.
[0004] As can be seen from the above, during the operation of the CTDB cluster, how to avoid the situation where the master node failure causes misjudgment, which may lead to cluster operation failure, is a problem to be solved in this field. Summary of the Invention
[0005] In view of this, the present invention aims to provide a cluster node failure handling method, apparatus, device, and medium. By modifying the internal processing logic of the CTDB, the method prevents the removal of healthy nodes when a master node misjudges the failure, maintaining the cluster's normal state and enabling normal service provision. Furthermore, when a node is determined to be faulty, the method can be quickly removed, ensuring high cluster availability and improving cluster reliability. The specific solution is as follows:
[0006] In a first aspect, the present application discloses a cluster node failure handling method, which is applied to a master node in a cluster trivial database, comprising:
[0007] If the heartbeat data and the inter-node heartbeat interaction data sent by the first target node in the cluster trivial database are not received within the first preset time, querying the inter-node heartbeat interaction data of the second target node from the node interaction record in which the inter-node interaction data of each node is locally stored; the node interaction record includes the heartbeat interaction record between the first target node and the second target node in the cluster trivial database;
[0008] Determine whether, in the inter-node heartbeat interaction data of the second target node, there is a heartbeat record between the first target node and the second target node within the first preset time;
[0009] If, in the inter-node heartbeat interaction data of the second target node, there is a heartbeat record between the first target node and the second target node within the first preset time, setting the current node state to a suspended state;
[0010] If there is no heartbeat record between the first target node and the second target node within the first preset time in the inter-node heartbeat interaction data of the second target node, the first target node is determined to be a faulty node, and the faulty node is removed from the cluster.
[0011] Optionally, before querying the inter-node heartbeat interaction data of the second target node from the node interaction record having the inter-node interaction data of each node stored locally, the method further includes:
[0012] Receive heartbeat data sent by other nodes in the cluster trivial database and inter-node heartbeat interaction data;
[0013] The received heartbeat data and the heartbeat interaction data between nodes are saved in the node interaction record corresponding to the local preset storage location.
[0014] Optionally, the step of saving the received heartbeat data and the inter-node heartbeat interaction data to a node interaction record corresponding to a local preset storage location includes:
[0015] The received heartbeat data and the heartbeat interaction data between nodes, wherein the data sending time is within the preset time period, are saved in the node interaction record corresponding to the local preset storage location.
[0016] Optionally, querying the inter-node heartbeat interaction data of the second target node from the node interaction record locally storing the inter-node interaction data of each node includes:
[0017] determining a second target node from the cluster trivial database;
[0018] The inter-node heartbeat interaction data of the second target node is queried from the node interaction record in which the inter-node interaction data of each node is stored locally.
[0019] Optionally, determining the second target node from the cluster trivial database includes:
[0020] Determine a target node from the cluster trivial database;
[0021] Determine whether the heartbeat data and the inter-node heartbeat interaction data sent by the target node are received within the preset heartbeat data sending time;
[0022] If the heartbeat data and the inter-node heartbeat interaction data sent by the target node are received within the preset heartbeat data sending time, the target node is determined as the second target node.
[0023] Optionally, after determining whether the heartbeat data and the inter-node heartbeat interaction data sent by the target node are received within the preset heartbeat data sending time, the method further includes:
[0024] If the heartbeat data sent by the target node and the heartbeat interaction data between nodes are not received within the preset heartbeat data sending time, the target node is re-determined from the cluster trivial database, and then the step of determining whether the heartbeat data sent by the target node and the heartbeat interaction data between nodes are received within the preset heartbeat data sending time is executed.
[0025] Optionally, after determining the target node, the method further includes:
[0026] Record the number of current historical target nodes;
[0027] Determining whether the number of the historical target nodes meets a preset threshold;
[0028] If the historical target node meets the preset threshold, the current node state is set to a suspended state.
[0029] In a second aspect, the present application discloses a cluster node fault handling device, which is applied to a master node in a cluster trivial database, comprising:
[0030] A record query module is configured to query the inter-node heartbeat interaction data of the second target node from the node interaction record in which the inter-node interaction data of each node is locally stored if the heartbeat data and the inter-node heartbeat interaction data sent by the first target node in the cluster trivial database are not received within a first preset time; the node interaction record includes the heartbeat interaction record between the first target node and the second target node in the cluster trivial database;
[0031] A record judgment module is used to judge whether there is a heartbeat record between the first target node and the second target node within the first preset time in the inter-node heartbeat interaction data of the second target node;
[0032] a record existence module for setting the current node state to a suspended state if, in the inter-node heartbeat interaction data of the second target node, there is a heartbeat record of the first target node and the second target node within the first preset time;
[0033] The record does not exist module is used to determine the first target node as a faulty node and remove the faulty node from the cluster if there is no heartbeat record between the first target node and the second target node in the inter-node heartbeat interaction data of the second target node within the first preset time.
[0034] In a third aspect, the present application discloses an electronic device, comprising:
[0035] Memory, used to store computer programs;
[0036] The processor is configured to execute the computer program to implement the aforementioned cluster node failure handling method.
[0037] In a fourth aspect, the present application discloses a computer storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the steps of the aforementioned disclosed cluster node failure handling method are implemented.
[0038] In this application, if the heartbeat data and inter-node heartbeat interaction data sent by the first target node in the cluster trivial database are not received within the first preset time, the inter-node heartbeat interaction data of the second target node is queried from the node interaction record of the inter-node interaction data of each node stored locally; the inter-node interaction data includes the heartbeat interaction record of the first target node and the second target node in the cluster trivial database; it is judged whether there is a heartbeat record of the first target node and the second target node in the inter-node heartbeat interaction data of the second target node within the first preset time; if there is a heartbeat record of the first target node and the second target node in the inter-node heartbeat interaction data of the second target node within the first preset time, the current node state is set to a suspended state; if there is no heartbeat record of the first target node and the second target node in the inter-node heartbeat interaction data of the second target node within the first preset time, the first target node is determined as a faulty node, and the faulty node is removed from the cluster. The present invention proposes a novel heartbeat sending mechanism, that is, when each node sends heartbeat data, it synchronously sends the heartbeat interaction data between the node and other nodes, so that each node can save the heartbeat interaction data between other nodes. When the master node cannot receive the heartbeat data of a certain node, it can judge which node is the faulty node through the heartbeat interaction data saved in each node. Finally, after the faulty node is determined, the faulty node is removed. In this way, the situation in the prior art where a normal node is mistakenly judged as a faulty node due to the failure of the master node and the normal node is removed from the cluster can be avoided, thereby ensuring the high availability of the cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0040] Figure 1 A flow chart of a cluster node failure handling method provided in this application;
[0041] Figure 2 A flowchart of a specific cluster node failure handling method provided in this application;
[0042] Figure 3 A schematic diagram of the structure of a cluster node fault handling device provided by this application;
[0043] Figure 4 This is a structural diagram of an electronic device provided in this application. DETAILED DESCRIPTION
[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0045] In the prior art, in a CTDB cluster, all nodes, including the master node, will periodically send heartbeat data to other nodes. When the master node A does not receive the heartbeat data from a certain node B within a preset time, the node B will be determined as a faulty node and removed from the cluster. However, the reason why the master node A does not receive the heartbeat data from B may be that the master node A is faulty. In order to avoid such a situation where the master node is misjudged due to its own fault, the present invention proposes a cluster node fault handling method, which can prevent the normal nodes from being removed when the master node is misjudged, maintain the normal state of the cluster, and provide services normally; when a node is determined to be faulty, it can quickly be removed to ensure the high availability of the cluster and improve the reliability of the cluster.
[0046] The embodiment of the present invention discloses a cluster node failure processing method, which is applied to the master node in the cluster trivial database. Figure 1 Said method comprises:
[0047] Step S11: If the heartbeat data and inter-node heartbeat interaction data sent by the first target node in the cluster trivial database are not received within the first preset time, the inter-node heartbeat interaction data of the second target node is queried from the node interaction record in which the inter-node interaction data of each node is locally stored; the inter-node interaction data includes the heartbeat interaction record between the first target node and the second target node in the cluster trivial database.
[0048] The cluster trivial database described in this embodiment is CTDB (i.e., Cluster Trivial Database). In the present invention, a new type of heartbeat data sending mechanism is pre-set for each node. When each node sends heartbeat data to a certain node, it will also send the heartbeat interaction data of the node and other nodes at the same time. For example, A is the master node in the CTDB cluster, and B, C, and D are all ordinary nodes. In the existing heartbeat mechanism, each node will send heartbeat data to other nodes in the cluster except itself, that is, node A sends heartbeat data to nodes B, C, and D; node B sends heartbeat data to nodes A, C, and D; node C sends heartbeat data to nodes A, B, and D; and node D sends heartbeat data to nodes A, B, and C. In the new heartbeat data sending mechanism proposed in this application, when a node sends heartbeat data to other nodes in the cluster except itself, it will synchronously send the heartbeat interaction data between itself and other nodes. That is, when node B sends heartbeat data to node A, the heartbeat interaction data between node B and node C, and node B and node D will also be sent to node A. When node A receives the heartbeat data sent by node B, it will also record the heartbeat interaction data between node B and node C, and node B and node D.
[0049] In this embodiment, before querying the inter-node heartbeat interaction data of the second target node from the node interaction record that locally stores the inter-node interaction data of each node, it can also include: receiving the heartbeat data and inter-node heartbeat interaction data sent by other nodes in the cluster trivial database; and saving the received heartbeat data and inter-node heartbeat interaction data to the node interaction record corresponding to the local preset storage location.
[0050] It is understood that the first preset time is greater than the frequency of heartbeat data transmission between nodes. In this embodiment, the value of the first preset time can be changed based on the specific real-time scenario. For example, when the heartbeat data transmission frequency is once per second, the first preset time can be set to three seconds. That is, each node will send data at a frequency of once per second. If the master node does not receive data sent by a node within three seconds, it indicates that the node or the master node has failed.
[0051] In the present embodiment, when each node sends heartbeat interaction data between nodes, it can select data within a preset time period to send, or it can send all historical heartbeat interaction data. In some preferred implementations, it can select historical heartbeat interaction data with other nodes within the past 32 seconds to send.
[0052] In this embodiment, the step of saving the received heartbeat data and the heartbeat interaction data between nodes to the node interaction record corresponding to the local preset storage location may include: saving the data in the received heartbeat data and the heartbeat interaction data between nodes whose data sending time is within the preset time period to the node interaction record corresponding to the local preset storage location.
[0053] In a specific embodiment, after node A receives the heartbeat data sent by node B and the heartbeat interaction data between node B and node C, and between node B and node D, it saves these heartbeat interaction data to the node interaction record stored in the local preset storage location. In this embodiment, when each node saves the inter-node heartbeat interaction data sent by the opposite node, it can select data within a preset time period for saving. In some preferred embodiments, the preset time period can be historical heartbeat interaction data within the past 32 seconds.
[0054] In a specific implementation, when each node receives heartbeat data sent by the peer node, it will also update the heartbeat record data corresponding to the peer node recorded in the node to ensure that the data recorded in the node always meets the preset time period.
[0055] In this embodiment, querying the inter-node heartbeat interaction data of the second target node from the node interaction records that locally store the inter-node interaction data of each node may include: determining the second target node from the cluster trivial database; querying the inter-node heartbeat interaction data of the second target node from the node interaction records that locally store the inter-node interaction data of each node.
[0056] In this embodiment, since the interaction records between nodes have been pre-saved in the node interaction records, when the main node A does not receive the heartbeat data sent by node B within the preset first time, it will check the locally saved node interaction records. If C is determined to be the second target node at this time, it will determine whether node B is operating normally from the node interaction data between node B and the second target node C in the node interaction record.
[0057] Step S12: Determine whether there is a heartbeat record between the first target node and the second target node within the first preset time in the inter-node heartbeat interaction data of the second target node.
[0058] In this embodiment, since the master node does not receive the heartbeat data of the first target node within the first preset time, when querying the node interaction record, it will determine whether the first target node has heartbeat interaction data within the first preset time to determine whether the first target node is abnormal.
[0059] Step S13: If, in the inter-node heartbeat interaction data of the second target node, there is a heartbeat record between the first target node and the second target node within the first preset time, the current node state is set to a suspended state.
[0060] In this embodiment, if the node interaction record contains heartbeat interaction data of the first target node and the second target node within the first preset time, it indicates that normal heartbeat interaction exists between the first target node and the second target node within the first preset time, and the first target node is in a normal state. At this time, it can be determined that the master node is in a sub-healthy state, and the master node will set its own state to suspended. After that, other modules in the cluster will shut down the network port of the master node and perform node isolation to ensure the normal operation of the cluster.
[0061] Step S14: If there is no heartbeat record between the first target node and the second target node in the inter-node heartbeat interaction data of the second target node within the first preset time, the first target node is determined as a faulty node, and the faulty node is removed from the cluster.
[0062] In this embodiment, if there is no heartbeat interaction data between the first target node and the second target node in the node interaction record within the first preset time, it indicates that there is no normal heartbeat interaction between the first target node and the second target node within the first preset event, and the first node is faulty. At this time, the master node will remove the faulty node from the cluster to ensure the normal operation of the cluster.
[0063] In this embodiment, if the heartbeat data and inter-node heartbeat interaction data sent by the first target node in the cluster trivial database are not received within the first preset time, the inter-node heartbeat interaction data of the second target node is queried from the node interaction record of the inter-node interaction data of each node stored locally; the inter-node interaction data includes the heartbeat interaction record of the first target node and the second target node in the cluster trivial database; it is determined whether there is a heartbeat record of the first target node and the second target node in the inter-node heartbeat interaction data of the second target node within the first preset time; if there is a heartbeat record of the first target node and the second target node in the inter-node heartbeat interaction data of the second target node within the first preset time, the current node state is set to a suspended state; if there is no heartbeat record of the first target node and the second target node in the inter-node heartbeat interaction data of the second target node within the first preset time, the first target node is determined as a faulty node, and the faulty node is removed from the cluster. The present invention proposes a novel heartbeat sending mechanism, that is, when each node sends heartbeat data, it synchronously sends the heartbeat interaction data between the node and other nodes, so that each node can save the heartbeat interaction data between other nodes. When the master node cannot receive the heartbeat data of a certain node, it can judge which node is the faulty node through the heartbeat interaction data saved in each node. Finally, after the faulty node is determined, the faulty node is removed. In this way, the situation in the prior art where a normal node is mistakenly judged as a faulty node due to the failure of the master node and the normal node is removed from the cluster can be avoided, thereby ensuring the high availability of the cluster.
[0064] Figure 2 This is a flow chart of a specific cluster node failure handling method provided by the embodiment of this application. Figure 2 As shown, the method includes:
[0065] Step S21: If the heartbeat data and inter-node heartbeat interaction data sent by the first target node in the cluster trivial database are not received within the first preset time, the second target node is determined from the cluster trivial database, and the inter-node heartbeat interaction data of the second target node is queried from the node interaction records of the inter-node interaction data of each node stored locally; the inter-node interaction data includes the heartbeat interaction records between the first target node and the second target node in the cluster trivial database.
[0066] In this embodiment, determining the second target node from the cluster trivial database may include: determining the target node from the cluster trivial database; judging whether the heartbeat data sent by the target node and the heartbeat interaction data between nodes are received within the preset heartbeat data sending time; if the heartbeat data sent by the target node and the heartbeat interaction data between nodes are received within the preset heartbeat data sending time, then determining the target node as the second target node.
[0067] In this embodiment, after determining whether the heartbeat data and the heartbeat interaction data between nodes are received within the preset heartbeat data sending time, the step may also include: if the heartbeat data and the heartbeat interaction data sent by the target node are not received within the preset heartbeat data sending time, re-determining the target node from the cluster trivial database, and then executing the step of determining whether the heartbeat data and the heartbeat interaction data sent by the target node are received within the preset heartbeat data sending time.
[0068] In this embodiment, when determining the second target node, the target node will be determined from the cluster first. For example, in some specific implementations, if the preset heartbeat data sending frequency in the cluster is to send heartbeat data once per second, the preset heartbeat data sending time is one second, the first preset time is three seconds, and when the master node A does not receive the heartbeat data and the heartbeat interaction data between nodes sent by the first target node B within three seconds, the master node will first determine the target node from the remaining nodes in the cluster. If the target node is determined to be C, it will be determined whether the heartbeat data and the heartbeat interaction data sent by the target node C are received within the previous second of the current time. If not, the target node will be re-determined to be D, and it will be determined whether the heartbeat data and the heartbeat interaction data sent by the target node D are received within the previous second of the current time. If received, it indicates that D is a node in a normal state, and D will be determined as the second target node at this time.
[0069] In this embodiment, after determining the target node, it can also include: recording the number of current historical target nodes; judging whether the number of historical target nodes meets the preset threshold; if the historical target nodes meet the preset threshold, setting the current node state to a suspended state.
[0070] In this embodiment, if the target node is determined multiple times, that is, in a possible implementation, if the heartbeat data and the heartbeat interaction data between the nodes are not received within the first second of the current time, the second target node will be re-determined to be E, and then it will be determined whether the heartbeat data and the heartbeat interaction data between the nodes are received within the first second of the current time, and finally it will be determined whether the target node needs to be re-determined. It should be pointed out that in this embodiment, in the process of repeatedly determining the target node, each time the target node is determined, the number of times the target node is determined will be recorded. If the number of times the target node is determined meets the preset threshold, the master node will be set to a suspended state. In some preferred implementations, in this case, after the master node is suspended, a cluster operation error message will be returned to the corresponding display interface. That is, in the process of determining the second target node, if the preset threshold nodes in the cluster do not send heartbeat data and the heartbeat interaction data between the nodes within the preset heartbeat sending time, it indicates that a large-scale node failure or a master node failure may have occurred in the cluster. At this time, the master node will be suspended and a cluster operation error message will be returned.
[0071] Step S22: Determine whether there is a heartbeat record between the first target node and the second target node within the first preset time in the inter-node heartbeat interaction data of the second target node.
[0072] In this embodiment, if the target node is in a normal heartbeat sending state, the target node will be determined as the second target node, and the operation state of the first target node will be determined using the interaction data between the second target node and the first target node.
[0073] Step S23: If, in the inter-node heartbeat interaction data of the second target node, there is a heartbeat record between the first target node and the second target node within the first preset time, the current node state is set to a suspended state.
[0074] For a more specific processing procedure of step S23, reference may be made to the corresponding contents disclosed in the aforementioned embodiment, which will not be described in detail here.
[0075] Step S24: If there is no heartbeat record between the first target node and the second target node in the inter-node heartbeat interaction data of the second target node within the first preset time, the first target node is determined as a faulty node, and the faulty node is removed from the cluster.
[0076] For a more specific processing procedure of step S24, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be elaborated here.
[0077] This embodiment proposes a detailed process for determining the second target node. Specifically, the target node is first determined. If no data is received from the target node within a preset heartbeat data transmission time, the target node is re-determined. If the number of determined target nodes meets a preset threshold, it indicates that a large-scale node failure or a master node failure may have occurred in the cluster. At this time, the master node will hang and return a cluster operation error message. In this embodiment, the target node in normal operation is ultimately determined as the second target node. The interactive data between the second target node and the first target node is then used to determine the operating status of the first target node, thereby improving the security of the cluster and ensuring its stable operation.
[0078] See also Figure 3 As shown, the embodiment of the present application discloses a cluster node fault processing device, which is applied to the master node in the cluster trivial database, and specifically may include:
[0079] A record query module 11 is configured to query the inter-node heartbeat interaction data of the second target node from the node interaction record locally storing the inter-node interaction data of each node if the heartbeat data and inter-node heartbeat interaction data sent by the first target node in the cluster trivial database are not received within a first preset time; the node interaction record includes the heartbeat interaction record between the first target node and the second target node in the cluster trivial database;
[0080] A record determination module 12 is configured to determine whether, in the inter-node heartbeat interaction data of the second target node, there is a heartbeat record between the first target node and the second target node within the first preset time;
[0081] A record existence module 13 is configured to set the current node state to a suspended state if, in the inter-node heartbeat interaction data of the second target node, there is a heartbeat record between the first target node and the second target node within the first preset time;
[0082] The record does not exist module 14 is used to determine the first target node as a faulty node and remove the faulty node from the cluster if there is no heartbeat record between the first target node and the second target node in the inter-node heartbeat interaction data of the second target node within the first preset time.
[0083] In the present invention, if the heartbeat data and the inter-node heartbeat interaction data sent by the first target node in the cluster trivial database are not received within the first preset time, the inter-node heartbeat interaction data of the second target node is queried from the node interaction record of the inter-node interaction data of each node stored locally; the inter-node interaction data includes the heartbeat interaction record of the first target node and the second target node in the cluster trivial database; it is judged whether there is a heartbeat record of the first target node and the second target node in the inter-node heartbeat interaction data of the second target node within the first preset time; if there is a heartbeat record of the first target node and the second target node in the inter-node heartbeat interaction data of the second target node within the first preset time, the current node state is set to a suspended state; if there is no heartbeat record of the first target node and the second target node in the inter-node heartbeat interaction data of the second target node within the first preset time, the first target node is determined as a faulty node, and the faulty node is removed from the cluster. The present invention proposes a novel heartbeat sending mechanism, that is, when each node sends heartbeat data, it synchronously sends the heartbeat interaction data between the node and other nodes, so that each node can save the heartbeat interaction data between other nodes. When the master node cannot receive the heartbeat data of a certain node, it can judge which node is the faulty node through the heartbeat interaction data saved in each node. Finally, after the faulty node is determined, the faulty node is removed. In this way, the situation in the prior art where a normal node is mistakenly judged as a faulty node due to the failure of the master node and the normal node is removed from the cluster can be avoided, thereby ensuring the high availability of the cluster.
[0084] In some specific embodiments, the cluster node failure handling device further includes:
[0085] A data receiving module is used to receive heartbeat data sent by other nodes in the cluster trivial database and heartbeat interaction data between nodes;
[0086] The data saving module is used to save the received heartbeat data and the heartbeat interaction data between nodes to the node interaction record corresponding to the local preset storage location.
[0087] In some specific embodiments, the data storage module specifically includes:
[0088] The data storage unit is used to save the received heartbeat data and the heartbeat interaction data between nodes, and the data whose data sending time is within the preset time period, into the node interaction record corresponding to the local preset storage location.
[0089] In some specific embodiments, the record query module 11 includes:
[0090] A second target node determining unit, configured to determine a second target node from the cluster trivial database;
[0091] The data query unit is used to query the inter-node heartbeat interaction data of the second target node from the node interaction record in which the inter-node interaction data of each node is stored locally.
[0092] In some specific embodiments, the second target node determination unit specifically includes:
[0093] a target node determining unit, configured to determine a target node from the cluster trivial database;
[0094] A data judgment unit is used to judge whether the heartbeat data sent by the target node and the heartbeat interaction data between nodes are received within a preset heartbeat data sending time;
[0095] The node determination unit is configured to determine the target node as a second target node if the heartbeat data and the inter-node heartbeat interaction data sent by the target node are received within a preset heartbeat data sending time.
[0096] In some specific embodiments, the second target node determination unit further includes:
[0097] The node unit is re-determined, and is used to re-determine the target node from the cluster trivial database if the heartbeat data and the heartbeat interaction data between nodes sent by the target node are not received within the preset heartbeat data sending time, and then execute the step of determining whether the heartbeat data and the heartbeat interaction data sent by the target node are received within the preset heartbeat data sending time.
[0098] In some specific embodiments, the cluster node fault handling device further includes:
[0099] Number recording unit, used to record the number of current historical target nodes;
[0100] A number determination unit, configured to determine whether the number of the historical target nodes meets a preset threshold;
[0101] The node abnormality unit is used to set the current node state to a suspended state if the historical target node meets a preset threshold.
[0102] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation to the scope of application of the present application.
[0103] Figure 4This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a display 24, an input / output interface 25, a communication interface 26, and a communication bus 27. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the cluster node failure handling method disclosed in any of the aforementioned embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0104] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 26 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0105] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, a magnetic disk, or an optical disk. The resources stored thereon may include an operating system 221, a computer program 222, and virtual machine data 223. The virtual machine data 223 may include various data. The storage method may be temporary storage or permanent storage.
[0106] The operating system 221 is used to manage and control the hardware devices on the electronic device 20 and the computer program 222, and can be Windows Server, NetWare, Unix, Linux, etc. In addition to including computer programs that can be used to implement the cluster node failure handling method executed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include computer programs that can be used to perform other specific tasks.
[0107] Furthermore, the present application also discloses a computer-readable storage medium, where the computer-readable storage medium includes a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a magnetic disk, or an optical disk, or any other form of storage medium known in the technical field. When the computer program is executed by a processor, the aforementioned cluster node fault handling method is implemented. For the specific steps of the method, reference can be made to the corresponding content disclosed in the aforementioned embodiments, and no further details will be given here.
[0108] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. Professionals can also further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0109] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0110] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0111] The above is a detailed introduction to the cluster node fault handling method, device, equipment, and storage medium provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for general technical personnel in this field, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A cluster node failure processing method, characterized in that: Applies to the master node in the cluster's trivial database, including: If the heartbeat data and the inter-node heartbeat interaction data sent by the first target node in the cluster trivial database are not received within the first preset time, query the inter-node heartbeat interaction data of the second target node from the node interaction record in which the inter-node interaction data of each node is locally stored; the node interaction record includes the heartbeat interaction record between the first target node and the second target node in the cluster trivial database; the first preset time is greater than the frequency of sending heartbeat data between nodes; Determine whether there is a heartbeat record between the first target node and the second target node within the first preset time in the inter-node heartbeat interaction data of the second target node; If, in the heartbeat interaction data between nodes of the second target node, there is a heartbeat record between the first target node and the second target node within the first preset time, the current node state is set to a suspended state, so that other modules in the cluster shut down the network port of the master node; If, in the inter-node heartbeat interaction data of the second target node, there is no heartbeat record between the first target node and the second target node within the first preset time, the first target node is determined as a faulty node, and the faulty node is removed from the cluster; The step of querying the node-to-node heartbeat interaction data of the second target node from the node interaction record in which the node-to-node interaction data of each node is stored locally includes: Determine a second target node from the cluster trivial database; Query the node-to-node heartbeat interaction data of the second target node from the node interaction record in which the node-to-node interaction data of each node is stored locally; Wherein, determining the second target node from the cluster trivial database includes: Determine a target node from the cluster trivial database; Determine whether the heartbeat data sent by the target node and the heartbeat interaction data between nodes are received within the preset heartbeat data sending time; If the heartbeat data and the heartbeat interaction data between nodes sent by the target node are received within the preset heartbeat data sending time, the target node is determined as the second target node; Wherein, after determining whether the heartbeat data sent by the target node and the heartbeat interaction data between nodes are received within the preset heartbeat data sending time, the method further includes: If the heartbeat data and the heartbeat interaction data between nodes sent by the target node are not received within the preset heartbeat data sending time, the target node is re-determined from the cluster trivial database, and then the step of determining whether the heartbeat data and the heartbeat interaction data sent by the target node are received within the preset heartbeat data sending time is performed; After determining the target node, the method further includes: Record the number of current historical target nodes; Determine whether the number of the historical target nodes meets a preset threshold; If the historical target node meets the preset threshold, the current node state is set to a suspended state, and cluster operation error information is returned to the display interface to indicate that a large-scale node failure or a master node failure has occurred in the cluster.
2. The cluster node failure handling method according to claim 1, characterized in that: Before querying the node-to-node heartbeat interaction data of the second target node from the node interaction record in which the node-to-node interaction data of each node is locally stored, the method further includes: Receive heartbeat data sent by other nodes in the cluster trivial database and heartbeat interaction data between nodes; The received heartbeat data and the heartbeat interaction data between nodes are saved in the node interaction record corresponding to the local preset storage location.
3. The cluster node failure handling method according to claim 2, characterized in that: The step of saving the received heartbeat data and the heartbeat interaction data between nodes to the node interaction record corresponding to the local preset storage location includes: The received heartbeat data and the heartbeat interaction data between nodes, wherein the data sending time is within a preset time period, are saved in the node interaction record corresponding to the local preset storage location.
4. A cluster node fault handling device, characterized in that: Applies to the master node in the cluster's trivial database, including: A record query module, for querying the inter-node heartbeat interaction data of the second target node from the node interaction record in which the inter-node interaction data of each node is locally stored if the heartbeat data and the inter-node heartbeat interaction data sent by the first target node in the cluster trivial database are not received within the first preset time; the node interaction record includes the heartbeat interaction record between the first target node and the second target node in the cluster trivial database; the first preset time is greater than the frequency of sending the inter-node heartbeat data; A record determination module, used to determine whether there is a heartbeat record between the first target node and the second target node within the first preset time in the inter-node heartbeat interaction data of the second target node; A record existence module is used to set the current node state to a suspended state if there is a heartbeat record between the first target node and the second target node in the inter-node heartbeat interaction data of the second target node within the first preset time, so that other modules in the cluster shut down the network port of the master node; A record non-existence module is used to determine the first target node as a faulty node and remove the faulty node from the cluster if there is no heartbeat record between the first target node and the second target node within the first preset time in the inter-node heartbeat interaction data of the second target node; Wherein, the record query module includes: A second target node determining unit, configured to determine a second target node from the cluster trivial database; A data query unit, used to query the inter-node heartbeat interaction data of the second target node from the node interaction record in which the inter-node interaction data of each node is locally stored; Wherein, the second target node determination unit includes: A target node determination unit, configured to determine a target node from the cluster trivial database; A data judgment unit, used to judge whether the heartbeat data sent by the target node and the heartbeat interaction data between nodes are received within a preset heartbeat data sending time; A node determination unit, configured to determine the target node as a second target node if the heartbeat data and the heartbeat interaction data between nodes sent by the target node are received within a preset heartbeat data sending time; Wherein, the second target node determination unit further includes: A node unit is re-determined, which is used to re-determine the target node from the cluster trivial database if the heartbeat data sent by the target node and the heartbeat interaction data between nodes are not received within the preset heartbeat data sending time, and then execute the step of determining whether the heartbeat data sent by the target node and the heartbeat interaction data between nodes are received within the preset heartbeat data sending time; Among them, it also includes: A number recording unit is used to record the number of current historical target nodes; A number determination unit, used to determine whether the number of the historical target nodes meets a preset threshold; The node abnormality unit is used to set the current node state to a suspended state if the historical target node meets a preset threshold, and return cluster operation error information to the display interface to indicate that a large-scale node failure or a master node failure has occurred in the cluster.
5. An electronic device, characterized in that: It comprises a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the cluster node failure processing method as described in any one of claims 1 to 3 is implemented.
6. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein, when the computer program is executed by a processor, the cluster node failure handling method according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Equipment detection method and device and communication equipment
CN114422412A