Node cluster exception recovery method and device, electronic equipment and storage medium
By obtaining the working status of the node cluster and performing targeted recovery operations, the problem of automatic master node election failure in abnormal situations of the node cluster is solved, and the recovery efficiency and availability of the node cluster are improved.
Patent Information
- Application Number
- CN202510412565.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-22
AI Technical Summary
When the node cluster is abnormal in master-slave sentry mode, it will automatically elect a new master node from multiple slave nodes, which will affect the node cluster's external services.
By obtaining the working status of all nodes in the node cluster, targeted abnormal recovery operations are carried out in response to different working statuses, including restarting the node, repairing the status, re-election of the master node, etc.
The operation instructions corresponding to different working states are realized to effectively restore the abnormal state of the node cluster, avoid blind operations, and improve recovery efficiency and availability.
Smart Images

Figure CN120353628A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to an abnormal recovery method and device for a node cluster, an electronic device, and a storage medium. Background Art
[0002] In the master-slave sentinel mode of a node cluster, a master node provides read-write services, multiple slave nodes provide read-only services, and multiple sentinel nodes monitor the status of the master node to ensure that a new master node can be elected from multiple slave nodes when the master node is abnormal.
[0003] An abnormality in the node cluster may cause the failure of automatically electing a new master node from multiple slave nodes, thereby affecting the services provided by the node cluster to the outside. Therefore, how to perform abnormal recovery on the node cluster is an urgent problem to be solved. Summary of the Invention
[0004] This application provides an abnormal recovery method and device for a node cluster, an electronic device, and a storage medium, so as to at least solve the problem of how to perform abnormal recovery on a node cluster in related technologies.
[0005] This application provides an abnormal recovery method for a node cluster, including:
[0006] When it is determined that the node cluster is abnormal, in response to an acquisition instruction for acquiring the working states of all nodes in the node cluster, acquiring the working states of all nodes;
[0007] In response to an operation instruction for performing abnormal recovery on the working states of all nodes, performing abnormal recovery on all nodes to obtain an abnormal recovery result; different operation instructions correspond to different working states of all nodes.
[0008] This application also provides an abnormal recovery device for a node cluster, including:
[0009] An acquisition unit, configured to, when it is determined that the node cluster is abnormal, in response to an acquisition instruction for acquiring the working states of all nodes in the node cluster, acquire the working states of all nodes;
[0010] A recovery unit, configured to, in response to an operation instruction for performing abnormal recovery on the working states of all nodes, perform abnormal recovery on all nodes to obtain an abnormal recovery result; different operation instructions correspond to different working states of all nodes.
[0011] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any one of the above-mentioned abnormal recovery methods for a node cluster when executing the computer program.
[0012] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any one of the above-mentioned node cluster exception recovery methods.
[0013] The present application also provides a computer program product including a computer program, which, when executed by a processor, implements the steps of any one of the above-mentioned node cluster exception recovery methods.
[0014] In the present application, since an exception in the node cluster may cause the failure of automatically electing a new master node from multiple slave nodes, in the case of determining that the node cluster has an exception, the working states of all nodes in the node cluster are obtained by responding to an acquisition instruction for obtaining the working states of all nodes in the node cluster, and an exception recovery operation instruction for the working states of all nodes is responded to, and exception recovery is performed on all nodes to obtain an exception recovery result; different working states of all nodes correspond to different operation instructions. The exception recovery of the node cluster is realized by responding to the operation instructions for performing exception recovery operations corresponding to different working states according to the different working states of all nodes. Therefore, the technical problem of how to perform exception recovery on the node cluster can be solved, and the technical effect of performing exception recovery on the node cluster according to the operation instructions corresponding to different working states can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a schematic flowchart of a method for exception recovery of a node cluster provided by an embodiment of the present application;
[0017] Figure 2 It is a schematic structural diagram of a node cluster provided by an embodiment of the present application;
[0018] Figure 3 It is a schematic flowchart of another method for exception recovery of a node cluster provided by an embodiment of the present application;
[0019] Figure 4 It is a schematic flowchart of the entire process of exception recovery of a node cluster provided by an embodiment of the present application;
[0020] Figure 5 It is a schematic structural diagram of an exception recovery device for a node cluster provided by an embodiment of the present application;
[0021] Figure 6 The structural schematic diagram of another abnormal recovery device for a node cluster provided by an embodiment of the present application. Specific embodiments
[0022] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0023] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusions, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0024] In order to enable those skilled in the art of this technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0025] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the abnormal recovery method of the node cluster depends, the specific application environment architecture or specific hardware architecture will be described herein.
[0026] An embodiment of the present application provides an abnormal recovery method for a node cluster. The method will be described in detail in combination with the execution process of the abnormal recovery method of the node cluster.
[0027] Figure 1 The flowchart of an abnormal recovery method for a node cluster provided by an embodiment of the present application.
[0028] As Figure 1 shown, the method includes the following steps:
[0029] Step 101, in the case of determining that there is an abnormality in the node cluster, in response to an acquisition instruction for acquiring the working states of all nodes in the node cluster, acquire the working states of all nodes.
[0030] A node refers to an independent computing unit or device, including but not limited to a physical server, virtual machine, or container. A node cluster refers to a collection of distributed systems composed of multiple nodes. For example, a cache middleware (Redis) cluster. The node cluster can include at least three nodes, and at least three nodes operate in a master-slave sentinel mode. To facilitate a better understanding of the structure of the node cluster, as Figure 2 shown, Figure 2 FIG. is a schematic structural diagram of a node cluster provided by an embodiment of the present application. Among them, one node serves as the master node, and the master node provides read and write services. At least one node serves as a slave node, and the slave node provides read-only services. At least one node serves as a sentinel node, and the sentinel node is used to monitor the master node.
[0031] An exception in the node cluster refers to an abnormal state that occurs in the node cluster. For example, the master node fails, the slave node cannot work properly, the sentinel node cannot correctly monitor the master node, etc., resulting in the cluster being unable to provide services normally. For the sake of understanding, an example is provided. The node cluster exception can be that after the node cluster is powered off and restarted, the node cluster cannot automatically select the master node.
[0032] The working state refers to the current running situation of the node, including the running state (working properly) and the non-running state (such as failure, offline, restarting, etc.). The acquisition instruction refers to an instruction issued by the user to query the current working state of all nodes in the node cluster, usually implemented through a management interface or a monitoring tool.
[0033] The state of the node cluster is monitored in real time through the sentinel node or other monitoring tools. When it is detected that the master node cannot respond, the slave node cannot respond, or the sentinel node cannot work properly, it is determined that the node cluster has an exception. After determining that the node cluster has an exception, the node cluster system automatically triggers the acquisition instruction, or the user manually triggers the acquisition instruction.
[0034] After obtaining the node status, the system can take targeted recovery measures according to the abnormal conditions of different nodes (such as restarting the node, re-electing the master node, etc.), avoiding blind operations and improving the recovery efficiency.
[0035] Step 102, in response to an operation instruction for abnormal recovery of the working states of all nodes, perform abnormal recovery on all nodes to obtain an abnormal recovery result; different working states of all nodes correspond to different operation instructions.
[0036] An operation instruction refers to a specific recovery command issued by a node cluster based on the working status of nodes, which is used to perform abnormal recovery operations on nodes. For example, restarting nodes, repairing node status, re-electing the master node, etc. Abnormal recovery refers to a process of taking a series of measures to make the node cluster resume normal operation in response to abnormal states in the node cluster (such as node failures, absence of the master node, inability of slave nodes to communicate, etc.). The abnormal recovery result refers to the result obtained after performing the abnormal recovery operation on the nodes, reflecting the recovery situation of the node cluster. For example, all nodes resume normal operation status, a new master node is successfully elected, etc.
[0037] By selecting the corresponding operation instruction according to the working status of the nodes, targeted abnormal recovery can be achieved, avoiding blind operations and improving the recovery efficiency.
[0038] With this application, since an abnormality in the node cluster will cause the automatic election of a new master node from multiple slave nodes to fail, when it is determined that there is an abnormality in the node cluster, the working status of all nodes is obtained by responding to an acquisition instruction for acquiring the working status of all nodes in the node cluster, and abnormal recovery is performed on all nodes in response to an operation instruction for abnormal recovery of the working status of all nodes, obtaining an abnormal recovery result; different working statuses of all nodes correspond to different operation instructions. It realizes the abnormal recovery of the node cluster by responding to the operation instruction for performing abnormal recovery operations corresponding to different working statuses according to the different working statuses of all nodes. Therefore, it can solve the technical problem of how to perform abnormal recovery on the node cluster and achieve the technical effect of performing abnormal recovery on the node cluster according to the operation instructions corresponding to different working statuses.
[0039] As a refinement of step 102, when performing abnormal recovery on all nodes in response to an operation instruction for abnormal recovery of the working status of all nodes and obtaining an abnormal recovery result, it can be implemented in but not limited to the following ways, such as Figure 3 shown Figure 3 is a schematic flowchart of a method for abnormal recovery of a node cluster provided by an embodiment of this application, including:
[0040] Step 201, if the working status of all nodes is the running state, then in response to a control instruction for restarting all nodes, all nodes are restarted to obtain the restart result of all nodes.
[0041] The running state refers to the state where the node is working properly, capable of responding to requests and executing tasks. In the running state, the node does not have any faults or abnormalities. A control instruction refers to a command used to instruct the node cluster system to restart all nodes. Restarting means performing an operation to restart the node. The restart operation usually includes stopping the currently running services, reloading the node's configuration and services, and finally restarting the node. The restart result refers to the state or feedback result returned by the node after performing the restart operation. It usually includes a successful restart (the node resumes the running state after restarting) or a failed restart (the node remains in a non-running state after restarting and fails to resume normal operation).
[0042] The restart operation can be used as a simple and effective means of troubleshooting. When the specific cause of the fault is uncertain, restarting all nodes can eliminate some temporary faults and avoid a complex troubleshooting process.
[0043] Step 202, if the restart result indicates that there is an abnormality in the node cluster, in response to a query instruction for querying whether there is a master node among all nodes, query whether there is a master node among all nodes.
[0044] A query instruction refers to a command used to instruct the node cluster system to check the node status, for confirming whether there is a master node in the node cluster. When an abnormality occurs in the node cluster, performing the master node query operation in a timely manner and responding with fault transfer or recovery measures can reduce the fault recovery time of the cluster. Especially when the master node is unavailable, it enables the node cluster system to select a new master node to ensure the continuous normal operation of the cluster.
[0045] Step 203, if there is no master node among all nodes, in response to a setting instruction for setting the target node as the master node, set the target node as the master node; the abnormal recovery result includes that there is a master node among all nodes, and the target node is any one of all nodes.
[0046] The target node refers to a node selected from all nodes, and the target node will be set as the new master node. The target node can be any slave node or sentinel node in the node cluster (in a specific configuration, the sentinel node can also have the ability to be converted into a master node), and the specific selection basis can be factors such as the performance, load, and data integrity of the node.
[0047] A setting instruction refers to an instruction used to instruct the node cluster system to set the target node as the master node when it is determined that there is no master node among all nodes. The setting instruction contains the identification information of the target node and the operation requirements for setting it as the master node, and can be transmitted in the form of specific system calls, command-line instructions, or network messages, etc.
[0048] When the original master node is missing, setting the target node as the master node in a timely manner can quickly restore the read and write service capabilities of the node cluster, reduce the service interruption time caused by the missing master node, and improve the availability of the node cluster.
[0049] In practical applications, after setting the target node as the master node, due to the change of the master node, it is necessary to update the relevant information of the master node. The following methods can be used but are not limited to: in response to the update instruction to update the identification information of the previous master node stored in the sentinel node among all nodes to the identification information of the target node, update the identification information of the previous master node to the identification information of the target node, so that the sentinel node can monitor the target node based on the identification information of the target node.
[0050] The identification information refers to the attributes used to uniquely identify a node, such as the Internet Protocol (IP) address, port number, or name of the node. The update instruction refers to the command issued by the node cluster system used to modify the master node identification information stored in the sentinel node.
[0051] By timely updating the master node identification information stored in the sentinel node, it can be ensured that the sentinel node always monitors the correct master node, avoiding monitoring failures caused by incorrect identification information.
[0052] In practical applications, after querying whether there is a master node among all nodes, there can only be one master node in the node cluster. To ensure that there is only one master node in the node cluster, the following methods can be used but are not limited to: if there are at least two master nodes among all nodes, in response to the modification instruction to modify at least two master nodes into one master node, modify at least two master nodes into one master node; the abnormal recovery result is that there is only one master node among all nodes.
[0053] The modification instruction refers to the instruction used to instruct the node cluster system to modify multiple master nodes into one master node when it is detected that there are at least two master nodes in the node cluster. The modification instruction clarifies the need for a merge operation and the target state of the merge, and can be transmitted in the form of specific system calls, command line instructions, or network messages, etc. Ensure that there is always one master node in the node cluster, avoid abnormalities in the node cluster caused by multiple master nodes, and improve the overall reliability of the node cluster system.
[0054] As a refinement of step 102, when performing an operation instruction for exception recovery in response to the working states of all nodes, and performing exception recovery on all nodes to obtain an exception recovery result, the following methods can be used but are not limited to: If there is at least one non-running state among the working states of all nodes, then in response to a repair instruction to repair the non-running state to a running state, repair the non-running state to a running state; the exception recovery result is that the working states of all nodes are running states.
[0055] The non-running state means that the node cannot work properly currently, including but not limited to states such as failure, offline, restarting, etc. The repair instruction is a command used to instruct the node cluster system to recover the node from the non-running state to the running state. Timely repairing the non-running state nodes can ensure that all nodes in the node cluster can work properly, improve the availability and reliability of the cluster, and reduce the service interruption time caused by node failures.
[0056] In practical applications, after updating the identification information of the previous master node to the identification information of the target node, after the master node changes, it is necessary to notify the slave nodes in the node cluster of the change of the master node. The following methods can be used but are not limited to: In response to a transmission instruction to transmit the identification information of the target node to the slave nodes among all nodes, transmit the identification information of the target node to the slave nodes, so that the slave nodes can communicate with the target node based on the identification information of the target node.
[0057] The transmission instruction is an instruction used to instruct the node cluster system to transmit the identification information of the target node to the slave nodes. The transmission instruction clarifies the transmission target (slave nodes) and the transmission content (identification information of the target node), and can be transmitted in the form of specific system calls, command line instructions, network messages, etc.
[0058] The slave nodes can accurately obtain the identification information of the target node and establish a communication connection with the target node, so as to ensure that the slave nodes can synchronize the data of the target node in a timely manner.
[0059] As a refinement of step 102, when performing an operation instruction for exception recovery in response to the working states of all nodes, and performing exception recovery on all nodes to obtain an exception recovery result, the following methods can be used but are not limited to: In response to a configuration instruction for permission configuration of the master node, slave nodes and sentinel nodes, perform permission configuration on the master node, slave nodes and sentinel nodes to obtain the respective permission configuration results of the master node, slave nodes and sentinel nodes, so that the node cluster can perform permission management based on the permission configuration results; the exception recovery result includes the permission configuration result.
[0060] A configuration instruction refers to an instruction used to indicate the permission configuration of the master node, slave node, and sentinel node in a node cluster. The configuration instruction contains the identification information of the nodes (such as IP address, node number) and the permission information to be configured. The permission configuration result refers to the permission status corresponding to each node after the permission configuration of the master node, slave node, and sentinel node. The permission configuration result records the operation permissions owned by the nodes, and the node cluster can perform permission management based on these results.
[0061] By means of reasonable permission configuration, the operation permissions of different nodes are restricted to prevent unauthorized access and data modification, thereby improving the security of the cluster.
[0062] In an implementable manner of the embodiments of the present disclosure, for a better understanding of the entire process of abnormal recovery of a node cluster, as Figure 4 shown, Figure 4 FIG. is a schematic flowchart of the entire process of abnormal recovery of a node cluster provided by an embodiment of the present application. For abnormal situations such as when the node cluster loses power, after recovery, it may affect the operation of stateful load components because the operation of the node cluster is stateful. Therefore, when encountering the above situation, it is necessary to check whether the nodes in the node cluster are running normally, that is, whether the containers in the container orchestration system are in a running state. If not, it means that there are problems with the startup of the nodes in the node cluster, such as file corruption or other reasons. This is what needs to be dealt with first. If the status of the nodes in the node cluster is all running, but the platform components or business components still report information about connection timeout failure to the node cluster, it is necessary to confirm whether the failure to form a cluster is caused by the inability of the node cluster to select a master. Generally, all nodes (replicas) of the node cluster can be restarted simultaneously to allow them to form a cluster again. If the automatic formation of the cluster still cannot be achieved after restarting, it is necessary to manually form the node cluster. This requires confirming whether there are multiple master nodes, that is, each node in the node cluster is in the role of the master node, resulting in the inability to form a cluster. It is necessary to query the working status information of all nodes in the node cluster, and each node needs to be queried. Based on the information returned by the query, determine whether the reason for the inability of the cluster to select a master is that all nodes are master nodes. If the number of connected slave nodes in the returned information is 0, it means that there is no master node in the current node cluster. It is necessary to select a certain node as the new master node, and at the same time modify the network address of the original master node in the stateful replica set of the sentinel node, update it to the new address and save it. Finally, modify the role, that is, modify the role of each node in the node cluster. Announce the network address and port information of the new master node to each node. This can avoid the impact caused by the inability of the node cluster to automatically recover.
[0063] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0064] An embodiment of the present application also provides an abnormal recovery device for a node cluster. Figure 5 As shown in the structural schematic diagram of an abnormal recovery device for a node cluster provided by an embodiment of the present application, Figure 5 as shown, it includes:
[0065] An acquisition unit 31, configured to, when it is determined that there is an abnormality in the node cluster, in response to an acquisition instruction for acquiring the working states of all nodes in the node cluster, acquire the working states of all nodes;
[0066] A recovery unit 32, configured to, in response to an operation instruction for performing abnormal recovery on the working states of all nodes, perform abnormal recovery on all nodes to obtain an abnormal recovery result; different operation instructions correspond to different working states of all nodes.
[0067] Through the present application, since an abnormality in the node cluster will cause the failure of automatically electing a new master node from multiple slave nodes, when it is determined that there is an abnormality in the node cluster, the working states of all nodes are acquired by responding to an acquisition instruction for acquiring the working states of all nodes in the node cluster, and abnormal recovery is performed on all nodes in response to an operation instruction for performing abnormal recovery on the working states of all nodes to obtain an abnormal recovery result; different operation instructions correspond to different working states of all nodes. The abnormal recovery of the node cluster is realized by responding to the operation instructions for performing abnormal recovery operations corresponding to different working states according to the different working states of all nodes. Therefore, the technical problem of how to perform abnormal recovery on the node cluster can be solved, and the technical effect of performing abnormal recovery on the node cluster according to the operation instructions corresponding to different working states can be achieved.
[0068] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the recovery unit 32 includes:
[0069] A restart module 321, configured to, when the working states of all nodes are all in the running state, in response to a control instruction for restarting all nodes, restart all nodes to obtain a restart result of all nodes;
[0070] A query module 322, configured to, when the restart result indicates that there is an abnormality in the node cluster, in response to a query instruction for querying whether there is a master node among all nodes, query whether there is a master node among all nodes;
[0071] A setting module 323, configured to set a target node as the master node in response to a setting instruction for setting the target node as the master node when there is no master node among all nodes; the exception recovery result includes that there is a master node among all nodes, and the target node is any one of all nodes.
[0072] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the exception recovery device of the node cluster further includes:
[0073] An update unit 33, configured to update the identification information of the previous master node to the identification information of the target node in response to an update instruction for updating the identification information of the previous master node stored in the sentinel nodes among all nodes to the identification information of the target node after setting the target node as the master node, so that the sentinel nodes monitor the target node based on the identification information of the target node.
[0074] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the exception recovery device of the node cluster further includes:
[0075] A modification unit 34, configured to modify at least two master nodes into one master node in response to a modification instruction for modifying at least two master nodes into one master node after querying whether there is a master node among all nodes; the exception recovery result is that there is only one master node among all nodes.
[0076] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the recovery unit 32 further includes:
[0077] A repair module 324, configured to repair at least one non-running state to a running state in response to a repair instruction for repairing the non-running state to a running state when there is at least one non-running state among the working states of all nodes; the exception recovery result is that the working states of all nodes are all running states.
[0078] Further, in a possible implementation manner of this embodiment, as Figure 6 shown, the exception recovery device of the node cluster further includes:
[0079] A transmission unit 35, configured to transmit the identification information of the target node to the slave nodes among all nodes in response to a transmission instruction for transmitting the identification information of the target node to the slave nodes after updating the identification information of the previous master node to the identification information of the target node, so that the slave nodes communicate with the target node based on the identification information of the target node.
[0080] Further, in a possible implementation manner of this embodiment, asFigure 6 As shown in Figure 6 , the recovery unit 32 further includes:
[0081] A configuration module 325, configured to perform permission configuration on the master node, slave node, and sentinel node in response to a configuration instruction for performing permission configuration on the master node, slave node, and sentinel node, so as to obtain the respective permission configuration results corresponding to the master node, slave node, and sentinel node, so that the node cluster performs permission management based on the permission configuration results;
[0082] The exception recovery result includes the permission configuration result.
[0083] For the description of the features in the corresponding embodiment of the exception recovery device of the node cluster, reference can be made to the relevant description in the corresponding embodiment of the exception recovery method of the node cluster, which will not be elaborated here one by one.
[0084] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the exception recovery method of the node cluster.
[0085] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the exception recovery method of the node cluster when running.
[0086] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.
[0087] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and the steps in any of the above embodiments of the exception recovery method of the node cluster are implemented when the computer program is executed by a processor.
[0088] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and the steps in any of the above embodiments of the exception recovery method of the node cluster are implemented when the computer program is executed by a processor.
[0089] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0090] The above has introduced in detail a method and device for abnormal recovery of a node cluster, an electronic device, and a storage medium provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. An abnormal recovery method for a node cluster, characterized in that, Including: When it is determined that there is an abnormality in the node cluster, in response to an acquisition instruction for acquiring the working states of all nodes in the node cluster, acquire the working states of all the nodes; In response to an operation instruction for performing abnormality recovery on the working states of all the nodes, perform abnormality recovery on all the nodes to obtain an abnormality recovery result; different operation instructions correspond to different working states of all the nodes.
2. The abnormal recovery method of the node cluster according to claim 1, characterized in that The performing abnormality recovery on all the nodes in response to the operation instruction for performing abnormality recovery on the working states of all the nodes to obtain an abnormality recovery result includes: If the working states of all the nodes are all running states, then in response to a control instruction for restarting all the nodes, restart all the nodes to obtain a restart result of all the nodes; If the restart result is that there is an abnormality in the node cluster, then in response to a query instruction for querying whether there is a master node among all the nodes, query whether there is the master node among all the nodes; If there is no such master node among all the nodes, then in response to a setting instruction for setting a target node as the master node, set the target node as the master node; the abnormality recovery result includes that there is the master node among all the nodes, and the target node is any one of all the nodes.
3. The method for abnormal recovery of a node cluster according to claim 2, wherein After setting the target node as the master node, the abnormality recovery method for the node cluster further includes: In response to an update instruction for updating the identification information of the previous master node stored in the sentinel node among all the nodes to the identification information of the target node, update the identification information of the previous master node to the identification information of the target node, so that the sentinel node monitors the target node based on the identification information of the target node.
4. The method for abnormal recovery of the node cluster according to claim 2, wherein After querying whether there is the master node among all the nodes, the abnormality recovery method for the node cluster further includes: If there are at least two such master nodes among all the nodes, then in response to a modification instruction for modifying the at least two master nodes into one master node, modify the at least two master nodes into one master node; the abnormality recovery result is that there is only one master node among all the nodes.
5. The method for abnormal recovery of a node cluster according to claim 2, wherein, The performing abnormality recovery on all the nodes in response to the operation instruction for performing abnormality recovery on the working states of all the nodes to obtain an abnormality recovery result further includes: If there is at least one non-running state among the working states of all the nodes, then in response to a repair instruction for repairing the non-running state to the running state, repair the non-running state to the running state; the abnormality recovery result is that the working states of all the nodes are all the running states.
6. The method for abnormal recovery of a node cluster according to claim 3, wherein After updating the identification information of the previous master node to the identification information of the target node, the abnormality recovery method for the node cluster further includes: In response to a transmission instruction for transmitting the identification information of the target node to the slave nodes among all the nodes, the identification information of the target node is transmitted to the slave nodes, so that the slave nodes can communicate with the target node based on the identification information of the target node.
7. The method for abnormal recovery of a node cluster according to claim 6, wherein, The performing an exception recovery operation on all the nodes in response to an operation instruction for exception recovery of the working states of all the nodes, and obtaining an exception recovery result further includes: In response to a configuration instruction for performing permission configuration on the master node, the slave nodes, and the sentinel nodes, performing permission configuration on the master node, the slave nodes, and the sentinel nodes to obtain respective permission configuration results of the master node, the slave nodes, and the sentinel nodes, so that the node cluster can perform permission management based on the permission configuration results; The exception recovery result includes the permission configuration result.
8. An abnormal recovery device for a node cluster, characterized in that, including: an obtaining unit, configured to, when it is determined that there is an exception in the node cluster, in response to an obtaining instruction for obtaining the working states of all the nodes in the node cluster, obtain the working states of all the nodes; a recovery unit, configured to, in response to an operation instruction for performing exception recovery on the working states of all the nodes, perform exception recovery on all the nodes to obtain an exception recovery result; Different operation instructions correspond to different working states of all the nodes.
9. An electronic device, characterized in that, including: a memory, configured to store a computer program; a processor, configured to, when executing the computer program, implement the steps of the method for exception recovery of the node cluster according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the method for exception recovery of the node cluster according to any one of claims 1 to 7 are implemented.