Backup exception handling method, system, device, computer equipment and storage medium
By detecting communication interruptions between multiple backup nodes on the backup server and initiating arbitration requests to the arbitration client group to obtain and summarize node status information, the problem of low data backup reliability in the prior art is solved, and higher data backup reliability and system stability are achieved.
Patent Information
- Application Number
- CN202411601592.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-11-11
AI Technical Summary
In the existing technology, in the user data protection scenarios such as finance and manufacturing, there is a problem of low data backup reliability, especially in the high availability of arbitration methods, which is prone to brain splitting.
By detecting communication interrupts between multiple backup nodes on the backup server, arbitration request is initiated to the arbitration client group, the status information of each node is obtained, and the node status summary information is summarized and returned. Backup exception processing is performed based on the summary information to ensure the reliability of data backup.
Improve the reliability of data backup and system stability, avoid split brain problems, and ensure business continuity.
Smart Images

Figure CN119377011B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of fault handling, and particularly to a method, system, device, computer device, computer-readable storage medium, and computer program product for handling node backup anomalies. Background Art
[0002] When providing data protection services for users, it is crucial to protect the data of critical and core business systems. In some scenarios of user data protection such as finance and manufacturing, situations like missing backups or incomplete backups are not allowed. Failing to back up in a timely manner will result in the loss of the recoverable time point of the business system and may even cause the database to stop due to the log filling up the storage space. Therefore, users have put forward requirements for the business continuity of backup systems such as high availability and redundancy. To meet these requirements, a high-availability mirror mode using two servers is proposed.
[0003] However, this arbitration method has low reliability and is prone to split-brain, resulting in low reliability of data backup. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a method, system, device, computer device, computer-readable storage medium, and computer program product for handling backup anomalies that can improve the reliability of data backup.
[0005] In a first aspect, the present application provides a method for handling backup anomalies, which is applied to any current backup node among multiple backup nodes included in a backup server. The backup server is connected to an arbitration client group, and the method includes:
[0006] When the current backup node detects a communication interruption with a target backup node, an arbitration request is initiated to the arbitration client group; the arbitration client group is used to respond to the arbitration request and return a node status acquisition command to each backup node included in the backup server; the target backup node is located among the multiple backup nodes included in the backup server;
[0007] Based on the node status acquisition command, the node status information of the current backup node is obtained, and the node status information of the current backup node is sent to the arbitration client group; the arbitration client group is used to summarize the node status information sent by each backup node, obtain node status summary information, and return it to each backup node included in the backup server;
[0008] The node status summary information is received, and backup anomaly handling is performed based on the node status summary information.
[0009] In one embodiment, when the current backup node is the master node among the multiple backup nodes, performing backup anomaly handling based on the node status summary information includes:
[0010] If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node, continue to use the current backup node as the primary node;
[0011] If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a faulty backup node, demote the current backup node to a backup node among multiple backup nodes.
[0012] In one embodiment, the method further includes:
[0013] In the case where the current backup node does not receive the node status summary information, continue to use the current backup node as the primary node.
[0014] In one embodiment, in the case where the current backup node is a backup node among multiple backup nodes, perform backup exception handling according to the node status summary information, including:
[0015] If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node, and the primary node among multiple backup nodes is a healthy backup node, continue to use the current backup node as the backup node;
[0016] If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node, and the primary node is a faulty backup node, send the static configuration information of the current backup node to the backup server; the backup server determines a new primary node from multiple backup nodes based on the static configuration information of each backup node, generates upgrade indication information and returns it to the current backup node;
[0017] Based on the upgrade indication information, determine whether to upgrade the current backup node to the new primary node.
[0018] In a second aspect, the present application further provides a backup exception handling method, which is applied to an arbitration client group. The arbitration client group is connected to a backup server, and the backup server includes multiple backup nodes, including:
[0019] Respond to the arbitration request initiated by the current backup node, and return a node status acquisition command to each backup node included in the backup server; the arbitration request is generated and returned to the arbitration client group when the current backup node detects a communication interruption with the target backup node; the current backup node is any one of the backup nodes, and the target backup node is located among multiple backup nodes included in the backup server;
[0020] Receive and summarize the node status information sent by each backup node to obtain the node status summary information; the node status information is obtained by each backup node in response to the node status acquisition command;
[0021] Send node status summary information to each backup node; the node status summary information is used for the current backup node to perform backup exception handling based on the node status summary information.
[0022] In a third aspect, the present application also provides a backup exception handling system, which includes any current backup node among the multiple backup nodes included in the backup server and an arbitration client group, and the backup server is connected to the arbitration client group;
[0023] The current backup node is used to initiate an arbitration request to the arbitration client group when the current backup node detects a communication interruption with the target backup node; the target backup node is located among the multiple backup nodes included in the backup server;
[0024] The arbitration client group is used to respond to the arbitration request and return a node status acquisition command to each backup node included in the backup server;
[0025] The current backup node is further used to receive the node status acquisition command, obtain the node status information of the current backup node based on the node status acquisition command, and send the node status information of the current backup node to the arbitration client group;
[0026] The arbitration client group is further used to summarize the node status information sent by each backup node, obtain the node status summary information and return it to each backup node included in the backup server;
[0027] The current backup node is further used to receive the node status summary information and perform backup exception handling according to the node status summary information.
[0028] In a fourth aspect, the present application also provides a backup exception handling device, which is applied to any current backup node among the multiple backup nodes included in the backup server, and the backup server is connected to the arbitration client group, and includes:
[0029] A request initiation module, which is used to initiate an arbitration request to the arbitration client group when the current backup node detects a communication interruption with the target backup node; the arbitration client group is used to respond to the arbitration request and return a node status acquisition command to each backup node included in the backup server; the target backup node is located among the multiple backup nodes included in the backup server;
[0030] An information acquisition module, which is used to obtain the node status information of the current backup node based on the node status acquisition command and send the node status information of the current backup node to the arbitration client group; the arbitration client group is used to summarize the node status information sent by each backup node, obtain the node status summary information and return it to each backup node included in the backup server;
[0031] An exception handling module, configured to receive node status summary information and perform backup exception handling based on the node status summary information.
[0032] In a fifth aspect, the present application further provides a backup exception handling device, which is applied to an arbitration client group. The arbitration client group is connected to a backup server, and the backup server includes multiple backup nodes, including:
[0033] A command sending module, configured to respond to an arbitration request initiated by a current backup node, and return a node status acquisition command to each backup node included in the backup server; the arbitration request is generated and returned to the arbitration client group when the current backup node detects a communication interruption with a target backup node; the current backup node is any one of the backup nodes, and the target backup node is located among the multiple backup nodes included in the backup server;
[0034] An information summarization module, configured to receive and summarize node status information sent by each backup node to obtain node status summary information; the node status information is obtained by each backup node in response to the node status acquisition command;
[0035] An information sending module, configured to send the node status summary information to each backup node; the node status summary information is used for the current backup node to perform backup exception handling based on the node status summary information.
[0036] In a sixth aspect, the present application further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0037] When the current backup node detects a communication interruption with a target backup node, initiate an arbitration request to the arbitration client group; the arbitration client group is configured to respond to the arbitration request and return a node status acquisition command to each backup node included in the backup server; the target backup node is located among the multiple backup nodes included in the backup server;
[0038] Obtain the node status information of the current backup node based on the node status acquisition command, and send the node status information of the current backup node to the arbitration client group; the arbitration client group is configured to summarize the node status information sent by each backup node to obtain node status summary information and return it to each backup node included in the backup server;
[0039] Receive the node status summary information, and perform backup exception handling based on the node status summary information.
[0040] In a seventh aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0041] When a communication interruption is detected between the current backup node and the target backup node, an arbitration request is initiated to the arbitration client group; the arbitration client group is used to respond to the arbitration request and return a node status acquisition command to each backup node included in the backup server; the target backup node is located among the multiple backup nodes included in the backup server;
[0042] Based on the node status acquisition command, obtain the node status information of the current backup node, and send the node status information of the current backup node to the arbitration client group; the arbitration client group is used to aggregate the node status information sent by each backup node, obtain the node status summary information and return it to each backup node included in the backup server;
[0043] Receive the node status summary information, and perform backup exception handling according to the node status summary information.
[0044] In an eighth aspect, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the following steps:
[0045] When a communication interruption is detected between the current backup node and the target backup node, an arbitration request is initiated to the arbitration client group; the arbitration client group is used to respond to the arbitration request and return a node status acquisition command to each backup node included in the backup server; the target backup node is located among the multiple backup nodes included in the backup server;
[0046] Based on the node status acquisition command, obtain the node status information of the current backup node, and send the node status information of the current backup node to the arbitration client group; the arbitration client group is used to aggregate the node status information sent by each backup node, obtain the node status summary information and return it to each backup node included in the backup server;
[0047] Receive the node status summary information, and perform backup exception handling according to the node status summary information.
[0048] The above backup exception handling method, system, device, computer equipment, computer-readable storage medium, and computer program product, when the current backup node in the backup server detects a communication interruption with the target backup node, initiate an arbitration request to the arbitration client group connected to the backup server. The arbitration client group responds to the arbitration request, returns a node status acquisition command to each backup node included in the backup server, the current backup node obtains node status information based on the node status acquisition command, and sends the node status information to the arbitration client group. The arbitration client group aggregates the node status information sent by each backup node, obtains the node status information and returns it to the backup node. The current backup node receives the node status summary information and performs backup exception handling according to the node status summary information. Each backup node in the backup server is used to implement data backup, and the backup nodes can communicate with each other to ensure data reliability. In the case of a communication interruption between the backup nodes, an arbitration request is initiated to the arbitration client group to handle backup exceptions, and the node status information of itself is sent under the instruction of the arbitration client group. The arbitration client group aggregates the node status information and returns the node status summary information to each backup node, enabling each backup node to know the node status of the other backup nodes. The current backup node performs backup exception handling based on the node status summary, improving the stability and availability of the backup system, and thus enhancing the reliability of data backup. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0050] Figure 1 It is an application environment diagram of the backup exception handling method in an embodiment;
[0051] Figure 2 It is a flowchart of the backup exception handling method in an embodiment;
[0052] Figure 3 It is a flowchart of the backup exception handling method in another embodiment;
[0053] Figure 4 It is a multi-network component diagram in an embodiment;
[0054] Figure 5 It is a flowchart of the switching service process when the master node fails in an embodiment;
[0055] Figure 6A schematic diagram of a switching service process when communication between a master node and a backup node fails in another embodiment;
[0056] Figure 7 A structural block diagram of a backup exception handling device in an embodiment;
[0057] Figure 8 It is a structural block diagram of a backup exception handling device in another embodiment;
[0058] Figure 9 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0060] The backup exception handling method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the backup server 102 communicates with the arbitration client group 104 through the network, and the backup server includes multiple backup nodes, and the current backup node and the target backup node are any one of the multiple backup nodes. The data storage system can arbitrate the data that the client group 104 needs to process. The data storage system can be integrated on the arbitration client group 104, or it can be placed on the cloud or other network servers. When the current backup node detects that the communication with the target backup node is interrupted, it initiates an arbitration request to the arbitration client group 104, and the arbitration client group 104 responds to the arbitration request and returns a node status acquisition command to each backup node included in the backup server 102. The current backup node obtains node status information based on the node status acquisition command, and sends the node status information of the current backup node to the arbitration client group 104. The arbitration client group 104 summarizes the node status information sent by each backup node, obtains the node status summary information and returns it to each backup node included in the backup server 102. The current backup node receives the node status summary information and performs backup exception processing according to the node status summary information. The backup server 102 and the arbitration client group 104 may be independent physical servers, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0061] In an exemplary embodiment, Figure 2 As shown, a backup exception handling method is provided, which is applied to Figure 1 The current backup node included in the backup server 102 is taken as an example to illustrate, including the following steps S201 to S203. Among them:
[0062] Step S201: When the current backup node detects a communication interruption with the target backup node, an arbitration request is sent to the arbitration client group; the arbitration client group is used to respond to the arbitration request and return a node status acquisition command to each backup node included in the backup server; the target backup node is located among the multiple backup nodes included in the backup server.
[0063] Among them, the situation of communication interruption may include three cases: backup node failure, heartbeat network failure between backup nodes, and network failure to which the backup node belongs. At this time, the states between the primary and backup nodes may conflict. Therefore, the backup request sends an arbitration request to the arbitration client group and waits for the relevant data to be returned to the backup node to determine the primary / backup state to which each should belong.
[0064] Optionally, when the current backup node in the backup server 102 detects its own failure, the failure of the target backup node, the heartbeat network failure between the current backup node and the target backup node, and the network failure to which the backup node belongs, an arbitration request is sent to the arbitration client group 104. The arbitration client group 104 responds to the arbitration request and returns a node status acquisition command to each backup node included in the backup server 102. When the current backup node does not detect a communication interruption with the target backup node, each backup node included in the backup server is in the process of normal operation and backup, and the arbitration module is in a dormant state. Only when relevant conditions are detected and triggered, the arbitration module starts to work and sends an arbitration request to the arbitration client group, so as to transmit arbitration-related information to and from the arbitration client group regularly or in real time later. By triggering arbitration in the above way, the operation cost of the backup server is reduced, and it is also ensured that abnormal situations can be handled in a timely manner when backup anomalies occur.
[0065] Step S202: Based on the node status acquisition command, obtain the node status information of the current backup node and send the node status information of the current backup node to the arbitration client group; the arbitration client group is used to summarize the node status information sent by each backup node, obtain the node status summary information and return it to each backup node included in the backup server.
[0066] Among them, the node status information may include the online status information of the backup node, the backup service status of the backup node, the original database status of the backup node, and the metadata database replication status.
[0067] Exemplarily, the current backup node obtains the node status information of the current backup node based on the node status acquisition command, and sends the node status information of the current backup node to the arbitration client group 104. The arbitration client group 104 aggregates the node status information sent by each backup node, obtains the aggregated node status information and returns it to each backup node included in the backup server 102. The current backup node can timely obtain its own status information and send it to the arbitration client group, ensuring the real-time update of the status information of each backup node in the backup server 102, which is beneficial for the arbitration client group to aggregate accurate node status information and return it to each backup node, helping each backup node know the status of the remaining backup nodes and facilitating timely exception handling.
[0068] Step S203: Receive the aggregated node status information and perform backup exception handling according to the aggregated node status information.
[0069] Optionally, when the current backup node is the primary node or the standby node among multiple backup nodes, the primary-standby switch is performed according to the node status of each backup node characterized by the aggregated node status information. By informing the current backup node of the aggregated information, the communication between the backup nodes with disconnected communication can be made interoperable, making the operations of each backup node during exception handling more accurate, avoiding the situation of split-brain due to competing for the primary node, and thus improving the reliability of data backup.
[0070] In the above backup exception handling method, when the current backup node in the backup server detects a communication interruption with the target backup node, an arbitration request is initiated to the arbitration client group connected to the backup server. The arbitration client group responds to the arbitration request, returns a node status acquisition command to each backup node included in the backup server. The current backup node obtains the node status information based on the node status acquisition command, and sends the node status information to the arbitration client group. The arbitration client group aggregates the node status information sent by each backup node, obtains the node status information and returns it to the backup node. The current backup node receives the aggregated node status information and performs backup exception handling according to the aggregated node status information. Each backup node in the backup server is used to implement data backup. The backup nodes can communicate with each other to ensure data reliability. In the case of communication interruption between the backup nodes, an arbitration request is initiated to the arbitration client group to handle backup exceptions, and the node status information of itself is sent under the instruction of the arbitration client group. The arbitration client group aggregates the node status information and returns the aggregated node status information to each backup node, enabling each backup node to know the node status of the remaining backup nodes. The current backup node performs backup exception handling according to the aggregated node status, improving the stability and availability of the backup system, and thus improving the reliability of data backup.
[0071] In one embodiment, when the current backup node is the primary node among multiple backup nodes, backup exception handling is performed according to the node status summary information, including:
[0072] If the current backup node receives the node status summary information and the node status summary information indicates that the current backup node is a healthy backup node, the current backup node continues to be the primary node;
[0073] If the current backup node receives the node status summary information and the node status summary information indicates that the current backup node is a faulty backup node, the current backup node is demoted to a standby node among the multiple backup nodes.
[0074] Among them, the primary node can be understood as the backup node responsible for managing backup tasks and coordinating other backup nodes, which helps improve the regularity of data backup; the standby node can be understood as an ordinary backup node other than the primary node.
[0075] Exemplarily, when the current backup node is the primary node, if the current backup node receives the node status summary information and the node status summary information indicates that the current backup node is a healthy backup node, the current backup node continues to be the primary node; in addition, if the current backup node receives the node status summary information and the node status summary information indicates that the current backup node is a faulty backup node, the current backup node is demoted to a standby node among the multiple backup nodes. On the premise that the current backup node assumes the primary node responsibility, if the current backup node is still a healthy backup node, it continues to assume this responsibility, and the other backup nodes cannot become the new primary node across the current backup node unless the current backup node is a faulty backup node and demotes itself to a standby node. The above master-slave switching method has the following advantages:
[0076] 1. Continuing to use the healthy current backup node as the primary node ensures the stability of the system in the normal running state and avoids service interruption or performance degradation caused by frequent primary node switching.
[0077] 2. If the current backup node fails and receives the corresponding status summary information, demoting it to a standby node can effectively prevent the faulty node from continuing to assume the primary node responsibility, thereby protecting the reliability of the system and data security.
[0078] 3. Restricting other backup nodes from competing to become the primary node when the current backup node is healthy ensures the orderliness of the system and the centralization of management, and avoids the split-brain problem caused by competing for the primary node role from affecting the business continuity of data backup.
[0079] In one of the embodiments, the method further includes: when the current backup node does not receive the node status summary information, the current backup node continues to be the primary node.
[0080] Optionally, when the current backup node is the primary node among multiple backup nodes, if the current backup node does not receive the node status summary information, no operation is performed, and the current backup node continues to be the primary node. When there is a lack of new status information, maintaining the role of the current primary node can reduce the management complexity of the system, avoid frequent adjustment of node roles due to status uncertainty, and ensure the stability of system operation. Choosing not to operate when no new status summary information is received helps prevent unnecessary switches caused by misjudgment or delayed status information, reducing the risk of potential service interruptions.
[0081] In an exemplary embodiment, when the current backup node is a secondary node among multiple backup nodes, backup exception handling is performed according to the node status summary information, including:
[0082] If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node and the primary node among multiple backup nodes is a healthy backup node, the current backup node continues to be the secondary node;
[0083] If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node and the primary node is a faulty backup node, the static configuration information of the current backup node is sent to the backup server; based on the static configuration information of each backup node, the backup server determines a new primary node from multiple secondary nodes, generates upgrade indication information and returns it to the current backup node;
[0084] Based on the upgrade indication information, determine whether to upgrade the current backup node to the new primary node.
[0085] Among them, the static configuration information can be understood as the set information that remains fixed in the system or network node, including network configuration, node identifier, hardware configuration information, etc.
[0086] Optionally, when the current backup node is a standby node among multiple backup nodes, if the current backup node receives node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node and the primary node among the multiple backup nodes is a healthy backup node, the current backup node continues to be used as a standby node. Additionally, if the current backup node receives node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node and the primary node is a faulty backup node, the static configuration information of the current backup node is sent to the backup server 102. The backup server 102 obtains the priorities of each backup node based on the static configuration information of each backup node, determines the backup node corresponding to the highest priority as the new primary node, generates upgrade indication information and returns it to the current backup node. If the upgrade indication information indicates an upgrade, the current backup node is upgraded to the new primary node; if the upgrade indication information does not indicate an upgrade, the current backup node continues to be used as a standby node.
[0087] The above method has the following advantages:
[0088] 1. When the primary node is healthy, the role of the current backup node is not promoted, ensuring system stability and consistency, and avoiding service interruptions caused by frequent status changes.
[0089] 2. When the primary node fails and the current backup node is healthy, the static configuration information is sent to the backup server in a timely manner. The backup server selects a new primary node based on the static configuration information, thereby improving the performance and efficiency of the backup system.
[0090] 3. By generating upgrade indication information, a feedback mechanism is established, enabling each backup node to make judgments according to the actual situation during role adjustment and reducing the risk of incorrect operations.
[0091] In an exemplary embodiment, as Figure 3 shown, a backup exception handling method is provided. Taking the current backup node included in the arbitration client group 104 in Figure 1 as an example, the method includes the following steps S301 to S303. Among them:
[0092] Step S301: In response to an arbitration request initiated by the current backup node, return a node status acquisition command to each backup node included in the backup server; the arbitration request is generated and returned by the arbitration client group when the current backup node detects a communication interruption with the target backup node; the current backup node is any one of the backup nodes, and the target backup node is located among the multiple backup nodes included in the backup server.
[0093] For example, when any current backup node in the backup server 102 detects a communication interruption with the target backup node, an arbitration request is generated and sent back to the arbitration client group 104. The arbitration client group responds to the arbitration request and returns a node status acquisition command to each backup node included in the backup server 102. Before receiving the arbitration request, the arbitration client group does not play an arbitration role and only runs tasks such as scheduled backup and recovery. There is no need to initiate a node status acquisition command to obtain the node status information of each backup node. The sleep-triggered operation mode reduces the operating cost of the arbitration client group.
[0094] Step S302: Receive and summarize the node status information sent by each backup node to obtain node status summary information; the node status information is obtained by each backup node in response to the node status acquisition command.
[0095] Step S303: Send the node status summary information to each backup node; the node status summary information is used for the current backup node to perform backup exception handling based on the node status summary information.
[0096] Optionally, each backup node responds to the node status acquisition command to obtain its own node status information and sends the node status information to the arbitration client group 104. The arbitration client group 104 receives and summarizes the node status information sent by each backup node to obtain node status summary information. Then, the node status summary information is sent to each backup node, and the current backup node performs backup exception handling based on the node status summary information. The above method achieves the following technical effects:
[0097] 1. By having each backup node actively report its own status, the system can understand the health status of each node in real time, improving the timeliness and accuracy of monitoring.
[0098] 2. The arbitration client group summarizes the status of each node, forming a perspective of centralized management, optimizing the scheduling of nodes and the allocation of resources, and improving the overall operation efficiency.
[0099] 3. Through information sharing among backup nodes, each node can make decisions based on the global status information, thereby enhancing the cooperation ability among nodes in the system.
[0100] In one embodiment, a backup exception handling system is provided. The system includes any current backup node among multiple backup nodes included in the backup server and an arbitration client group. The backup server is connected to the arbitration client group;
[0101] The current backup node is used to initiate an arbitration request to the arbitration client group when the current backup node detects a communication interruption with the target backup node; the target backup node is located among multiple backup nodes included in the backup server; the arbitration client group is used to respond to the arbitration request and return a node status acquisition command to each backup node included in the backup server; the current backup node is further used to receive the node status acquisition command, obtain the node status information of the current backup node based on the node status acquisition command, and send the node status information of the current backup node to the arbitration client group; the arbitration client group is further used to aggregate the node status information sent by each backup node, obtain the node status aggregation information and return it to each backup node included in the backup server; the current backup node is further used to receive the node status aggregation information and perform backup exception handling according to the node status aggregation information.
[0102] Exemplarily, the backup exception handling system includes any current backup node among multiple backup nodes included in the backup server and the arbitration client group, and the backup server is connected to the arbitration client group;
[0103] When the current backup node detects a communication interruption with the target backup node, the current backup node initiates an arbitration request to the arbitration client group, the target backup node is located among multiple backup nodes included in the backup server, the arbitration client group responds to the arbitration request and returns a node status acquisition command to each backup node included in the backup server, the current backup node receives the node status acquisition command, obtains the node status information of the current backup node based on the node status acquisition command, and sends the node status information of the current backup node to the arbitration client group, the arbitration client group aggregates the node status information sent by each backup node, obtains the node status aggregation information and returns it to each backup node included in the backup server, and the current backup node receives the node status aggregation information.
[0104] When the current backup node is the master node, if the current backup node receives the node status aggregation information and the node status aggregation information indicates that the current backup node is a healthy backup node, or the current backup node does not receive the node status aggregation information, the current backup node continues to be used as the master node; in addition, if the current backup node receives the node status aggregation information and the node status aggregation information indicates that the current backup node is a faulty backup node, the current backup node is demoted to a standby node among multiple backup nodes.
[0105] When the current backup node is a standby node among multiple backup nodes, if the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node and the master node among the multiple backup nodes is a healthy backup node, the current backup node continues to be a standby node. In addition, if the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node and the master node is a faulty backup node, the static configuration information of the current backup node is sent to the backup server. The backup server obtains the priorities of each backup node based on the static configuration information of each backup node, determines the backup node corresponding to the highest priority as the new master node, generates an upgrade indication message and returns it to the current backup node. If the upgrade indication message indicates an upgrade, the current backup node is upgraded to the new master node; if the upgrade indication message does not indicate an upgrade, the current backup node continues to be a standby node.
[0106] The current backup node can detect the communication interruption with the target backup node in a timely manner and initiate an arbitration request. The arbitration client group responds to the arbitration request and summarizes the node status information of each backup node to obtain the node status summary information, providing a global perspective for the current backup node, realizing a reliable adjustment of the backup role of the current backup node, ensuring the business continuity of data backup, and improving the reliability and availability of the entire backup system.
[0107] In an exemplary embodiment, a specific implementation manner for implementing the backup exception handling method is provided:
[0108] 1. Component description:
[0109] As Figure 4 shown, the backup management architecture consists of four components, namely: Other Agent (other clients), Agent Group (client group), Master (backup server), Storage (storage server), Master_A (master node), Master_B (standby node).
[0110] The backup server contains a master node and a standby node. The master node manages all nodes of Other Agent, Agent Group, and Storage. Due to different business network regions, in the actual application scenario, the nodes of Agent Group are connected to different networks, but the backup server is configured with multiple network cards, enabling the clients (Agents) to connect to it. The following is a detailed description:
[0111] Backup server: It consists of a backup management service and a metadata database, providing functions such as backup system user and user group management, job configuration and scheduling management for backup, recovery, remote replication, etc., global configuration management of the backup console, and recording the corresponding relationship between resource information, backup set information, and storage media.
[0112] Storage server: It consists of a storage management service and storage media, providing interfaces for writing and reading backup set data, automatically clearing expired or invalid data according to the configuration, and providing management functions for storage media such as disks, tapes, optical discs, and objects.
[0113] Backup client: A backup proxy host installed on the business host, used for job acquisition, backup data reading, and transmitting data to the storage server.
[0114] Client group (i.e., the arbitration client group): Mark multiple clients as arbitration nodes to form a client group, providing an arbitration function for the high availability of the backup server.
[0115] Heartbeat line: When the backup server becomes highly available, it is the heartbeat network connection between two server nodes, used as a link to detect the online status of the nodes.
[0116] 2. Operating principle:
[0117] 1. The clients marked as arbitration join the client group and are always managed by the backup server. When a client goes offline, it is automatically removed from the group.
[0118] 2. The client information records of the client group, including client host name, address, connection port, backup proxy type, etc., are recorded in the metadata database of the backup server. And in the high-availability cluster nodes, through the synchronization of the metadata database, the client group information records in the cluster nodes are kept consistent.
[0119] 3. The arbitration information includes: the online status information of the backup server nodes, the backup service status within the backup server nodes, the metadata database status within the backup server nodes, and the metadata database replication status.
[0120] 4. When no node failure occurs in the backup server cluster, the clients within the client group only run scheduled backup, recovery, and other tasks, without the need to transmit high-availability arbitration-related information to the backup server regularly or in real time.
[0121] 5. When the primary node of the backup server cluster fails, the arbitration mechanism is triggered. The standby node of the cluster reads the client group list from the metadata database and sends arbitration request information to all clients within the group.
[0122] 6. The standby node that issues an arbitration request needs to wait for the arbitration information of all clients to be returned. It can only perform information aggregation and merging until it times out.
[0123] 7. The standby node determines whether to switch by summarizing the information:
[0124] a). All the detection items of the master node are normal, all the detection items of the standby node are normal, and the replication delay of the metadata database is lower than the threshold. Switching is not allowed.
[0125] b). The master node cannot be accessed, all the detection items of the standby node are normal, and the replication delay of the metadata database is lower than the threshold. Switching is allowed.
[0126] c). The backup service of the master node fails, all the detection items of the standby node are normal, and the replication delay of the metadata database is lower than the threshold. Switching is allowed.
[0127] d). The metadata database of the master node fails, all the detection items of the standby node are normal, and the replication delay of the metadata database is lower than the threshold. Switching is allowed.
[0128] e). There are abnormalities or the master node cannot be accessed in the detection items of the master node, the backup service of the standby node fails, and the replication delay of the metadata database is lower than the threshold. Switching is not allowed.
[0129] f). There are abnormalities or the master node cannot be accessed in the detection items of the master node, the metadata database of the standby node fails, and the replication delay of the metadata database is higher than the threshold. Switching is not allowed.
[0130] g). There are abnormalities or the master node cannot be accessed in the detection items of the master node, all the detection items of the standby node are normal, and the replication delay of the metadata database is higher than the threshold. Switching is not allowed.
[0131] 3. Example:
[0132] 1. As Figure 5 shown, the switching service process when Master_A fails:
[0133] (1) The master node of Master_A fails. Master_B finds that it cannot connect to Master_A or receives a failure notification from Master_A.
[0134] (2) Master_B sends an arbitration request to all nodes in the Agent Group.
[0135] (3) All nodes (C1\2\…\n) in the Agent Group send a check command to check the status of Master_A and Master_B (host status, service status, metadata database status, etc.).
[0136] (4) Master_A and Master_B respectively return status information (Master_A fails, Master_B is healthy).
[0137] (5) C1\2\…\n merges the status information of Master_A and Master_B.
[0138] (6) The Agent Group sends the status information to Master_A and Master_B.
[0139] (7) If Master_A can receive the status information, it executes the downgrade of Master_A to Standby; if it cannot receive it, it skips.
[0140] (8) When Master_B receives that the return results of all nodes in the Agent Group are consistent, both indicating that Master_A fails and Master_B is healthy, Master_B is upgraded to Active (primary).
[0141] 2. As Figure 6 shown, the switching service process 1 when Master_A and Master_B communication fails:
[0142] (1) In case of a heartbeat network failure, Master_B finds that it cannot connect to Master_A, and at the same time, Master_A finds that it cannot connect to Master_B.
[0143] (2) Master_A and Master_B send arbitration requests to all nodes in the Agent Group.
[0144] (3) C1\2\…\n sends a check command to check the status of Master_A and Master_B (host status, service status, metadatabase status, etc.).
[0145] (4) Master_A and Master_B respectively return status information (Master_A is healthy, Master_B is healthy).
[0146] (5) C1\2\…\n merges the status information of Master_A and Master_B.
[0147] (6) The Agent Group sends the status information to Master_A and Master_B.
[0148] (7) Master_A receives the information that both Master_A and Master_B are healthy and remains in the Active (primary) state.
[0149] (8) Master_B receives the information that both Master_A and Master_B are healthy, remains in the Standby state, and no split-brain occurs.
[0150] 3. As Figure 6 shown, the switching service process 2 for communication failure between Master_A and Master_B:
[0151] (1) Network n fails. Master_B finds that it cannot connect to Master_A, and at the same time, Master_A finds that it cannot connect to Master_B.
[0152] (2) Master_A and Master_B send arbitration requests to all nodes in the Agent Group.
[0153] (3) C1 on the non-failed network 1 sends a check command to check the status of Master_A and Master_B (host status, service status, metadatabase status, etc.).
[0154] (4) Master_A\Master_B respectively return status information (Master_A is healthy, Master_B is healthy).
[0155] (5) C1 merges the status information of Master_A and Master_B.
[0156] (6) C1 sends the status information to Master_A and Master_B.
[0157] (7) Master_A receives the information that both Master_A and Master_B are healthy and remains in the Active state.
[0158] (8) Master_B receives the information that both Master_A and Master_B are healthy, remains in the Standby state, and no split-brain occurs.
[0159] Compared with the prior art, the present application has the following advantages:
[0160] 1. The method of the present application has strong compatibility, can adapt to various backup application scenarios, and the solution has strong stability.
[0161] 2. The present application can ensure that when a failure occurs in the distributed cluster components, the backup system is still available, guaranteeing the high availability requirements of data storage.
[0162] 3. The present application provides a technology for automatically detecting the host status, avoiding the occurrence of split-brain problems.
[0163] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0164] Based on the same inventive concept, an embodiment of the present application further provides a backup exception handling device for implementing the backup exception handling method described above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the backup exception handling device provided below can refer to the limitations on the backup exception handling method in the foregoing, and will not be repeated here.
[0165] In an exemplary embodiment, as Figure 7 shown, a backup exception handling device is provided, which is applied to any current backup node among multiple backup nodes included in a backup server. The backup server is connected to an arbitration client group and includes: a request initiation module 701, an information acquisition module 702, and an exception handling module 703, where:
[0166] The request initiation module 701 is configured to initiate an arbitration request to the arbitration client group when a communication interruption is detected between the current backup node and a target backup node; the arbitration client group is configured to respond to the arbitration request and return a node status acquisition command to each backup node included in the backup server; the target backup node is located among the multiple backup nodes included in the backup server;
[0167] The information acquisition module 702 is configured to acquire the node status information of the current backup node based on the node status acquisition command and send the node status information of the current backup node to the arbitration client group; the arbitration client group is configured to summarize the node status information sent by each backup node, obtain node status summary information and return it to each backup node included in the backup server;
[0168] The exception handling module 703 is configured to receive the node status summary information and perform backup exception handling according to the node status summary information.
[0169] In one embodiment, when the current backup node is the primary node among multiple backup nodes, the exception handling module 703 is further configured to continue to use the current backup node as the primary node if the current backup node receives the node status summary information and the node status summary information indicates that the current backup node is a healthy backup node; and to demote the current backup node to a standby node among multiple backup nodes if the current backup node receives the node status summary information and the node status summary information indicates that the current backup node is a faulty backup node.
[0170] In one of the embodiments, the exception handling module 703 is further configured to continue to use the current backup node as the primary node when the current backup node does not receive the node status summary information.
[0171] In an exemplary embodiment, when the current backup node is a standby node among multiple backup nodes, the exception handling module 703 further includes:
[0172] A first processing sub-module, configured to continue to use the current backup node as a standby node if the current backup node receives the node status summary information, the node status summary information indicates that the current backup node is a healthy backup node, and the primary node among multiple backup nodes is a healthy backup node.
[0173] A second processing sub-module, configured to send the static configuration information of the current backup node to the backup server if the current backup node receives the node status summary information, the node status summary information indicates that the current backup node is a healthy backup node, and the primary node is a faulty backup node; the backup server determines a new primary node from multiple standby nodes based on the static configuration information of each backup node, generates upgrade indication information and returns it to the current backup node; and determines whether to upgrade the current backup node to the new primary node based on the upgrade indication information.
[0174] In an exemplary embodiment, as Figure 8 shown, a backup exception handling device is provided, which is applied to an arbitration client group. The arbitration client group is connected to a backup server, and the backup server includes multiple backup nodes, including: a command issuing module 801, an information summarizing module 802, and an information sending module, where:
[0175] The command issuing module is configured to respond to an arbitration request initiated by the current backup node and return a node status acquisition command to each backup node included in the backup server; the arbitration request is generated and returned to the arbitration client group when the current backup node detects a communication interruption with a target backup node; the current backup node is any one of the backup nodes, and the target backup node is located among the multiple backup nodes included in the backup server;
[0176] An information summarization module, configured to receive and summarize the node status information sent by each backup node to obtain the node status summary information; the node status information is obtained by each backup node in response to the node status acquisition command.
[0177] An information sending module, configured to send the node status summary information to each backup node; the node status summary information is used for the current backup node to perform backup exception handling based on the node status summary information.
[0178] Each module in the above backup exception handling device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0179] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 9 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the node status information and the node status summary information data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a backup exception handling method.
[0180] Those skilled in the art can understand that Figure 9 the structure shown in
[0181] is only a block diagram of a part of the structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0182] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the backup exception handling method of the above embodiment is implemented.
[0183] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the backup exception handling method of the above embodiment is implemented.
[0184] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0185] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0186] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in the present application.
[0187] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A backup exception handling method, characterized in that: Applied to any current backup node among a plurality of backup nodes included in a backup server, the backup server being connected to an arbitration client group, the method comprising: In the case where the current backup node detects that the communication with the target backup node is interrupted, an arbitration request is initiated to the arbitration client group; the arbitration client group is used to respond to the arbitration request and return a node status acquisition command to each backup node included in the backup server; the target backup node is located in the multiple backup nodes included in the backup server; Based on the node status acquisition command, the node status information of the current backup node is acquired, and the node status information of the current backup node is sent to the arbitration client group; the arbitration client group is used to aggregate the node status information sent by each of the backup nodes, obtain the node status aggregate information and return it to each of the backup nodes included in the backup server; Receiving the node status summary information, and when the current backup node is a backup node among the multiple backup nodes, performing backup exception processing according to the node status summary information, including: If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node, and the master node among the multiple backup nodes is a healthy backup node, the current backup node continues to be used as a backup node; If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node and the master node is a faulty backup node, the static configuration information of the current backup node is sent to the backup server; the backup server determines a new master node from the multiple backup nodes based on the static configuration information of each backup node, generates upgrade indication information and returns it to the current backup node; Based on the upgrade indication information, determine whether to upgrade the current backup node to a new master node.
2. The method according to claim 1, characterized in that In the case where the current backup node is a master node among the multiple backup nodes, performing backup exception processing according to the node status summary information includes: If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node, the current backup node continues to serve as the master node; If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a faulty backup node, the current backup node is downgraded to a backup node among the multiple backup nodes.
3. The method according to claim 2, characterized in that The method further comprises: In the case that the current backup node does not receive the node status summary information, the current backup node continues to serve as the master node.
4. A backup exception handling method, characterized in that: Applied to an arbitration client group, the arbitration client group is connected to a backup server, the backup server includes a plurality of backup nodes, the method includes: In response to an arbitration request initiated by the current backup node, a node status acquisition command is returned to each backup node included in the backup server; the arbitration request is generated and returned to the arbitration client group when the current backup node detects that the communication with the target backup node is interrupted; the current backup node is any one of the backup nodes, and the target backup node is located among the multiple backup nodes included in the backup server; Receiving and summarizing the node status information sent by each backup node to obtain node status summary information; the node status information is obtained by each backup node in response to the node status acquisition command; Sending the node status summary information to each of the backup nodes; the node status summary information is used to, when the current backup node is a backup node among the multiple backup nodes, if the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node, and the master node among the multiple backup nodes is a healthy backup node, continue to use the current backup node as a backup node; If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node and the master node is a faulty backup node, the static configuration information of the current backup node is sent to the backup server; the backup server determines a new master node from the multiple backup nodes based on the static configuration information of each backup node, generates upgrade indication information and returns it to the current backup node; Based on the upgrade indication information, determine whether to upgrade the current backup node to a new master node.
5. A backup exception handling system, characterized in that: The system includes any current backup node among a plurality of backup nodes included in a backup server and an arbitration client group, wherein the backup server is connected to the arbitration client group; The current backup node is configured to initiate an arbitration request to the arbitration client group when the current backup node detects that the communication with the target backup node is interrupted; The target backup node is located in a plurality of backup nodes included in the backup server; The arbitration client group is used to respond to the arbitration request and return a node status acquisition command to each backup node included in the backup server; The current backup node is further configured to receive the node status acquisition command, acquire node status information of the current backup node based on the node status acquisition command, and send the node status information of the current backup node to the arbitration client group; The arbitration client group is further used to summarize the node status information sent by each backup node, obtain node status summary information and return it to each backup node included in the backup server; The current backup node is further configured to receive the node status summary information, and when the current backup node is a backup node among the multiple backup nodes, perform backup exception processing according to the node status summary information, including: If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node, and the master node among the multiple backup nodes is a healthy backup node, the current backup node continues to be used as a backup node; If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node and the master node is a faulty backup node, the static configuration information of the current backup node is sent to the backup server; the backup server determines a new master node from the multiple backup nodes based on the static configuration information of each backup node, generates upgrade indication information and returns it to the current backup node; Based on the upgrade indication information, determine whether to upgrade the current backup node to a new master node.
6. A backup exception handling device, characterized in that: Applied to any current backup node among a plurality of backup nodes included in a backup server, the backup server being connected to an arbitration client group, the device comprising: a request initiation module, configured to initiate an arbitration request to the arbitration client group when the current backup node detects that the communication with the target backup node is interrupted; the arbitration client group is configured to respond to the arbitration request and return a node status acquisition command to each backup node included in the backup server; the target backup node is located among the multiple backup nodes included in the backup server; An information acquisition module, configured to acquire the node status information of the current backup node based on the node status acquisition command, and send the node status information of the current backup node to the arbitration client group; the arbitration client group is configured to aggregate the node status information sent by each of the backup nodes, obtain the node status aggregate information and return it to each of the backup nodes included in the backup server; The exception handling module is used to receive the node status summary information, and when the current backup node is a backup node among the multiple backup nodes, perform backup exception handling according to the node status summary information, including: If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node, and the master node among the multiple backup nodes is a healthy backup node, the current backup node continues to be used as a backup node; If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node and the master node is a faulty backup node, the static configuration information of the current backup node is sent to the backup server; the backup server determines a new master node from the multiple backup nodes based on the static configuration information of each backup node, generates upgrade indication information and returns it to the current backup node; Based on the upgrade indication information, determine whether to upgrade the current backup node to a new master node.
7. A backup exception handling device, characterized in that: Applied to an arbitration client group, the arbitration client group is connected to a backup server, the backup server includes a plurality of backup nodes, and the device includes: A command issuing module, used for responding to the arbitration request initiated by the current backup node, and returning a node status acquisition command to each backup node included in the backup server; the arbitration request is generated and returned to the arbitration client group when the current backup node detects that the communication with the target backup node is interrupted; the current backup node is any one of the backup nodes, and the target backup node is located in the multiple backup nodes included in the backup server; An information aggregation module, used to receive and aggregate the node status information sent by each backup node to obtain node status aggregation information; the node status information is obtained by each backup node in response to the node status acquisition command; An information sending module, configured to send the node status summary information to each of the backup nodes; the node status summary information is used to, when the current backup node is a backup node among the multiple backup nodes, if the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node, and the primary node among the multiple backup nodes is a healthy backup node, continue to use the current backup node as a backup node; If the current backup node receives the node status summary information, and the node status summary information indicates that the current backup node is a healthy backup node and the master node is a faulty backup node, the static configuration information of the current backup node is sent to the backup server; the backup server determines a new master node from the multiple backup nodes based on the static configuration information of each backup node, generates upgrade indication information and returns it to the current backup node; Based on the upgrade indication information, determine whether to upgrade the current backup node to a new master node.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Server switching method, MooseFS system and storage medium
CN115145782A