Multi-cluster MySQL instance high-availability management system based on cloud platform

By designing a multi-cluster MySQL instance high-availability management system on the cloud platform, using the heartbeat detection module and failover module to realize real-time monitoring and automatic fault handling of the MySQL database cluster, the problem of how to ensure stable operation and high availability when managing massive MySQL database clusters on the cloud platform is solved, and seamless switching and high availability of database services are achieved.

CN120029825APending Publication Date: 2025-05-23SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510097540.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When managing massive MySQL database clusters on a cloud platform, how to ensure the stable operation and high availability of the database cluster, especially how to respond and deal with failures quickly.

Method used

A cloud-based multi-cluster MySQL instance high-availability management system is designed, including a heartbeat detection module and a failover module. The heartbeat detection module regularly monitors the health status of the cluster, checks the cluster operation status in real time, and starts a voting mechanism to decide whether to perform the failover process when an exception occurs. After receiving the fault handling notification, the failover module automatically performs the failover operation, selects the appropriate failover policy, and transfers the load of the failed node to other normal nodes.

Benefits of technology

Real-time monitoring and automatic fault handling of the cloud platform MySQL database cluster is realized, ensuring seamless switching and high availability of database services, improving the efficiency of fault handling, and reducing the risk of human errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029825A_ABST
    Figure CN120029825A_ABST
Patent Text Reader

Abstract

The invention provides a multi-cluster MySQL instance high-availability management system based on a cloud platform, and belongs to the technical field of cloud computing, and the system comprises a heartbeat detection module which regularly collects the operation condition information of each MySQL database high-availability cluster on the cloud platform to monitor the overall operation condition of the cluster in real time. Once it is monitored that cluster operation is abnormal, a voting mechanism is started, and whether a failover module is notified to execute a failover process or not is determined according to a voting result; and the failover module is responsible for automatically executing failover operation when the cluster has a fault so as to guarantee seamless switching of database services. And the failover module selects the most suitable failover strategy according to the property and severity of the fault, and transfers the load of the fault node to other normal nodes. All the operations can ensure continuity and stability of database services, and service interruption caused by cluster faults is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing, and in particular to a multi-cluster MySQL instance high-availability management system based on a cloud platform. Background Art

[0002] In the current era of rapid digital development, data has become a key asset for enterprise operations and development. As a core component of data storage and management, the stability, reliability and performance of databases are crucial. However, with the continuous expansion of enterprise business and the rapid growth of data, the scale of MySQL database clusters on cloud platforms is also expanding, and the difficulty and complexity of management are also increasing. Especially in the face of massive MySQL clusters, how to ensure the stable operation and high availability of database clusters has become an important challenge facing cloud platforms. Summary of the invention

[0003] In order to solve the above technical problems, the present invention provides a multi-cluster MySQL instance high-availability management system based on a cloud platform, which realizes real-time monitoring and failover of a large number of MySQL database clusters on the cloud platform, ensures that it can respond quickly and take effective measures when a cluster fails, and guarantees the continuity and stability of database services.

[0004] The technical solution of the present invention is:

[0005] A multi-cluster MySQL instance high-availability management system based on a cloud platform includes a heartbeat detection module and a failover module.

[0006] in

[0007] The heartbeat detection module is used to periodically monitor the health status of the MySQL database high-availability cluster. This mechanism verifies the overall operation of the cluster in real time by regularly collecting the operation status information of the MySQL database high-availability cluster. Once an abnormality is detected in the cluster operation, the voting mechanism will be started to decide whether to execute the failover process based on the voting results to ensure seamless switching of database services and high availability of data.

[0008] Furthermore,

[0009] The heartbeat detection module includes several heartbeat detection nodes. Each node registers a scheduled heartbeat task. These tasks periodically initiate heartbeat detection to the MySQL cluster to determine the health status of the cluster.

[0010] If the heartbeat detection of the heartbeat detection module is abnormal, a voting process will be initiated. During this process, the heartbeat detection module creates a voting item in ETCD. Before voting, each heartbeat detection module will obtain the voting information of the current abnormal MySQL cluster to ensure that each module only casts one vote to avoid repeated voting.

[0011] The changes in the number of votes are monitored in real time through the ETCD watch API. When the total number of votes is greater than half, the current MySQL cluster is determined to be abnormal, and the failover module is notified to handle the abnormal MySQL cluster.

[0012] The failover module is another important component in the present invention. It is responsible for automatically performing failover operations when a cluster failure occurs to ensure seamless switching of database services. The failover module will select the most appropriate failover strategy based on the nature and severity of the failure to transfer the load of the failed node to other normal nodes. These operations can ensure the continuity and stability of database services and avoid service interruptions caused by cluster failures.

[0013] Furthermore,

[0014] The failover strategy is to automatically select a standby server as the new primary server when a MySQL cluster fails, and reconfigure and optimize the MySQL high-availability cluster accordingly.

[0015] After receiving the fault handling notification from the heartbeat detection module, the fault transfer module automatically starts the fault handling process.

[0016] The failover module first obtains the log information of all slaves and selects the slave with the most complete log as the new master; if there are several slaves with the most complete logs, their relay logs are further compared, and finally an optimal slave is randomly selected as the new master.

[0017] Next, the failover module logs into the failed host, saves its binary logs, and determines whether there are binary log differences between the new host and the original failed host.

[0018] If differences exist, the new master applies them to ensure data consistency.

[0019] Then, the failover module disconnects the failed master from all slaves, resets the slave information and reestablishes the connection between the slaves and the new master, and finally transfers the front-end request to the new master.

[0020] The beneficial effects of the present invention are

[0021] The present invention realizes real-time monitoring of a large number of MySQL database clusters on a cloud platform through a heartbeat detection module, and can promptly detect abnormal situations in cluster operation. Compared with traditional manual inspections or simple timed monitoring methods, the real-time monitoring of the present invention is more accurate and efficient, and can ensure rapid response when a cluster failure occurs. Secondly, the present invention introduces a voting mechanism and a failover module, which can automatically failover when a cluster failure occurs, ensuring seamless switching of database services. Compared with traditional manual switching or simple backup and recovery methods, this automated failover method not only improves the efficiency of fault handling, but also reduces the risk of human error, further enhancing the high availability of the database. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is the structural block diagram of the heartbeat detection module;

[0023] Figure 2 It is a flowchart of fault handling of the fault handling module. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0025] The present invention provides a multi-cluster MySQL instance high-availability management system based on a cloud platform, which realizes real-time monitoring and automatic fault handling of the MySQL database cluster through a heartbeat detection module and a fault transfer module, thereby ensuring seamless switching and high availability of database services.

[0026] The heartbeat detection module is used to periodically monitor the health status of the MySQL database high-availability cluster. This mechanism verifies the overall operation of the cluster in real time by regularly collecting the operation status information of the MySQL database high-availability cluster. Once an abnormality is detected in the cluster operation, a voting mechanism will be started to decide whether to execute the failover process based on the voting results to ensure seamless switching of database services and high availability of data.

[0027] Specifically, the heartbeat detection module includes multiple heartbeat detection nodes, and each node registers a scheduled heartbeat task. These tasks determine the health status of the cluster by periodically initiating heartbeat detection to the MySQL cluster. If the heartbeat detection of the heartbeat detection module is abnormal, a voting process will be initiated. In this process, the heartbeat detection module creates a voting item in ETCD, for example, setting the key to ` / vote / instance-1`, and the value is a key-value pair of the current heartbeat detection module name. Before voting, each heartbeat detection module will obtain the voting information of the current abnormal MySQL cluster to ensure that each module only casts one vote to avoid repeated voting. The changes in the number of votes are monitored in real time through the watch API of ETCD. When the total number of votes is greater than the number of votes required by the majority strategy (for example, two out of three votes), the current MySQL cluster is determined to be abnormal, thereby notifying the failover module to handle the abnormal MySQL cluster.

[0028] The failover module, after receiving the fault handling notification from the heartbeat detection module, automatically starts the fault handling process. It first obtains the log information of all slaves and selects the slave with the most complete log as the new master. If there are multiple slaves with the most complete logs, their relay logs are further compared, and finally an optimal slave is randomly selected as the new master. Then, the failover module will log in to the failed host, save its binary log, and determine whether there is a binary log difference between the new host and the original failed host. If there is a difference, the new host will apply these differences to ensure data consistency. Then, the failover module disconnects the failed host from all slaves, resets the slave information and reestablishes the connection between the slave and the new host, and finally transfers the front-end request to the new host.

[0029] A multi-cluster MySQL instance high-availability management system based on a cloud platform includes the following steps:

[0030] S1. The heartbeat detection module determines whether the MySQL cluster fails. The number of deployed heartbeat detection modules is an odd number. Taking three heartbeat detection modules as an example, the steps are as follows:

[0031] S1.1 Each heartbeat detection module registers a scheduled heartbeat task for each MySQL cluster on the cloud platform. S1.2 Each heartbeat detection module relies on the scheduled heartbeat task to initiate periodic heartbeat detection to the MySQL cluster.

[0032] S1.2.1 Define cluster node information variables: sshReachable (SSH service unreachable), containsVipNodes (contains VIP nodes), aliveNodes (surviving nodes), deadNodes (dead nodes), nodeRunningStatus (node ​​running information)

[0033] S1.2.2 Check the statistics of cluster node information:

[0034] ●SSH connection detection: Connect to each node via SSH to check whether the node’s SSH connection is reachable, count unreachable nodes in sshReachable, and count reachable nodes in aliveNodes

[0035] ●VIP status detection: If the node is not a read-only instance, execute the Shell command through SSH to check whether the node contains VIP, and count the nodes containing VIP into containsVipNodes

[0036] ●MySQL information query: Use SSH to connect to the node IP, query the node's MySQL master-slave status information through the database connection, and count the information in nodeRunningStatus S1.2.3 Summarize the cluster status based on the statistical node information:

[0037] ●Judge whether there is a master node (Master). If not, return MASTER_DB_DOWN.

[0038] ●If there are multiple master nodes, MYSQL_CLUSTER_BRAIN_CLEFT is returned.

[0039] ● Check whether the slave node (Slave) points to the same master node. If not, return SLAVE_POINT_MASTER_NOT_CONSISTENT.

[0040] ● Check if there is a slave MySQL service that cannot be connected. If so, SLAVE_DB_DOWN is returned.

[0041] ● Check if the slave node is out of sync with the master node. If not, return SLAVE_UNSYNCED.

[0042] If none of the above conditions are met, the cluster status is normal and MYSQL_CLUSTER_NORMAL is returned.

[0043] S1.3 If the heartbeat detection of the heartbeat detection module is not MYSQL_CLUSTER_NORMAL, the heartbeat detection module initiates a voting process. Specifically, create a voting item in ETCD, for example, set the key to / vote / instance-1, and the value to the key-value pair of the current heartbeat detection module name.

[0044] S1.4 Each heartbeat detection module obtains the voting information of the current abnormal MySQL cluster before voting to ensure that each heartbeat detection module only casts one vote to avoid repeated voting.

[0045] S1.5 uses the watch API of ETCD to monitor the changes in the number of votes in real time. When the total number of votes is greater than two (majority strategy), it is determined that the current MySQL cluster is abnormal, and the failover module is notified to handle the abnormal MySQL cluster.

[0046] S2. After the failover module receives the notification from the heartbeat detection module that a fault needs to be handled, the failover module starts the fault handling process for the MySQL cluster with the problem. The steps are as follows:

[0047] S2.1 obtains the latest slave, specifically the Master_Log_File (file name of the master server binary log that the current I / O thread of the slave server is reading) and Read_Master_Log_Pos (the position that the current I / O thread of the slave server has read from the master server binary log file specified by Master_Log_File) values ​​output by the SHOW SLAVE STATUS (displays the replication status of the slave server) command, compares the Master_Log_File and Read_Master_Log_Pos values ​​of all slaves, and selects the slave with the most complete log as the latest slaves (latest slave). At this time, there may be multiple or only one latest slave selected. If there are multiple latest slaves selected, select the latest slave with the most complete relay log as the new host. If there are multiple latest slaves with the most complete relay log, randomly select a latest slave as the new host. If there is only one latest slave selected, use the latest slave as the new host.

[0048] S2.2 logs in to the host where the failure occurs, and saves the binary log of the host where the failure occurs.

[0049] Determine whether there is a binary log difference between the latest slave server and the original failed host, and if so, instruct the new host to apply the binary log difference.

[0050] S2.3 disconnects the failed host from all corresponding slaves.

[0051] S2.4 Other slaves synchronize the relay log from the latest slave server. Specifically, all relay log files are traversed in each other slave, and the CRC32 value of the last valid transaction is found. The CRC32 value is used to obtain the differential relay log from the latest slave server through the SSH command, and the generated relay log difference files are copied to other slaves. These relay log difference files are applied on other slaves to ensure that data is not lost.

[0052] S2.5 instructs the latest slave server to execute reset slave all to clear the original replication status and reconfigure it as a new master server.

[0053] S2.6 establishes a connection between the slave machine and the new master server.

[0054] S2.7 transfers the front-end request to the new primary server.

[0055] The above description is only a preferred embodiment of the present invention, which is only used to illustrate the technical solution of the present invention, and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A multi-cluster MySQL instance high-availability management system based on a cloud platform, characterized in that: include: The heartbeat detection module is used to periodically monitor the health status of the MySQL database high-availability cluster. It monitors the overall operation of the cluster in real time by regularly collecting the operation status information of each MySQL database high-availability cluster on the cloud platform. Once an abnormality is detected in the cluster operation, the voting mechanism will be started to decide whether to notify the failover module to execute the failover process based on the voting results. The failover module is responsible for automatically performing failover operations when a cluster failure occurs to ensure seamless switching of database services.

2. The system according to claim 1, characterized in that The heartbeat detection module includes several heartbeat detection nodes. Each node registers a scheduled heartbeat task. These tasks periodically initiate heartbeat detection to the MySQL cluster to determine the health status of the cluster.

3. The system according to claim 2, characterized in that If the heartbeat detection of the heartbeat detection module is abnormal, a voting process will be initiated. During this process, the heartbeat detection module creates a voting item in ETCD. Before voting, each heartbeat detection module will obtain the voting information of the current abnormal MySQL cluster to ensure that each module only casts one vote to avoid repeated voting.

4. The system according to claim 3, characterized in that The changes in the number of votes are monitored in real time through the ETCD watch API. When the total number of votes is greater than half, the current MySQL cluster is determined to be abnormal, and the failover module is notified to handle the abnormal MySQL cluster.

5. The system according to claim 1, characterized in that The failover module will select the most appropriate failover strategy based on the nature and severity of the failure and transfer the load of the failed node to other normal nodes.

6. The system according to claim 5, characterized in that The failover strategy is to automatically select a standby server as the new primary server when a MySQL cluster fails, and reconfigure and optimize the MySQL high-availability cluster accordingly.

7. The system according to claim 4, characterized in that After receiving the fault handling notification from the heartbeat detection module, the fault transfer module automatically starts the fault handling process.

8. The system according to claim 7, characterized in that The failover module first obtains the log information of all slaves and selects the slave with the most complete log as the new master. If there are several slaves with the most complete logs, their relay logs are further compared and finally the best slave is randomly selected as the new master. Next, the failover module will log in to the failed host, save its binary log, and determine whether there is a binary log difference between the new host and the original failed host; If there are differences, the new master applies them to ensure data consistency; Then, the failover module disconnects the failed master from all slaves, resets the slave information and reestablishes the connection between the slaves and the new master, and finally transfers the front-end request to the new master.