Database state perception-based main / standby automatic fault switching system and method
By designing a master-standby automatic failover system based on database state awareness, the existing system's dependence on zookeeper service is solved, safe, accurate and automatic switching in case of main library failure is achieved, and the stability and reliability of the database system are improved.
Patent Information
- Application Number
- CN202510112403.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing database automatic failover system relies on the zookeeper service, which leads to a large limit on the number of available master and spare nodes, which affects the accuracy and security of database handover.
A main and standby automatic failover system based on database state awareness is designed, including a daemon module, a synchronization program module, a main library query interface module and a coordination program module. By real-time monitoring of the database status, differentiated synchronization data, sensing the main and standby node status and automatically switching the main library, the automatic switching of the main and standby database is realized.
By reducing dependence on zookeeper services, the system improves the flexibility of the master and standby nodes, realizes safe, accurate and automatic switching in the event of a master library failure, reduces the need for manual intervention, and improves the stability and reliability of the database system.
Smart Images

Figure CN120045547A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of database technology, and in particular to a primary-standby automatic fault switching system and method based on database state perception. Background Art
[0002] The database is an important guarantee for the normal operation of enterprise application systems. The database is responsible for the data storage and processing of various functions in the application system. The availability of the database directly affects the availability of the application system, so the importance of the database is self-evident. For application systems with high availability requirements, in order to deal with database failures, a master-slave architecture is usually used to deploy the database. The existing technologies provide systems and methods for switching the backup database to the new master database to provide data operation services after the master-slave fails. With the advancement of information technology, how to ensure the accuracy and security of database switching has always been a high concern in the industry.
[0003] When the primary database fails, switching to the standby database is a key step to ensure business availability. According to the traditional manual switching steps, first, you need to confirm that the primary database has indeed encountered a failure and cannot resume normal service through regular restart or other quick recovery methods. Next, before starting the switching process, you need to ensure that the standby database data is up-to-date and in good condition, which is usually done by checking the synchronization status and network latency of the standby database. Subsequently, before the switch, relevant personnel, including operation and maintenance, development, and management, should be notified of the upcoming database switch operation. Based on business needs and disaster recovery plans, decide whether to perform a planned switch or an emergency switch. When performing the switch, use database management tools or scripts to perform operations, which may involve deactivating the primary database service, starting the standby database service, adjusting DNS or load balancer settings, etc. For systems that support automatic failover, you may only need to execute commands or trigger scripts. After the switch is completed, the new primary database needs to be verified to ensure that it can provide services normally, including checking the database status, performance indicators, and application connectivity. Ensure that all read and write operations are normal and that applications can access the new primary database smoothly. Next, update all relevant documents and configurations to reflect the new database role assignments, and adjust the monitoring system and related alarm settings. After the original primary database is repaired, it should be reconfigured as a backup database, and data may need to be synchronized from the current primary database to the repaired backup database. Finally, perform a fault analysis and process review, record the problems and solutions during the switchover process, and evaluate and optimize the failover strategy based on this experience. If conditions permit, conduct a switchover drill in a non-production environment to ensure smooth operation in actual operation.
[0004] For example, Chinese patent CN111338767A discloses a PostgreSQL master-slave database automatic switching system and method, which uses the characteristics of the zookeeper open source distributed coordination service to automatically elect a leader, each node observes the database status in real time, and multiple observation points vote on the database status. The leader, as a vote monitor, decides whether to trigger the automatic switching of the PostgreSQL master-slave database during a fault period based on the number of votes from each observation point. However, this method relies on the zookeeper service, which uses a voting mechanism. In order to avoid the "brain split" phenomenon, the number of available nodes is required to be greater than half of the total number of nodes. At the same time, in order to prevent the cluster from being unable to provide services due to network isolation, and to save cluster resources under the same fault tolerance capability, the number of nodes in the zookeeper cluster in the actual production environment is generally an odd number and at least 3. Therefore, when the zookeeper service is applied to the field of automatic database switching, there is a great restriction on the number of available master and standby nodes. Summary of the invention
[0005] Based on this, it is necessary to provide a master-slave automatic fault switching system and method based on database status perception to address the above technical problems.
[0006] In a first aspect, the present invention provides a primary-standby automatic fault switching system based on database state perception, the system comprising:
[0007] The daemon module is used to monitor the operation status of the local database in real time, periodically publish and update the local database status information, and distinguish and monitor the master database node and the standby database node according to the system command, introduce the master database offline alarm mechanism, and check the actual availability of the master database node;
[0008] The synchronization program module is used to regularly connect the main database and the local database, execute differentiated synchronization strategies according to different synchronization requirements and environmental factors, synchronize the updated data of the main database to the backup database, and regularly optimize the performance and storage usage of various databases;
[0009] The main database query interface module is used to send transmission control protocol messages to the daemon module, and realize the safe positioning and automatic identification of the application connected to the main database through query and response detection;
[0010] The coordination program module is used to make decisions on the roles of the primary and standby databases, manage and publish primary database information, regularly update primary and database information, and sense the status of primary and standby database nodes. When the primary database needs to be switched, it screens the standby databases with available status and executes automatic switching of the primary and standby databases.
[0011] Further, the daemon module includes: a joint monitoring unit, a status service unit, a main database configuration unit, a database parsing unit and an offline alarm unit;
[0012] The joint monitoring unit is used to monitor the local database in real time, and comprehensively obtain the configuration information, performance statistics and real-time status of the local database as the local database status information;
[0013] The state service unit is used to execute the state service, create a key-value store that can be synchronized between specified database nodes, and periodically update the database state information by continuously running the state service;
[0014] The master database configuration unit is used to set the master database node according to the master-slave database role decision command issued by the coordination program module, publish the master database version information of the current master database node, and locally retain the master database version information of the master database node;
[0015] The database parsing unit is used to parse the corresponding database type in the database configuration file according to the local IP address. Each database node runs a daemon process independently.
[0016] The offline alarm unit is used to check the connection status of the local database according to the main database offline alarm mechanism, determine the actual availability of the main database node, and trigger the alarm information.
[0017] Further, according to the master-slave role decision command issued by the coordination program module, the master library node is set, the master library version information of the current master library node is released, and the master library version information of the master library node is locally retained, including:
[0018] Before the master database node is specified, the first database server node is used as the master database node by default. During use, the master and standby database role decision command issued by the coordination program module is obtained. If the database configuration of any machine is selected or switched to the master database server, the machine will be used as the master database node, and the communication between the master database node and the standby database node will be maintained through the network;
[0019] Use the daemon process of the main library node to call the status service to publish the main library version information, which includes the main library serial number, whether it is a center, the configuration library version number, and the submission library version number;
[0020] The daemon process of the main library node is used to locally retain the main library version information and store it in specific files and election tables. When the daemon process is started again, the latest retained content is selected from the specific storage file and election table as the local recorded main library version information.
[0021] Furthermore, according to the main database offline alarm mechanism, the connection status of the local database is checked to determine the actual availability of the main database node. The triggered alarm information includes:
[0022] Use the daemon process corresponding to the master database node to check the connection status of the local database. If the master database can be accessed normally, no alarm is triggered. If the master database cannot be accessed, the alarm information is published to the general alarm node through the status service to trigger the alarm information.
[0023] When the application initiates a connection request to the main database, the request is sent to the daemon process corresponding to the main database node first, and the daemon process is used to verify the communication status between itself and the coordination program module. If the communication status is normal, the application can access the main database normally. If the communication status is abnormal, no response is given.
[0024] In the standby database node, the daemon process of the standby database node is used to simulate the database access application, and periodically attempts to establish a connection through the main database access interface, and then determines whether the main database can be accessed normally. If the main database can be accessed normally, no alarm is triggered. If the main database cannot be accessed, the alarm information is published to the general alarm node through the status service to trigger the alarm information.
[0025] Furthermore, the synchronization program module includes: a full synchronization unit, an incremental synchronization unit, a regular maintenance unit and a resource protection unit;
[0026] Among them, the full synchronization unit is used to regularly connect the main database and the local database. If the local database is the main database, no operation is performed. If the local database is the backup database, the main and backup synchronization operations are performed. When the initial synchronization or the version difference is greater than the set threshold, the full synchronization strategy is adopted to perform data synchronization between the main database and the backup database.
[0027] The incremental synchronization unit is used to regularly connect the master database and the local database. If the local database is the master database, no operation is performed. If the local database is the slave database, the master-slave synchronization operation is performed. When the version difference is less than the set threshold, the incremental synchronization strategy is adopted to perform the master-slave synchronization operation.
[0028] Regular maintenance unit to optimize system performance and storage usage;
[0029] The resource protection unit is used to equip the locking mechanism during the data import and export, configuration library update and incremental data cleanup process, limiting only one thread to perform data import and export, configuration library update and incremental data cleanup at any time.
[0030] Furthermore, when the synchronization is initialized or the version difference is greater than the set threshold, the full synchronization strategy is adopted to perform data synchronization between the primary database and the standby database, including:
[0031] Compare the submitted version numbers of the primary database and the standby database. If the submitted version numbers of the primary database and the standby database are inconsistent, export the entire primary database to the standby database. The entire database includes the configuration database and the submitted database.
[0032] Compare the maximum version numbers of the libraries submitted by the master database and the library submitted by the standby database. If the maximum version number of the library submitted by the master database is greater than the maximum version number of the library submitted by the standby database, export the entire library from the master database to the standby database.
[0033] If the maximum version number of the standby database's configuration library is less than 0, or the maximum version number of the standby database's configuration library is greater than the maximum version number of its own submitted library, or the maximum version number of the standby database's configuration library is less than the maximum version number of its own submitted library and the difference is greater than the set threshold, the entire primary library will be exported to the standby library.
[0034] Furthermore, when the version difference is less than the set threshold, the incremental synchronization strategy is adopted to perform data synchronization between the primary database and the standby database, including:
[0035] Compare the commit version numbers of the master and slave databases. If the commit version numbers of the master and slave databases are different and the difference is less than the set threshold, use the end of each incremental commit as the trigger point. Whenever an incremental commit ends, synchronize the slave database to the same section as the end of the incremental commit of the master database.
[0036] Further, the regular maintenance unit includes: an incremental cleaning subunit, a file cleaning subunit, a check update subunit and a verification update subunit;
[0037] Among them, the incremental cleaning subunit is used to regularly clean up the old incremental data in the submission library;
[0038] The file cleaning subunit is used to regularly clean up local storage files;
[0039] The check and update subunit is used to use the standby database node to periodically check and import new incremental data or compressed packages to the local submission database;
[0040] The check and update subunit is used to regularly check the version of the configuration library and update the configuration library from the local submission library.
[0041] Further, the coordination program module includes: an update publishing unit, a check and judgment unit, a node configuration unit and a master-slave switching unit;
[0042] Among them, the update publishing unit is used to regularly check the main library version information retained by the daemon module. If the information is inconsistent, the latest main library version information is published and the daemon module is called for updating; database information and topics about database information are published regularly;
[0043] The checking and judging unit is used to regularly check the connection status with the daemon module, sense the status of the primary and standby database nodes, and judge the availability of the database;
[0044] The node configuration unit is used to select a database that meets the requirements of the main database as the main database according to the availability of the database, and determine whether to perform the main database switch;
[0045] The master-slave switching unit is used to screen the available slave databases, use the synchronization program module to check the configuration version and submission version of the slave database, find the slave database for master-slave database switching, and after the search is completed, use the notification mechanism to trigger the automatic switching of the master-slave database.
[0046] Furthermore, the connection with the daemon module is checked regularly to sense the status of the primary and standby database nodes and determine the availability of the database, including:
[0047] The coordination process uses socket technology to communicate with the daemon process of each database node. When the socket connection fails, the database status of the corresponding database node is determined to be unknown.
[0048] When the socket connection is successful, the daemon process is used to verify the connection with the local database. If the verification is successful, the database status of the corresponding database node is determined to be available. If the verification fails, the database status of the corresponding database node is determined to be unavailable.
[0049] Further, according to the availability of the database, a database that meets the requirements of the main database is selected to be configured as the main database, and determining whether to perform the main database switch includes:
[0050] When the database status is unavailable, the corresponding database node is determined to be abnormal. When the main database is detected to be abnormal, it is determined that the main database is inaccessible at this time and the main and standby databases need to be switched;
[0051] When the database status of the primary database is unknown and there is no primary database offline alarm, it is determined that the primary database is accessible and the network is connected, only the daemon process cannot communicate, and there is no need to perform primary and standby database switching;
[0052] When the database status of the primary database is unknown and there is a primary database offline alarm, it is determined that the primary database is inaccessible and the network is not connected, the daemon process cannot communicate, and the primary and standby databases need to be switched.
[0053] Further, the available standby databases are screened, the configuration version and the submitted version of the standby database are checked by using the synchronization program module, and the standby database is found for switching between the primary and standby databases. After the search is completed, the notification mechanism is used to trigger the automatic switching between the primary and standby databases, including:
[0054] Call the daemon process to filter out available standby databases, check whether the configuration version and commit version of the standby database are consistent with the original master database, and retain the consistent standby database. If there are multiple available and synchronized standby databases, select one as the new master database to be switched according to the order of the database configuration files;
[0055] The coordination process is called to notify the application to disconnect before the primary and standby databases switch and wait for the application's response within a limited time. After the time limit expires, the coordination process is used to trigger automatic switching.
[0056] In a second aspect, the present invention further provides a method for automatic master-slave failover based on database status perception, the method comprising the following steps:
[0057] S1. Monitor the operation status of the local database in real time, periodically publish and update the local database status information, and distinguish and monitor the master database node and the backup database node according to system commands, introduce the master database offline alarm mechanism, and check the actual availability of the master database node;
[0058] S2. Regularly connect the main database and the local database, implement differentiated synchronization strategies based on different synchronization requirements and environmental factors, synchronize the updated data of the main database to the standby database, and regularly optimize the performance and storage usage of various databases;
[0059] S3, sending a transmission control protocol message to the daemon module, and realizing the secure positioning and automatic identification of the application connection to the main database through inquiry and response detection;
[0060] S4. Execute decisions on the roles of the primary and standby databases, manage and publish information about the primary database, regularly update information about the primary and standby databases, and sense the status of the primary and standby database nodes. When the primary database needs to be switched, select the standby databases with available status and execute automatic switching of the primary and standby databases.
[0061] The beneficial effects of the present invention are as follows: by integrating a multi-node daemon, a multi-node synchronization program, a master database query interface and a single instance coordination program, a database fault automatic switching system is jointly formed, aiming to realize safe and accurate automatic switching when a master database fails; the daemon not only regularly monitors and publishes the local database status, but also adapts and is compatible with a variety of database types, and introduces a master database offline alarm mechanism to avoid misjudgment of master database failures; the synchronization program is responsible for synchronizing the master database data changes to the standby database in real time, adding a lock mechanism during the synchronization process to ensure data consistency, flexibly adopting data synchronization strategies according to the size of the difference between versions, and timely optimizing system performance through regular maintenance tasks; the application realizes a secure connection with the master database through the master database query interface; the coordination program is not only responsible for the management and publication of the master database information, but also can detect the master database status and identify the fault type, and promptly notify the application to disconnect, thereby realizing accurate anomaly detection and safe automatic switching when a master database fails, reducing the need for manual intervention, ensuring the safety of the master-standby automatic switching process, and effectively improving the stability and reliability of the database system. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0063] Figure 1 This is a system principle block diagram of a master-slave automatic fault switching system based on database status perception according to an embodiment of the present invention;
[0064] Figure 2 is a flow chart of a primary-standby automatic failover method based on database status awareness according to an embodiment of the present invention;
[0065] Figure 3 2 is a schematic diagram of a table storage structure for storing information in a main database according to an embodiment of the present invention;
[0066] Figure 4 is a schematic diagram of a standby database node synchronization process according to an embodiment of the present invention;
[0067] Figure 5 is a schematic diagram of an automatic switching triggering process according to an embodiment of the present invention;
[0068] Figure 6 It is a schematic diagram of visual display of the coordination program according to an embodiment of the present invention.
[0069] Figure numbers: 1. Guardian module; 2. Synchronization module; 3. Main database query interface module; 4. Coordination module. DETAILED DESCRIPTION
[0070] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0071] See also Figure 1 , provides a master-slave automatic failover system based on database status perception, the system comprising:
[0072] Daemon module 1 is used to monitor the operation status of the local database in real time, periodically publish and update the local database status information, and distinguish and monitor the main database node and the backup database node according to the system command, introduce the main database offline alarm mechanism, and check the actual availability of the main database node.
[0073] In the description of the present invention, the daemon module 1 includes: a joint monitoring unit (not marked in the figure), a status service unit (not marked in the figure), a main library configuration unit (not marked in the figure), a database parsing unit (not marked in the figure) and an offline alarm unit (not marked in the figure).
[0074] The joint monitoring unit is used to monitor the local database in real time, and comprehensively obtain the configuration information, performance statistics and real-time status of the local database as the local database status information.
[0075] The state service unit is used to execute the state service, create a key-value store that can be synchronized between specified database nodes, and periodically update the database state information by continuously running the state service.
[0076] Specifically, the state service is a self-developed distributed service, which is mainly used to solve the coordination problem between multiple nodes in distributed applications. It maintains a key-value data structure similar to the file system internally, and can synchronize the key-value library between specified nodes. Applications can access the data stored in it through key values with specific rules. The daemon runs continuously and periodically updates the database status information through the state service, publishing the topic " / domain name / Machine / node id / DatabaseState".
[0077] The master database configuration unit is used to set the master database node according to the master-slave database role decision command issued by the coordination program module 4, publish the master database version information of the current master database node, and locally retain the master database version information of the master database node.
[0078] In the description of the present invention, according to the master-slave database role decision command issued by the coordination program module 4, the master database node is set, the master database version information of the current master database node is released, and the master database version information of the master database node is locally retained, including:
[0079] Step 101: Before the master database node is specified, the first database server node is used as the master database node by default. During use, the master-slave database role decision command issued by the coordination program module is obtained. If the database configuration of any machine is selected or switched to the master database server, the machine will be used as the master database node, and the communication between the master database node and the standby database node is maintained through the network.
[0080] Step 102: Use the daemon process of the master library node to call the status service to publish the master library version information, which includes the master library serial number, whether it is a center, the configuration library version number, and the submission library version number.
[0081] Step 103: Use the daemon process of the master library node to locally retain the master library version information and store it in a specific file and an election table. When the daemon process is started again, the latest retained content is selected from the specific storage file and the election table as the master library version information recorded locally.
[0082] The database parsing unit is used to parse the corresponding database type in the database configuration file according to the local IP address. Each database node runs a daemon process independently.
[0083] Specifically, to ensure compatibility under different database environments, the daemon of the present invention can parse the corresponding database type in the database configuration file according to the local IP, and can currently support the adaptation of five mainstream database systems: oracle, mysql, kingbase, postgresql, and dm. Secondly, each database node independently runs a daemon process, but the daemon itself does not make decisions on the main and standby databases. Whether the local database of a certain machine is the main database or the standby database is notified to the daemon after the coordination program makes a decision. When there is no coordination program involved, the daemon on each machine is only responsible for monitoring and publishing the database status of its own node. With the participation of the coordination program, it is distinguished whether the local database monitored by each daemon is the main database or the standby database. At this time, if the daemon of the main database fails, the function of monitoring the node status of the main database will be affected, resulting in the inability to determine whether the main database is actually available, which may cause erroneous automatic switching decisions.
[0084] The offline alarm unit is used to check the connection status of the local database according to the main database offline alarm mechanism, determine the actual availability of the main database node, and trigger the alarm information.
[0085] In the description of the present invention, according to the main database offline alarm mechanism, the connection status of the local database is checked to determine the actual availability of the main database node, and the triggered alarm information includes:
[0086] Step 111: Use the daemon process corresponding to the main database node to check the connection status of the local database. If the main database can be accessed normally, no alarm is triggered. If the main database cannot be accessed, the alarm information is published to the general alarm node through the status service to trigger the alarm information.
[0087] Step 112, when the application initiates a connection request to the main library, the request is sent to the daemon process corresponding to the main library node first, and the daemon process is used to verify the communication status between itself and the coordination program module 4. If the communication status is normal, the application accesses the main library normally. If the communication status is abnormal, no response is given.
[0088] Step 113: In the standby database node, use the daemon process of the standby database node to simulate the database access application, regularly try to establish a connection through the main database access interface, and then determine whether the main database can be accessed normally. If the main database can be accessed normally, no alarm is triggered. If the main database cannot be accessed, the alarm information is published to the general alarm node through the status service to trigger the alarm information.
[0089] The synchronization program module 2 is used to regularly connect the main database and the local database, execute differentiated synchronization strategies according to different synchronization requirements and environmental factors, synchronize the updated data of the main database to the backup database, and regularly optimize the performance and storage usage of various databases.
[0090] In the description of the present invention, the synchronization program module 2 includes: a full synchronization unit (not marked in the figure), an incremental synchronization unit (not marked in the figure), a regular maintenance unit (not marked in the figure) and a resource protection unit (not marked in the figure).
[0091] Among them, the full synchronization unit is used to regularly connect the main database and the local database. If the local database is the main database, no operation is performed. If the local database is the backup database, the main and backup synchronization operations are performed. When the initial synchronization or the version difference is greater than the set threshold, the full synchronization strategy is adopted to perform data synchronization between the main database and the backup database.
[0092] In the description of the present invention, when the synchronization is initialized or the version difference is greater than the set threshold, the full synchronization strategy is adopted to perform data synchronization between the primary database and the standby database, including:
[0093] Step 201: Compare the submitted version numbers of the main database and the standby database. If the submitted version numbers of the main database and the standby database are inconsistent, export the entire main database to the standby database, where the entire database includes the configuration database and the submitted database.
[0094] Step 202: compare the maximum version numbers of the library submitted by the main library and the library submitted by the standby library. If the maximum version number of the library submitted by the main library is greater than the maximum version number of the library submitted by the standby library, export the entire library from the main library to the standby library.
[0095] Step 203: If the maximum version number of the standby database's configuration library is less than 0, or the maximum version number of the standby database's configuration library is greater than the maximum version number of its own submitted library, or the maximum version number of the standby database's configuration library is less than the maximum version number of its own submitted library and the difference is greater than a set threshold, then export the entire main library to the standby library.
[0096] The incremental synchronization unit is used to regularly connect the main database and the local database. If the local database is the main database, no operation is performed. If the local database is the backup database, the main and backup synchronization operations are performed. When the version difference is less than the set threshold, the incremental synchronization strategy is adopted to perform the main and backup synchronization operations.
[0097] In the description of the present invention, when the version difference is less than the set threshold, an incremental synchronization strategy is adopted to perform data synchronization between the primary database and the standby database, including:
[0098] Step 211: Compare the commit version numbers of the master database and the slave database. If the commit version numbers of the master database and the slave database are different and the difference is less than the set threshold, use the end of each incremental commit as the trigger point. Whenever an incremental commit ends, synchronize the slave database to be consistent with the section at the end of the incremental commit of the master database.
[0099] Periodic maintenance unit to optimize system performance and storage usage.
[0100] In the description of the present invention, the regular maintenance unit includes: an incremental cleaning subunit, a file cleaning subunit, a check update subunit and a verification update subunit.
[0101] Among them, the incremental cleaning subunit is used to regularly clear the old incremental data submitted to the library.
[0102] The file cleaning subunit is used to regularly clean up local storage files.
[0103] The check and update subunit is used to use the standby database node to periodically check and import new incremental data or compressed packages to the local submission library.
[0104] The check and update subunit is used to regularly check the version of the configuration library and update the configuration library from the local submission library.
[0105] The resource protection unit is used to equip the locking mechanism during the data import and export, configuration library update and incremental data cleanup process, limiting only one thread to perform data import and export, configuration library update and incremental data cleanup at any time.
[0106] The main database query interface module 3 is used to send transmission control protocol messages to the daemon module, and realize the safe positioning and automatic identification of the application connected to the main database through query and response detection.
[0107] Specifically, the main database information structure in the query interface is as follows Figure 3 As shown, in order to realize the safe positioning and automatic identification of the application connection to the main library, the query function of the present invention is realized by sending a TCP (Transmission Control Protocol) message to the daemon program through the interface, which will connect to the daemon program of each database node. When the first daemon program responds, the query ends. The interface is called by the database interface, and the main library query function is automatically called internally when connecting to the configuration library.
[0108] The coordination program module 4 is used to make decisions on the roles of the primary and standby databases, manage and publish the primary database information, regularly update the primary database information and database information, and sense the status of the primary and standby database nodes. When the primary database needs to be switched, it screens the standby databases with available status and executes automatic switching of the primary and standby databases.
[0109] In the description of the present invention, the coordination program module 4 includes: an update publishing unit (not marked in the figure), a checking and judging unit (not marked in the figure), a node configuration unit (not marked in the figure) and a master-slave switching unit (not marked in the figure).
[0110] The update publishing unit is used to regularly check the main database version information retained by the daemon module. If the information is inconsistent, the latest main database version information is published and the daemon module is called for updating. Database information and topics about database information are published regularly.
[0111] The checking and judging unit is used to periodically check the connection status with the daemon module, sense the status of the primary and standby database nodes, and judge the availability of the database.
[0112] In the description of the present invention, regularly checking the connection status with the daemon module, sensing the status of the primary and standby database nodes, and judging the availability of the database includes:
[0113] Step 301: The coordination process uses the socket technology to communicate with the daemon process of each database node over the network. When the socket connection fails, the database state of the corresponding database node is determined to be unknown.
[0114] Step 302: When the socket connection is successful, the daemon process is used to verify the connection with the local database. If the verification is successful, the database status of the corresponding database node is determined to be available. If the verification fails, the database status of the corresponding database node is determined to be unavailable.
[0115] The node configuration unit is used to select a database that meets the requirements of the main database as the main database according to the availability of the database, and determine whether to perform the main database switch.
[0116] In the description of the present invention, according to the availability of the database, selecting a database that meets the requirements of the main database to be configured as the main database, and determining whether to perform the main database switch includes:
[0117] Step 311: When the database status is unavailable, the corresponding database node is determined to be abnormal. When the main database is detected to be in an abnormal state, it is determined that the main database is inaccessible at this time and a master-slave database switch needs to be performed.
[0118] Step 312: When the database status of the primary database is unknown and there is no primary database offline alarm, it is determined that the primary database is accessible and the network is connected, only the daemon process cannot communicate, and there is no need to perform primary-standby database switching.
[0119] Step 313: When the database status of the main database is unknown and there is a main database offline alarm, it is determined that the main database is inaccessible and the network is not connected, the daemon process cannot communicate, and the main and standby databases need to be switched.
[0120] The master-slave switching unit is used to screen the available slave databases, use the synchronization program module to check the configuration version and submission version of the slave database, find the slave database for master-slave database switching, and after the search is completed, use the notification mechanism to trigger the automatic switching of the master-slave database.
[0121] In the description of the present invention, screening available standby databases, using a synchronization program module to check the configuration version and the submitted version of the standby database, finding the standby database for master-slave database switching, and after the search is completed, using a notification mechanism to trigger the automatic switch of the master-slave database includes:
[0122] Step 321: Call the daemon process to filter the available standby databases, check whether the configuration version and the submitted version of the standby database are consistent with the original master database, retain the consistent standby databases, and if there are multiple available and synchronized standby databases, select one as the new master database to be switched according to the order of the database configuration files.
[0123] Step 322: Call the coordination process to notify the application to disconnect before the primary and standby databases switch and wait for the application's response within a limited time. After the time limit expires, use the coordination process to trigger automatic switching.
[0124] Specifically, after determining that the main database needs to be switched, in order to find a suitable standby database for switching, the coordination program of the present invention screens the available standby databases based on the daemon program. For each available standby database node, the synchronization program checks whether its configuration version and submission version are consistent with the original main database to ensure that the standby database data is the latest. If there are multiple standby databases that are available and synchronized, one is selected as the new main database to be switched according to the order of the database configuration files. To ensure the data security and consistency of the system during automatic switching, the coordination program of the present invention notifies the application to disconnect before switching and waits for the application's response within a limited time. The coordination program does not directly cause the application to be disconnected, but adopts a notification mechanism to let the application or user know the main database failure event in a timely manner through the status service within a limited time, and records this notification event in the log. After receiving the notification of the main database failure, it will decide whether to disconnect from the original main database based on the running status and business needs of the application. After the time limit ends, the coordination program triggers automatic switching.
[0125] The above-mentioned master-slave automatic fault switching system based on database status perception integrates a daemon, a synchronization program, a master database query interface and a coordination program to perform automatic switching of database faults. The daemon of the present invention monitors and publishes the status of the local database, and issues a master database offline alarm when the master database fails, thereby avoiding the phenomenon of erroneous switching; the synchronization program synchronizes the updated data of the master database to the standby database in real time, flexibly adopts data synchronization strategies according to the differences between versions, and timely optimizes system performance through regular maintenance tasks; the application realizes a secure connection with the master database through the master database query interface; the coordination program publishes the master database information, detects the status of the master database and identifies the fault type, and promptly notifies the application to disconnect, thereby realizing accurate anomaly detection and safe automatic switching when the master database fails, reducing the need for manual intervention, ensuring the safety of the master-slave automatic switching process, and effectively improving the stability and reliability of the database system.
[0126] See also Figure 2 , and also provides a primary-standby automatic failover method based on database status perception, the method comprising the following steps:
[0127] S1. Monitor the operation status of the local database in real time, publish and update the local database status information periodically, and distinguish and monitor the main database node and the backup database node according to system commands, introduce the main database offline alarm mechanism, and check the actual availability of the main database node.
[0128] S2. Regularly connect the main database and the local database, implement differentiated synchronization strategies based on different synchronization requirements and environmental factors, synchronize the updated data of the main database to the backup database, and regularly optimize the performance and storage usage of various databases.
[0129] S3. Send a transmission control protocol message to the daemon module, and through inquiry and response detection, realize the secure positioning and automatic identification of the application connection to the main database.
[0130] S4. Execute decisions on the roles of the primary and standby databases, manage and publish information about the primary database, regularly update information about the primary and standby databases, and sense the status of the primary and standby database nodes. When the primary database needs to be switched, select the standby databases with available status and execute automatic switching of the primary and standby databases.
[0131] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments.
[0132] The daemon module, synchronization module, master database query interface module and coordination program module in the master-slave automatic fault switching system based on database status perception of the present invention can all be applied on computers or machine nodes in the form of independent programs or interfaces. An introduction to each program or interface is as follows.
[0133] 1. Daemon
[0134] The daemon has the basic functions of regularly monitoring and publishing the status of the local database. In order to monitor the operation status of the local database in real time, query the system view to obtain the database instance information, listener status, table space usage and number of database connections, etc., and execute system commands to obtain the database size, occupancy rate, CPU usage rate and memory occupancy rate, etc. Through the joint monitoring strategy based on SQL instructions (Structured Query Language instructions) and shell scripts, the configuration information, performance statistics and real-time status of the database are fully grasped to ensure the stable operation and efficient performance of the database.
[0135] The state service is a self-developed distributed service, mainly used to solve the coordination problem between multiple nodes in distributed applications. It maintains a key-value data structure (key-value library) similar to the file system internally, and can synchronize the key-value library between specified nodes. Applications can access the data stored in it through key values with specific rules. The daemon runs continuously and periodically updates the database status information through the state service, publishing the topic " / domain name / Machine / node id / DatabaseState", and its topic content is shown in the following code:
[0136] message DatabaseState
[0137] {
[0138] uint32 instance_mode=1; instance operation mode
[0139] uint32 instance_status=2; instance running status
[0140] bool listener_available=3; listener is available
[0141] uint32 database_size=4; database disk usage (MB)
[0142] float database_ratio=5;Database disk occupancy rate (%)
[0143] uint32 cpstbs_size=6; cps table space usage (MB)
[0144] float cpstbs_ratio=7; cps table space usage (%)
[0145] uint32 hdstbs_size=8; hds table space usage (MB)
[0146] float hdstbs_ratio=9; hds table space usage (%)
[0147] uint64 device_id=10; Collection node ID
[0148] }
[0149] The local database, i.e. the local database, is a database instance installed and run on a certain machine. The decision of the primary and standby database roles is made by the coordinator. The initial role depends on the configuration in the system. The primary database is specified when the automation system is installed. By default, it is the database of the first database server node. The subsequent switching of the primary and standby database roles depends on the coordinator. In the database environment, the primary and standby databases are of the same database type and can run in the x86 and arm architectures. Based on the minimum configuration required by the state service, the processor: Intel(R) Atom(TM) CPU Z510@1.10GHz or ARM Cortex-A7(ARMv7)
[0150] 1.2GHz, memory: 512MB, network: 100M Ethernet, storage: 8GB. The system currently runs on several operating systems including Linux (Ubuntu 16.04, Ubuntu 20.04), NeXT, Kylin, etc.
[0151] The machine node name is obtained based on the system command. If the database of the machine is configured or switched to the master database server, the machine assumes the role of the master database node, and the master and standby database nodes maintain communication based on the network. The daemon of the master database node will call the status service to publish the master database version information with the subject " / config / master_db_info" so that other system components can obtain these key configuration information. The subject content includes the master database serial number, whether it is a center, the configuration database version number, and the submission database version number. At the same time, in order to quickly restore the master database status when the daemon is restarted or fails, the daemon will locally retain the master database information, and store the master database information string obtained from the coordinator in a specific file and dbelection table (election table). The table structure is as follows: Figure 3 As shown, when the daemon is started again, it will select the most recent one from the two storage locations as the master database information for local records.
[0152] To ensure compatibility under different database environments, the daemon of the present invention can parse the corresponding database type in the database configuration file according to the local IP, and can currently support the adaptation of oracle, mysql, kingbase, postgresql, dm five mainstream database systems. Secondly, each database node independently runs a daemon process (a process independently running in each database node in the daemon, in the present invention, the daemon can refer to all daemon processes), but the daemon itself does not make a decision on the main standby library. Whether the local database of a certain machine is the main library or the standby library is notified by the coordination program after the decision. When there is no coordination program involved, the daemon on each machine is only responsible for monitoring and publishing the database status of its own node. With the participation of the coordination program, it is distinguished whether the local database monitored by each daemon is the main library or the standby library. At this time, if the daemon of the main library fails, the function of monitoring the node status of the main library will be affected, resulting in the inability to determine whether the main library is actually available, which may cause an erroneous automatic switching decision.
[0153] In order to accurately judge the actual availability of the main database, the present invention introduces a main database offline alarm mechanism in the daemon. The daemon of the main database node will make a judgment by checking the connection status of the local database; in the standby database node, the daemon simulates a database access application and regularly attempts to establish a connection through the main database query interface. If it is found that the main database is inaccessible, the alarm information is published to the general alarm node " / global / system / alarm" through the status service to notify the system administrator. In addition, in order to ensure the security of the application connecting to the main database, when the application initiates a main database connection request, the request will first be sent to the daemon process. Before responding to the connection request, the daemon will verify the TCP (Transmission Control Protocol) communication status between itself and the coordinator. Through this verification mechanism, the daemon ensures that the main database information will only be passed to the inquiring party when the system status is normal, otherwise no response will be made.
[0154] 2. Synchronization Program
[0155] In order to make the data of the backup database consistent with the main database, the synchronization program has the basic function of synchronizing the updated data of the main database to the backup database, and each database node will run a synchronization program instance. The synchronization program will periodically connect to the main database and the local database (i.e., the local database). The role of the local database may be the main database or the backup database. At this time, if the local database is the main database, no operation will be performed. If the local database is the backup database, the synchronization program will perform the main-backup synchronization operation. In order to improve the efficiency of the synchronization process and reduce the amount of data transmission, the synchronization program of the present invention is set to use full synchronization when initializing synchronization or when there is a large version difference, otherwise incremental synchronization is used.
[0156] If the complete commit version numbers of the master and slave databases are inconsistent, there are two cases of inconsistency. The first case is that there is no record of complete commit in the historical version table of the slave database's commit database. The second case is that there is a complete commit in the historical version table of the slave database's commit database, but the version number or commit time corresponding to the most recent complete commit is inconsistent with that of the master database. In this case, the synchronization program will export the entire database (configuration database, commit database) from the master database to the slave database. If the maximum version number of the master database's commit database is greater than the maximum version number of the slave database's commit database (full commit and incremental commit are not distinguished at this time), the entire database will also be exported. If there is an inconsistency problem in the slave database itself, the maximum version number of the slave database's configuration database is less than 0, or the maximum version number of the slave database's configuration database is greater than the maximum version number of its own commit database, or the maximum version number of the slave database's configuration database is less than the maximum version number of its own commit database and exceeds the set threshold, the entire database will also be exported.
[0157] Incremental synchronization is generally used when the difference in commit numbers is less than the set threshold. The incremental synchronization work is triggered by the end of each incremental commit. That is, every time an incremental commit ends, the synchronization program will synchronize the standby database to be consistent with the section at the end of the incremental commit of the primary database. Its essence is that the standby database packages the configuration data modified between two adjacent incremental commits of the primary database and completes it in a one-time batch update.
[0158] The synchronization program of the present invention selects appropriate synchronization strategies according to different synchronization requirements and environments, thereby improving the flexibility and reliability of synchronization operations.
[0159] At the same time, in order to optimize system performance and storage usage, the synchronization program of the present invention includes a series of maintenance tasks, including regularly clearing old incremental data from the submission library to prevent unlimited storage space growth, and regularly clearing local storage files to prevent excessive disk space occupation. The standby library node regularly checks and imports new incremental data or compressed packages to the local submission library, and regularly checks the version of the configuration library, and updates the configuration library from the local submission library when necessary.
[0160] In addition, to ensure the reliability of the synchronization process and prevent potential data conflicts and errors caused by repeated operations, the present invention is equipped with a lock mechanism in the data import and export, configuration library update, and incremental data cleaning processes in the synchronization program. The lock mechanism is implemented by the QMutex class of the Qt framework. Each lock protects the corresponding operation or resource, which ensures that only one thread can perform data import and export, configuration library update, and incremental data cleaning at any time, avoiding data inconsistency problems and realizing serialization of operations.
[0161] like Figure 4 The figure shows the workflow of the synchronization program in the standby database node. The synchronization program monitors the incremental commit operation of the primary database, synchronously reads the commit information of the primary database when the incremental commit of the primary database is completed, determines the full or incremental synchronization method, and uses the backup tool to export the full database or incremental content of the primary database as a .dump file, transfers the .dump file to the standby database, and modifies the standby database configuration library to be consistent with the primary database.
[0162] 3. Main database query interface
[0163] The main database information structure in the query interface is as follows Figure 4 As shown, in order to realize the safe positioning and automatic identification of the application connection to the main library, the query function of the present invention is realized by sending a TCP (Transmission Control Protocol) message to the daemon program through the interface, which will connect to the daemon program of each database node. When the first daemon program responds, the query ends. The interface is called by the database interface, and the main library query function is automatically called internally when connecting to the configuration library.
[0164] 4. Coordination procedures
[0165] The coordination program starts in fault-tolerant mode and has the basic functions of main database information management and publishing. Each node has a daemon program. The decision of the main and standby database roles is executed by the coordination program. The initial role depends on the configuration in the system. The main database is specified when the automation system is installed. The default is the database of the first database server node. The subsequent role switching depends on the coordination program. In order to enable the daemons of different nodes to obtain consistent main database information, the coordination program regularly checks the main database information retained by each daemon program. If there is inconsistency, the latest main database information will be published, and each daemon program will be updated according to the published main database information.
[0166] At the same time, the coordination program regularly publishes database information, and publishes the main database information topic " / global / center / DBCoordinator / dbswitch", and the topic content is: main database serial number, station serial number, whether it is a center, priority URL serial number, main database host name, elector host name, and main election timestamp. When the new main database is switched, the topic will be published once at the same time. In addition, the coordination program regularly publishes the topic " / config / all_config_db_info" about all database information, and the topic content is: [main database serial number, main database host name, total number of databases, elector host name, election time]; [fault identification, standby database host name, configuration database version number, submission database version number, last synchronization time];...
[0167] In order to perceive the status of the master and standby database nodes, the coordinator of the present invention periodically checks the connection status with each daemon program to determine the availability of the database. The coordinator communicates with the daemon program of each machine node based on the socket technology of the Qt framework. When the socket connection fails, the coordinator will determine that the database status of the corresponding node is unknown. If the connection is successfully established, the daemon process will verify the connection status with the local database. If the verification is passed, the coordinator believes that the database of the node is in an available state; if the verification fails, it is determined to be unavailable. Therefore, the coordinator can perceive the status of the multi-node database as available, unknown or unavailable.
[0168] When the status is unavailable, it is identified as an abnormality. In the initial state, the coordination program determines that the main database is inaccessible according to the system configuration and needs to perform a master-slave database switch. Figure 5 As shown, when the coordination program of the present invention senses that the main database is abnormal, it makes a logical judgment on whether to switch, including main database status detection and fault type judgment. Figure 6As shown in the figure, the Qt-based coordination program interface visualizes the roles, status, and fault types of multi-node databases. When the status of the main database is unavailable, the main database is inaccessible. At this time, the network of the main database node is connected, the daemon can communicate, and the fault type is a service layer fault related to the original main database node database; when the status of the main database is unknown and there is no main database offline alarm, the main database is accessible, the network of the main database node is connected, and only the daemon cannot communicate, then there is no need to switch, and the fault type is a business layer fault related to the original main database daemon; if the status of the main database is unknown and there is a main database offline alarm issued by the standby database node, the main database is inaccessible, the network of the main database node is not connected, and the daemon cannot communicate. The fault type is a hardware layer fault related to the machine and network in the main database node.
[0169] After determining that the main database needs to be switched, in order to find a suitable standby database for switching, the coordinator of the present invention screens the available standby databases based on the daemon program. For each available standby database node, the synchronization program checks whether its configuration version and submission version are consistent with the original main database to ensure that the standby database data is the latest. If there are multiple standby databases that are available and synchronized, one is selected as the new main database to be switched according to the order of the database configuration files.
[0170] To ensure the data security and consistency of the system during automatic switching, the coordination program of the present invention notifies the application to disconnect before switching and waits for the application's response within a limited time. The coordination program does not directly cause the application to disconnect, but adopts a notification mechanism to let the application or user know the main database failure event in a timely manner through the status service within a limited time, and records this notification event in the log. After receiving the notification of the main database failure, it will decide whether to disconnect from the original main database based on the operating status and business needs of the application. After the time limit ends, the coordination program triggers automatic switching.
[0171] In summary, with the help of the above technical solution of the present invention, by integrating a multi-node daemon, a multi-node synchronization program, a main database query interface and a single instance coordination program, a database fault automatic switching system is jointly formed, aiming to achieve safe and accurate automatic switching when the main database fails; the daemon not only monitors and publishes the local database status regularly, but also adapts and is compatible with a variety of database types, and introduces a main database offline alarm mechanism to avoid misjudgment of the main database failure; the synchronization program is responsible for synchronizing the main database data changes to the standby database in real time, adding a lock mechanism during the synchronization process to ensure data consistency, flexibly adopting data synchronization strategies according to the size of the difference between versions, and timely optimizing system performance through regular maintenance tasks; the application realizes a secure connection with the main database through the main database query interface; the coordination program is not only responsible for the management and release of the main database information, but also can detect the main database status and identify the fault type, and promptly notify the application to disconnect, thereby achieving accurate anomaly detection and safe automatic switching when the main database fails, reducing the need for manual intervention, ensuring the safety of the main-standby automatic switching process, and effectively improving the stability and reliability of the database system.
[0172] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
Claims
1. A master-slave automatic fault switching system based on database status perception, characterized in that: The system includes: The daemon module is used to monitor the operation status of the local database in real time, periodically publish and update the local database status information, and distinguish and monitor the master database node and the standby database node according to the system command, introduce the master database offline alarm mechanism, and check the actual availability of the master database node; The synchronization program module is used to regularly connect the main database and the local database, execute differentiated synchronization strategies according to different synchronization requirements and environmental factors, synchronize the updated data of the main database to the backup database, and regularly optimize the performance and storage usage of various databases; The main database query interface module is used to send transmission control protocol messages to the daemon module, and realize the safe positioning and automatic identification of the application connection to the main database through query and response detection; The coordination program module is used to make decisions on the roles of the primary and standby databases, manage and publish primary database information, regularly update primary and database information, and sense the status of primary and standby database nodes. When the primary database needs to be switched, it screens the standby databases with available status and executes automatic switching of the primary and standby databases.
2. The active / standby automatic fault switching system based on database status perception according to claim 1 is characterized in that: The daemon module includes: a joint monitoring unit, a status service unit, a main database configuration unit, a database analysis unit and an offline alarm unit; The joint monitoring unit is used to monitor the local database in real time, and comprehensively obtain the configuration information, performance statistics and real-time status of the local database as the local database status information; The state service unit is used to execute the state service, create a key-value library that can be synchronized between designated database nodes, and periodically update the database state information by continuously running the state service; The master library configuration unit is used to set the master library node according to the master-standby library role decision command issued by the coordination program module, publish the master library version information of the current master library node, and locally retain the master library version information of the master library node; The database parsing unit is used to parse the corresponding database type in the database configuration file according to the local IP address, and each database node independently runs a daemon process; The offline alarm unit is used to check the connection status of the local database according to the main database offline alarm mechanism, determine the actual availability of the main database node, and trigger the alarm information.
3. The active / standby automatic fault switching system based on database status perception according to claim 2 is characterized in that: The step of setting the master library node according to the master-slave library role decision command issued by the coordination program module, publishing the master library version information of the current master library node, and locally retaining the master library version information of the master library node includes: Before the master database node is specified, the first database server node is used as the master database node by default. During use, the master and standby database role decision command issued by the coordination program module is obtained. If the database configuration of any machine is selected or switched to the master database server, the machine will be used as the master database node, and the communication between the master database node and the standby database node will be maintained through the network; Use the daemon process of the master library node to call the status service to publish the master library version information, which includes the master library serial number, whether it is a center, the configuration library version number and the submission library version number; The daemon process of the main library node is used to locally retain the main library version information and store it in specific files and election tables. When the daemon process is started again, the latest retained content is selected from the specific storage file and election table as the local recorded main library version information.
4. The active / standby automatic fault switching system based on database status perception according to claim 3 is characterized in that: According to the main database offline alarm mechanism, the connection status of the local database is checked to determine the actual availability of the main database node, and the triggering alarm information includes: Use the daemon process corresponding to the master database node to check the connection status of the local database. If the master database can be accessed normally, no alarm is triggered. If the master database cannot be accessed, the alarm information is published to the general alarm node through the status service to trigger the alarm information. When the application initiates a connection request to the main database, the request is sent to the daemon process corresponding to the main database node first, and the daemon process is used to verify the communication status between itself and the coordination program module. If the communication status is normal, the application accesses the main database normally. If the communication status is abnormal, no response is given; In the standby database node, the daemon process of the standby database node is used to simulate the database access application, and periodically attempts to establish a connection through the main database access interface, and then determines whether the main database can be accessed normally. If the main database can be accessed normally, no alarm is triggered. If the main database cannot be accessed, the alarm information is published to the general alarm node through the status service to trigger the alarm information.
5. The active / standby automatic fault switching system based on database status perception according to claim 1 is characterized in that: The synchronization program module includes: a full synchronization unit, an incremental synchronization unit, a regular maintenance unit and a resource protection unit; The full synchronization unit is used to regularly connect the main database and the local database. If the local database is the main database, no operation is performed. If the local database is the backup database, the main and backup synchronization operations are performed. When the synchronization is initialized or the version difference is greater than the set threshold, the full synchronization strategy is adopted to perform data synchronization between the main database and the backup database. The incremental synchronization unit is used to regularly connect the main database and the local database. If the local database is the main database, no operation is performed. If the local database is the backup database, the main and backup synchronization operations are performed. When the version difference is less than a set threshold, the incremental synchronization strategy is adopted to perform the main and backup synchronization operations. The periodic maintenance unit is used to optimize system performance and storage usage; The resource protection unit is used to equip the locking mechanism during the data import and export, configuration library update and incremental data cleaning process, so as to limit only one thread to perform the data import and export, configuration library update and incremental data cleaning work at any time.
6. The active / standby automatic fault switching system based on database status perception according to claim 5 is characterized in that: When the synchronization is initialized or the version difference is greater than the set threshold, the full synchronization strategy is adopted to perform data synchronization between the primary database and the standby database, including: Compare the submitted version numbers of the main database and the standby database. If the submitted version numbers of the main database and the standby database are inconsistent, export the entire main database to the standby database, where the entire database includes the configuration database and the submitted database. Compare the maximum version numbers of the libraries submitted by the master database and the library submitted by the standby database. If the maximum version number of the library submitted by the master database is greater than the maximum version number of the library submitted by the standby database, export the entire library from the master database to the standby database. If the maximum version number of the standby database's configuration library is less than 0, or the maximum version number of the standby database's configuration library is greater than the maximum version number of its own submitted library, or the maximum version number of the standby database's configuration library is less than the maximum version number of its own submitted library and the difference is greater than the set threshold, the entire primary library will be exported to the standby library.
7. The active / standby automatic fault switching system based on database status perception according to claim 6 is characterized in that: When the version difference is less than the set threshold, the incremental synchronization strategy is adopted to perform data synchronization between the primary database and the standby database, including: Compare the commit version numbers of the master and slave databases. If the commit version numbers of the master and slave databases are different and the difference is less than the set threshold, use the end of each incremental commit as the trigger point. Whenever an incremental commit ends, synchronize the slave database to the same section as the end of the incremental commit of the master database.
8. The active / standby automatic fault switching system based on database status perception according to claim 7 is characterized in that: The regular maintenance unit includes: an incremental cleaning subunit, a file cleaning subunit, a check update subunit and a verification update subunit; The incremental cleaning subunit is used to regularly clean up old incremental data submitted to the repository; The file cleaning subunit is used to regularly clean up local storage files; The check and update subunit is used to use the standby database node to periodically check and import new incremental data or compressed packages into the local submission database; The checking and updating subunit is used to regularly check the version of the configuration library and update the configuration library from the local submission library.
9. The active / standby automatic fault switching system based on database status perception according to claim 2, characterized in that: The coordination program module includes: an update publishing unit, a check and judgment unit, a node configuration unit and a master-slave switching unit; The update publishing unit is used to regularly check the main library version information retained by the daemon module, and if the information is inconsistent, publish the latest main library version information and call the daemon module to update; regularly publish database information and topics about database information; The checking and judging unit is used to regularly check the connection status with the daemon module, sense the status of the primary and standby database nodes, and judge the availability of the database; The node configuration unit is used to select a database that meets the requirements of the main database as the main database according to the availability of the database, and determine whether to perform the main database switch; The master-slave switching unit is used to screen the available slave libraries, use the synchronization program module to check the configuration version and the submission version of the slave library, find the slave library to switch the master library, and after the search is completed, use the notification mechanism to trigger the automatic switching of the master library.
10. The active / standby automatic fault switching system based on database status perception according to claim 9, characterized in that: The periodic checking of the connection with the daemon module, sensing the status of the primary and standby database nodes, and judging the availability of the database includes: The coordination process uses socket technology to communicate with the daemon process of each database node. When the socket connection fails, the database status of the corresponding database node is determined to be unknown. When the socket connection is successful, the daemon process is used to verify the connection with the local database. If the verification is successful, the database status of the corresponding database node is determined to be available. If the verification fails, the database status of the corresponding database node is determined to be unavailable.
11. The active / standby automatic fault switching system based on database status perception according to claim 10, characterized in that: The step of selecting a database that meets the requirements of the master database as the master database according to the availability of the database, and determining whether to perform the master database switch includes: When the database status is unavailable, the corresponding database node is determined to be abnormal. When the main database is detected to be abnormal, it is determined that the main database is inaccessible at this time and the main and standby databases need to be switched; When the database status of the primary database is unknown and there is no primary database offline alarm, it is determined that the primary database is accessible and the network is connected, only the daemon process cannot communicate, and there is no need to perform primary and standby database switching; When the database status of the primary database is unknown and there is a primary database offline alarm, it is determined that the primary database is inaccessible and the network is not connected, the daemon process cannot communicate, and the primary and standby databases need to be switched.
12. The active / standby automatic fault switching system based on database status perception according to claim 11, characterized in that: The screening of available standby databases, using the synchronization program module to check the configuration version and the submitted version of the standby database, finding the standby database for switching between the primary and standby databases, and after the search is completed, using the notification mechanism to trigger the automatic switching between the primary and standby databases include: Call the daemon process to filter out available standby databases, check whether the configuration version and commit version of the standby database are consistent with the original master database, and retain the consistent standby database. If there are multiple available and synchronized standby databases, select one as the new master database to be switched according to the order of the database configuration files; The coordination process is called to notify the application to disconnect before the primary and standby databases switch and wait for the application's response within a limited time. After the time limit expires, the coordination process is used to trigger automatic switching.
13. A method for automatic master-slave fault switching based on database status perception, using the automatic master-slave fault switching system based on database status perception according to any one of claims 1 to 12, characterized in that: The method comprises the following steps: S1. Monitor the operation status of the local database in real time, periodically publish and update the local database status information, and distinguish and monitor the master database node and the backup database node according to system commands, introduce the master database offline alarm mechanism, and check the actual availability of the master database node; S2. Regularly connect the main database and the local database, implement differentiated synchronization strategies based on different synchronization requirements and environmental factors, synchronize the updated data of the main database to the standby database, and regularly optimize the performance and storage usage of various databases; S3, sending a transmission control protocol message to the daemon module, and realizing the safe positioning and automatic identification of the application connection to the main library through inquiry and response detection; S4. Execute decisions on the roles of the primary and standby databases, manage and publish information about the primary database, regularly update information about the primary and standby databases, and sense the status of the primary and standby database nodes. When the primary database needs to be switched, select the standby databases with available status and execute automatic switching of the primary and standby databases.
Citation Information
Patent Citations
PostgreSQL master-slave database automatic switching system and PostgreSQL master-slave database automatic switching method
CN111338767A
Cited By
Synchronization and switching method and device for multi-level heterogeneous data sources and electronic equipment
CN121255862A
Intelligent switching method and device for main and standby fragmentation library nodes in distributed memory database
CN121412072A
Multi-disk deployment method, device and storage medium of MySQL database
CN122507725A