Management method, device, medium and system for multi-center database cluster
By dynamically monitoring the database cluster status and switching in real-time, combining data synchronization and caching services, the problem of insufficient cluster switching flexibility in multi-center database cluster management is solved, and management efficiency is improved.
Patent Information
- Application Number
- CN202510328710.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has poor flexibility in cluster switching in multi-center database cluster management, resulting in low management efficiency.
By calling the real-time exploration service to dynamically monitor the main database status of the database cluster, and switch to other available clusters when the main database is unavailable; calling the data synchronization service to synchronize the data of multiple database clusters in real time; calling the data cache service to temporarily store flow data during database cluster switching or network jitter, ensuring that the data can be written after switching or recovery.
It realizes effective switching during cluster switching, improves the management efficiency of database clusters, and solves the problem of insufficient flexibility in cluster switching in existing solutions.
Smart Images

Figure CN120144675A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of database cluster management. Specifically, it relates to a management method, device, medium, and system for a multi-center database cluster. Background Technique
[0002] Multi-center database cluster management is a common problem faced in the process of financial data processing. Currently, there is no standard solution, and all are in the exploratory stage.
[0003] The existing feasible solutions mainly include the following two:
[0004] (1) Manage multi-cluster data sources based on the configuration file method: Define the information of multiple data sources in the system configuration file, and select the corresponding data source as needed during runtime. This method will have a greater impact on the upper-layer services, is not flexible and efficient enough, and is only applicable to small-scale applications.
[0005] (2) Manage multi-cluster data sources based on the connection pool method: Use the connection pool to manage database connections and perform data source switching at the connection pool level. The connection pool can be configured with multiple data sources, and the application can obtain database connections from the connection pool through the API of the connection pool. This method has high configuration complexity, high requirements for server resources, and needs to establish connections with all database clusters in advance, resulting in resource waste.
[0006] That is, the existing solutions have poor flexibility during cluster switching in multi-center database cluster management. Summary of the Invention
[0007] The main purpose of this application is to provide a management method, device, medium, and system for a multi-center database cluster, so as to at least solve the problem that the existing solutions have poor flexibility during cluster switching in multi-center database cluster management.
[0008] To achieve the above purpose, according to one aspect of this application, a management method for a multi-center database cluster is provided. The method includes: invoking a real-time liveness detection service to dynamically monitor the status of the primary databases of multiple database clusters, and when it is detected that the current primary database is unavailable, selecting one of the other available database clusters as the new primary database; invoking a data synchronization service to synchronize the data of multiple database clusters in real time; invoking a data caching service to temporarily store the streaming data when the database cluster switches or there is network jitter, so as to complete the data writing after the database cluster switches or the network recovers.
[0009] Optionally, call the data synchronization service to synchronize the data of multiple database clusters in real time, including: when the master database receives a write operation, use the master database to record the write operation in the WAL log and mark the write operation in the WAL log as the committed state; asynchronously replicate the WAL log in the committed state to the slave database through the network, where the slave database continues to replicate the parsed data to other slave databases through the asynchronous streaming replication mechanism, and the parsed data is the data obtained after parsing the WAL log in the committed state.
[0010] Optionally, during the process of calling the real-time liveness detection service to dynamically monitor the status of the master databases of multiple database clusters, the method further includes: obtaining the priorities of each database cluster; performing liveness detection processing on the database clusters based on the priorities of each database cluster.
[0011] Optionally, during the process of calling the real-time liveness detection service to dynamically monitor the status of the master databases of multiple database clusters, the method further includes: updating the failure count once every time an unavailable database cluster is detected; generating an alarm message when the current failure count is greater than a preset number, and the alarm message is used to prompt that there is no available database cluster currently.
[0012] Optionally, call the data caching service to temporarily store the streaming data when the database cluster switches or there is network jitter, so as to complete the data writing after the database cluster switches or the network recovers, including: when the database cluster switches or there is network jitter, process the streaming data and write the message into the message cache middleware; when it is necessary to read the message, read the message in the message cache middleware and write the message into the database; when the database write is successful, perform ACK confirmation on the message and remove the message from the queue; when the database write fails, call the real-time liveness detection service again to dynamically monitor the status of the master databases of multiple database clusters, and when it is detected that the current master database is unavailable, use other available database clusters as the new master database and retry reading the message until the database write is successful.
[0013] Optionally, after writing the message into the database, the method further includes: when all database clusters are unavailable, continuously retain the message in the message cache middleware.
[0014] Optionally, calling the real-time liveness detection service includes: calling the real-time liveness detection service using a standard Rest-style HTTP interface or a WebSocket interface.
[0015] According to another aspect of the present application, there is provided a management device for a multi-center database cluster, the device comprising: a first processing unit, configured to dynamically monitor the status of the primary databases of multiple database clusters by invoking a real-time probing service, and in the case where the current primary database is detected to be unavailable, use one of the other available database clusters as the new primary database; a second processing unit, configured to call a data synchronization service to synchronize the data of multiple database clusters in real time; a third processing unit, configured to call a data caching service, and when a database cluster switch or network jitter occurs, temporarily store the streaming data to complete data writing after the database cluster switch or network recovery.
[0016] According to another aspect of the present application, there is provided a computer-readable storage medium, the computer-readable storage medium comprising a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute any of the above methods.
[0017] According to another aspect of the present application, there is provided a management system for a multi-center database cluster, the system comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include those for executing any of the above methods.
[0018] Applying the technical solution of the present application, by dynamically monitoring the status of the primary databases of multiple database clusters with a real-time probing service, and in the case where the current primary database is detected to be unavailable, using one of the other available database clusters as the new primary database, calling a data synchronization service to synchronize the data of multiple database clusters in real time, and calling a data caching service to temporarily store the streaming data when a database cluster switch or network jitter occurs, it can perform an effective switch during cluster switching compared with the existing solution, improving the management efficiency of the database cluster, and thus solving the problem of poor flexibility during cluster switching in the management of multi-center database clusters in the existing solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The specification drawings forming a part of the present application are used to provide a further understanding of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0020] Figure 1 It shows a schematic flowchart of a management method for a multi-center database cluster provided by an embodiment of the present application;
[0021] Figure 2 It shows a schematic diagram of the principle of a management method for a multi-center database cluster provided by an embodiment of the present application;
[0022] Figure 3 Shows a schematic flowchart of the operation of a data synchronization module provided according to an embodiment of the present application;
[0023] Figure 4 Shows a schematic flowchart of the operation of a real-time probing module provided according to an embodiment of the present application;
[0024] Figure 5 Shows a schematic flowchart of the operation of a data caching module provided according to an embodiment of the present application;
[0025] Figure 6 Shows a structural block diagram of a management device for a multi-center database cluster provided according to an embodiment of the present application. Detailed implementation manners
[0026] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0027] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data used may be interchanged under appropriate circumstances so as to describe the embodiments of the present application here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0029] As introduced in the background art, multi - center database cluster management is a common problem faced in the process of financial data processing. Currently, there is no standard solution, and all are in the exploratory stage. The existing feasible solutions mainly include the following two: (1) Managing multi - cluster data sources based on configuration files: Define the information of multiple data sources in the system configuration file, and select the corresponding data source as needed during runtime. This method has a greater impact on the upper - layer services, is not flexible and efficient enough, and is only applicable to small - scale applications. (2) Managing multi - cluster data sources based on connection pools: Use connection pools to manage database connections and switch data sources at the connection pool level. Multiple data sources can be configured in the connection pool, and the application can obtain database connections from the connection pool through the connection pool's API. This method has high configuration complexity, high requirements for server resources, requires establishing connections with all database clusters in advance, and there is a situation of resource waste. To solve the problem that the existing solutions have poor flexibility during cluster switching in multi - center database cluster management, the embodiments of this application provide a management method, device, medium, and system for multi - center database clusters.
[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0031] In this embodiment, a management method for a multi - center database cluster running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer - executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0032] Figure 1 It is a schematic flowchart of a management method for a multi - center database cluster provided according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:
[0033] Step S101, call the real - time probing service to dynamically monitor the status of the primary databases of multiple database clusters, and in the case where the current primary database is detected to be unavailable, use one of the other available database clusters as the new primary database;
[0034] Among them, during the process of calling the real - time probing service to dynamically monitor the status of the primary databases of multiple database clusters, the above - mentioned method further includes: obtaining the priorities of each database cluster; performing probing processing on the above - mentioned database clusters based on the priorities of each database cluster.
[0035] Specifically, database clusters with higher priorities usually have better performance or more resources. Probing these clusters first can ensure that when switching, the system can select the database with the best performance as much as possible, thereby improving the service quality and response speed of the overall system. In the case of a failure in the current primary database cluster, the live probing based on priorities can quickly locate the next available cluster as the new primary database in the preset order, which reduces the system recovery time and improves the efficiency of failover. The priority mechanism can be combined with the load balancing strategy to direct write operations to clusters with lower loads, avoiding all clusters from bearing high loads simultaneously, which helps to balance the workloads of each database cluster and improve the stability and efficiency of the entire system. In the extreme case where all database clusters are unavailable, the priority strategy can guide the system to attempt to connect to the last available database cluster. Even if its performance is poor or resources are limited, it can ensure that the system does not completely interrupt and maintains a minimum level of availability.
[0036] In addition, during the process of calling the live probing service to dynamically monitor the status of the primary databases of multiple database clusters, the above method further includes: updating the failure count once every time an unavailable database cluster is detected; generating an alarm message when the current failure count is greater than the preset number of times, and the above alarm message is used to prompt that there is no available database cluster at present. Thus, it can quickly identify and isolate the faulty database cluster, reduce the impact time of the failure on the entire system, and prevent the business service from continuously attempting to connect to the faulty database, thereby avoiding unnecessary resource consumption and performance loss. Generating an alarm message when it is detected that all database clusters have reached the preset number of failure times indicates that the system is currently facing serious data resource availability problems. At this time, the system can take emergency measures, such as starting a backup data source, reducing the concurrency of the service, or pausing non-core services, to ensure the continuity and stability of critical services. The generation of the alarm message helps the operation and maintenance personnel quickly locate the problem and take corresponding measures for troubleshooting or system recovery. At the data center level, this may include resource reconfiguration, emergency repair or expansion of the faulty cluster to restore the availability of the data service. The alarm message provides the system administrator with the health status information of the database cluster, which helps them analyze the cause of the failure, formulate preventive measures, or adjust the database usage strategy, such as reallocating priorities, optimizing the cluster configuration, or increasing cluster redundancy, to improve the overall availability and fault tolerance of the system.
[0037] The call to the live probing service in step S101 includes: calling the live probing service using a standard Rest-style HTTP interface or a WebSocket interface.
[0038] The standard RESTful HTTP interface is an interface designed based on the REST (Representational State Transfer) architectural style. It conforms to some design principles of REST, such as the unique identification of resources, the use of HTTP verbs for operations, and statelessness. This interface design is characterized by simplicity, flexibility, ease of understanding and extension, which can improve the maintainability and scalability of the system. The standard RESTful HTTP interface usually includes the URI of resources, HTTP methods (GET, POST, PUT, DELETE, etc.), data formats (such as JSON, XML), etc., enabling the client to communicate with the server via HTTP requests and perform operations on resources.
[0039] The REST (Representational State Transfer) style HTTP interface is widely used in web services due to its simple and stateless characteristics. Almost all programming languages and platforms support the invocation of RESTful APIs. WebSocket provides a protocol for full-duplex communication over a single TCP connection. It provides a persistent connection between the client and the server and is suitable for scenarios that require real-time communication. The adoption of these two interface standards ensures the seamless integration of the data source management device and the upper-layer business application, improving the compatibility and portability of the system. The WebSocket interface is based on a persistent connection, allowing the server to actively push information to the client. This enables the real-time liveness detection service to immediately notify relevant services when the state of the database cluster changes, greatly improving the speed of fault detection and response, and thus shortening the service recovery time. Although RESTful APIs may not be as real-time as WebSocket, their statelessness makes requests and responses faster and is suitable for frequent health check requests. RESTful APIs use the HTTP protocol and transmit data through status codes and lightweight data formats such as JSON / XML, which makes them relatively low in resource consumption, especially in a high-concurrency environment. Because of its full-duplex communication characteristics, WebSocket can reduce the round-trip time of requests, improve data transmission efficiency, and reduce network latency, and is suitable for scenarios that require high-speed and low-latency data transmission. RESTful APIs can be easily integrated with existing web security mechanisms such as HTTPS, OAuth, etc. to ensure the security of data transmission. Although WebSocket does not provide encryption by default, it can run on top of TLS / SSL to achieve secure real-time communication. Both of these interface standards support multiple authentication and authorization mechanisms to ensure secure communication between the data source management device and the upper-layer service.
[0040] Step S102: Invoke the data synchronization service to synchronize the data of multiple database clusters in real time;
[0041] Among them, invoking the data synchronization service to synchronize the data of multiple database clusters in real time includes: when the master database receives a write operation, use the above master database to record the write operation in the WAL log and mark the write operation in the WAL log as the committed state; asynchronously replicate the WAL log in the committed state to the slave database through the network. Among them, the slave database continues to replicate the parsed data to other slave databases through the asynchronous streaming replication mechanism, and the parsed data is the data obtained by parsing the WAL log in the committed state.
[0042] Specifically, the WAL log ensures that before any write operation is executed in the master database, these operations are pre-recorded in the log. Even if a failure occurs during the writing process, the integrity of the data can be guaranteed. The WAL log marked as the committed state ensures that all data received by the slave databases are completed transactions, avoiding the risk of data inconsistency. By asynchronously replicating the WAL log through the network, real-time or near-real-time data synchronization can be achieved, enabling the slave database to quickly obtain the latest state of the master database. The asynchronous streaming replication mechanism not only improves the speed of data synchronization but also reduces the real-time write latency of the master database, enabling the entire data source management system to quickly respond to the failure of the master database and achieve distributed disaster tolerance of data. The method of asynchronously replicating the WAL log reduces the write latency of the master database because the write operation and data synchronization can be carried out concurrently. At the same time, since the replication operation is non-blocking, the master database can continue to process new write requests, while the slave database can independently process the received WAL log, improving the resource utilization rate and response speed of the entire system. The existence of the WAL log provides a solid foundation for fault recovery. If the master database or the slave database fails, the system can recover the data through the WAL log to ensure the integrity of the data and the consistency of the system. The asynchronous streaming replication mechanism also ensures the continuity of data synchronization. Even if a certain slave database is temporarily unavailable, other slave databases can still continue to replicate the data, thereby reducing the impact of a single point of failure.
[0043] Among them, the real-time synchronization of data from multiple database clusters by invoking the data synchronization service can be applied in the following scenarios: When the data source management device starts up, initialize the data synchronization service, including setting synchronization policies (such as full synchronization, incremental synchronization), synchronization frequency, conflict resolution mechanisms, data format conversion rules, etc. Configure the connection information of each database cluster, including but not limited to database type, version, address, port, username, password, and specific configuration parameters of the database. Obtain the current master-slave status of each database cluster through the real-time liveness detection service, and automatically identify the master database and slave databases. When a failure or switch occurs in the master database, the data synchronization service can automatically identify the new master database and adjust the direction and strategy of data synchronization to ensure data consistency and correctness. The data synchronization service adopts a multi-threaded or asynchronous processing method to synchronize multiple database clusters in parallel, improving the speed and efficiency of data synchronization. Dynamically allocate synchronization threads according to the load conditions of each database to avoid additional load caused by synchronization operations and affect database performance. During the data synchronization process, record any synchronization errors and automatically or manually retry according to the preset retry policy to ensure that the data can be successfully synchronized eventually. Introduce data verification and error recovery mechanisms, such as using the Write-Ahead Log (WAL) and streaming replication mechanisms of the database for data consistency verification to ensure that the synchronized data is exactly the same as the source data. In the extreme case where all database clusters are unavailable, the data synchronization service can start the local cache mechanism to temporarily store the data that fails to be synchronized. After any cluster recovers, immediately perform data synchronization or rollback. Implement logging and monitoring of the data synchronization status to ensure that during fault recovery, the last synchronization time point can be quickly located, reducing the complexity and time of data recovery. Optimize the data synchronization algorithm to reduce the latency of data synchronization and ensure that data can be synchronized in real-time or near real-time between database clusters. Dynamically adjust the synchronization frequency and rate according to business requirements and network conditions to balance real-time performance and resource consumption. Provide standard interfaces such as RESTful HTTP or WebSocket to allow upper-layer applications or external services to subscribe to data synchronization status updates or trigger specific data synchronization operations. Integrate closely with other key components such as the real-time liveness detection service and data cache service to form a closed-loop system for data source management to ensure the stability and reliability of data synchronization.
[0044] The beneficial effects applied to the above scenarios are as follows: Through real-time data synchronization, it can ensure that the data of all database clusters is consistent. Even in a distributed environment, data differences can be eliminated, which is crucial for industries such as finance and e-commerce that have extremely high requirements for data accuracy and consistency. The multi-threaded processing and load balancing features of the data synchronization service reduce the load on a single database cluster and avoid system performance bottlenecks caused by data synchronization operations, thereby improving the overall stability and availability of the system. Dynamically adjusting the frequency and rate of data synchronization, as well as allocating synchronization threads according to the database load conditions, makes resource utilization more efficient, reduces resource waste caused by unnecessary data synchronization, and improves the overall operation efficiency of the system. When all database clusters are unavailable, the local cache mechanism can ensure that data is not lost. Once any cluster is restored, data synchronization is immediately performed, improving data integrity and security. At the same time, the fast data recovery mechanism reduces the impact time of faults on the business.
[0045] Step S103, call the data cache service. When the database cluster switches or there is network jitter, temporarily store the streaming data to complete the data writing after the database cluster switches or the network is restored.
[0046] In the above steps, by dynamically monitoring the status of the primary database of multiple database clusters with a real-time probing service, and in the case of detecting that the current primary database is unavailable, taking one of the other available database clusters as the new primary database, calling the data synchronization service to synchronize the data of multiple database clusters in real time, and calling the data cache service to temporarily store the streaming data when the database cluster switches or there is network jitter, it can perform an effective switch during cluster switching compared with the existing solution, improving the management efficiency of the database cluster, thus solving the problem of poor flexibility in cluster switching in the management of multi-center database clusters in the existing solution.
[0047] In an embodiment of the present application, calling the data cache service to temporarily store the streaming data when the database cluster switches or there is network jitter to complete the data writing after the database cluster switches or the network is restored includes: when the database cluster switches or there is network jitter, process the streaming data and write the message into the message cache middleware; in the case of needing to read the above message, read the above message from the above message cache middleware and write the above message into the database; in the case of successful writing to the above database, perform an ACK confirmation on the above message and remove the above message from the queue; in the case of failed writing to the above database, call the real-time probing service again to dynamically monitor the status of the primary database of multiple database clusters, and in the case of detecting that the current primary database is unavailable, take the other available database clusters as the new primary database and re-attempt to read the above message until the writing to the above database is successful.
[0048] ACK confirmation refers to the confirmation message sent by the receiving end after receiving a data packet in computer network communication. ACK is the abbreviation of the English word Acknowledgement, indicating that the data packet sent by the sending end has been successfully received. ACK confirmation is usually used in the TCP protocol to ensure the reliable transmission of data packets. Through ACK confirmation, the sending end can know that the data packet has successfully reached the receiving end, and thus can continue to send the next data packet. ACK confirmation is an important mechanism in network communication, ensuring the reliable transmission of data and the stability of communication.
[0049] After writing the above message into the database, the above method further includes: in the case where all database clusters are unavailable, continuously retaining the above message in the above message caching middleware.
[0050] Specifically, in the case of database cluster switching or unreliable network, the data caching middleware becomes a temporary buffer for storing streaming data. Even when the primary database is unavailable, the write operation can still be successfully executed, ensuring data persistence and non-loss, which is particularly important for critical business scenarios such as financial transactions and order processing. By writing the streaming data into the message cache first, even during the process of primary database switching, the application can continue to process the business logic without waiting for the database switching to complete, which greatly improves the service continuity and user experience. When the database write fails, by calling the real-time liveness detection service again and dynamically monitoring and switching to the new primary database, it can ensure that the data can ultimately be successfully written into the database. This mechanism enhances the fault recovery ability of the data source management system, and can ensure data correctness and system stability even in extreme cases. The message caching middleware supports asynchronous processing, which means that data writing and reading can be performed non-blockingly, improving the system throughput and response speed, and is particularly suitable for high-concurrency business scenarios. The use of the message caching middleware can balance the load between database clusters, because the write operation is first processed in the cache and then dynamically allocated to different database clusters according to the database availability. In this way, even if a certain cluster is temporarily unavailable, other clusters can continue to process data, avoiding the impact of single-point failures.
[0051] This application implements a data source management device that encapsulates and manages database-related operations such as database connection, query, data synchronization, and switching, thereby achieving decoupling between the business system and database operations. The business system only needs to access this device and does not need to concern itself with database-related configurations and operations. By uniformly configuring and managing the database clusters of multiple data centers, this device ensures the high availability of the database and the security of data. Using the data synchronization module inside the device, real-time synchronization of data between multiple clusters is achieved; using the real-time probing module inside the device, dynamic switching between master and slave of the database clusters is achieved; using the data caching module inside the device, zero data loss during network jitter and master-slave switching is achieved. As an intermediate layer, this device provides a large number of extensible interfaces to provide additional functional enhancements. For example, functions such as data caching, data conversion, and data encryption can be implemented to provide more services and security guarantees. At the same time, this device supports access from multiple types of business services. Whether it is a file system or a third-party API, unified access and management can be performed through a proxy. This scalability and compatibility enable the system to adapt to different business types and access requirements.
[0052] As Figure 2 shown, the data source management device is located between the business application and the database cluster. It is internally divided into an interface layer, a business logic layer, and an infrastructure layer. The OpenAPI interface layer exposes standard Rest-style HTTP interfaces for upper-layer business applications to call. The business logic layer is divided into a data synchronization module, a real-time probing module, and a data caching module according to logical functions. The infrastructure layer is divided into a configuration management module and a core logic module. Among them, the configuration management module is responsible for managing information related to the configuration of multi-center database clusters; the core logic module is responsible for providing basic function support to the business logic layer. Figure 2 Each of the database clusters in Library. The data source management device is used to centrally manage the connection information of data sources, maintain and optimize the connection of data sources, and ensure that only authorized users or applications can access specific data through permission and authentication mechanisms. Configuration management is used to centrally manage and maintain the configuration information of systems, applications, or devices to ensure consistency, traceability, and efficiency, mainly including configuration distribution, change management, and configuration verification.
[0053] As Figure 3As shown in the figure, the operation process of the data synchronization module is as follows: After the implementation device is started, it probes the liveness of the database clusters in sequence according to the configured order, and selects the first cluster with successful connection as the primary cluster. After the primary cluster is successfully selected, all write transactions will be routed to the primary database cluster. The primary database cluster uses the WAL log of the database and the asynchronous streaming replication mechanism to synchronize data to other database clusters. The data with synchronization failure is written into the synchronization failure record file of the corresponding cluster. The missing transaction compensation service will regularly scan the synchronization failure record file and synchronize the data to the corresponding database cluster to achieve the final data consistency.
[0054] As Figure 4 shown in the figure, the operation process of the real-time liveness probe module is described as follows: After the device is normally started, the service will start a background thread to probe the liveness of the database, and regularly detect whether the primary database cluster is available; when the primary database cluster is unavailable, it is necessary to determine whether the automatic data source switching function is enabled. To prevent accidental data source switching caused by network jitter, it can be configured to turn off the automatic data source switching. When the automatic data source switching is turned off, the device will continuously probe the liveness of the existing primary database, and after reaching the configured number of failure times, the service will issue an alarm. When the automatic data source switching is turned on, when the number of liveness probe failures of the current primary database reaches the configured threshold, the device will probe the liveness of other central database clusters in sequence, and select the first cluster with successful liveness probe as the new primary database.
[0055] As Figure 5 shown in the figure, the operation process of the data caching module is described as follows: After the application service generates the transaction data that needs to be stored in the database, it calls the interface of this device; after processing the transaction information, it writes the message into the message caching middleware; the transaction data consumption thread reads the message from the message caching middleware and writes it into the database; after the database write is successful, it acknowledges the message and removes the message from the queue; if the database write fails, it will call the dynamic data source switching process to reselect the primary database. After the primary database is successfully selected, it will try to consume the cached data again; when all clusters are unavailable, the data will be cached in the message caching middleware for a long time to ensure that the application service can continue to provide services externally and the data will not be lost.
[0056] Advantages of the present application: Under the microservices architecture, by establishing an intermediate layer between the business service and the underlying data source, the business service is effectively decoupled from the specific data source, eliminating the strong dependence of the service on the database. The device can complete the switching of the underlying database without the application service being aware of it while ensuring a high degree of synchronization of the data in each cluster database. By combining the write-ahead log, streaming replication function of the database, and the message caching middleware, quasi-real-time synchronization of data between clusters is achieved, and a complete data recovery mechanism is realized. When the primary data source fails, the present application can automatically select another available data source as the primary data source to ensure the stability and reliability of the application program; in the extreme case where all database clusters are unavailable, the data can still be ensured not to be lost and the service can still be available through the caching service. Through the present application, isolation between the data source and the application program can be achieved. By controlling and managing the permissions of the device API interface, the security of the data and the privacy of the data source are improved.
[0057] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the management method for a multi-center database cluster of the present application will be described in detail below in conjunction with specific embodiments.
[0058] This embodiment relates to a specific management method for a multi-center database cluster, including:
[0059] Call the real-time liveness detection service using a standard Rest-style HTTP interface or a WebSocket interface to dynamically monitor the status of the primary databases of multiple database clusters, and in the case where the current primary database is detected to be unavailable, use one of the other available database clusters as the new primary database;
[0060] During the process of calling the real-time liveness detection service to dynamically monitor the status of the primary databases of multiple database clusters, obtain the priorities of each database cluster; perform liveness detection processing on the above-mentioned database clusters based on the priorities of each database cluster; update the failure count once each time an unavailable database cluster is detected; in the case where the current failure count is greater than a preset number of times, generate an alarm message, and the above-mentioned alarm message is used to prompt that there is no available database cluster currently;
[0061] In the case where the primary database receives a write operation, use the above-mentioned primary database to record the above-mentioned write operation in the WAL log, and mark the above-mentioned write operation in the above-mentioned WAL log as the committed state;
[0062] Asynchronously replicate the WAL log in the committed state to the slave databases through the network, where the slave databases continue to replicate the parsed data to other slave databases through the asynchronous streaming replication mechanism, and the parsed data is the data obtained after parsing the WAL log in the committed state;
[0063] Call the data caching service. When the database cluster switches or there is network jitter, temporarily store the streaming data to complete data writing after the database cluster switches or the network recovers.
[0064] Specifically, when the database cluster switches or there is network jitter, process the streaming data and write the message into the message caching middleware; in the case of needing to read the above message, read the above message in the above message caching middleware and write the above message into the database; in the case of successful writing to the above database, perform ACK confirmation on the above message and remove the above message from the queue; in the case of failed writing to the above database, call the real-time liveness detection service again to dynamically monitor the status of the primary databases of multiple database clusters, and in the case of detecting that the current primary database is unavailable, use other available database clusters as the new primary database and retry reading the above message until the above database writing is successful.
[0065] After writing the above message into the database, the above method further includes: in the case that all database clusters are unavailable, continuously retain the above message in the above message caching middleware.
[0066] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0067] The embodiment of the present application also provides a management device for a multi-center database cluster. It should be noted that the management device for the multi-center database cluster in the embodiment of the present application can be used to execute the management method for the multi-center database cluster provided by the embodiment of the present application. The device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0068] The following introduces the management device for the multi-center database cluster provided by the embodiment of the present application.
[0069] Figure 6 It is a structural block diagram of a management device for a multi-center database cluster provided according to an embodiment of the present application. As Figure 6 shown, the device includes:
[0070] The first processing unit 61 is used to call the real-time liveness detection service to dynamically monitor the status of the primary databases of multiple database clusters, and in the case of detecting that the current primary database is unavailable, use one of the other available database clusters as the new primary database;
[0071] The second processing unit 62 is used to call the data synchronization service to synchronize the data of multiple database clusters in real time;
[0072] The third processing unit 63 is used to call the data caching service. When the database cluster undergoes a switch or network jitter, it temporarily stores the streaming data to complete the data writing after the database cluster switch or network recovery.
[0073] In the above device, by using the real-time liveness detection service to dynamically monitor the status of the primary databases of multiple database clusters, and in the case of detecting that the current primary database is unavailable, using one of the other available database clusters as the new primary database, calling the data synchronization service to synchronize the data of multiple database clusters in real time, and calling the data caching service to temporarily store the streaming data when the database cluster undergoes a switch or network jitter, it can perform an effective switch during the cluster switch compared with the existing solution, improving the management efficiency of the database cluster, thereby solving the problem of poor flexibility in the cluster switch during the management of the multi-center database cluster in the existing solution.
[0074] In an embodiment of the present application, the second processing unit includes a first processing module and a second processing module. The first processing module is used to, when the primary database receives a write operation, use the above primary database to record the write operation in the WAL log and mark the write operation in the above WAL log as a committed state; the second processing module asynchronously replicates the WAL log in the committed state to the slave database through the network, where the slave database continues to replicate the parsed data to other slave databases through the asynchronous stream replication mechanism, and the parsed data is the data obtained after parsing the WAL log in the committed state.
[0075] In an embodiment of the present application, the first processing unit includes an acquisition module and a third processing module. The acquisition module is used to obtain the priority of each database cluster during the process of calling the real-time liveness detection service to dynamically monitor the status of the primary databases of multiple database clusters; the third processing module is used to perform liveness detection processing on the above database clusters based on the priority of each database cluster.
[0076] In an embodiment of the present application, the first processing unit includes a fourth processing module and a fifth processing module. The fourth processing module is configured to update the failure count once every time an unavailable database cluster is detected during the process of dynamically monitoring the status of the primary databases of multiple database clusters by invoking the real-time liveness detection service. The fifth processing module is configured to generate an alarm message when the current failure count is greater than a preset number, and the alarm message is used to prompt that there is no available database cluster currently.
[0077] In an embodiment of the present application, the third processing unit includes a sixth processing module, a seventh processing module, an eighth processing module, and a ninth processing module. The sixth processing module is configured to process the streaming data and write the message into the message cache middleware when a database cluster switches or there is network jitter. The seventh processing module is configured to read the message from the message cache middleware when the message needs to be read and write the message into the database. The eighth processing module is configured to perform an ACK confirmation on the message and remove the message from the queue when the writing into the database is successful. The ninth processing module is configured to call the real-time liveness detection service again to dynamically monitor the status of the primary databases of multiple database clusters when the writing into the database fails, and when it is detected that the current primary database is unavailable, use other available database clusters as the new primary database and retry reading the message until the writing into the database is successful.
[0078] In an embodiment of the present application, the third processing unit includes a tenth processing module, which is configured to continuously retain the message in the message cache middleware when all database clusters are unavailable after the message is written into the database.
[0079] In an embodiment of the present application, the first processing unit includes an eleventh processing module, which is configured to call the real-time liveness detection service by using a standard Rest-style HTTP interface or a WebSocket interface.
[0080] The above multi-center database cluster management device includes a processor and a memory. The first processing unit, the second processing unit, the third processing unit, etc. are all stored in the memory as program units, and the processor executes the program units stored in the memory to implement corresponding functions. The above modules are all located in the same processor; or, the above modules are respectively located in different processors in any combination form.
[0081] The processor contains a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem of poor flexibility in cluster switching in the existing solution for multi-center database cluster management can be solved.
[0082] The memory may include non-permanent memory in the form of computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0083] An embodiment of the present invention provides a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls a device where the computer-readable storage medium is located to execute the management method of the multi-center database cluster.
[0084] An embodiment of the present invention provides a processor, where the processor is used to run a program, and when the program runs, it executes the management method of the multi-center database cluster.
[0085] An embodiment of the present invention provides a device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements at least the following steps: calling a real-time probing service to dynamically monitor the status of the primary databases of multiple database clusters, and when it is detected that the current primary database is unavailable, taking one of the other available database clusters as the new primary database; calling a data synchronization service to synchronize the data of multiple database clusters in real time; calling a data caching service to temporarily store the streaming data when the database cluster switches or there is network jitter, so as to complete the data writing after the database cluster switches or the network recovers. The device herein may be a server, a PC, a PAD, a mobile phone, etc.
[0086] Optionally, calling a data synchronization service to synchronize the data of multiple database clusters in real time includes: when the primary database receives a write operation, using the primary database to record the write operation in the WAL log and marking the write operation in the WAL log as a committed state; asynchronously replicating the WAL log with the committed state to the slave databases through the network, where the slave databases continue to replicate the parsed data to other slave databases through an asynchronous streaming replication mechanism, and the parsed data is the data obtained by parsing the WAL log with the committed state.
[0087] Optionally, during the process of calling a real-time probing service to dynamically monitor the status of the primary databases of multiple database clusters, the method further includes: obtaining the priorities of each database cluster; performing a probing process on the database clusters based on the priorities of each database cluster.
[0088] Optionally, during the process of calling the real-time liveness detection service to dynamically monitor the status of the primary databases of multiple database clusters, the above method further includes: updating the failure count once for each detected unavailable database cluster; generating an alarm message when the current failure count is greater than a preset number, where the alarm message is used to indicate that there is no available database cluster currently.
[0089] Optionally, call the data caching service. When a database cluster switches or there is network jitter, temporarily store the streaming data to complete data writing after the database cluster switches or the network resumes, including: when a database cluster switches or there is network jitter, process the streaming data and write the message into the message caching middleware; when it is necessary to read the above message, read the above message from the above message caching middleware and write the above message into the database; when the above database write is successful, perform an ACK confirmation on the above message and remove the above message from the queue; when the above database write fails, call the real-time liveness detection service again to dynamically monitor the status of the primary databases of multiple database clusters, and when it is detected that the current primary database is unavailable, use another available database cluster as the new primary database and retry reading the above message until the above database write is successful.
[0090] Optionally, after writing the above message into the database, the above method further includes: when all database clusters are unavailable, continuously store the above message in the above message caching middleware.
[0091] Optionally, calling the real-time liveness detection service includes: calling the real-time liveness detection service using a standard Rest-style HTTP interface or a WebSocket interface.
[0092] This application also provides a computer program product, which when executed on a data processing device, is adapted to execute a program initialized with at least the following method steps: calling the real-time liveness detection service to dynamically monitor the status of the primary databases of multiple database clusters, and when it is detected that the current primary database is unavailable, using one of the other available database clusters as the new primary database; calling the data synchronization service to synchronize the data of multiple database clusters in real time; calling the data caching service, when a database cluster switches or there is network jitter, temporarily store the streaming data to complete data writing after the database cluster switches or the network resumes.
[0093] Optionally, call the data synchronization service to synchronize the data of multiple database clusters in real time, including: when the master database receives a write operation, use the above master database to record the write operation in the WAL log, and mark the write operation in the WAL log as the committed state; asynchronously replicate the WAL log in the committed state to the slave database through the network, where the slave database continues to replicate the parsed data to other slave databases through the asynchronous stream replication mechanism, and the parsed data is the data obtained after parsing the WAL log in the committed state.
[0094] Optionally, during the process of calling the real-time liveness detection service to dynamically monitor the status of the master databases of multiple database clusters, the above method further includes: obtaining the priorities of each database cluster; performing liveness detection processing on the above database clusters based on the priorities of each database cluster.
[0095] Optionally, during the process of calling the real-time liveness detection service to dynamically monitor the status of the master databases of multiple database clusters, the above method further includes: updating the failure count once every time an unavailable database cluster is detected; generating an alarm message when the current failure count is greater than the preset number of times, and the above alarm message is used to prompt that there is no available database cluster currently.
[0096] Optionally, call the data caching service to temporarily store the streaming data when the database cluster switches or there is network jitter, so as to complete the data writing after the database cluster switches or the network recovers, including: when the database cluster switches or there is network jitter, process the streaming data and write the message into the message caching middleware; when it is necessary to read the above message, read the above message in the message caching middleware and write the above message into the database; when the above database write is successful, perform ACK confirmation on the above message and remove the above message from the queue; when the above database write fails, call the real-time liveness detection service again to dynamically monitor the status of the master databases of multiple database clusters, and when it is detected that the current master database is unavailable, use other available database clusters as the new master database and retry reading the above message until the above database write is successful.
[0097] Optionally, after writing the above message into the database, the above method further includes: when all database clusters are unavailable, continuously retain the above message in the above message caching middleware.
[0098] Optionally, calling the real-time liveness detection service includes: calling the real-time liveness detection service using a standard Rest-style HTTP interface or a WebSocket interface.
[0099] The present application also provides a management system for a multi-center database cluster. The system includes: one or more processors, a memory, and one or more programs. Among them, the above one or more programs are stored in the above memory and are configured to be executed by the above one or more processors. The above one or more programs include those for executing any one of the above methods. By dynamically monitoring the status of the primary databases of multiple database clusters with a real-time probing service, and in the case of detecting that the current primary database is unavailable, taking one of the other available database clusters as the new primary database, invoking a data synchronization service to synchronize the data of multiple database clusters in real time, and invoking a data caching service to temporarily store the streaming data when there is a cluster switch or network jitter, it can perform an effective switch during cluster switching compared with the existing solutions, improving the management efficiency of the database cluster, and thus solving the problem of poor flexibility during cluster switching in the management of multi-center database clusters in the existing solutions.
[0100] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to be implemented. In this way, the present invention is not limited to any specific combination of hardware and software.
[0101] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0102] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate for implementing in the processFigure 1 means for the functions specified in one or more processes and / or blocks Figure 1 or multiple blocks.
[0103] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions in the process Figure 1 one or more processes and / or blocks Figure 1 or the functions specified in multiple blocks.
[0104] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions in the process Figure 1 one or more processes and / or blocks Figure 1 or the functions specified in multiple blocks.
[0105] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0106] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0107] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0108] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.
[0109] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0110] 1) The management method of the multi-center database cluster of the present application dynamically monitors the status of the primary databases of multiple database clusters through real-time probing services. And in the case where the current primary database is detected to be unavailable, one of the other available database clusters is used as the new primary database, and the data synchronization service is called to synchronize the data of multiple database clusters in real time. The data caching service is called to temporarily store the streaming data when the database cluster switches or there is network jitter. Compared with the existing solutions, it can perform effective switching during cluster switching, improving the management efficiency of the database cluster, thereby solving the problem of poor flexibility in cluster switching in the management of multi-center database clusters in the existing solutions.
[0111] 2) The management device of the multi-center database cluster of the present application dynamically monitors the status of the primary databases of multiple database clusters through real-time probing services. And in the case where the current primary database is detected to be unavailable, one of the other available database clusters is used as the new primary database, and the data synchronization service is called to synchronize the data of multiple database clusters in real time. The data caching service is called to temporarily store the streaming data when the database cluster switches or there is network jitter. Compared with the existing solutions, it can perform effective switching during cluster switching, improving the management efficiency of the database cluster, thereby solving the problem of poor flexibility in cluster switching in the management of multi-center database clusters in the existing solutions.
[0112] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A management method for a multi-center database cluster, characterized in that: include: Call the real-time liveness detection service to dynamically monitor the status of the master databases of multiple database clusters, and if it is detected that the current master database is unavailable, use one of the other available database clusters as the new master database; Call the data synchronization service to synchronize data of multiple database clusters in real time; Call the data cache service to temporarily store the streaming data when the database cluster switches or the network jitters, so as to complete the data writing after the database cluster switches or the network recovers.
2. The method according to claim 1, characterized in that Call the data synchronization service to synchronize data of multiple database clusters in real time, including: When the primary database receives a write operation, the primary database records the write operation in a WAL log, and marks the write operation in the WAL log as a committed state; The WAL log in the committed state is asynchronously copied to a slave database through a network, wherein the slave database continues to copy the parsed data to other slave databases through an asynchronous streaming replication mechanism, and the parsed data is the data after parsing the WAL log in the committed state.
3. The method according to claim 1, characterized in that In the process of calling the real-time live detection service to dynamically monitor the status of the master databases of the multiple database clusters, the method further includes: Get the priority of each database cluster; Perform activation processing on the database cluster based on the priority of each database cluster.
4. The method according to claim 1, characterized in that: In the process of calling the real-time live detection service to dynamically monitor the status of the master databases of the multiple database clusters, the method further includes: The failure count is updated every time an unavailable database cluster is detected; When the current failure count is greater than a preset number of times, an alarm message is generated, where the alarm message is used to prompt that there is currently no available database cluster.
5. The method according to claim 1, characterized in that Call the data cache service to temporarily store the streaming data when the database cluster switches or the network jitters, so as to complete the data writing after the database cluster switches or the network recovers, including: When the database cluster switches or the network jitters, the flow data is processed and the messages are written into the message cache middleware; When the message needs to be read, read the message in the message cache middleware and write the message into a database; If the database is written successfully, ACK the message and remove the message from the queue; In the event that the database write fails, the real-time liveness detection service is called again to dynamically monitor the status of the primary databases of multiple database clusters. If it is detected that the current primary database is unavailable, other available database clusters are used as new primary databases and the message is retried to be read until the database write is successful.
6. The method according to claim 5, characterized in that After writing the message into the database, the method further comprises: In the case that all database clusters are unavailable, the message is continuously stored in the message cache middleware.
7. The method according to any one of claims 1 to 6, characterized in that Call the real-time detection service, including: Use standard REST-style HTTP interface or WebSocket interface to call real-time detection service.
8. A management device for a multi-center database cluster, characterized in that: include: The first processing unit is used to call the real-time live detection service to dynamically monitor the status of the master databases of the multiple database clusters, and when it is detected that the current master database is unavailable, use one of the other available database clusters as a new master database; The second processing unit is used to call the data synchronization service to synchronize the data of multiple database clusters in real time; The third processing unit is used to call the data cache service, and when the database cluster switches or the network jitters, the flow data is temporarily stored to complete the data writing after the database cluster switches or the network recovers.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.
10. A management system for a multi-center database cluster, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of claims 1 to 7.