A database recovery method, device, equipment and medium

CN122777367APending Publication Date: 2026-09-18SHENZHEN INST OF COMPUTING SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610901269.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0004]本发明实施例提供一种数据库的恢复方法、装置、设备及介质,以解决如何对数据库的恢复过程进行优化,以降低恢复过程对数据库业务连续性的影响的问题

Benefits of technology

[0010] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned database recovery method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122777367A_ABST
    Figure CN122777367A_ABST
Patent Text Reader

Abstract

The application discloses a database recovery method, device and equipment and a medium, and comprises the following steps: the method is applied to a master cluster, logs unsynchronized by a fault instance are sent to a backup cluster based on a recovery instance, a recovery set is constructed, data pages in the recovery set are marked to obtain a marking result, when a session receives a page request for a target data page, if the target data page is a data page in the recovery set and there is a marking to be obtained, a page request is sent to the backup cluster, and the latest version of the target data page sent by the backup cluster is sent to the session. The method is applied to the backup cluster, after receiving the logs unsynchronized by the fault instance sent by the master cluster, playback is performed, when a page request is received, the latest version of the target data page is sent to the master cluster according to a target log sequence number of a log corresponding to the target data page and a latest log sequence number of a log played back at the current time. The influence of the recovery process on the continuity of database services is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of database technology, and in particular to a database recovery method, apparatus, device, and medium. Background Technology

[0002] In a master-slave database system, when an instance in the master cluster fails, the remaining instances need to recover the failed instance to continue providing services. Current technologies typically involve the following stages for online instance recovery: Fault detection and instance eviction: Quickly and accurately identify the failed instance and remove it from the cluster. Global resource recovery: Take over the global resources previously managed by the failed instance and restore their consistency. Recovery set confirmation: Read the failed instance's redo logs, construct a recovery set, and filter the recovery set. Log replay: Perform sequential log replay based on the recovery set. After these stages, the impact of the failed instance is completely eliminated, and the remaining online instances on the master node can continue providing services without obstruction. However, log replay is usually the most time-consuming step, and its duration is closely related to the size of the replay log data. When the data volume is extremely large, the log replay stage time may increase linearly, directly extending the overall database recovery window and thus causing a more severe impact on business continuity.

[0003] Therefore, optimizing the database recovery process to reduce its impact on database business continuity has become an urgent problem to be solved. Summary of the Invention

[0004] This invention provides a database recovery method, apparatus, device, and medium to address the problem of how to optimize the database recovery process to reduce its impact on database business continuity.

[0005] A database recovery method is provided, wherein the recovery method is applied to a primary cluster, the primary cluster and a backup cluster forming a primary-backup database system, the primary cluster comprising at least two database instances, including: When a database instance in the primary cluster fails, a database instance other than the one that failed is selected as a recovery instance. Based on the recovery instance, logs of the failed instance that were not synchronized to the backup cluster are sent to the backup cluster. The backup cluster is used to replay all the unsynchronized logs. The recovery instance constructs a recovery set based on the logs of the faulty instance, and marks each data page according to the status in the global resource information corresponding to each data page in the recovery set, to obtain the marking result; When a page request for a target data page is received from a session, if the target data page is a data page in the recovery set and the tagging result of the target data page contains a tag to be acquired, then the page request is sent to the backup cluster. The backup cluster is used to determine the target log sequence number of the log corresponding to the target data page from all unsynchronized logs when it receives the page request, and to determine the latest log sequence number of the logs that have been replayed for all unsynchronized logs at the current time. Based on the target log sequence number and the latest log sequence number, the backup cluster sends the latest version of the target data page to the primary cluster. Receive the latest version of the target data page sent by the backup cluster, and send the latest version of the target data page to the session.

[0006] A database recovery method, wherein the recovery method is applied to a standby cluster, and the primary cluster and the standby cluster form a primary-standby database system, comprising: After receiving the logs from the main cluster indicating that a faulty instance in the main cluster has not synchronized, all the unsynchronized logs are replayed. Upon receiving a page request for a target data page from the main cluster, the target log sequence number of the log corresponding to the target data page is determined from all unsynchronized logs, and the latest log sequence number of the logs that have been replayed to all unsynchronized logs at the current moment is determined. Based on the target log sequence number and the latest log sequence number, the latest version of the target data page is determined, and the latest version of the target data page is sent to the main cluster.

[0007] A database recovery device is applied to a primary cluster, wherein the primary cluster and a backup cluster form a primary-backup database system, and the primary cluster includes at least two database instances, including: The log synchronization module is used to select a database instance other than the one that failed as a recovery instance when a database instance in the primary cluster fails, and to send the logs of the failed instance that were not synchronized to the backup cluster to the backup cluster based on the recovery instance. The backup cluster is used to replay all the unsynchronized logs. The page marking module is used by the recovery instance to construct a recovery set based on the logs of the fault instance, and to mark each data page according to the status in the global resource information corresponding to each data page in the recovery set, so as to obtain the marking result; The page request module is used to send the page request to the backup cluster when it receives a page request for a target data page from a session, if the target data page is a data page in the recovery set and the tagging result of the target data page contains a tag to be obtained. The backup cluster is used to determine the target log sequence number of the log corresponding to the target data page from all unsynchronized logs when it receives the page request, and to determine the latest log sequence number of the logs that have been replayed for all unsynchronized logs at the current time. Based on the target log sequence number and the latest log sequence number, the backup cluster sends the latest version of the target data page to the primary cluster. The page receiving module is used to receive the latest version of the target data page sent by the backup cluster, and send the latest version of the target data page to the session.

[0008] A database recovery device, applied to a standby cluster, wherein the primary cluster and the standby cluster form a primary-standby database system, comprising: The log replay module is used to replay all unsynchronized logs after receiving logs from the faulty instance that failed in the main cluster. The sequence number determination module is used to determine the target log sequence number of the log corresponding to the target data page from all unsynchronized logs when receiving a page request for the target data page sent by the main cluster, and to determine the latest log sequence number of the log that has been replayed for all unsynchronized logs at the current time. The page sending module is used to determine the latest version of the target data page based on the target log sequence number and the latest log sequence number, and send the latest version of the target data page to the main cluster.

[0009] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the database recovery method described above.

[0010] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned database recovery method.

[0011] A database system with a primary cluster and a backup cluster forming a primary-backup architecture, wherein the primary cluster includes at least two database instances, the above-mentioned database recovery method, when applied to the primary cluster, selects the database instance other than the one that failed as the recovery instance when one database instance in the primary cluster fails. Based on the recovery instance, the logs of the failed instance that were not synchronized to the backup cluster are sent to the backup cluster. The recovery instance constructs a recovery set based on the logs of the failed instance. According to the status of each data page in the global resource information corresponding to each data page in the recovery set, each data page is marked to obtain the marking result. When a page request for a target data page is received from a session, if the target data page is a data page in the recovery set and the marking result of the target data page contains a mark to be obtained, the page request is sent to the backup cluster. The latest version of the target data page sent by the backup cluster is received and sent to the session. When the above database recovery method is applied to the standby cluster, after receiving the unsynchronized logs of the failed instance in the primary cluster sent by the primary cluster, all unsynchronized logs are replayed. When a page request for a target data page is received from the primary cluster, the target log sequence number corresponding to the target data page is determined from all unsynchronized logs, and the latest log sequence number of the logs that have been replayed from all unsynchronized logs at the current moment is determined. Based on the target log sequence number and the latest log sequence number, the latest version of the target data page is determined, and the latest version of the target data page is sent to the primary cluster.

[0012] In this approach, when an instance in the primary cluster fails, log replay is not performed on the primary cluster. Instead, the recovery instance sends the out-of-sync logs of the failed instance to the backup cluster for replay. The backup cluster can quickly replay these logs, ensuring that the data in the backup cluster quickly becomes consistent with the primary cluster during the online recovery of the primary cluster. This guarantees data consistency for subsequent page requests. When the primary cluster receives a page request from a business application, if the latest version of the data page is not available, it can obtain the latest version from the backup cluster. Since the recovery instance in the primary cluster can quickly provide services without log replay, the log replay process during online recovery in the primary cluster is eliminated, thus shortening the recovery time and reducing the impact of the recovery process on database business continuity. Attached Figure Description

[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1This is a flowchart of a database recovery method according to an embodiment of the present invention; Figure 2 This is another flowchart of a database recovery method according to one embodiment of the present invention; Figure 3 This is a flowchart of a database recovery method according to an embodiment of the present invention; Figure 4 This is a schematic diagram of a database recovery device according to an embodiment of the present invention; Figure 5 This is a schematic diagram of a database recovery device according to an embodiment of the present invention; Figure 6 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] In one embodiment, such as Figure 1 As shown, a database recovery method is provided, applied to a primary cluster, where the primary cluster and a backup cluster form a primary-backup database architecture. The primary cluster includes at least two database instances, and includes the following steps: Step S101: When a database instance in the primary cluster fails, select the database instance other than the one that failed as the recovery instance, and send the logs of the failed instance that were not synchronized to the backup cluster to the backup cluster based on the recovery instance.

[0017] In this embodiment, a server node in the primary cluster can be a database instance. A recovery instance can refer to a normally functioning database instance in the primary cluster that is selected to perform fault recovery tasks, excluding the database instance that has failed. A fault instance can refer to a database instance in the primary cluster that has failed due to an abnormal crash. Logs of a fault instance that have not been synchronized to the backup cluster can refer to database redo logs that were generated by the fault instance before the failure but have not yet been transmitted to the backup cluster.

[0018] Specifically, when any database instance in the primary cluster fails, a working database instance other than the one that failed is selected as the recovery instance. Based on the recovery instance, the logs that the failed instance did not synchronize to the backup cluster are determined and sent to the backup cluster so that the backup cluster can replay the logs that the failed instance did not synchronize.

[0019] Step S102: The recovery instance constructs a recovery set based on the logs of the faulty instance, and marks each data page according to the status in the global resource information corresponding to each data page in the recovery set, thus obtaining the marking result.

[0020] In this embodiment, the recovery set can refer to the set of data pages that the faulty instance updated before the fault, as determined by the logs of the faulty instance. The global resource information can refer to the unified data structure maintained in the primary cluster, which records the current status and version information of the data pages managed by all database instances. The status of the data page in the global resource information can refer to the version status of a specific data page recorded in the global resource information, used to identify whether the data page is the latest version. The marking result can refer to the record information formed by adding a pending acquisition identifier to the data page that needs to be obtained from the backup cluster based on the version status of the data page in the global resource information.

[0021] Specifically, the recovery instance determines the data pages that the faulty instance updated before the failure based on the logs of the faulty instance, constructs a recovery set based on the data pages that the faulty instance updated before the failure, and determines the version status of the data page from the global resource information corresponding to the data page for any data page in the recovery set. If the version status of the data page is not the latest status, the data page is marked as pending acquisition in the global resource information corresponding to the data page to obtain the marking result of the data page. If the version status of the data page is the latest status, the data page is not marked as pending acquisition.

[0022] Step S103: When a page request for a target data page is received from a session, if the target data page is a data page in the recovery set and there is a tag to be acquired in the tagging result of the target data page, then a page request is sent to the standby cluster.

[0023] Step S104: Receive the latest version of the target data page sent by the backup cluster, and send the latest version of the target data page to the session.

[0024] In this embodiment, the target data page can refer to the specific data page that the session wants to access when it initiates a page request. The page request can refer to the instruction that the session sends to the database system to read or access a specific data page. The tag to be acquired can refer to the tag information added to the global resource information of the data page based on the tag result, which indicates that the data page needs to obtain the latest version from the backup cluster.

[0025] Specifically, when a page request for a target data page is received from a session, if the target data page is a data page in the recovery set and the target data page's tagging result contains a tag to be acquired, that is, the page request is for a page updated by the faulty instance, but due to the fault of the faulty instance, the latest version of the data page does not exist in the primary cluster, then the page request is sent to the backup cluster. This allows the backup cluster to determine the target log sequence number of the log corresponding to the target data page from all unsynchronized logs when it receives the page request, as well as the latest log sequence number of the logs that have been replayed for all unsynchronized logs at the current moment. Based on the target log sequence number and the latest log sequence number, the backup cluster sends the replayed latest version of the target data page to the primary cluster. The primary cluster receives the latest version of the target data page sent by the backup cluster and sends the latest version of the target data page to the session.

[0026] For example, in a master-slave shared cluster environment, the master cluster consists of two instances, instance 1 and instance 2, and the slave cluster consists of one instance. Instance 1 of the master cluster modifies data items B1, B2, and B3. Then, instance 2 of the master cluster modifies data item B3. Before the logs of the modifications to data items B1, B2, and B3 by instance 1 of the master cluster are sent to the slave cluster, instance 1 of the master cluster fails. The recovery process in this embodiment can be as follows: a. The primary cluster instance 2, which is also the recovery instance, takes over the logs that the sending instance 1 did not send to the backup cluster and asynchronously waits for the replay completion notification. After the backup cluster finishes replaying, it notifies the primary cluster's recovery instance.

[0027] b. The recovery instance of the primary cluster performs global resource recovery.

[0028] c. The primary cluster recovery instance reads the logs of the failed instance 1 and constructs a recovery set, which contains data items B1, B2, and B3.

[0029] d. Quickly filter the recovery set based on global resource information. At this time, the global resource statuses corresponding to data items B1, B2, and B3 are: None, None, and Latest Version, respectively (the status of None can mean that only the faulty instance 1 has updated this data item, and the status is updated to None after the faulty instance crashes). Mark the data items B1 and B2 as pending acquisition in the global resource information, add data items B1 and B2 to the pending acquisition queue, and add data item B3 to the pending write queue.

[0030] e. At this time, the primary cluster can provide normal read and write services. When a session needs to access data item B2, since B2 is in a pending state, the data page is retrieved from the backup cluster. If the log sequence number of the log corresponding to data item B2 is less than or equal to the maximum log sequence number of the currently replayed log, it indicates that the backup cluster has replayed the latest version of data item B2. Otherwise, it waits until the latest version of data item B2 is replayed. After that, the backup cluster sends the latest version of data item B2 to the primary cluster. The recovery instance clears the pending retrieval mark for data item B2 on the global resources, removes data item B2 from the pending retrieval queue, adds the latest version of data item B2 to the pending disk queue, writes it to disk, and then removes it from the pending disk queue.

[0031] f. The background thread asynchronously traverses the queue to be retrieved. For data item B1, it retrieves the latest version of the data page from the backup cluster, clears the pending retrieval mark for data item B1 on the global resource, removes data item B1 from the queue to be retrieved, adds the latest version of data item B1 to the disk-to-write queue, writes it to disk, and then removes it from the disk-to-write queue.

[0032] g. After the background thread asynchronously writes data item B3 to disk, it is removed from the disk write queue.

[0033] In this embodiment, when an instance in the primary cluster fails, log replay is not performed in the primary cluster. Instead, the logs that were not synchronized by the failed instance are sent to the backup cluster for replay based on the recovery instance. The backup cluster can quickly replay these logs, so during the online recovery of the primary cluster, the data in the backup cluster can quickly become consistent with the primary cluster. This ensures data consistency for subsequent page sending. Therefore, when the primary cluster receives a page request from a business, if the latest version of the data page is not available, it can obtain the latest version of the data page from the backup cluster. Since the recovery instance of the primary cluster can quickly provide services without log replay, the log replay process during online recovery in the primary cluster is eliminated, thus shortening the recovery time and reducing the impact of the recovery process on the continuity of database services.

[0034] In one embodiment, such as Figure 2 As shown, a database recovery method is provided. In step S102 above, the recovery instance constructs a recovery set based on the logs of the failed instance. Each data page is marked according to its status in the global resource information corresponding to each data page in the recovery set. After obtaining the marking results, the method further includes the following steps: Step S201: The recovery instance adds the data pages with pending tags in all the tag results of the recovery set to the pending retrieval queue.

[0035] Step S202: The background asynchronously and in parallel traverses the data pages in the queue to be retrieved. For any data page in the queue to be retrieved, a page request for the data page is sent to the backup cluster, and the latest version of the data page sent by the backup cluster is received.

[0036] Step S203: Clear the pending tags in the tagging results of the data page and remove the data page from the pending retrieval queue.

[0037] In this embodiment, the queue to be retrieved can refer to a buffer queue created by the recovery instance, used to temporarily store all data pages that contain tags to be retrieved in the tagging results and need to retrieve the latest version from the backup cluster.

[0038] Specifically, the recovery instance can also add data pages with pending acquisition marks from all the marked results in the recovery set to the acquisition queue. The background asynchronously and in parallel traverses the acquisition queue. For any data page in the acquisition queue, it sends a page request for that data page to the backup cluster. When the backup cluster receives the page request for that data page, if the log sequence number of the currently replayed log is greater than the log sequence number of the log corresponding to that data page, it indicates that the latest version of that data page has been replayed. If the sequence number of the currently replayed log is less than or equal to the log sequence number of the log corresponding to that data page, it indicates that the latest version of that data page has not yet been replayed, and it waits until the latest version of that data page is replayed. Then, it sends the latest version of that data page to the backup cluster. After receiving the latest version of that data page, the backup cluster clears the pending acquisition mark in the global resource information corresponding to that data page and removes that data page from the acquisition queue.

[0039] Optionally, after removing the data page from the fetch queue, the process also includes: The recovery instance adds the latest version of the data page corresponding to the data page to the write queue; The background asynchronously and in parallel traverses the latest version data pages in the disk write queue. For any latest version data page in the disk write queue, the latest version data page is written to disk. After the latest version data page is written to disk, the latest version data page is removed from the disk write queue.

[0040] The write queue can refer to a buffer queue created by the recovery instance for temporarily storing the latest version of data pages that are to be persisted.

[0041] That is, after obtaining the latest version of the data page and removing the data page from the queue to be obtained, the recovery instance can also add the latest version of the data page to the queue to be written to disk. The background asynchronously and in parallel traverses the latest version of the data pages in the queue to be written to disk, performs a write operation on each latest version of the data page, and removes the latest version of the data page after writing to disk from the queue to be written to disk after the write operation is completed.

[0042] Optionally, in step S102 above, during the process of marking each data page according to the status in the global resource information corresponding to each data page in the recovery set, for any data page in the recovery set, the version status of the data page is determined from the global resource information corresponding to the data page. If the version status of the data page is the latest status, but no disk write operation has been performed, then the data page with the latest version status is also added to the disk write queue, so that when the recovery instance asynchronously traverses the disk write queue in the background, it performs a disk write operation on the data page with the latest status.

[0043] In this embodiment, by introducing a queue to be acquired and a queue to be written to disk, asynchronous parallel processing of data acquisition and persistence operations in the background is realized during the fault recovery process. This provides a timely data response channel for business requests, minimizes blocking delays caused by waiting for data pages to be acquired or written to disk, and reduces the impact of the recovery process on business continuity.

[0044] In one embodiment, such as Figure 3 As shown, a database recovery method is provided, applied to a standby cluster, in a database system where the primary cluster and standby cluster form a primary-standby architecture. The method includes the following steps: Step S301: After receiving the logs of the failed instance in the main cluster that has not been synchronized, the logs of all the failed instances are replayed.

[0045] Step S302: Upon receiving a page request for the target data page from the master cluster, determine the target log sequence number of the log corresponding to the target data page from all unsynchronized logs, and determine the latest log sequence number of the logs that have been replayed to all unsynchronized logs at the current moment.

[0046] Step S303: Based on the target log sequence number and the latest log sequence number, determine the target data page of the latest version and send the target data page of the latest version to the main cluster.

[0047] In this embodiment, the target log sequence number can refer to the log sequence number of the log that records the operation of the faulty instance on the target data page, and the latest log sequence number can refer to the maximum log sequence number of the log that has been replayed in the standby cluster.

[0048] Specifically, after receiving the logs from the primary cluster regarding the unsynchronized logs of the failed instance in the primary cluster, the standby cluster replays all the unsynchronized logs. When it receives a page request from the primary cluster for the target data page, it determines the log sequence number (target log sequence number) of the logs recording the operation of the failed instance on the target data page, and determines the maximum log sequence number (latest log sequence number) of the logs currently being replayed. Based on the target log sequence number and the latest log sequence number, it determines the latest version of the target data page to be replayed and sends the latest version of the target data page to the primary cluster.

[0049] Optionally, based on the target log sequence number and the latest log sequence number, the latest version of the target data page is determined, and the latest version of the target data page is sent to the main cluster, including: If the target log sequence number is less than or equal to the latest log sequence number, it is determined that the latest version of the target data page has been replayed at the current moment, and the latest version of the target data page is sent to the main cluster. If the target log sequence number is greater than the latest log sequence number, it is determined that the latest version of the target data page has not been replayed at the current time. All unsynchronized logs are replayed until the target log sequence number is less than or equal to the latest log sequence number. Then, the latest version of the target data page is sent to the main cluster.

[0050] That is, if the target log sequence number is less than or equal to the latest log sequence number, it means that the latest version of the target data page has been replayed, and the latest version of the target data page can be directly obtained and sent to the main cluster. If the target log sequence number is greater than the latest log sequence number, it means that the latest version of the target data page has not yet been replayed, and it is necessary to wait for the log to be replayed to the latest version of the target data page before sending the latest version of the target data page to the main cluster.

[0051] In this embodiment, when an instance in the primary cluster fails, log replay is not performed in the primary cluster. Instead, the logs that were not synchronized by the failed instance are sent to the backup cluster for replay based on the recovery instance. The backup cluster can quickly replay these logs, so during the online recovery of the primary cluster, the data in the backup cluster can quickly become consistent with the primary cluster. This ensures data consistency for subsequent page sending. Therefore, when the primary cluster receives a page request from a business, if the latest version of the data page is not available, it can obtain the latest version of the data page from the backup cluster. Since the recovery instance of the primary cluster can quickly provide services without log replay, the log replay process during online recovery in the primary cluster is eliminated, thus shortening the recovery time and reducing the impact of the recovery process on the continuity of database services.

[0052] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0053] In one embodiment, a database recovery device is provided, applied to a primary cluster, wherein the primary cluster and a backup cluster form a primary-backup database system. The primary cluster includes at least two database instances, and this database recovery device corresponds one-to-one with the database recovery methods described in the above embodiments. Figure 4 As shown, the database recovery device includes a log synchronization module 41, a page marking module 42, a page request module 43, and a page receiving module 44. Detailed descriptions of each functional module are as follows: The log synchronization module 41 is used to select a database instance other than the one that failed as a recovery instance when a database instance in the primary cluster fails, and to send the logs of the failed instance that were not synchronized to the backup cluster to the backup cluster based on the recovery instance. The backup cluster is used to replay all the unsynchronized logs. The page marking module 42 is used by the recovery instance to construct a recovery set based on the logs of the fault instance, and to mark each data page according to the status in the global resource information corresponding to each data page in the recovery set, so as to obtain the marking result; The page request module 43 is used to send the page request to the backup cluster when it receives a page request for a target data page from a session, if the target data page is a data page in the recovery set and the tagging result of the target data page contains a tag to be obtained. The backup cluster is used to determine the target log sequence number of the log corresponding to the target data page from all unsynchronized logs when it receives the page request, and to determine the latest log sequence number of the logs that have been replayed for all unsynchronized logs at the current time. Based on the target log sequence number and the latest log sequence number, the backup cluster sends the latest version of the target data page to the primary cluster. The page receiving module 44 is used to receive the latest version of the target data page sent by the backup cluster, and send the latest version of the target data page to the session.

[0054] Optionally, the above-mentioned page markup module 42 includes: A construction unit is used by the recovery instance to determine the data pages that the faulty instance updated before the failure based on the logs of the faulty instance, and to construct the recovery set based on the data pages that the faulty instance updated before the failure. The pending acquisition marking unit is used to determine the version status of any data page in the recovery set from the global resource information corresponding to the data page. If the version status of the data page is not the latest status, the data page is marked as pending acquisition in the global resource information corresponding to the data page to obtain the marking result of the data page.

[0055] Optionally, the database recovery device further includes: The first addition module is used by the recovery instance to add the data pages containing the tag to be acquired in all the tag results of the recovery set to the acquisition queue; The first traversal module is used to asynchronously and in parallel traverse the data pages in the queue to be acquired in the background. For any data page in the queue to be acquired, it sends a page request for the data page to the backup cluster and receives the latest version of the data page sent by the backup cluster. The mark clearing module is used to clear the unacquired marks in the mark results of the data page and remove the data page from the unacquired queue.

[0056] Optionally, the database recovery device further includes: The second addition module is used by the recovery instance to add the latest version of the data page corresponding to the data page to the disk to be written queue; The disk writing module is used to asynchronously and in parallel traverse the latest version data pages in the disk writing queue in the background. For any latest version data page in the disk writing queue, the module writes the latest version data page to disk. After the disk writing of the latest version data page is completed, the latest version data page is removed from the disk writing queue.

[0057] Specific limitations regarding the database recovery device can be found in the limitations of the database recovery method described above, and will not be repeated here. Each module in the aforementioned database recovery device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0058] In one embodiment, a database recovery device is provided, applied to a standby cluster, in a database system where the primary cluster and the standby cluster form a primary-standby architecture. This database recovery device corresponds one-to-one with the database recovery methods described in the above embodiments. Figure 5 As shown, the database recovery device includes a log playback module 51, a sequence number determination module 52, and a page sending module 53. Detailed descriptions of each functional module are as follows: The log replay module 51 is used to replay all the unsynchronized logs after receiving the logs of the faulty instance that failed in the main cluster sent by the main cluster. The sequence number determination module 52 is used to determine the target log sequence number of the log corresponding to the target data page from all unsynchronized logs when receiving a page request for the target data page sent by the main cluster, and to determine the latest log sequence number of the log that has been replayed for all unsynchronized logs at the current time. The page sending module 53 is used to determine the latest version of the target data page based on the target log sequence number and the latest log sequence number, and send the latest version of the target data page to the main cluster.

[0059] Optionally, the above-mentioned page sending module 53 includes: The first judgment unit is used to determine that the current time has been replayed to the latest version of the target data page if the target log sequence number is less than or equal to the latest log sequence number, and to send the latest version of the target data page to the main cluster. The second determination unit is used to determine that the latest version of the target data page has not been replayed at the current time if the target log sequence number is greater than the latest log sequence number, and to continue replaying all unsynchronized logs until the target log sequence number is less than or equal to the latest log sequence number, and then send the latest version of the target data page to the main cluster.

[0060] Specific limitations regarding the database recovery device can be found in the limitations of the database recovery method described above, and will not be repeated here. Each module in the aforementioned database recovery device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0061] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores outdated logs. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a database recovery method.

[0062] In one embodiment, a computer device is provided, including a memory, a first processor, a second processor, and a computer program stored in the memory and executable on the first and second processors. When the first processor executes the computer program, it implements the database recovery method applied to the master cluster described in the above embodiments, for example... Figure 1 As shown in S101-S104, or Figure 2 As shown, to avoid repetition, it will not be described again here. Alternatively, when the processor executes a computer program, it implements the functions of each module / unit in this embodiment of a recovery device for the database applied to the main cluster, for example, Figure 4 The functions of the log synchronization module 41, page marking module 42, page request module 43, and page receiving module 44 shown are not described again here to avoid repetition. When the second processor executes the computer program, it implements the database recovery method applied to the standby cluster in the above embodiments, for example... Figure 3 S301-S303, as shown, will not be described again here to avoid repetition. Alternatively, the processor executes a computer program to implement the functions of each module / unit in this embodiment of a database recovery device applied to the main cluster, for example, Figure 5 The functions of the log playback module 51, the sequence number determination module 52, and the page sending module 53 shown are not described again here to avoid duplication.

[0063] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a first processor, the computer program implements the database recovery method applied to the master cluster in the above embodiments, for example... Figure 1 As shown in S101-S104, or Figure 2 As shown, to avoid repetition, it will not be described again here. Alternatively, when the processor executes a computer program, it implements the functions of each module / unit in this embodiment of a recovery device for the database applied to the main cluster, for example, Figure 4The functions of the log synchronization module 41, page marking module 42, page request module 43, and page receiving module 44 shown are not described again here to avoid repetition. When this computer program is executed by the second processor, it implements the database recovery method applied to the standby cluster in the above embodiments, for example... Figure 3 S301-S303, as shown, will not be described again here to avoid repetition. Alternatively, the processor executes a computer program to implement the functions of each module / unit in this embodiment of a database recovery device applied to the main cluster, for example, Figure 5 The functions of the log playback module 51, the sequence number determination module 52, and the page sending module 53 shown are not described again here to avoid duplication.

[0064] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0065] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0066] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A database recovery method, characterized in that, The recovery method is applied to the primary cluster, which, together with the backup cluster, forms a primary-backup database system. The primary cluster includes at least two database instances, including: When a database instance in the primary cluster fails, a database instance other than the one that failed is selected as a recovery instance. Based on the recovery instance, logs of the failed instance that were not synchronized to the backup cluster are sent to the backup cluster. The backup cluster is used to replay all the unsynchronized logs. The recovery instance constructs a recovery set based on the logs of the faulty instance, and marks each data page according to the status in the global resource information corresponding to each data page in the recovery set, to obtain the marking result; When a page request for a target data page is received from a session, if the target data page is a data page in the recovery set and the tagging result of the target data page contains a tag to be acquired, then the page request is sent to the backup cluster. The backup cluster is used to determine the target log sequence number of the log corresponding to the target data page from all unsynchronized logs when it receives the page request, and to determine the latest log sequence number of the logs that have been replayed for all unsynchronized logs at the current time. Based on the target log sequence number and the latest log sequence number, the backup cluster sends the latest version of the target data page to the primary cluster. Receive the latest version of the target data page sent by the backup cluster, and send the latest version of the target data page to the session.

2. The database recovery method according to claim 1, characterized in that, The recovery instance constructs a recovery set based on the logs of the faulty instance, and marks each data page according to the status in the global resource information corresponding to each data page in the recovery set, obtaining a marking result, including: The recovery instance determines the data pages that the faulty instance updated before the failure based on the logs of the faulty instance, and constructs the recovery set based on the data pages that the faulty instance updated before the failure. For any data page in the recovery set, the version status of the data page is determined from the global resource information corresponding to the data page. If the version status of the data page is not the latest, the data page is marked as to be acquired in the global resource information corresponding to the data page, and the marking result of the data page is obtained.

3. The database recovery method according to claim 1, characterized in that, After the recovery instance constructs a recovery set based on the logs of the faulty instance, and marks each data page according to the status in the global resource information corresponding to each data page in the recovery set, and obtains the marking results, the process further includes: The recovery instance adds the data pages containing the tag to be retrieved in all the tag results of the recovery set to the retrieval queue; The background asynchronously and in parallel traverses the data pages in the queue to be retrieved. For any data page in the queue to be retrieved, a page request for the data page is sent to the backup cluster, and the latest version of the data page is received from the backup cluster. Remove the pending tags from the tagging results of the data page and remove the data page from the pending retrieval queue.

4. The database recovery method according to claim 3, characterized in that, After removing the data page from the queue to be retrieved, the process further includes: The recovery instance adds the latest version of the data page corresponding to the data page to the write queue; The background asynchronously and in parallel traverses the latest version data pages in the disk write queue. For any latest version data page in the disk write queue, the latest version data page is written to disk. After the write operation of the latest version data page is completed, the latest version data page is removed from the disk write queue.

5. A database recovery method, characterized in that, The recovery method is applied to a standby cluster, and the primary cluster and the standby cluster form a primary-standby database system, including: After receiving the logs from the main cluster indicating that a faulty instance in the main cluster has not synchronized, all the unsynchronized logs are replayed. Upon receiving a page request for a target data page from the main cluster, the target log sequence number of the log corresponding to the target data page is determined from all unsynchronized logs, and the latest log sequence number of the logs that have been replayed to all unsynchronized logs at the current moment is determined. Based on the target log sequence number and the latest log sequence number, the latest version of the target data page is determined, and the latest version of the target data page is sent to the main cluster.

6. The database recovery method according to claim 5, characterized in that, The step of determining the latest version of the target data page based on the target log sequence number and the latest log sequence number, and sending the latest version of the target data page to the main cluster, includes: If the target log sequence number is less than or equal to the latest log sequence number, it is determined that the current time has been replayed to the latest version of the target data page, and the latest version of the target data page is sent to the main cluster; If the target log sequence number is greater than the latest log sequence number, it is determined that the latest version of the target data page has not been replayed at the current time. All unsynchronized logs are replayed until the target log sequence number is less than or equal to the latest log sequence number. Then, the latest version of the target data page is sent to the main cluster.

7. A database recovery device, characterized in that, The recovery device is applied to the primary cluster, which, together with the backup cluster, forms a primary-backup database system. The primary cluster includes at least two database instances, including: The log synchronization module is used to select a database instance other than the one that failed as a recovery instance when a database instance in the primary cluster fails, and to send the logs of the failed instance that were not synchronized to the backup cluster to the backup cluster based on the recovery instance. The backup cluster is used to replay all the unsynchronized logs. The page marking module is used by the recovery instance to construct a recovery set based on the logs of the fault instance, and to mark each data page according to the status in the global resource information corresponding to each data page in the recovery set, so as to obtain the marking result; The page request module is used to send the page request to the backup cluster when it receives a page request for a target data page from a session, if the target data page is a data page in the recovery set and the tagging result of the target data page contains a tag to be obtained. The backup cluster is used to determine the target log sequence number of the log corresponding to the target data page from all unsynchronized logs when it receives the page request, and to determine the latest log sequence number of the logs that have been replayed for all unsynchronized logs at the current time. Based on the target log sequence number and the latest log sequence number, the backup cluster sends the latest version of the target data page to the primary cluster. The page receiving module is used to receive the latest version of the target data page sent by the backup cluster, and send the latest version of the target data page to the session.

8. A database recovery device, characterized in that, The recovery device is applied to the backup cluster, and the primary cluster and the backup cluster form a primary-backup database system, including: The log replay module is used to replay all unsynchronized logs after receiving logs from the faulty instance that failed in the main cluster. The sequence number determination module is used to determine the target log sequence number of the log corresponding to the target data page from all unsynchronized logs when receiving a page request for the target data page sent by the main cluster, and to determine the latest log sequence number of the log that has been replayed for all unsynchronized logs at the current time. The page sending module is used to determine the latest version of the target data page based on the target log sequence number and the latest log sequence number, and send the latest version of the target data page to the main cluster.

9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the database recovery method as described in any one of claims 1 to 4, or the database recovery method as described in any one of claims 5 to 6.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the database recovery method as described in any one of claims 1 to 4, or the database recovery method as described in any one of claims 5 to 6.