Data synchronization method and device and electronic equipment
By introducing specific processes and logging mechanisms into the primary database, the problem of taking into account performance and differences in data synchronization between primary and secondary databases is solved, and the performance improvement of primary database and the controllability of differences is achieved.
Patent Information
- Application Number
- CN202411983388.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
When synchronizing data from the primary and secondary databases, how to take into account the performance of the primary and secondary databases?
By introducing the write log process, send process and session process in the primary database, log information of the redo log not synchronized to the backup database, including timestamps and serial numbers. When the session process receives a transaction commit request, it obtains the earliest timestamp and the current timestamp in the pre-recorded log information, and performs transaction commit when the time interval is less than or equal to the preset time threshold.
While ensuring the performance of the primary database, the differences between the primary and secondary databases are controlled within a certain range, reducing the impact on the performance of the primary database, and making the differences between the primary and secondary databases controllable.
Smart Images

Figure CN119938402A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to database technology, and in particular to a data synchronization method, device and electronic device. Background Art
[0002] In a database cluster, the primary and standby databases work together. If the primary database fails, the standby database takes over and continues providing services. Therefore, while the primary database is running, data must be backed up to the standby database.
[0003] Currently, data synchronization between primary and standby databases mainly includes synchronous and asynchronous modes. In synchronous mode, after executing a transaction, the primary database waits for the standby database to fully synchronize its data before committing the transaction. In asynchronous mode, the primary database commits the transaction immediately after executing the transaction, without waiting for the standby database to fully synchronize its data.
[0004] However, synchronous mode can affect the performance of the primary database, while asynchronous mode can lead to significant data discrepancies between the primary and standby databases. Therefore, when synchronizing data between the primary and standby databases, it is crucial to balance the performance of the primary database and the discrepancies between them. Summary of the Invention
[0005] The embodiments of the present application provide a data synchronization method, apparatus, and electronic device, which can control the differences between the primary and backup databases within a certain range while ensuring the performance of the primary database.
[0006] In a first aspect, an embodiment of the present application provides a data synchronization method. A server cluster includes a primary database and a standby database. The method is applied to the primary database, where the primary database runs a log writing process, a sending process, and a session process. The method includes:
[0007] During the execution of a transaction, the sending process sends the redo logs generated during the execution of the transaction to the standby database;
[0008] The log writing process detects whether the current redo log is synchronized to the standby database, and if not, records log information of the current redo log, the log information including a timestamp;
[0009] When the session process receives a transaction commit request, it obtains the earliest first timestamp and the current second timestamp from at least one pre-recorded log information;
[0010] When the session process determines that the time interval between the first timestamp and the second timestamp is less than or equal to a preset time threshold, the transaction commit is executed.
[0011] In a possible implementation, the log information further includes the first sequence number of the current redo log;
[0012] The log writing process detects whether the current redo log is synchronized to the standby database. If not, it records the log information of the current redo log, including:
[0013] The log writing process obtains the second sequence number of the redo log synchronized to the standby database, the first sequence number of the current redo log, and the current timestamp;
[0014] If the first sequence number is greater than the second sequence number, the log writing process determines that there is a redo log that has not been synchronized to the standby database;
[0015] The log writing process records the first sequence number and the current timestamp as log information.
[0016] In a possible implementation, the log writing process records the first sequence number and the current timestamp as log information, including:
[0017] The log writing process locks the log queue and records the first sequence number and the current timestamp as log information at the end of the log queue; the log information in the log queue is arranged in chronological order;
[0018] The log writing process unlocks the log queue;
[0019] The session process obtains the earliest first timestamp from at least one pre-recorded log message, including:
[0020] The session process obtains a first timestamp in the log information at the head of the log queue.
[0021] In one possible implementation, the method further includes:
[0022] The sending process receives a feedback message returned by the standby database, where the feedback message is returned after the standby database completes synchronization of the redo log sent by the sending process;
[0023] The sending process obtains the second sequence number of the redo log synchronized to the standby database carried in the feedback message;
[0024] The sending process locks the log queue and determines whether the log queue is empty;
[0025] If the log queue is not empty, the sending process obtains the third sequence number in the log information at the head of the log queue;
[0026] The sending process compares the third sequence number with the second sequence number;
[0027] If the third sequence number is less than or equal to the second sequence number, the sending process deletes the log information in the log queue whose sequence number is less than or equal to the second sequence number;
[0028] The sending process unlocks the log queue.
[0029] In one possible implementation, the method further includes:
[0030] If the log queue is empty, the sending process sends a wake-up signal to all session processes in the waiting queue; the session processes in the waiting queue are session processes waiting for transaction submission;
[0031] The session process performs transaction commit based on the received wake-up signal.
[0032] In one possible implementation, the method further includes:
[0033] If the third sequence number is greater than the second sequence number, the sending process obtains the third timestamp in the log information at the head of the log queue and the fourth timestamp of each session process in the waiting queue; the session process in the waiting queue is the session process waiting for transaction submission;
[0034] The sending process calculates the time interval between the third timestamp and each fourth timestamp;
[0035] The sending process sends a wake-up signal to the target conversation process in the waiting queue; the target conversation process is a conversation process whose corresponding time interval is less than or equal to the preset time threshold;
[0036] The target session process performs transaction commit based on the received wake-up signal.
[0037] In a possible implementation, when the session process receives a transaction commit request, before obtaining the earliest first timestamp and the current second timestamp in at least one pre-recorded log information, the method further includes:
[0038] The session process detects whether the current redo log is synchronized to the standby database. If not, the log queue is locked and the log information of the current redo log is recorded at the end of the log queue.
[0039] The method further comprises:
[0040] When the session process determines that the time interval between the first timestamp and the second timestamp is greater than the preset time threshold, the transaction commit is not performed, and the session process and the current timestamp are recorded in the waiting queue;
[0041] The session process unlocks the log queue.
[0042] In a second aspect, an embodiment of the present application provides a data synchronization device, including:
[0043] A sending process is used to send redo logs generated during the execution of a transaction to the standby database;
[0044] A log writing process is used to detect whether the current redo log is synchronized to the standby database. If not, the log information of the current redo log is recorded, and the log information includes a timestamp.
[0045] A session process is used to obtain the earliest first timestamp and the current second timestamp in at least one pre-recorded log information when receiving a transaction commit request; when the session process determines that the time interval between the first timestamp and the second timestamp is less than or equal to a preset time threshold, execute transaction commit.
[0046] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor;
[0047] The memory stores computer-executable instructions;
[0048] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.
[0049] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementation methods of the first aspect.
[0050] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.
[0051] The data synchronization method, device and electronic device provided by the embodiments of the present application. The server cluster includes a primary database and a standby database, and the method is applied to the primary database, and the primary database runs a log writing process, a sending process and a session process; the method includes: in the process of executing a transaction, the sending process sends the redo log generated during the transaction execution process to the standby database; the log writing process detects whether the current redo log is synchronized to the standby database, and if not, records the log information of the current redo log, and the log information includes a timestamp; when the session process receives a transaction commit request, it obtains the earliest first timestamp and the current second timestamp in at least one pre-recorded log information; when the session process determines that the time interval between the first timestamp and the second timestamp is less than or equal to the preset time threshold, the transaction commit is executed. In this way, when the primary and standby databases are synchronized, the standby database is allowed to lose data for a period of time, which can reduce the impact on the performance of the primary database and control the difference between the primary and standby databases within a certain range. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0053] Figure 1 A flowchart of the data synchronization process provided for this application;
[0054] Figure 2 A schematic diagram of a log queue provided for this application;
[0055] Figure 3 A flowchart of the operations performed by a sending process provided by this application;
[0056] Figure 4 A schematic diagram of the relationship between a process and a log queue provided for this application;
[0057] Figure 5 A flow chart of a log writing process for recording log information provided by this application;
[0058] Figure 6 A flowchart of the operations performed by a sending process provided by this application;
[0059] Figure 7 A flowchart of operations performed by a session process provided by this application;
[0060] Figure 8 A schematic diagram of the structure of a data synchronization device provided by this application;
[0061] Figure 9 This is a schematic diagram of the structure of an electronic device provided in this application.
[0062] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0063] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0064] First, let’s explain the terms involved in this application:
[0065] Recovery Point Objective (RPO): A metric that measures the amount of production data lost after a disaster. It can be simply described as the maximum amount of data loss a facility can tolerate. In this document, this refers to the data gap between the primary and standby databases. If the primary database fails, the standby database will lose data within that gap.
[0066] Redo log (Write-Ahead Log, referred to as WAL): used to record data changes in the database. Each time a log is recorded, a corresponding WAL sequence number is generated, and the WAL sequence number is incremental.
[0067] Log Sequence Number (LSN): Monotonically increasing, it refers to the sequence number of the WAL. Each data operation generates a WAL, and each corresponding LSN has a unique LSN.
[0068] Primary database: A database that can be read and written. In a primary-slave database cluster, only one node is the primary database.
[0069] Standby database: A read-only database. In a master-slave database cluster, the standby database synchronizes data from the master database.
[0070] A database cluster consists of a primary and standby database. The primary database handles write operations and synchronizes changes to the standby database. The standby database replicates data from the primary database and primarily handles read operations, ensuring data consistency and availability. For example, in a PostgreSQL-based master-slave replication architecture, the primary database sends write operation logs to the standby database, which then updates its own data based on these logs, achieving data synchronization. Data synchronization between the primary and standby databases ensures that, if the primary database fails, the standby database can quickly take over its responsibilities and continue providing services, ensuring normal system operation.
[0071] In some implementations, data synchronization between the primary and standby databases can be performed in synchronization mode. Specifically, after the primary database records the data modification operation in the redo log, it sends the redo log to the standby database. Upon receiving the redo log, the standby database analyzes and processes the redo log. After the analysis and processing are complete, it sends a confirmation message to the primary database. Only after receiving the confirmation message from the standby database does the primary database determine that the redo log write was successful and feedback the successful data modification operation to the application.
[0072] However, in this possible implementation, if the data in the primary database has not been synchronized to the standby database due to network problems or other reasons, or if the primary database cannot receive confirmation information sent by the standby database, the primary database will not be able to commit transactions, which will affect the performance of the primary database.
[0073] In some implementations, data synchronization between the primary and standby databases can be performed asynchronously. Specifically, after the primary database records the data modification operation in the redo log, it determines that the redo log write is successful and reports the success of the data modification operation to the application. Of course, after recording the redo log, the primary database also sends the redo log to the standby database.
[0074] However, in this possible implementation, the difference between the primary and standby databases is unpredictable, and there may be a situation where the difference between the primary and standby databases is large. In this way, when the primary database fails, it may not be possible to upgrade the standby database to the primary database.
[0075] Based on this, the present application provides a data synchronization method that pre-sets a time threshold. When the primary database is preparing to commit a transaction, the method determines the time between the redo log synchronized to the standby database and the earliest time of the unsynchronized redo log. If the time interval between the two times is within the preset time threshold, the transaction is committed. This eliminates the need for the primary database to wait until the redo log is fully synchronized to the standby database before committing a transaction, improving the performance of the primary database. Furthermore, the difference between the primary and standby databases is controlled within a certain timeframe, making the difference between the primary and standby databases manageable.
[0076] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0077] Figure 1 A flowchart of the data synchronization process provided for this application.
[0078] It should be noted that the data synchronization method provided in this application is applied to the main database in the server cluster, and the main database runs a log writing process, a sending process, and a session process. Of course, the server cluster also includes a standby database.
[0079] like Figure 1 As shown, the method includes:
[0080] S101. During the execution of a transaction, a sending process sends redo logs generated during the transaction execution to a standby database.
[0081] A redo log is a log file used by databases to record changes made to data by transactions. During a transaction, whenever data in the database is modified (such as inserts, updates, and deletes), the database system records these changes in the redo log while modifying the data pages in memory.
[0082] The standby database has a dedicated process that receives redo logs from the sending process. After receiving the redo logs, the standby database stores them in local redo log files. Based on the received redo logs, the standby database performs the same operations in its local database to synchronize data with the primary database.
[0083] It is understood that after the sending process sends redo logs to the standby database, the standby database receives and synchronizes the redo logs. After synchronization is complete, the standby database sends a feedback message to the sending process of the primary database to inform the sending process of the redo log reception and processing status. Therefore, the sending process can receive feedback messages returned by the standby database.
[0084] For example, the feedback message carries the sequence number of the redo log that has been synchronized to the standby database. The feedback message may also carry status information of the standby database applying the redo log. The embodiments of the present application do not specifically limit the content of the feedback message.
[0085] S102: The log writing process detects whether the current redo log is synchronized to the standby database. If not, the log information of the current redo log is recorded, and the log information includes a timestamp.
[0086] Exemplarily, the log writing process can determine the second sequence number of the redo log synchronized to the standby database, the first sequence number of the current redo log, and the current timestamp based on the message fed back by the standby database; and compare the first sequence number with the second sequence number.
[0087] If the first sequence number is greater than the second sequence number, the log writer determines that there are redo logs that have not been synchronized to the standby database. The log writer records the first sequence number and the current timestamp as log information.
[0088] It is understandable that the first sequence number is greater than the second sequence number, indicating that the current redo log has not yet been synchronized to the standby database. Therefore, the log writing process needs to record the sequence number of the transaction log and the current timestamp.
[0089] In the embodiment of the present application, if the first sequence number is less than or equal to the second sequence number, it indicates that the current redo log has been synchronized to the standby database, and there is no need to record the first sequence number. The log writing process can continue to obtain the sequence number of the new redo log.
[0090] For example, the log writing process can obtain the serial number of the current redo log once every preset time period, that is, every preset time period, it can detect whether there is a redo log that has not been synchronized to the standby database, and record the serial number and timestamp of the redo log, so as to determine the time difference between the primary and standby databases during the synchronization process based on the timestamp corresponding to the recorded redo log.
[0091] In this way, by comparing the serial numbers, it is determined whether there is a redo log that has not been synchronized to the standby database. If so, the serial number and timestamp corresponding to the redo log are recorded. This allows the subsequent determination of the time difference between the primary and standby databases based on the recorded timestamp, thereby determining whether the transaction corresponding to the redo log can be committed.
[0092] In this application, a log queue can be used to record redo logs that have not been synchronized to the standby data. Figure 2 A schematic diagram of a log queue provided for this application.
[0093] like Figure 2 As shown, the log queue can be a first-in-first-out queue, and each item in the queue consists of a timestamp and a sequence number of the redo log corresponding to the time, and the timestamp and sequence number of each item are one-to-one corresponding.
[0094] In the log queue, the timestamp and sequence number are monotonically increasing, the minimum value is stored at the head of the queue, and the maximum value is stored at the tail of the queue.
[0095] Therefore, the process of the log writing process recording log information can be as follows: the log writing process locks the log queue, and records the first sequence number and the current timestamp as log information to the end of the log queue; the log information in the log queue is arranged in chronological order.
[0096] It is understandable that after the log writing process locks the log queue, the log queue can only be modified by the log writing process and cannot be accessed and modified by other processes.
[0097] After the log writing process finishes recording the log information, the log queue is unlocked.
[0098] In this way, the log writing process records the log information into the log queue after the log writing process records the redo log, which can avoid the situation where other processes modify the log queue while the log writing process records the redo log, causing the log queue to be more chaotic.
[0099] S103: When the session process receives the transaction commit request, it obtains the earliest first timestamp and the current second timestamp from at least one pre-recorded log information.
[0100] Exemplarily, the transaction commit request received by the session process may be sent by an application program.
[0101] Since redo logs are recorded in the log queue, and the timestamps of log information stored in the log queue are incremented, the session process can obtain the first timestamp in the log information at the head of the log queue, that is, the earliest first timestamp in at least one pre-recorded log information.
[0102] S104: When the session process determines that the time interval between the first timestamp and the second timestamp is less than or equal to a preset time threshold, execute transaction commit.
[0103] The preset time threshold can be set by the user based on actual conditions such as the purpose of the database cluster. For example, the preset interval can be 5 hours or 1 day. The embodiment of the present application does not specifically limit the preset time threshold.
[0104] It can be understood that the time interval between the first timestamp and the second timestamp is less than or equal to the preset time threshold, indicating that the time difference between the primary database and the standby database is within the preset time threshold, that is, within the range allowed by the user, and therefore, the transaction commit can be executed.
[0105] The data synchronization method provided in an embodiment of the present application pre-sets a time threshold. Before the primary database's session process commits data, it determines whether the time interval between the time of the redo log synchronized to the standby database and the earliest time of the unsynchronized redo log is less than or equal to the preset time threshold. If so, the transaction commit is performed. This eliminates the need for the primary database to wait until the redo log is fully synchronized to the standby database before executing a transaction commit, improving the performance of the primary database. Furthermore, the difference between the primary and standby databases is controlled within a certain timeframe, making the difference between the primary and standby databases manageable.
[0106] In this application, the session process can also record log information into the log queue.
[0107] For example, when the session process receives a transaction commit request, before obtaining the earliest first timestamp and the current second timestamp in at least one pre-recorded log information, the session process can detect whether the current redo log is synchronized to the standby database. If not, the log queue is locked and the log information of the current redo log is recorded at the tail of the log queue.
[0108] The method by which the session process can detect whether the current redo log is synchronized to the standby database is similar to the method by which the log writing process detects whether the current redo log is synchronized to the standby database. Please refer to the above embodiment and will not be repeated here.
[0109] It's understandable that both the log writer and the session process can log redo log information that hasn't been synchronized to the standby database into the queue. In other words, a single redo log message might be written to the log queue by both the log writer and the session process. Consequently, two log messages with the same or different sequence numbers might exist in the log queue.
[0110] In this way, the session process can also record the log information of the redo log that has not been synchronized to the standby database into the queue, which can avoid the situation where the log writing process misses the unsynchronized redo log records.
[0111] In a possible implementation, when the session process determines that the time interval between the first timestamp and the second timestamp is greater than a preset time threshold, the transaction commit is not performed, and the session process and the current timestamp are recorded in a waiting queue.
[0112] The time interval between the first timestamp and the second timestamp is greater than the preset time threshold, indicating that the time difference between the primary database and the standby database has exceeded the preset time threshold, that is, exceeds the allowable range. Therefore, the transaction cannot be committed and can enter the waiting state.
[0113] Exemplarily, when recording a session process into a waiting queue, log information of the redo log of the transaction corresponding to the session process can be added to the waiting queue, so that when the session process is awakened, the redo log can be sent directly to the standby database based on the log queue.
[0114] In this way, when the difference between the primary and standby databases exceeds a preset time threshold, the transaction commit is not executed, so that the difference between the primary and standby databases is controlled within a certain time, making the difference between the primary and standby databases controllable.
[0115] In this application, after the sending process sends the redo logs generated during the transaction execution to the standby database, it can receive a feedback message returned by the standby database, and based on the feedback information, execute the operation of deleting the log information of the redo logs that have been backed up to the standby database in the log queue, and waking up the session process in the waiting queue.
[0116] The following describes the process of sending the process to delete the log information of the redo log that has been backed up to the standby database in the log queue. Figure 3 As shown, Figure 3 A flowchart of the operations performed by a sending process provided in this application.
[0117] like Figure 3 As shown, the sending process can perform the following operations:
[0118] S301. Receive a feedback message returned by the standby database. The feedback message is returned after the standby database completes synchronization of the redo log sent by the sending process.
[0119] After the sending process sends the redo logs generated during transaction execution to the standby database, it can receive feedback messages returned by the standby database.
[0120] S302: The sending process obtains the second sequence number of the redo log synchronized to the standby database carried in the feedback message.
[0121] S303: The sending process locks the log queue and determines whether the log queue is empty.
[0122] After the sending process locks the log queue, it can prevent other processes from accessing and modifying the log queue.
[0123] Determine whether the log queue is empty, that is, determine whether there is log information in the log queue.
[0124] Exemplarily, when there is no log information in the log queue, an identifier indicating that the log queue is empty can be generated for the log queue. When the sending process detects the identifier, it can be determined that the log queue is empty. When the sending process cannot detect the identifier, it can be determined that the log queue is not empty.
[0125] S304: If the log queue is not empty, the sending process obtains the third sequence number in the log information at the head of the log queue.
[0126] If the queue is not empty, it means that there may be redo logs that have not been synchronized to the standby database. Therefore, the sending process obtains the third sequence number from the log information at the head of the log queue.
[0127] S305: The sending process compares the third sequence number with the second sequence number.
[0128] S306: If the third sequence number is less than or equal to the second sequence number, the sending process deletes the log information with a sequence number less than or equal to the second sequence number in the log queue.
[0129] If the third sequence number is less than or equal to the second sequence number, it means that the redo logs with the second sequence number and before the second sequence number have been synchronized to the standby database. Therefore, the sending process can delete the log information with a sequence number less than or equal to the second sequence number from the log queue.
[0130] S307: The sending process unlocks the log queue.
[0131] The sending process unlocks the log queue, allowing other processes to access and modify the log queue.
[0132] In this way, the log information in the log queue that has been synchronized to the redo log of the standby database is deleted by sending the log, so that the log information contained in the log queue is all log information of the redo log that has not been synchronized to the standby database. In addition, the sending process locks the log queue so that only the sending process can access and modify the log queue at the same time, which can prevent the log queue from being cluttered when multiple processes modify the log queue.
[0133] In this application, the sending process can wake up the session process in the waiting queue. The sending process waking up the session process can include the following two possible implementations:
[0134] In a possible implementation, if the log queue is empty, the sending process sends a wake-up signal to all session processes in the waiting queue; the session processes in the waiting queue are session processes waiting for transaction submission.
[0135] Accordingly, the session process executes transaction commit based on the received wake-up signal.
[0136] In this way, when the log queue is empty, it indicates that the current transaction has no redo logs that have not been synchronized to the standby database. At this time, the primary database has sufficient capacity to synchronize the redo logs of other transactions. Therefore, a wake-up signal can be sent to all session processes in the waiting queue to enable the waiting session processes to execute transaction commits. This can reduce the impact on primary database performance caused by session processes being in a waiting state, which occupies primary database resources. Compared with the synchronous mode, this application can improve primary database performance.
[0137] In another possible implementation, if the third sequence number is greater than the second sequence number, the sending process obtains the third timestamp in the log information at the head of the log queue and the fourth timestamp of each session process in the waiting queue; the session processes in the waiting queue are session processes waiting for transaction commit; and the sending process calculates the time interval between the third timestamp and each fourth timestamp.
[0138] The sending process sends a wake-up signal to a target session process in a waiting queue; the target session process is a session process whose corresponding time interval is less than or equal to a preset time threshold.
[0139] When the third sequence number is greater than the second sequence number, it indicates that the redo log corresponding to the log information at the head of the log queue has not yet been synchronized to the standby database. At this time, the primary database does not have sufficient capacity to synchronize the redo logs of other transactions. Therefore, it is possible to wake up the session process that meets the conditions instead of waking up all session processes.
[0140] Accordingly, the target session process executes transaction commit based on the received wake-up signal.
[0141] In this way, the sending process wakes up the session process in the waiting state, so that the session process in the waiting state can execute transaction submission, which can reduce the situation where the session process is always waiting and occupies the main database resources, affecting the performance of the main database. Compared with the synchronous mode, this application can improve the performance of the main database.
[0142] Based on the above embodiment, the relationship between the process running in the main database and the log queue can be as follows: Figure 4 shown. Figure 4 A schematic diagram of the relationship between a process and a log queue provided for this application.
[0143] like Figure 4 As shown, the log queue can be accessed and modified by multiple processes, including at least the log writer, the sender, and the session process. Items in the log queue are ordered, with the latest timestamp and the largest sequence number at the end of the queue and the oldest timestamp and the smallest sequence number at the head of the queue. If all redo logs have been synchronized to the standby database, the log queue is empty.
[0144] The log writer process can append the sequence number and timestamp of the current redo log to the end of the queue.
[0145] Exemplarily, the log writing process can record the current timestamp and sequence number at every preset time or when the transaction is committed, and append the recorded content to the end of the queue. The specific method of appending the sequence number and timestamp by the log writing process can be found in the above embodiment and will not be repeated here.
[0146] The sending process can remove the log information that has been synchronized to the standby database from the head of the queue and wake up the session process that meets the conditions.
[0147] The specific method of sending process removal log information and waking up the session process can be found in the above embodiment and will not be repeated here.
[0148] The session process records the timestamp T corresponding to the current transaction submission and obtains T1 from the head of the queue. If there is no T1 or T-T1 is less than the preset time threshold, the transaction can be submitted directly. Otherwise, the session process sleeps and waits.
[0149] Among them, T is recorded when the session process receives the transaction commit.
[0150] It should be noted that if there is no T1, the log queue is empty, that is, the transaction's redo logs have all been synchronized to the standby database, so the transaction can be committed directly. If T - T1 is less than the preset time threshold, the difference between the primary and standby databases is considered to be within a controllable range, and the transaction can be committed directly.
[0151] In order to further illustrate the data synchronization method of the present application, the operations performed by the log writing process, the sending process and the session process during the data synchronization process of the primary and standby databases are described below.
[0152] First, the operations performed by the log writing process are described. Figure 5 As shown, Figure 5 A flow chart of a log writing process for recording log information provided in this application.
[0153] like Figure 5 As shown, the log information recorded by the log writing process may include:
[0154] Step 1: Write log information or sleep.
[0155] Step 2: Get the current timestamp and the sequence number curLSN of the latest redo log at preset intervals.
[0156] The preset time can be 1 second or 2 seconds, and can be set according to actual conditions. The embodiment of this application does not make any specific limitations.
[0157] Step 3: Get the sequence number syncLSN of the redo log that has been synchronized to the standby database, compare curLSN with syncLSN, and determine whether the current redo log has been synchronized to the standby database.
[0158] If curLSN is less than or equal to syncLSN, the latest redo log has been synchronized to the standby database. You can proceed to step 6 below.
[0159] If curLSN is greater than syncLSN, the latest redo log has not been synchronized to the standby database. In this case, perform step 4.
[0160] Step 4: Lock the log queue and record the current timestamp and curLSN at the end of the log queue.
[0161] Step 5: Unlock the log queue.
[0162] Step 6: No operation, return to step 1.
[0163] Next, we will explain the operations performed by the sending process. Figure 6 As shown, Figure 6 A flowchart of the operations performed by a sending process provided in this application.
[0164] like Figure 6 As shown, the operations performed by the sending process may include:
[0165] Step 1: Send redo logs to the standby database or put it into hibernation.
[0166] Step 2: When receiving the message returned by the standby database, obtain the sequence number that has been synchronized to the standby database and update it to syncLSN.
[0167] Step 3: Lock the log queue.
[0168] Step 4: Determine whether the log queue is empty.
[0169] If the log queue is empty, proceed to step 9 below.
[0170] If the log queue is not empty, proceed to step 5 below.
[0171] Step 5: Get the timestamp Tn and sequence number Ln of the log queue header record.
[0172] Specifically, obtain the item at the head of the log queue and obtain the timestamp Tn and sequence number Ln recorded therein.
[0173] Step 6: Determine whether Ln is less than or equal to syncLSN.
[0174] If Ln <= syncLSN, the redo log corresponding to this item has been synchronized to the standby database, and you can proceed to step 7 below.
[0175] If Ln>syncLSN, the redo log corresponding to this item has not yet been synchronized to the standby database. You can perform the following step 8.
[0176] Step 7: Delete the current item from the head of the queue.
[0177] Step 8: Get the timestamp Tx of the session processes in the waiting queue, compare Tx with Tn in sequence, and wake up all session processes whose Tx-Tn is less than a preset time threshold.
[0178] Step 9: Wake up all waiting session processes.
[0179] Step 10: Unlock the log queue.
[0180] Finally, the operations performed by the session process are described in Figure 7 As shown, Figure 7 A flowchart of the operations performed by a session process provided in this application.
[0181] like Figure 7 As shown, the operations performed by the session process may include:
[0182] Step 1: Start a transaction and perform data operations.
[0183] Step 2: After receiving the transaction commit request from the application, the master and slave nodes commit the transaction.
[0184] Step 3. Obtain the current timestamp curTime, the sequence number curLSN of the current redo log, and the sequence number syncLSN synchronized to the standby database.
[0185] Step 4: Compare the curLSN and syncLSN to determine whether the current redo log has been synchronized to the standby database.
[0186] If curLSN ≤ syncLSN, the latest redo log has been synchronized to the standby database, and you can proceed to step 12.
[0187] If curLSN > syncLSN, the current redo log has not been synchronized to the standby database. You need to record the sequence number that has not been synchronized to the standby database in the log queue. You can perform the following step 5.
[0188] Step 5: Lock the log queue.
[0189] Step 6: Record the current timestamp curTime and sequence number curLSN to the end of the log queue.
[0190] Step 7: Get the timestamp Tn of the log queue header item record.
[0191] Step 8: Compare curTime with Tn to determine whether the difference between the primary and standby databases is within the preset range.
[0192] If curTime-Tn is less than the preset time threshold, it means that the difference between the primary and standby databases is within the preset range. The current transaction can be committed directly, and the following steps 11-12 can be executed.
[0193] If curTime-Tn is greater than or equal to the preset time threshold, it indicates that the current primary and standby databases differ too much, exceeding the preset range. The current transaction cannot be committed directly and needs to wait. In this case, execute step 9 below.
[0194] Step 9: Add the current queue to the waiting queue and record the current timestamp curTime in the waiting queue.
[0195] Step 10: Unlock the log queue and enter a dormant state, waiting to be awakened by the sending process.
[0196] If the time difference between the timestamp curTime of the current session process and the head of the log queue is less than the preset time threshold, it will be awakened by the sending process.
[0197] The method for the sending process to wake up the session process in the waiting state can be referred to the above embodiment, which will not be repeated here.
[0198] If the session process is awakened, the following step 12 may be executed.
[0199] Step 11: Unlock the log queue.
[0200] Step 12: Commit the transaction.
[0201] Step 13: When the transaction is committed, reply to the application with a message "transaction completed and committed".
[0202] In summary, the data synchronization method provided by the embodiment of the present application adds a new mode in addition to the traditional synchronous mode and asynchronous mode: the user can specify a preset time threshold. The preset time threshold is a value that quantifies the maximum time for which data is lost in units of time. The preset time threshold value ranges from 0 to 86400, and the corresponding unit is seconds. During the synchronization process between the primary and standby databases, the difference between the primary and standby databases can always be controlled to be less than or equal to the preset time threshold. When the user configures the preset time threshold to 0, it is equivalent to the synchronous mode, that is, no data can be lost; when it is configured to be greater than 0 and less than or equal to 86400 (seconds), it has higher performance than the synchronous mode. Compared with the asynchronous mode, when the primary database fails, the standby database will lose data of the preset time threshold at most.
[0203] Figure 8 A schematic diagram of the structure of a data synchronization device provided by this application is shown as follows: Figure 8 As shown, the data synchronization device 80 provided in this embodiment includes:
[0204] The sending process 801 is used to send the redo logs generated during the execution of the transaction to the standby database.
[0205] The log writing process 802 is used to detect whether the current redo log is synchronized to the standby database. If not, the log information of the current redo log is recorded. The log information includes a timestamp.
[0206] The session process 803 is used to obtain the earliest first timestamp and the current second timestamp in at least one pre-recorded log information when receiving a transaction commit request; when the session process determines that the time interval between the first timestamp and the second timestamp is less than or equal to the preset time threshold, execute the transaction commit.
[0207] In a possible implementation, the log information further includes the first sequence number of the current redo log.
[0208] The log write process 802 is specifically used to obtain the second sequence number of the redo log that has been synchronized to the standby database, the first sequence number of the current redo log, and the current timestamp; if the first sequence number is greater than the second sequence number, the log write process determines that there is a redo log that has not been synchronized to the standby database; and records the first sequence number and the current timestamp as log information.
[0209] In a possible implementation, the log writing process 802 is specifically configured to lock the log queue and record the first sequence number and the current timestamp as log information at the end of the log queue; the log information in the log queue is arranged in chronological order.
[0210] The log writing process 802 is specifically used to unlock the log queue.
[0211] The session process 803 is specifically used to obtain the first timestamp in the log information at the head of the log queue.
[0212] In one possible implementation, the sending process 801 is also used to receive a feedback message returned by the standby database, which is returned after the standby database completes synchronization of the redo log sent by the sending process; obtain the second sequence number of the redo log synchronized to the standby database carried in the feedback message; lock the log queue and determine whether the log queue is empty; if the log queue is not empty, the sending process obtains the third sequence number in the log information at the head of the log queue; compares the third sequence number with the second sequence number; if the third sequence number is less than or equal to the second sequence number, the sending process deletes the log information in the log queue with a sequence number less than or equal to the second sequence number; and unlocks the log queue.
[0213] In a possible implementation, the sending process 801 is further configured to send a wake-up signal to all session processes in the waiting queue if the log queue is empty; the session processes in the waiting queue are session processes waiting for transaction submission.
[0214] The session process 803 is further configured to execute transaction submission based on the received wake-up signal.
[0215] In one possible implementation, the sending process 801 is further configured to obtain, if the third sequence number is greater than the second sequence number, the third timestamp in the log information at the head of the log queue and the fourth timestamp of each session process in the waiting queue; the session processes in the waiting queue are session processes waiting for transaction submission; calculate the time interval between the third timestamp and each fourth timestamp; send a wake-up signal to the target session process in the waiting queue; the target session process is a session process whose corresponding time interval is less than or equal to a preset time threshold; the target session process executes transaction submission based on the received wake-up signal.
[0216] In one possible implementation, the session process 803 is further configured to detect whether the current redo log has been synchronized to the standby database. If not, the session process locks the log queue and records the log information of the current redo log to the end of the log queue. If the session process determines that the time interval between the first timestamp and the second timestamp is greater than a preset time threshold, the transaction is not committed, and the session process and the current timestamp are recorded in the waiting queue. The log queue is then unlocked.
[0217] The data synchronization device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.
[0218] Figure 9 This is a schematic diagram of the structure of an electronic device provided by this application. Figure 9 As shown, the electronic device 90 provided in this embodiment includes: at least one processor 901 and a memory 902. Optionally, the device 90 further includes a communication component 903. The processor 901, the memory 902, and the communication component 903 are connected via a bus 904.
[0219] During the specific implementation process, at least one processor 901 executes the computer-executable instructions stored in the memory 902, so that the at least one processor 901 performs the above method.
[0220] The specific implementation process of the processor 901 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.
[0221] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly implemented by a hardware processor or implemented by a combination of hardware and software modules in the processor.
[0222] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.
[0223] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be classified into address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.
[0224] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0225] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.
[0226] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0227] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.
[0228] The division of units is merely a logical functional division; actual implementations may employ alternative divisions, such as combining or integrating multiple units or components into another system, or omitting or disabling certain features. Furthermore, any direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units, either through an interface, electrical, mechanical, or other means.
[0229] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0230] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0231] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0232] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0233] Finally, it should be noted that those skilled in the art will readily identify other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The present invention is not limited to the precise structure described above and illustrated in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof. The scope of the present invention is limited solely by the appended claims.
Claims
1. A data synchronization method, characterized in that: The server cluster includes a main database and a standby database. The method is applied to the main database, which runs a log writing process, a sending process, and a session process. The method includes: During the execution of the transaction, the sending process sends the redo log generated during the execution of the transaction to the standby database; The log writing process detects whether the current redo log is synchronized to the standby database, and if not, records log information of the current redo log, wherein the log information includes a timestamp; When the session process receives a transaction commit request, it obtains the earliest first timestamp and the current second timestamp in at least one pre-recorded log information; When the session process determines that the time interval between the first timestamp and the second timestamp is less than or equal to a preset time threshold, the transaction commit is executed.
2. The method according to claim 1, characterized in that The log information also includes the first sequence number of the current redo log; The log writing process detects whether the current redo log is synchronized to the standby database. If not, the log information of the current redo log is recorded, including: The log writing process obtains the second sequence number of the redo log synchronized to the standby database, the first sequence number of the current redo log, and the current timestamp; If the first sequence number is greater than the second sequence number, the log writing process determines that there is a redo log that has not been synchronized to the standby database; The log writing process records the first sequence number and the current timestamp as log information.
3. The method according to claim 2, characterized in that The log writing process records the first sequence number and the current timestamp as log information, including: The log writing process locks the log queue, and records the first sequence number and the current timestamp as log information at the end of the log queue; the log information in the log queue is arranged in chronological order; The log writing process unlocks the log queue; The session process obtains the earliest first timestamp in at least one pre-recorded log information, including: The session process obtains a first timestamp in the log information at the head of the log queue.
4. The method according to claim 3, characterized in that The method further comprises: The sending process receives a feedback message returned by the standby database, where the feedback message is returned after the standby database completes synchronization of the redo log sent by the sending process; The sending process obtains the second sequence number of the redo log synchronized to the standby database and carried in the feedback message; The sending process locks the log queue and determines whether the log queue is empty; If the log queue is not empty, the sending process obtains the third sequence number in the log information at the head of the log queue; The sending process compares the third sequence number with the second sequence number; If the third sequence number is less than or equal to the second sequence number, the sending process deletes the log information in the log queue whose sequence number is less than or equal to the second sequence number; The sending process unlocks the log queue.
5. The method according to claim 4, characterized in that The method further comprises: If the log queue is empty, the sending process sends a wake-up signal to all session processes in the waiting queue; the session processes in the waiting queue are session processes waiting for transaction submission; The session process performs transaction submission based on the received wake-up signal.
6. The method according to claim 4, characterized in that The method further comprises: If the third sequence number is greater than the second sequence number, the sending process obtains the third timestamp in the log information at the head of the log queue and the fourth timestamp of each session process in the waiting queue; the session process in the waiting queue is the session process waiting for transaction submission; The sending process calculates the time interval between the third timestamp and each fourth timestamp; The sending process sends a wake-up signal to the target conversation process in the waiting queue; the target conversation process is a conversation process whose corresponding time interval is less than or equal to the preset time threshold; The target session process performs transaction submission based on the received wake-up signal.
7. The method according to claim 5 or 6, characterized in that: When the session process receives a transaction commit request, before obtaining the earliest first timestamp and the current second timestamp in at least one pre-recorded log information, the method further includes: The session process detects whether the current redo log is synchronized to the standby database, and if not, locks the log queue and records the log information of the current redo log to the tail of the log queue; The method further comprises: When the conversation process determines that the time interval between the first timestamp and the second timestamp is greater than the preset time threshold, the transaction commit is not executed, and the conversation process and the current timestamp are recorded in the waiting queue; The session process unlocks the log queue.
8. A data synchronization device, characterized in that: include: A sending process is used to send the redo logs generated in the process of executing the transaction to the standby database during the execution of the transaction; A log writing process is used to detect whether the current redo log is synchronized to the standby database, and if not, to record the log information of the current redo log, wherein the log information includes a timestamp; The session process is used to obtain the earliest first timestamp and the current second timestamp in at least one pre-recorded log information when receiving a transaction commit request; When the session process determines that the time interval between the first timestamp and the second timestamp is less than or equal to a preset time threshold, the transaction commit is executed.
9. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.