WAL-based time sequence database high-availability main and standby replication method and device

Through a dual-node master-slave architecture, log cache optimization, and batch transmission, it solves the node redundancy, synchronization inefficiency, and IO bottleneck problems of time series databases, and achieves highly available data synchronization and fault recovery, making it suitable for lightweight deployment and high-frequency write scenarios.

CN120704947APending Publication Date: 2025-09-26上海沄熹科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510804872.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Traditional master-slave replication solutions in time series databases suffer from high node redundancy, inefficient data synchronization, insufficient fault recovery reliability, and disk IO bottlenecks, making them difficult to meet the needs of lightweight deployment and high-frequency write scenarios.

Method used

It adopts a dual-node master-slave architecture, combined with log cache optimization and batch WAL transmission, and optimizes data synchronization and fault recovery processes through consistency check tables and breakpoint recovery mechanisms.

Benefits of technology

It reduces the complexity of distributed deployment and hardware costs, improves network transmission efficiency and system throughput, ensures zero data loss and zero inconsistency, and meets the reliability requirements of industrial-grade time series data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704947A_ABST
    Figure CN120704947A_ABST
Patent Text Reader

Abstract

The invention relates to the field of databases, and particularly provides a time sequence database high-availability master-slave replication method and device based on WAL, and the method comprises the following steps: S1, replacing a traditional three-node architecture with a master-slave dual-node system; s2, performing log cache optimization; s3, carrying out batch WAL transmission and dynamic control; and S4, recovering the consistency verification table and the breakpoint. Compared with the prior art, the method has the advantages that the problems of node redundancy, IO bottleneck and consistency of a traditional scheme can be solved through batch WAL transmission, dynamic cache optimization and strong consistency verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of database technology, and specifically provides a high-availability master-slave replication method and device for a time series database based on WAL. Background Art

[0002] Traditional master-slave replication solutions are commonly used in relational databases. Their distributed high-availability architecture typically relies on at least three nodes (e.g., master, slave, and arbitration nodes), resulting in high hardware costs and complex deployment. New time series databases, however, are centered around time series data, characterized by high-frequency writes and sequential access, and require lightweight deployment in distributed scenarios such as the Internet of Things and industrial monitoring. Existing solutions have the following drawbacks:

[0003] High node redundancy: The traditional three-node architecture is not suitable for the lightweight deployment requirements of time series databases, increasing the deployment cost of edge computing and embedded scenarios. Furthermore, the data volume of time series databases is much larger than that of relational databases. Traditional multi-node distributed solutions have difficulty scheduling large-scale data, and maintaining distributed metadata will result in many performance bottlenecks.

[0004] Inefficient data synchronization: Real-time log transmission or full replication does not utilize the batch characteristics of time series data. High-frequency I / O operations lead to delays in master-slave synchronization.

[0005] Insufficient fault recovery reliability: The lack of a systematic consistency verification mechanism makes LSN breakpoints and data loss more likely to occur after multiple failovers.

[0006] Disk I / O bottleneck: When the master database frequently reads WAL files, disk seek latency becomes a performance bottleneck, significantly affecting batch generation efficiency, especially in high-frequency write scenarios.

[0007] In summary, in response to the distributed high availability requirements of time series databases, there is an urgent need in this field to solve the node redundancy, IO bottleneck, and consistency problems of traditional solutions through batched WAL transmission, dynamic cache optimization, and strong consistency verification. Summary of the Invention

[0008] The present invention aims to address the deficiencies of the above-mentioned prior art and provides a highly practical WAL-based high-availability master-slave replication method for time series databases.

[0009] A further technical task of the present invention is to provide a reasonably designed, safe and applicable WAL-based time series database high-availability master-slave replication device.

[0010] The technical solution adopted by the present invention to solve its technical problem is:

[0011] A highly available master-slave replication method for a time series database based on WAL has the following steps:

[0012] S1: A dual-node system with a master database and a slave database replaces the traditional three-node architecture.

[0013] S2. Optimize log cache;

[0014] S3, batch WAL transmission and dynamic control;

[0015] S4. Consistency check table and breakpoint recovery.

[0016] Furthermore, in step S1, the master database receives read and write requests, generates WAL records and maintains the WAL log cache, and reads and transmits them to the slave database in batches;

[0017] The standby database synchronizes data with the primary database in real time, and seamlessly switches to a hot standby node when the primary database fails. It supports read-only queries to share the read pressure.

[0018] Furthermore, in step S2, the master database introduces a memory cache queue. When initializing the cache, the cache is initialized when the master and slave databases are enabled. During the initialization phase, the master database needs to obtain the current consistency check table from the slave database. The consistency check table records the LSN position of each table through WAL playback, indicating that the data before the current LSN of the table has been received and the playback is complete. The master database obtains the WALMgr of all tables based on the consistency check table and sets the current_read_lsn of the WALMgr to the consistency table LSN, indicating that the master database has completed reading the table before this LSN and will start reading from this LSN in the future.

[0019] First, set up a master-slave connection on the slave database. The slave database sends its own consistency check table to the master database. If it is the first time, the consistency check table is empty. Otherwise, it is a disconnection and reconnection. After receiving the consistency check table, the master database obtains the WAL_Mgr of all tables in the consistency check table and sets its own Current_read_lsn to the LSN of the consistency check table.

[0020] Furthermore, when the dual-write mechanism is implemented, when a new WAL record is generated, it is written to the disk file and the cache queue at the same time. The cache is a global cache with a configured maximum capacity. The cache write order is consistent with the WAL record generation order.

[0021] When performing priority cache reads, continuous LSN segment data is first extracted from the cache when generating batches. Only when LSN discontinuity is detected, a disk read is triggered to complete the missing records and update the cache;

[0022] The LSN continuous detection strategy is:

[0023] After the WALMgr corresponding to each table reads the WAL, it updates its current_read_lsn, obtains the latest LogEntry in the cache queue. If its start_lsn is equal to the current_read_lsn of the current WALMgr, it means it is continuous, and the LogEntry is read. If start_lsn < current_read_lsn, it means the LogEntry has been read and the LogEntry is discarded;

[0024] If start_lsn > current_read_lsn, it means that the LogEntries that have been read are not continuous with the LogEntries in the cache. At this time, the missing LogEntries need to be read from the WAL file, where the starting LSN of the missing part is the current_read_lsn of the current WALMgr, and the ending LSN is the start_lsn of the first LogEntry in the current buffer.

[0025] Furthermore, the cache eviction policy is as follows:

[0026] After the standby database confirms the batch application, the standby database updates the consistency check table in memory. The standby database persists the consistency check table to disk every certain period of time. After the master database reads the cache, it deletes the records within the corresponding LSN range in the cache, releases the memory space, and ensures that the cache only retains the latest unconfirmed log data;

[0027] When the cache of the master database is full and the standby database is not enough to consume the writes of the master database, the master database only writes to the WAL file after a new write, and does not write to the cache until the data in the cache is consumed and then the cache write is restarted;

[0028] When performing fault detection, a fault determination is triggered after a certain number of consecutive heartbeat timeouts. When the standby database detects a master database fault, it traverses the consistency check table and performs a WAL replay for all unapplied batches to ensure that the local database is updated to the latest LSN;

[0029] Switch to the read-write mode, generate a new version number for the check table, and broadcast the new master database address and the current LSN synchronization point;

[0030] When the node recovers and the original master database is repaired and connected as a new standby database, the breakpoint is determined by comparing the check table, and only the unconfirmed batches are synchronized, and the data integrity is ensured by verifying the check code.

[0031] Furthermore, in step S2, the master database encapsulates the WAL batches according to the double-threshold condition, that is, the cumulative log data reaches the preset size or the number of records reaches a certain number on the capacity threshold;

[0032] The time threshold exceeds the preset interval from the last batch generation. If either of the two conditions is met, batch packaging is triggered, and a batch unit containing the start LSN, end LSN, log data and check code is generated;

[0033] The master database dynamically adjusts batch parameters based on real-time network status and backup database load.

[0034] Furthermore, when the standby database is idle, the transmission frequency is increased to shorten the synchronization delay. When the primary database reads the batch process, the log cache contains logEntry1, LogEntry2, ..., LogEntryN. The primary database reads from the cache queue. At this time, the queue head is LogEntry1, and its start_lsn is 100, which is equal to the current_read_lsn of WAL_MgrA. Therefore, LogEntry1 is read, dequeued, and the current_read_lsn is set to the end_lsn of LogEntry1, 150.

[0035] The master database continues to read LogEntry2. At this time, the start_lsn of LogEntry2 is 100, which is less than the current_read_lsn of WAL_MgrB. This indicates that the log has already been read and does not need to be read again, so it is discarded directly.

[0036] The master database reads LogEntry3. At this time, the start_lsn of LogEntry3 is 500, while the current_read_lsn of WAL_MgrA is 150. This indicates that a termination occurred between reading and caching. The next LogEntry to be read should be 150, but there are no logs with LSNs from 150 to 500 in the cache. Therefore, the master database reads LogEntries with LSNs from 150 to 500 from the WAL file, and then reads LogEntry3. Then the current_read_lsn is set to the end_lsn of LogEntry3, which is 600.

[0037] When the number or volume of read LogEntries reaches the set batch capacity threshold, the primary database sends the read batch of logs to the secondary database for the next read.

[0038] Because the cache cannot be written to after it is full, the number of logs in the cache lags behind the actual WAL files. When all LogEntries in the cache are read, that is, when the cache is empty, the master database reads from the WAL files in each table, assembles the read logs into a batch request, and sends it to the slave database. At this time, the capacity threshold management policy still applies to reading from files.

[0039] Furthermore, in step S4, after receiving the batch, the standby database verifies the data integrity through the checksum and requests retransmission if it fails;

[0040] The primary and standby databases regularly compare and verify the LSN ranges in the tables. When the primary database recovers from a failure and becomes the standby database, the start_lsn is used to quickly locate the breakpoint, skipping the synchronized batches and achieving incremental catch-up.

[0041] A new checksum version number is generated for each master-slave switch to ensure the uniqueness and validity of the checkpoints after multiple rounds of failures. When the master database crashes and restarts, the slave database sends the consistency checksum. The master database reinitializes the log cache based on this consistency checksum and sets the current_read_lsn of all WAL_Mgrs.

[0042] A highly available master-slave replication device for a time series database based on WAL, comprising: at least one memory and at least one processor;

[0043] The at least one memory is configured to store a machine-readable program;

[0044] The at least one processor is used to call the machine-readable program to execute a high-availability master-slave replication method for a time series database based on WAL.

[0045] Compared with the prior art, the WAL-based time series database high-availability master-slave replication method and device of the present invention have the following outstanding beneficial effects:

[0046] This paper uses a dual-node active-standby architecture to adapt to time series databases, breaking the traditional three-node dependency and reducing the complexity and hardware costs of distributed deployments. A dynamic batch mechanism is designed based on the batch characteristics of time series data, combined with log caching to reduce disk I / O, achieving both improved network transmission efficiency and host performance.

[0047] A bidirectional checksum table records LSN synchronization points and checksums to ensure zero data loss and inconsistency after multiple rounds of failover, meeting the reliability requirements of industrial-grade time series data. The log cache mechanism specifically addresses the IO bottleneck of WAL files, significantly reducing the master-slave synchronization delay in high-frequency write scenarios in the IoT, thereby improving system throughput. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0049] Figure 1This is a schematic diagram of the cache initialization process in the high-availability master-slave replication method of a WAL-based time series database;

[0050] Figure 2 This is a flowchart of the master database reading LogEntry in the high-availability master-slave replication method of a WAL-based time series database. DETAILED DESCRIPTION

[0051] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention will be further described in detail below in conjunction with specific embodiments. Obviously, the embodiments described are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0052] A best embodiment is given below:

[0053] In this embodiment, a high-availability master-slave replication method for a time series database based on WAL has the following steps:

[0054] S1: A dual-node system with a master database and a slave database replaces the traditional three-node architecture.

[0055] The master database receives read and write requests, generates WAL records, maintains a WAL log cache, and reads and transmits them to the slave database in batches. This database synchronizes data with the master database in real time, seamlessly switching to a hot standby node in the event of a master failure. It supports read-only queries to offload read traffic. This architecture eliminates the need for arbitration nodes and monitors node status through bidirectional heartbeat detection (including LSN consistency), reducing node deployment costs by 50%.

[0056] S2. Optimize log cache;

[0057] The main library introduces a memory cache queue to reduce disk IO interactions for WAL files.

[0058] When the cache is initialized, the cache is initialized when the master and slave are turned on. During the initialization phase, the current consistency check table needs to be obtained from the slave database. The consistency check table records the LSN position of each table through WAL playback, indicating that the data before the current LSN of the table has been received and the playback is complete. The master database obtains the WALMgr of all tables based on the consistency check table and sets the current_read_lsn of WALMgr to the consistency table LSN, indicating that the master database has completed reading the table before this LSN and will start reading from this LSN subsequently.

[0059] The cache initialization process is as follows Figure 1As shown in the figure, first, set up the master-slave connection on the standby database. The standby database sends its own consistency check table to the master database. If it is the first time, the consistency check table is empty; otherwise, it is a reconnect after disconnection. After receiving the consistency check table, the master database obtains the WAL_Mgr of all Tables in the consistency check table and sets its own Current_read_lsn to the LSN of the consistency check table.

[0060] Double-write mechanism: When a new WAL record is generated, it is written to the disk file and the cache queue simultaneously. The cache is a global cache with a configurable maximum capacity (such as 1GB). The cache write order is the same as the generation order of the WAL record.

[0061] Priority cache reading: When generating batches, preferentially extract continuous LSN segment data from the cache. Only when it is detected that the LSN is discontinuous (such as the cache missing LSN = 1500), trigger disk reading to complete the missing records and update the cache.

[0062] LSN continuous detection strategy: After reading the WAL, the WALMgr corresponding to each table updates its own current_read_lsn. Obtain the latest LogEntry in the cache queue. If its start_lsn is equal to the current_read_lsn of the current WALMgr, it means it is continuous, and read this LogEntry.

[0063] If start_lsn < current_read_lsn, it means this LogEntry has been read and discard this LogEntry; if start_lsn > current_read_lsn, it means that the LogEntry that has been read is not continuous with the LogEntry in the cache. At this time, it is necessary to read the missing LogEntry from the WAL file, where the starting LSN of the missing part is the current_read_lsn of the current WALMgr, and the ending LSN is the start_lsn of the first LogEntry in the current buffer.

[0064] Cache eviction strategy: After the standby database confirms the batch application, the standby database updates the consistency check table in its memory. The standby database persists the consistency check table to disk every certain period of time. After the master database reads from the cache, it deletes the records within the corresponding LSN range in the cache to release memory space, ensuring that the cache only retains the latest unconfirmed log data. This mechanism can reduce disk I / O operations by 60%-80% in high-frequency write scenarios, and the batch generation efficiency is increased by more than 40%. When the cache of the master database is full and the standby database is not enough to consume the writes of the master database, the master database only writes to the WAL file after a new write, and does not write to the cache until the data in the cache is consumed and then resumes writing to the cache.

[0065] Fault detection: Three consecutive heartbeat timeouts (3 seconds by default) trigger a fault determination. When the standby database detects a primary database failure:

[0066] Traverse the consistency check table and perform WAL replay on all unapplied batches to ensure that the local database is updated to the latest LSN;

[0067] Switch to read-write mode, generate a new checksum version number, and broadcast the new master database address and current LSN synchronization point;

[0068] Node recovery: When the original primary database is repaired and connected as a new standby database, the breakpoint is determined by comparing the checksum table (for example, the original maximum LSN is 10000, and the new primary database's current LSN is 12000). Only unconfirmed batches with start_lsn>10000 are synchronized, and data integrity is ensured through checksum verification.

[0069] S3, batch WAL transmission and dynamic control;

[0070] Batch generation strategy: The master database encapsulates WAL batches according to dual threshold conditions;

[0071] Capacity threshold: The accumulated log data reaches a preset size (e.g., 16 MiB) or the number of records reaches 1,000;

[0072] Time threshold: The time since the last batch generation exceeds the preset interval (such as 50ms).

[0073] When either the capacity threshold or the time threshold is met, batch encapsulation is triggered, and a batch unit including a start LSN, an end LSN, log data, and a check code is generated.

[0074] Dynamic tuning mechanism: The master database dynamically adjusts batch parameters based on real-time network status (TCP throughput) and standby database load (heartbeat response time):

[0075] Increase the transmission frequency when the standby database is idle to shorten the synchronization delay.

[0076] The process of the main database reading a batch is as follows Figure 2 As shown, at this time, the log cache contains logEntry1, LogEntry2, ..., LogEntryN and other logs.

[0077] Among them, the tableID corresponding to LogEntry1 is 1, the corresponding table is TableA, and the current_read_lsn of its WAL_MgrA is 100. The tableID corresponding to LogEntry2 is 2, the corresponding table is TableB, and the current_read_lsn of its WAL_MgrA is 200. The tableID corresponding to LogEntry3 is 1, the corresponding table is TableA, and the current_read_lsn of its WAL_MgrA is 100.

[0078] The master reads from the cache queue. At this time, the queue head is LogEntry1, whose start_lsn is 100, which is equal to the current_read_lsn of WAL_MgrA. So LogEntry1 is read, dequeued, and current_read_lsn is set to the end_lsn of LogEntry1, 150.

[0079] The master database continues to read LogEntry2. At this time, the start_lsn of LogEntry2 is 100, which is less than the current_read_lsn of WAL_MgrB. This indicates that the log has already been read and does not need to be read again, so it is discarded directly.

[0080] The master database reads LogEntry3. At this time, the start_lsn of LogEntry3 is 500, and the current_read_lsn of its WAL_MgrA is 150. This indicates that a termination occurred between reading and caching. The next LogEntry to be read should be 150, but there are no logs with LSNs from 150 to 500 in the cache. Therefore, the master database reads LogEntries with LSNs from 150 to 500 from the WAL file, and then reads LogEntry3. Then, the current_read_lsn is set to the end_lsn of LogEntry3, which is 600.

[0081] When the number or capacity of read LogEntries reaches the set batch capacity threshold, the primary database sends the read batch of logs to the secondary database for the next read.

[0082] Because the cache cannot be written to once it is full, the number of log entries in the cache lags behind the actual WAL files. When all the LogEntries in the cache are read, that is, when the cache is empty, the master database reads from the WAL files in each table, assembles the read logs into a batch request, and sends it to the slave database. At this time, reading from files still applies to the capacity threshold management policy described above.

[0083] S4, consistency check table and breakpoint recovery;

[0084] Integrity check: After receiving a batch, the standby database verifies data integrity using a checksum. If a failure occurs, a retransmission request is made.

[0085] Breakpoint synchronization: The primary and standby databases regularly compare and verify the LSN ranges in the tables. When the primary database recovers from a failure and becomes the standby database, the start_lsn is used to quickly locate the breakpoint, skipping the synchronized batches and achieving incremental catch-up.

[0086] Version number management: A new checksum version number (such as VERSION=3) is generated for each active / standby switchover to ensure the uniqueness and validity of the checkpoints after multiple rounds of failures.

[0087] When the primary database crashes and restarts, the standby database sends its consistency check table. The primary database reinitializes the log cache based on this consistency check table and sets the current_read_lsn of all WAL_Mgrs.

[0088] Based on the above method, a highly available master-slave replication device for a time series database based on WAL in this embodiment includes: at least one memory and at least one processor;

[0089] The at least one memory is configured to store a machine-readable program;

[0090] The at least one processor is used to call the machine-readable program to execute a high-availability master-slave replication method for a time series database based on WAL.

[0091] The above-mentioned specific implementation methods are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above-mentioned specific implementation methods. Any technical solutions that conform to the above-mentioned specific implementation methods of the present invention and any appropriate changes or substitutions made thereto by ordinary technicians in the relevant technical field shall fall within the patent protection scope of the present invention.

[0092] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A high-availability master-slave replication method for a time series database based on WAL, characterized in that: It has the following steps: S1. Replace the traditional three - node architecture with a two - node system of the master database Master and the slave database Slave; S2. Optimize the log cache; S3. Conduct batch - based WAL transmission and dynamic control; S4. Perform consistency check table and breakpoint recovery.

2. A high-availability master-slave replication method for a time series database based on WAL according to claim 1, characterized in that: In step S1, the master database receives read - write requests, generates WAL records, and maintains the WAL log cache, and transmits them to the slave database through batch - based reading; The slave database synchronizes the master database data in real - time, serves as a hot - standby node for seamless switching in case of master database failure, and supports read - only queries to share the read pressure.

3. A high-availability master-slave replication method for a time series database based on WAL according to claim 2, characterized in that: In step S2, the master database introduces an in - memory cache queue. When initializing the cache, the cache is initialized when the master - slave is enabled. During the initialization phase, the current consistency check table needs to be obtained from the slave database. The consistency check table records the LSN positions of each Table through WAL replay, indicating that the data of this table before the current LSN has been received and the replay has been completed. The master database obtains the WALMgr of all Tables according to the consistency check table, and sets the current_read_lsn of the WALMgr to the LSN of the consistency table, indicating that the master database has completed reading this table before this LSN, and subsequent reading starts from this LSN; First, establish a master - slave connection on the slave database. The slave database sends its own consistency check table to the master database. If it is the first time, the consistency check table is empty; otherwise, it is a reconnect after disconnection. After receiving the consistency check table, the master database obtains the WAL_Mgr of all Tables in the consistency check table and sets its own Current_read_lsn to the LSN of the consistency check table.

4. A high-availability master-slave replication method for a time series database based on WAL according to claim 3, characterized in that: When implementing the double - write mechanism, when a new WAL record is generated, it is written into the disk file and the cache queue simultaneously. The cache is a global cache, and its maximum capacity is configured. The cache writing order is the same as the generation order of WAL records; When performing priority cache reading, when generating batches, data of continuous LSN segments is preferentially extracted from the cache. Only when it is detected that the LSN is discontinuous, disk reading is triggered to complete the missing records and update the cache; The LSN continuous detection strategy is: After reading the WAL, the WALMgr corresponding to each table updates its own current_read_lsn. Obtain the latest LogEntry in the cache queue. If its start_lsn is equal to the current_read_lsn of the current WALMgr, it means it is continuous, and read the LogEntry. If start_lsn < current_read_lsn, it means the LogEntry has been read, and discard the LogEntry; If start_lsn > current_read_lsn, it means that the read LogEntry is not continuous with the LogEntry in the cache. In this case, the missing LogEntry needs to be read from the WAL file. The starting LSN of the missing part is the current_read_lsn of the current WALMgr, and the ending LSN is the start_lsn of the first LogEntry in the current buffer.

5. A high-availability master-slave replication method for a time series database based on WAL according to claim 4, characterized in that: The cache elimination strategy is: After the standby database confirms the batch applies, it updates the consistency check table in memory and persists it to disk at regular intervals. After the primary database cache reads the records in the corresponding LSN range, it deletes them from the cache to free up memory space and ensure that the cache only retains the latest unconfirmed log data. When the master database's cache is full and the slave database is insufficient to consume the master's writes, the master database writes only to the WAL file after the new write, without writing to the cache. It waits until all the data in the cache is consumed, and then re-enables cache writes. During fault detection, a certain number of consecutive heartbeat timeouts trigger a fault determination. When the standby database detects a primary database failure, it traverses the consistency check table and performs a WAL replay on all unapplied batches to ensure that the local database is updated to the latest LSN. Switch to read-write mode, generate a new checksum version number, and broadcast the new master database address and current LSN synchronization point; When the node is restored and the original primary database is repaired and connected as the new backup database, the breakpoint is determined by comparing the checksum table, and only unconfirmed batches are synchronized, and data integrity is ensured through verification of the checksum.

6. A high-availability master-slave replication method for a time series database based on WAL according to claim 5, characterized in that: In step S3, the master database packages the WAL batch according to the dual threshold conditions. The accumulated log data reaches the preset size or the number of records reaches a certain number above the capacity threshold. The time threshold exceeds the preset interval from the last batch generation. If either of the two conditions is met, batch packaging is triggered, and a batch unit containing the start LSN, end LSN, log data and check code is generated; The master database dynamically adjusts batch parameters based on real-time network status and backup database load.

7. A high-availability master-slave replication method for a time series database based on WAL according to claim 6, characterized in that: When the standby database is idle, the transmission frequency is increased to shorten the synchronization delay. When the primary database reads the batch process, the log cache contains logEntry1, LogEntry2, ..., LogEntryN. The primary database reads from the cache queue. At this time, the queue head is LogEntry1, and its start_lsn is 100, which is equal to the current_read_lsn of WAL_MgrA. Therefore, LogEntry1 is read, dequeued, and the current_read_lsn is set to the end_lsn of LogEntry1, 150. The master database continues to read LogEntry2. At this time, the start_lsn of LogEntry2 is 100, which is less than the current_read_lsn of WAL_MgrB. This indicates that the log has already been read and does not need to be read again, so it is discarded directly. The master database reads LogEntry3. At this time, the start_lsn of LogEntry3 is 500, while the current_read_lsn of WAL_MgrA is 150. This indicates that a termination occurred between reading and caching. The next LogEntry to be read should be 150, but there are no logs with LSNs from 150 to 500 in the cache. Therefore, the master database reads LogEntries with LSNs from 150 to 500 from the WAL file, and then reads LogEntry3. Then the current_read_lsn is set to the end_lsn of LogEntry3, which is 600. When the number or volume of read LogEntries reaches the set batch capacity threshold, the primary database sends the read batch of logs to the secondary database for the next read. Because the cache cannot be written to after it is full, the number of logs in the cache lags behind the actual WAL files. When all LogEntries in the cache are read, that is, when the cache is empty, the master database reads from the WAL files in each table, assembles the read logs into a batch request, and sends it to the slave database. At this time, the capacity threshold management policy still applies to reading from files.

8. A high-availability master-slave replication method for a time series database based on WAL according to claim 7, characterized in that: In step S4, after receiving the batch, the standby database verifies the data integrity using the checksum. If the checksum fails, a retransmission request is made. The primary and standby databases regularly compare and verify the LSN ranges in the tables. When the primary database recovers from a failure and becomes the standby database, the start_lsn is used to quickly locate the breakpoint, skipping the synchronized batches and achieving incremental catch-up. A new checksum version number is generated for each master-slave switch to ensure the uniqueness and validity of the checkpoints after multiple rounds of failures. When the master database crashes and restarts, the slave database sends the consistency checksum. The master database reinitializes the log cache based on this consistency checksum and sets the current_read_lsn of all WAL_Mgrs.

9. A highly available master-slave replication device for a time series database based on WAL, characterized in that: include: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Intelligent optimization system and method based on one-master multi-standby database index

    CN121092547A