Database node restart processing method and device based on delayed playback and medium

By using stream replication and delay queue asynchronous processing technology when the database backup node is restarted, real-time synchronization of the master and backup nodes is achieved, solving the problem of business interruption when the delayed replay backup node is restarted, and improving the system's reliability and resource utilization rate.

CN120162387AActive Publication Date: 2025-06-17HIGHGO SOFTWARE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510645602.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-17
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

When a database backup node configured for delayed playback occurs when a new transaction generated by the primary node cannot be synchronized to the backup node in time, resulting in business interruption and discontinuity.

Method used

When the standby node is restarted, real-time synchronization of the main and standby nodes is established in seconds through stream replication, and delay queue asynchronous processing, dynamic sorting and asynchronous sharding technology are used to ensure that the main and standby nodes are synchronized in real time without affecting each other.

Benefits of technology

It greatly shortens the stream replication recovery time, realizes real-time synchronization of the master and spare nodes, avoids the risk of WAL log accumulation and interruption of the master node, reduces storage pressure, improves resource utilization, and ensures business continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162387A_ABST
    Figure CN120162387A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a database node restart processing method and device based on delayed replay and a medium, belongs to the technical field of databases, and solves the problem that when a database standby node configured with delayed replay is restarted, services are discontinuous. Comprising the following steps: respectively carrying out restart detection and state capture on a main node and a standby node; determining a peak-valley period corresponding to the current time window, and adjusting the parameter configuration file of the standby node based on the peak-valley period; after the adjustment is completed, performing real-time synchronization on the latest WAL log pulled from the main node in a local corresponding real-time synchronization area through stream replication; in a local corresponding delay replay area, delay queues are constructed based on the historical WAL logs which are not applied locally, the delay queues are dynamically sorted, and the sorted delay queues are replayed through asynchronous fragmentation; and when it is detected that the real-time synchronization region is in a flow replication stable state, performing recovery adjustment on the parameter configuration file to complete database node restart processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of database technologies, and in particular, to a method, device, and medium for restarting a database node based on delayed replay. Background Art

[0002] In modern enterprise-level database systems, to ensure high data availability and disaster tolerance capabilities, the master-slave architecture has become the mainstream deployment mode. The slave node uses streaming replication technology to synchronize the transaction logs of the master node in real time, ensuring that it can quickly take over the service in case of a master node failure and maintain business continuity.

[0003] However, when a database slave node configured with delayed replay restarts, the system will face serious performance bottlenecks and service interruption risks. After restarting, the slave node needs to fully replay the locally accumulated WAL logs. This process may take up to several hours due to factors such as a large amount of log data and high transaction complexity. Before the replay is completed, the slave node cannot establish a streaming replication connection with the master node, resulting in data synchronization interruption. During this period, new transactions generated by the master node cannot be synchronized to the slave node in a timely manner. If the master node fails at this time, the slave node cannot fully take over the service due to data loss, which will cause business interruption and thus result in business discontinuity. Summary of the Invention

[0004] Embodiments of this application provide a method, device, and medium for restarting a database node based on delayed replay, which are used to solve the following technical problems: When a database slave node configured with delayed replay restarts, new transactions generated by the master node cannot be synchronized to the slave node in a timely manner, which will cause business interruption and thus result in business discontinuity.

[0005] Embodiments of this application adopt the following technical solutions: Embodiments of this application provide a method for restarting a database node based on delayed replay. The method includes: when a restart signal of the slave node is obtained, performing restart detection and status capture on the master node and the slave node respectively; when the detection passes, determining the peak-valley period corresponding to the current time window, and adjusting the parameter configuration file of the slave node based on the peak-valley period; after the adjustment is completed, based on the captured status information, in the corresponding local real-time synchronization area, performing real-time synchronization of the latest WAL logs pulled from the master node through streaming replication; and, in the corresponding local delayed replay area, constructing a delay queue based on the locally unapplied historical WAL logs, and dynamically sorting the delay queue, so as to perform replay on the sorted delay queue through asynchronous sharding; when it is detected that the real-time synchronization area is in a stable state of streaming replication, performing restoration adjustment on the parameter configuration file to complete the restart processing of the database node.

[0006] The embodiment of the present application is established within seconds through stream replication, and the stream replication recovery time is greatly shortened. Through asynchronous processing of the delay queue, real-time synchronization between the master node and the standby node is achieved without mutual influence. Secondly, the rapid establishment of real-time stream replication ensures that the master node continuously sends the latest WAL logs, avoiding the risks of WAL log accumulation and interruption on the master node. In addition, the asynchronous replay of local historical WAL logs in the embodiment of the present application reduces the resource contention for real-time stream replication. The master node does not need to retain a large number of historical WAL logs for a long time, reducing the storage pressure and improving the resource utilization rate. Even if the delay queue is not completed, the standby node can still provide read-only queries of the latest data, supporting failover and load balancing to ensure business continuity.

[0007] In an implementation manner of the present application, when a restart signal of the standby node is obtained, restart detection and status capture are respectively performed on the master node and the standby node, which specifically includes: obtaining the restart signal of the standby node through the database log or the system daemon process; calling the information acquisition instruction to obtain the control information corresponding to the current node, and detecting the last consistent log sequence number before restart; when the detection passes, saving the delay parameter value, the physical replication slot name, and the master node before restart; and recording the list of local unapplied WAL log files and the corresponding log sequence number range.

[0008] In an implementation manner of the present application, detecting the last consistent log sequence number before restart specifically includes: obtaining the reference log sequence number in the preset control file; where the reference log sequence number is the position of the WAL log that has completed consistency processing at the preset historical moment; comparing the reference log sequence number with the WAL log corresponding to the standby node; determining the continuity between the prev-LSN field of the current WAL log and the end log sequence number of the previous WAL log; and performing matching verification on the WAL log through the cyclic redundancy check method; when the WAL log is continuous and the matching verification passes, it is determined that the detection of the last consistent log sequence number before restart passes.

[0009] In an implementation manner of the present application, determining the peak-valley period corresponding to the current time window and adjusting the parameter configuration file of the standby node based on the peak-valley period specifically includes: determining the peak-valley period corresponding to the current time window; when the current time window belongs to the low valley period, configuring to retain some delay parameters; when the current time window belongs to the peak period, configuring to completely disable the delay parameters, and determining the delay parameters in the parameter configuration file; setting the delay parameters to 0; after the adjustment of the parameter configuration file is completed, sending the SIGHUP signal to reload the configuration to make the delay parameters take effect.

[0010] In an implementation manner of the present application, dynamic sorting is performed on the delay queue, specifically including: determining corresponding priority data in a preset load status table based on the load status corresponding to peak and valley periods; performing a first weight assignment on the delay queue according to the priority data; performing associated data mining on historical transaction data through federated learning based on the delay queue to determine associated data in the historical transaction data, and performing a second weight assignment on the associated data according to a preset association degree; and performing dynamic sorting on the delay queue through the first weight assignment and the second weight assignment.

[0011] In an implementation manner of the present application, asynchronous sharding is used to replay the sorted delay queue, specifically including: obtaining network status parameters, and determining multiple replay policies based on the network status parameters; where the network status parameters at least include one of network bandwidth, delay parameters, and packet loss rate; inputting the multiple replay policies into a preset digital twin model to output the replay efficiency corresponding to each replay policy through the preset digital twin model; determining a reference policy among the multiple replay policies based on the replay efficiency, and replaying the sorted delay queue through the reference policy and asynchronous sharding.

[0012] In an implementation manner of the present application, after replaying the sorted delay queue through asynchronous sharding, the method further includes: calling the pg_wal_replay_resume function when it is detected that there is a blocking replay of the local WAL log; through the pg_wal_replay_resume function, skipping the blocking replay of the local WAL log and directly establishing a primary-standby connection based on the sorted delay queue and real-time streaming replication data.

[0013] In an implementation manner of the present application, when it is detected that the real-time synchronization area is in a stable state of streaming replication, the parameter configuration file is restored and adjusted to complete the restart process of the database node, specifically including: starting a background thread to periodically query the difference in log sequence numbers between the primary node and the standby node; determining that the streaming replication has been successfully established when the delay time corresponding to the real-time synchronization area is less than a preset delay time and the continuous stable duration is greater than a preset running duration; restoring the adjusted parameters in the parameter configuration file to their initial values and reloading the configuration; asynchronously replaying the local historical WAL log in the background.

[0014] An embodiment of the present application provides a database node restart processing device based on delayed replay, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: when a standby node restart signal is obtained, perform restart detection and status capture on the primary node and the standby node respectively; when the detection passes, determine the peak-valley period corresponding to the current time window, and adjust the parameter configuration file of the standby node based on the peak-valley period; after the adjustment is completed, based on the captured status information, in the corresponding local real-time synchronization area, perform real-time synchronization of the latest WAL log pulled from the primary node through stream replication; and, in the corresponding local delayed replay area, construct a delay queue based on the local unapplied historical WAL log, and perform dynamic sorting on the delay queue to replay the sorted delay queue through asynchronous sharding; when it is detected that the real-time synchronization area is in a stable state of stream replication, perform recovery adjustment on the parameter configuration file to complete the database node restart processing.

[0015] A non-volatile computer storage medium provided by an embodiment of the present application stores computer-executable instructions, and the computer-executable instructions are set to: when a standby node restart signal is obtained, perform restart detection and status capture on the primary node and the standby node respectively; when the detection passes, determine the peak-valley period corresponding to the current time window, and adjust the parameter configuration file of the standby node based on the peak-valley period; after the adjustment is completed, based on the captured status information, in the corresponding local real-time synchronization area, perform real-time synchronization of the latest WAL log pulled from the primary node through stream replication; and, in the corresponding local delayed replay area, construct a delay queue based on the local unapplied historical WAL log, and perform dynamic sorting on the delay queue to replay the sorted delay queue through asynchronous sharding; when it is detected that the real-time synchronization area is in a stable state of stream replication, perform recovery adjustment on the parameter configuration file to complete the database node restart processing.

[0016] The above at least one technical solution adopted by the embodiment of the present application can achieve the following beneficial effects: In the embodiment of the present application, stream replication is established within seconds, and the stream replication recovery time is greatly shortened. Through asynchronous processing of the delay queue, real-time synchronization between the primary node and the standby node is realized without mutual influence. Secondly, the rapid establishment of real-time stream replication ensures that the primary node continuously sends the latest WAL log, avoiding the risks of WAL log accumulation and interruption on the primary node. In addition, the asynchronous replay of the local historical WAL log in the embodiment of the present application reduces the resource contention for real-time stream replication. The primary node does not need to retain a large amount of historical WAL logs for a long time, reducing the storage pressure and improving the resource utilization rate. Even if the delay queue is not completed, the standby node can still provide read-only queries of the latest data, supporting failover and load balancing, and ensuring business continuity. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings: Figure 1 It is a flowchart of a method for restarting a database node based on delayed replay provided by an embodiment of the present application; Figure 2 It is a schematic diagram of the process of node restart and status capture provided by an embodiment of the present application; Figure 3 It is a schematic diagram of the process of adjusting delay parameters and streaming replication priorities provided by an embodiment of the present application; Figure 4 It is a schematic diagram of the process of monitoring the stability of streaming replication and restoring parameters provided by an embodiment of the present application; Figure 5 It is a schematic diagram of the structure of a device for restarting a database node based on delayed replay provided by an embodiment of the present application.

[0018] Reference Numerals: 500: Device for restarting a database node based on delayed replay, 501: Processor, 502: Memory. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Embodiments of the present application provide a method, device, and medium for restarting a database node based on delayed replay.

[0020] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0021] The following will detail the technical solutions proposed in the embodiments of the present invention through the drawings.

[0022] Figure 1 It is a flowchart of a method for restarting a database node based on delayed replay provided by an embodiment of the present application. As Figure 1 shown, the method for restarting a database node based on delayed replay includes the following steps: Step 101: When a restart signal of the standby node is obtained, perform restart detection and status capture on the primary node and the standby node respectively.

[0023] In an implementation manner of the present application, a restart signal of the standby node is obtained through a database log or a system daemon process. An information acquisition instruction is called to obtain control information corresponding to the current node, and the last consistent log sequence number before restart is detected. When the detection passes, the delay parameter value, physical replication slot name, and primary node before restart are saved. In addition, a list of WAL log files not applied locally and the corresponding log sequence number range are recorded.

[0024] Specifically, Figure 2 is a schematic flow diagram of node restart and status capture provided by an embodiment of the present application. As Figure 2 shown, node restart and status capture include the following processes: Step 201: Capture a restart signal of the standby node through a database log or a system daemon process.

[0025] Step 202: Call the pg_controldata command to obtain control information of the current node, and confirm the last consistent LSN (log sequence number) before restart.

[0026] Step 203: Save the delay parameter value (recovery_min_apply_delay), physical replication slot name (primary_slot_name), and primary node before restart.

[0027] Step 204: Check the local WAL log directory, and record the list of unapplied WAL files and their corresponding LSN ranges.

[0028] For example: (0000000100000001000000A0 to 000000010000000100000500).

[0029] In an implementation manner of the present application, confirming the last consistent LSN (log sequence number) before restart includes obtaining a reference log sequence number in a preset control file; wherein, the reference log sequence number is the position of the WAL log that has completed consistency processing at a preset historical moment. Compare the reference log sequence number with the WAL log corresponding to the standby node. Determine the continuity between the prev-LSN field of the current WAL log and the end log sequence number of the previous WAL log, and perform matching verification on the WAL log through a cyclic redundancy check method. When the WAL log is continuous and the matching verification passes, it is determined that the verification of the last consistent log sequence number before restart passes.

[0030] Specifically, obtain the base LSN from pg_control. Here, pg_control is a binary file in the PostgreSQL database, which is used to store the core metadata of the database cluster. Compare it with the standby node's WAL log, introduce CRC (Cyclic Redundancy Check) verification to enhance reliability, read the metadata of the WAL file header, verify whether its prev-LSN field is continuous with the end LSN of the previous file, and verify the CRC32 value. If it is continuous and the CRC matches, the log integrity is confirmed.

[0031] Furthermore, the pg_control file includes the base LSN (Log Sequence Number). The base LSN represents the WAL log position that the database system recognizes as having completed consistency processing, and it is an important identifier for measuring the data consistency state of the database. The standby node obtains the WAL log from the primary node through streaming replication to achieve data synchronization. Compare the base LSN obtained from pg_control with the WAL log of the standby node to check whether the WAL log on the standby node is complete and consistent with the data state of the primary node. If the WAL log of the standby node does not match the base LSN after a certain LSN position, it indicates that there may be missing, damaged, or replication errors in the WAL log of the standby node.

[0032] To further ensure that no data errors occur during the storage and transmission of the WAL log, introduce CRC technology. At the receiving end, perform the same calculation on the data again, and compare the obtained CRC value with the previously stored or transmitted CRC value. If the two CRC values are consistent, it means that the data has not changed during transmission or storage; otherwise, it indicates that the data may be damaged or tampered with.

[0033] Furthermore, the header of each WAL file contains important metadata information. The prev-LSN field records the end LSN of the previous WAL file of the current WAL file. By reading the prev-LSN field of the WAL file header and verifying whether it is continuous with the end LSN of the previous WAL file, the logical order between WAL files can be determined. The WAL log is generated and recorded sequentially in chronological order, and the continuity between files is crucial for ensuring the integrity and consistency of transaction processing. If the prev-LSN field is not continuous with the end LSN of the previous file, it means that there may be missing or incorrect entries in the WAL file sequence, and further inspection and repair are required to ensure the continuity and integrity of the log.

[0034] Further, after verifying the continuity of the prev-LSN field, it is also necessary to check the CRC32 value of the content of the WAL file. By calculating the CRC32 value of the WAL file content and comparing it with the CRC32 value stored in the file, if the two match, it can be confirmed that the WAL file has not suffered data corruption during storage or transmission, and the content is complete and accurate. Only when the prev-LSN field is continuous and the CRC32 value matches can the integrity of the WAL file be finally confirmed. By verifying all WAL files, the integrity of the entire WAL log can be ensured, guaranteeing the stable operation and data consistency of the database system.

[0035] Step 102: In the case of passing the detection, determine the peak-valley period corresponding to the current time window, and adjust the parameter configuration file of the standby node based on the peak-valley period.

[0036] Figure 3 It is a schematic flow diagram of adjusting delay parameters and streaming replication priorities provided by an embodiment of the present application, as Figure 3 shown. The delay parameter adjustment includes the following steps 301: Step 301: Before modification, check the current time window (such as peak and valley periods), and dynamically decide whether to completely disable the delay parameter or partially retain it.

[0037] Specifically, modify the postgresql.auto.conf file of the standby node. This file is a configuration file. Set recovery_min_apply_delay to 0 and send a SIGHUP signal to reload the configuration.

[0038] In an implementation manner of the present application, determine the peak-valley period corresponding to the current time window. In the case where the current time window belongs to the valley period, configure to retain some delay parameters. In the case where the current time window belongs to the peak period, configure to completely disable the delay parameters, and determine the delay parameters in the parameter configuration file and set the delay parameters to 0. After adjusting the parameter configuration file, send a SIGHUP signal to reload the configuration to make the delay parameters take effect.

[0039] Specifically, during the low-load period, since the system load is low and the real-time requirement for data synchronization is not so urgent, in order to reduce the processing pressure on the standby node and avoid high-intensity data processing on the standby node under low-load conditions, some delay parameters can be selected to be retained, and the data synchronization speed can be appropriately slowed down, allowing the standby node to process data in a more relaxed state. During the high-load period, in order to ensure that the data on the primary node can be quickly synchronized to the standby node, meet the strict requirements for data consistency in high-concurrency scenarios, and avoid data lag caused by delays affecting the normal operation of the business, the delay parameters are usually completely disabled.

[0040] Furthermore, postgresql.auto.conf is an important file for the PostgreSQL database to store automatic configuration parameters, which records many key parameters affecting the operation of the database. After determining the adjustment strategy of the delay parameters, corresponding modifications need to be made in this file. The recovery_min_apply_delay parameter is used to control the minimum delay time for the standby node to apply transaction logs. Setting it to 0 means that after receiving the transaction logs sent by the primary node, the standby node will not wait for any delay and will immediately apply these logs. During the high-load period, it can minimize the time interval of data synchronization to ensure that the data on the standby node quickly catches up with the changes on the primary node and achieve data real-time and consistency. The role of sending the SIGHUP signal is to notify the database process to reread the configuration file to make the just-modified setting of the recovery_min_apply_delay parameter take effect.

[0041] Step 103: After the adjustment is completed, based on the captured status information, in the corresponding local real-time synchronization area, through streaming replication, the latest WAL logs pulled from the primary node are synchronously replicated in real time.

[0042] In one implementation manner of this application, Figure 3 is a schematic flowchart of the adjustment of delay parameters and streaming replication priority provided by the embodiment of this application. As Figure 3 shown, the streaming replication in the real-time synchronization area includes the following steps 302: Step 302: In the real-time synchronization area, directly pull the latest WAL logs (LSN greater than the locally applied LSN) from the primary node and synchronously replicate them in real time through the streaming replication process.

[0043] Step 104: And, in the corresponding local delay replay area, build a delay queue based on the historical WAL logs that have not been applied locally, and dynamically sort the delay queue to replay the sorted delay queue through asynchronous sharding.

[0044] In one implementation manner of this application, Figure 3This is a schematic flowchart of the adjustment of delay parameters and stream replication priorities provided by the embodiments of the present application. As Figure 3 shown, the stream replication in the delay replay area includes the content of steps 303 to 304 as follows: Step 303: In the delay replay area, mark the unapplied old WAL logs locally as the "delay queue", dynamically sort them based on the transaction type and data importance (e.g., payment transactions have priority), and start asynchronous shard replay (adhering to the original delay parameters), while ensuring exclusive access to the real-time stream replication channel resources.

[0045] In an implementation manner of the present application, based on the load status corresponding to the peak and valley periods, the corresponding priority data is determined in the preset load status table. According to the priority data, the first weight assignment is performed on the delay queue. Based on the delay queue, the historical transaction data is mined for associated data through federated learning to determine the associated data in the historical transaction data, and the second weight assignment is performed on the associated data according to the preset association degree. Through the first weight assignment and the second weight assignment, the delay queue is dynamically sorted.

[0046] Specifically, the preset load status table is a pre-set rule table that corresponds different load statuses to the corresponding priority data one by one. Through real-time monitoring, the specific load status during the peak and valley periods is judged, and based on the monitoring results, accurate positioning is performed in the load status table to find the matching priority data. The delay queue stores the unapplied old WAL logs locally, and these logs record the transaction change information of the database. The first weight assignment is to assign a weight value to each log in the delay queue based on the previously determined priority data.

[0047] Furthermore, taking the transactions in the delay queue as the entry point, the historical transaction data of multiple data sources is included in the analysis scope. During the data processing process, each data source does not need to directly transmit and share the original data, but performs calculations locally and then securely aggregates the calculation results through encrypted communication. In this process, the hidden association relationships between transactions are deeply mined through the federated learning algorithm. For example, it is found that certain transactions often occur after specific transactions, or there are close data interaction relationships between the transactions of certain business lines. These mined transaction data with association relationships are the determined associated data.

[0048] Furthermore, the preset correlation degree is a criterion preset to measure the tightness of the correlation between transactions. For the correlated data mined through federated learning, according to the preset correlation degree rules, weight values are assigned to these data again, that is, the second weight assignment. Transactions with a high correlation degree will be assigned a higher weight, while those with a low correlation degree will have a lower weight. The results of the first weight assignment and the second weight assignment are comprehensively considered to calculate the final weight value of each transaction. According to the final weight value, the delay queue is dynamically sorted, with transactions with a high weight ranked at the front of the queue and given priority for processing; transactions with a low weight are ranked behind. Through this dynamic sorting method, during the database delay replay process, transactions that are more critical to the system operation and business process can be preferentially processed, the system resources can be reasonably allocated, the database processing efficiency and data consistency can be improved, and it is ensured that the core business can operate stably and efficiently under different load conditions.

[0049] In an implementation manner of the present application, network state parameters are obtained, and multiple replay strategies are determined based on the network state parameters; wherein, the network state parameters at least include one of network bandwidth, delay parameters, and packet loss rate. The multiple replay strategies are input into a preset digital twin model to output the replay efficiency corresponding to each replay strategy through the preset digital twin model. A reference strategy is determined from the multiple replay strategies based on the replay efficiency, and the sorted delay queue is replayed through the reference strategy and asynchronous sharding.

[0050] Specifically, the network state parameters are obtained in real time through a monitoring mechanism, and these parameters include key indicators such as network bandwidth, delay parameters, and packet loss rate. Based on these parameters, multiple different replay strategies are generated according to preset rules and algorithms. For example, when the network bandwidth is low and the packet loss rate is high, a "sharding compression transmission" strategy is generated, and the WAL log data is sharded and compressed before transmission to reduce the amount of data to adapt to the network conditions. Different strategies target different network environment characteristics and provide multiple options for efficient replay.

[0051] Furthermore, the digital twin model in the embodiments of the present application is a virtual mapping of the real database replay process. Through a large amount of data training and optimization in advance, it can accurately simulate the actual operation conditions under different replay strategies. The multiple generated replay strategies are input into this preset digital twin model. During the simulation process, the model will comprehensively consider various factors such as network state, delay queue data characteristics, and system resources, and calculate the replay efficiency corresponding to each replay strategy.

[0052] Further, after obtaining the replay efficiency corresponding to each replay policy, these efficiency values are compared and analyzed. The policy with the highest replay efficiency is determined as the reference policy. After determining the reference policy, combined with the asynchronous sharding technology, the replay operation is performed on the already sorted delay queue. Asynchronous sharding can divide the WAL log data in the delay queue into multiple segments and process these segments in parallel, further improving the replay efficiency. The combination of the reference policy and asynchronous sharding enables the database to select the optimal way to replay the delay queue according to the network environment and data characteristics, ensuring that the data replay task can be efficiently and stably completed under different network conditions, and guaranteeing the normal operation of the database and data consistency.

[0053] Step 304: Call the pg_wal_replay_resume function to skip the blocking replay of the local WAL log and directly establish a primary-standby connection based on the dynamically sorted delay queue status and real-time streaming replication data.

[0054] In an implementation manner of this application, when it is detected that there is a blocking replay of the local WAL log, the pg_wal_replay_resume function is called. This function is an internal function for resuming the WAL log replay. Through the pg_wal_replay_resume function, the blocking replay of the local WAL log is skipped, and a primary-standby connection is directly established based on the sorted delay queue and real-time streaming replication data.

[0055] Step 105: When it is detected that the real-time synchronization area is in a stable state of streaming replication, the parameter configuration file is restored and adjusted to complete the restart process of the database node.

[0056] In an implementation manner of this application, a background thread is started to periodically query the difference in the log sequence numbers between the primary node and the standby node. When the delay in the real-time synchronization area is less than the preset delay time and the continuous stable duration is greater than the preset running duration, it is determined that the streaming replication has been successfully established. The adjusted parameters in the parameter configuration file are restored to their initial values, and the configuration is reloaded. The local historical WAL log is replayed asynchronously in the background.

[0057] Figure 4 This is a schematic diagram of the process for monitoring the stability of streaming replication and restoring parameters provided by an embodiment of this application. As Figure 4 shown, the monitoring of the stability of streaming replication and parameter restoration includes the following steps: Step 401: Start a background thread to periodically query the LSN difference between the primary and standby nodes.

[0058] Step 402: When the delay in the real-time synchronization area is lower than the threshold (such as 5 seconds) and remains stable for 5 minutes, it is determined that the streaming replication has been successfully established.

[0059] Step 403: Restore recovery_min_apply_delay to its initial value (such as 12h), and reload the configuration.

[0060] Step 404: Asynchronously replay the local old WAL logs in the background to ensure data eventual consistency without affecting real-time streaming replication.

[0061] Figure 5 This is a schematic structural diagram of a database node restart processing device based on delayed replay provided by an embodiment of the present application. As Figure 5 shown, the database node restart processing device 500 based on delayed replay includes: at least one processor 501; and a memory 502 communicatively connected to the at least one processor 501; wherein, the memory 502 stores instructions executable by the at least one processor 501, and when the instructions are executed by the at least one processor 501, the at least one processor 501 is capable of: performing restart detection and status capture on the primary node and the standby node respectively when a standby node restart signal is obtained; determining the peak-valley period corresponding to the current time window when the detection passes, and adjusting the parameter configuration file of the standby node based on the peak-valley period; after the adjustment is completed, based on the captured status information, in the corresponding local real-time synchronization area, perform real-time synchronization of the latest WAL logs pulled from the primary node through streaming replication; and in the corresponding local delayed replay area, construct a delay queue based on the local unapplied historical WAL logs, and dynamically sort the delay queue, so as to replay the sorted delay queue through asynchronous sharding; when it is detected that the real-time synchronization area is in a stable state of streaming replication, perform recovery adjustment on the parameter configuration file to complete the database node restart processing.

[0062] A non-volatile computer storage medium provided by an embodiment of the present application stores computer-executable instructions, and the computer-executable instructions are set to: perform restart detection and status capture on the primary node and the standby node respectively when a standby node restart signal is obtained; determine the peak-valley period corresponding to the current time window when the detection passes, and adjust the parameter configuration file of the standby node based on the peak-valley period; after the adjustment is completed, based on the captured status information, in the corresponding local real-time synchronization area, perform real-time synchronization of the latest WAL logs pulled from the primary node through streaming replication; and in the corresponding local delayed replay area, construct a delay queue based on the local unapplied historical WAL logs, and dynamically sort the delay queue, so as to replay the sorted delay queue through asynchronous sharding; when it is detected that the real-time synchronization area is in a stable state of streaming replication, perform recovery adjustment on the parameter configuration file to complete the database node restart processing.

[0063] The embodiments in the present application are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0064] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the embodiments of the present application. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A database node restart processing method based on delayed replay, characterized in that: The method comprises: When the restart signal of the standby node is obtained, restart detection and status capture are performed on the primary node and the standby node respectively; If the detection passes, determine the peak and valley time periods corresponding to the current time window, and adjust the parameter configuration file of the standby node based on the peak and valley time periods; After the adjustment is completed, based on the captured status information, the latest WAL log pulled from the master node is synchronized in real time in the local corresponding real-time synchronization area through streaming replication; And, in the locally corresponding delayed replay area, a delay queue is constructed based on the locally unapplied historical WAL log, and the delay queue is dynamically sorted to replay the sorted delay queue through asynchronous sharding; When it is detected that the real-time synchronization area is in a stable state of stream replication, the parameter configuration file is restored and adjusted to complete the database node restart process.

2. The database node restart processing method based on delayed replay according to claim 1 is characterized in that: When a restart signal of the standby node is obtained, restart detection and status capture are performed on the master node and the standby node respectively, specifically including: Obtain the restart signal of the standby node through the database log or system daemon process; Call the information acquisition instruction to obtain the control information corresponding to the current node, and detect the last consistency log sequence number before the restart; If the detection passes, the delay parameter value, physical replication slot name, and master node before the restart are saved; Also, the list of locally unapplied WAL log files and the corresponding log sequence number ranges are recorded.

3. The database node restart processing method based on delayed replay according to claim 2 is characterized in that: The detecting of the last consistent log sequence number before the restart specifically includes: In the preset control file, obtain the base log sequence number; wherein the base log sequence number is the position of the WAL log that has completed the consistency processing corresponding to the preset historical moment; Compare the benchmark log sequence number with the WAL log corresponding to the standby node; Determine the continuity between the prev-LSN field of the current WAL log and the end log sequence number of the previous WAL log; And, performing a matching check on the WAL log by a cyclic redundancy check method; When the WAL log is continuous and the matching check passes, it is determined that the last consistency log sequence number check before the restart passes.

4. The database node restart processing method based on delayed replay according to claim 1, characterized in that: The determining the peak and valley time periods corresponding to the current time window, and adjusting the parameter configuration file of the standby node based on the peak and valley time periods specifically includes: Determine the peak and valley periods corresponding to the current time window; When the current time window belongs to a valley period, it is configured to retain some delay parameters; In the case where the current time window belongs to a peak period, the delay parameter is configured to be completely disabled, and the delay parameter is determined in a parameter configuration file; Setting the delay parameter to 0; After the parameter configuration file is adjusted, a SIGHUP signal is sent to reload the configuration so that the delay parameter takes effect.

5. The database node restart processing method based on delayed replay according to claim 1, characterized in that: The dynamically sorting the delay queue specifically includes: Based on the load status corresponding to the peak and valley time periods, determining corresponding priority data in the preset load status table; According to the priority data, assigning a first weight to the delay queue; Based on the delay queue, performing associated data mining on the historical transaction data through federated learning to determine associated data in the historical transaction data, and assigning a second weight to the associated data according to a preset association degree; The delay queue is dynamically sorted by means of the first weight assignment and the second weight assignment.

6. The database node restart processing method based on delayed replay according to claim 1, characterized in that: The replaying of the sorted delay queue by asynchronous sharding specifically includes: Acquire network status parameters, and determine multiple replay strategies based on the network status parameters; wherein the network status parameters include at least one of network bandwidth, delay parameter, and packet loss rate; Inputting the plurality of replay strategies into a preset digital twin model, so as to output the replay efficiencies corresponding to the respective replay strategies through the preset digital twin model; A reference strategy is determined from among the plurality of replay strategies based on the replay efficiency, so as to replay the sorted delay queue through the reference strategy and asynchronous sharding.

7. The database node restart processing method based on delayed replay according to claim 1, characterized in that: After replaying the sorted delay queue by asynchronous sharding, the method further includes: When blocking replay of local WAL logs is detected, the pg_wal_replay_resume function is called; Through the pg_wal_replay_resume function, the blocking replay of the local WAL log is skipped, and a master-slave connection is established directly based on the sorted delay queue and the real-time streaming replication data.

8. The database node restart processing method based on delayed replay according to claim 1, characterized in that: When it is detected that the real-time synchronization area is in a stable state of stream replication, the parameter configuration file is restored and adjusted to complete the database node restart process, specifically including: Start a background thread to regularly query the difference in log sequence numbers between the primary node and the backup node; When the delay time corresponding to the real-time synchronization area is less than the preset delay time and the continuous stability time is greater than the preset running time, it is determined that the stream replication has been established normally; Restoring the adjusted parameters in the parameter configuration file to their initial values ​​and reloading the configuration; Replay local historical WAL logs asynchronously in the background.

9. A database node restart processing device based on delayed replay, characterized in that: The device comprises a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute the method according to any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions can execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data processing method and device adopted after node restart

    CN106909473A

  • Database stream replication method and device

    CN114090339A

  • Data delay statistical method and device, equipment and storage medium

    CN116775727A

  • Restoring database consistency integrity

    US20150254298A1

  • Asynchronous Garbage Collection in Database Redo Log Replay

    US20180239676A1