A distributed log synchronization method and system based on a window pipeline delegation mechanism

CN122673282APending Publication Date: 2026-09-01BEIJING XINSHU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611003888.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0004](1)主服务高可用机制效率低下:在分布式系统中,主节点的选举和切换效率对监控系统的容错能力和连续性至关重要

Benefits of technology

[0038] Regarding high availability management of the primary service, this invention introduces a dynamic log priority synchronization mechanism based on the RAFT protocol. By calculating the degree of log differences among candidate nodes, the synchronization priority is dynamically adjusted, prioritizing the synchronization of critical nodes and critical log segments. Compared to existing technologies where all slave nodes participate in election equally without priority distinction, this effectively reduces log synchronization waiting time during primary node switching, enabling monitoring tasks to recover faster and improving system fault tolerance and service continuity. Simultaneously, this invention employs an incremental update strategy instead of the traditional full or batch synchronization method. During primary node switching and fault recovery, only differing log data is synchronized, significantly reducing the amount of synchronized data and lowering log synchronization latency, thereby accelerating the overall system recovery speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure SMS_37
    Figure SMS_37
Patent Text Reader

Abstract

To address the problems of low master node election efficiency, large log synchronization latency, single point of failure in the collection engine, and difficulty in storage expansion in existing distributed database monitoring systems, this invention discloses a distributed log synchronization method and system based on a window pipeline delegation mechanism. The main service high availability management module introduces a dynamic log priority synchronization mechanism and incremental update strategy based on the RAFT protocol to achieve rapid master node election and switching. The log distribution optimization module adopts a window pipeline delegation mechanism, significantly reducing log distribution latency through dynamic grouping, dynamic batch windows, load-aware weight matching, and timeout retry strategies. The collection engine load balancing module adopts a paired deployment method to achieve master-slave load sharing and millisecond-level fault takeover. The storage engine horizontal scaling module automatically completes new node synchronization and hierarchical archiving and cleanup through a sharding strategy. This invention effectively improves the reliability, scalability, and real-time performance of the distributed database monitoring system under high-load scenarios, enabling 24 / 7 uninterrupted operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a distributed log synchronization method and system based on a window pipeline delegation mechanism, belonging to the field of database monitoring. Background Technology

[0002] In the field of database operations and maintenance, there is currently no reliable solution for high availability of distributed monitoring systems. As database size and complexity increase, there is a growing demand for distributed monitoring systems to operate 24 / 7 and efficiently manage the main service, data collection engine, and log distribution module. However, existing technologies have significant shortcomings in key areas such as high availability management of the main service, load balancing of the data collection engine, and log distribution performance, directly impacting the reliability, scalability, and performance optimization of the monitoring system.

[0003] Specifically, these shortcomings are mainly reflected in the following two aspects:

[0004] (1) Inefficient master service high availability mechanism: In a distributed system, the efficiency of master node election and switching is crucial to the fault tolerance and continuity of the monitoring system. However, existing master service high availability mechanisms often suffer from low master node election efficiency and large log synchronization delays after master node failure. This situation leads to a long master node switching process, affecting the real-time performance of monitoring tasks and the overall stability of the system.

[0005] (2) Low log synchronization efficiency affects fault recovery speed: The log distribution module is a key link to ensure system data consistency, but the existing log distribution mechanism often faces the problem of high log synchronization latency under high load scenarios. Especially during master node switching or fault recovery, the log synchronization process is usually affected by latency, resulting in increased system recovery time. This not only affects the real-time performance of the system, but also limits the scalability of the system in a large-scale distributed environment.

[0006] The aforementioned shortcomings limit the reliability and performance optimization potential of distributed monitoring systems under high-load scenarios. Therefore, targeted solutions are proposed to address these issues. Attached Figure Description

[0007] Figure 1 This is a diagram of the architecture of a high-availability database monitoring system.

[0008] Figure 2 This is the algorithm flowchart. Summary of the Invention

[0009] A distributed log synchronization method based on a window pipeline delegation mechanism, characterized by comprising:

[0010] (1) High availability management of the main service: The status of the main service node of the database monitoring system is monitored based on the RAFT protocol. When the main node fails, the election process is triggered. The new main node must meet the consistency of the latest log status. During the log synchronization process, the synchronization priority is dynamically adjusted by calculating the log difference of the candidate nodes, prioritizing the synchronization of key nodes and key log segments, and using an incremental update strategy for log synchronization.

[0011] (2) Log distribution optimization: A window pipeline-based delegation mechanism is adopted to dynamically divide nodes into high-priority groups and low-priority groups, and the grouping strategy is adjusted according to real-time latency; a dynamic window control mechanism is introduced to dynamically adjust the batch size of distributed data according to network bandwidth and node processing capacity; segmented synchronization of logs is achieved through a pipeline model; and the grouping status is updated after each round of distribution.

[0012] (3) Load balancing of acquisition engines: Each pair of acquisition engines is configured with dual modes of load sharing and disaster recovery. The primary and backup engines share the acquisition tasks, and the backup engine retains some redundant resources. When the primary acquisition engine fails, the backup engine automatically detects the anomaly through the health check mechanism and takes over the task. After the task is completed, the acquired data is synchronized to the primary engine storage.

[0013] (4) Storage engine horizontal scaling steps: Add new nodes to the storage cluster through sharding strategy. After the new node is added, the historical log data within the range is automatically synchronized and the access is completed within the preset time. Perform automatic cleanup on historical logs regularly, prioritize the deletion of archived data to free up storage space, and transfer the archived data to low-cost media for long-term storage according to the set compression ratio and storage strategy.

[0014] Furthermore, the log distribution optimization step further includes:

[0015] (a) Initialization: Define the delegator node group senderVec and the receiver node group receiverVec. The senderVec consists of all slave nodes that have completed log synchronization; define the lag threshold. ,in The `nextIndexSortVec` is a list of lag values, where `nextIndexSortVec` represents an array of node indices sorted in ascending order of `nextIndex`, `nextIndex` represents the next log position that the follower needs to receive, `k` is an adjustment parameter, `median(ΔL)` is the median lag value of all nodes, and `MAD(ΔL)` is the median absolute deviation of the lag values. The batch distribution window size is defined. R is the network bandwidth, and C is the log entry size. The hysteresis of the receiver node;

[0016] (b) Principal-Recipient Matching: The matching weight formula is defined as follows ,in , Let be the lags for principal i and receiver j, respectively. Given the current load of the delegator node i; calculate the weights for all possible matching pairs, sort the weights in descending order, and prioritize matching them.

[0017] (c) Log distribution: For each receiver node j, find its matching delegator node i, and the leader distributes logs based on the lag of receiver node j and the batch window size. Calculate the log range to be distributed: N j ), , and These are the starting and ending indices for distribution; the leader updates the replication window of the receiver node and initiates a log distribution request to the delegator node;

[0018] (d) Timeout handling: For tasks that time out due to network latency or other issues, (supplement the timeout threshold formula or provide a fixed value), the leader will reassign the task to the idle delegator node with the lowest load;

[0019] (e) When the lag of all nodes satisfies When the system reaches consistency, the log distribution process ends.

[0020] Furthermore, in the main service high availability management steps, the dynamic priority synchronization mechanism specifically involves: calculating the log difference degree of candidate nodes, sorting the nodes from largest to smallest lag, and prioritizing the synchronization of nodes whose lag is greater than the current average lag multiplied by a preset coefficient.

[0021] Furthermore, in the load balancing step of the acquisition engine, the health check mechanism uses heartbeat detection, and the task takeover time is in the millisecond range.

[0022] Furthermore, in the horizontal scaling step of the storage engine, new nodes are added to the storage cluster through a sharding strategy. After the new node is added, historical log data within the shard range is automatically synchronized, and the access is completed within a preset time threshold. Historical logs are automatically cleaned up periodically, and archived data in the storage is deleted first to free up storage space. The archived data is compressed according to the set compression ratio and transferred to low-cost media for long-term storage according to the storage strategy to ensure the availability of data queries.

[0023] Based on the above method, this invention proposes a distributed database monitoring system based on a high availability mechanism, which consists of the following modules;

[0024] The main service high availability management module monitors the status of the main service node of the database monitoring system based on the RAFT protocol. When the main node fails, an election process is triggered, and the new main node must meet the requirement of consistency with the latest log status. During the log synchronization process, the synchronization priority is dynamically adjusted by calculating the degree of log difference between candidate nodes, prioritizing the synchronization of key nodes and key log segments, and using an incremental update strategy for log synchronization.

[0025] Log distribution optimization module: It adopts a window pipeline-based delegation mechanism to dynamically divide nodes into high-priority groups and low-priority groups, and adjusts the grouping strategy according to real-time latency; it introduces a dynamic window control mechanism to dynamically adjust the batch size of distributed data according to network bandwidth and node processing capacity; it realizes segmented synchronization of logs through a pipeline model; and it updates the grouping status after each round of distribution.

[0026] Data Acquisition Engine Load Balancing Module: Each pair of acquisition engines is configured with dual modes of load sharing and disaster recovery. The primary and backup engines share the acquisition tasks, and the backup engine retains some redundant resources. When the primary acquisition engine fails, the backup engine automatically detects the anomaly through a health check mechanism and takes over the task. After the task is completed, the acquired data is synchronized to the primary engine's storage.

[0027] Storage engine horizontal scaling module: Adds new nodes to the storage cluster through sharding strategy. After the new node is added, it automatically synchronizes the historical log data within the range and completes the access within a preset time. It also performs automatic cleanup of historical logs on a regular basis, prioritizing the deletion of archived data to free up storage space, and transferring the archived data to low-cost media for long-term storage according to the set compression ratio and storage strategy.

[0028] Furthermore, the log distribution optimization module completes log distribution optimization using the following steps:

[0029] (a) Initialization: Define the delegator node group senderVec and the receiver node group receiverVec. The senderVec consists of all slave nodes that have completed log synchronization; define the lag threshold. ,in The `nextIndexSortVec` is a list of lag values, where `nextIndexSortVec` represents an array of node indices sorted in ascending order of `nextIndex`, `nextIndex` represents the next log position that the follower needs to receive, `k` is an adjustment parameter, `median(ΔL)` is the median lag value of all nodes, and `MAD(ΔL)` is the median absolute deviation of the lag values. The batch distribution window size is defined. R is the network bandwidth, and C is the log entry size. The hysteresis of the receiver node;

[0030] (b) Principal-Recipient Matching: The matching weight formula is defined as follows ,in and Let be the lags for principal i and receiver j, respectively. Given the current load of the delegator node i; calculate the weights for all possible matching pairs, sort the weights in descending order, and prioritize matching them.

[0031] (c) Log distribution: For each receiver node j, find its matching delegator node i, and the leader distributes logs based on the lag of receiver node j and the batch window size. Calculate the log range to be distributed: N j ), , and These are the starting and ending indices for distribution; the leader updates the replication window of the receiver node and initiates a log distribution request to the delegator node;

[0032] (d) Timeout handling: For tasks that time out due to network latency or other issues, (supplement the timeout threshold formula or provide a fixed value), the leader will reassign the task to the idle delegator node with the lowest load;

[0033] (e) When the lag of all nodes satisfies When the system reaches consistency, the log distribution process ends.

[0034] Furthermore, in the main service high availability management module, the dynamic priority synchronization mechanism is as follows: by calculating the log difference degree of candidate nodes, the nodes are sorted from largest to smallest according to their lag, and nodes with a lag greater than the current average lag multiplied by a preset coefficient are synchronized first.

[0035] Furthermore, in the data acquisition engine load balancing module, the health check mechanism uses heartbeat detection, and the task takeover time is in the millisecond range.

[0036] Furthermore, the horizontal scaling module of the storage engine includes the following sub-modules: a new node expansion sub-module, used to add new nodes to the storage cluster through a sharding strategy. After a new node is added, it automatically synchronizes historical log data within the shard range and completes the access within a preset time threshold; an automatic archive cleanup sub-module, used to periodically clean up historical logs and prioritize deleting archived data in storage to free up storage space; and a tiered storage sub-module, used to compress archived data according to a set compression ratio and transfer it to low-cost media for long-term storage according to a storage strategy to ensure data query availability.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] Regarding high availability management of the primary service, this invention introduces a dynamic log priority synchronization mechanism based on the RAFT protocol. By calculating the degree of log differences among candidate nodes, the synchronization priority is dynamically adjusted, prioritizing the synchronization of critical nodes and critical log segments. Compared to existing technologies where all slave nodes participate in election equally without priority distinction, this effectively reduces log synchronization waiting time during primary node switching, enabling monitoring tasks to recover faster and improving system fault tolerance and service continuity. Simultaneously, this invention employs an incremental update strategy instead of the traditional full or batch synchronization method. During primary node switching and fault recovery, only differing log data is synchronized, significantly reducing the amount of synchronized data and lowering log synchronization latency, thereby accelerating the overall system recovery speed.

[0039] Regarding log distribution optimization, this invention proposes a delegation mechanism based on window pipelines. This mechanism dynamically groups nodes into high-priority and low-priority groups and adjusts the grouping strategy based on real-time latency, prioritizing nodes with large lags. This avoids the problem in existing technologies where the overall synchronization time is dragged down by the slowest node. This invention introduces a dynamic window control mechanism, based on network bandwidth R, log entry size C, and receiver node lag. Dynamically calculate the batch distribution window size This avoids network congestion or bandwidth waste caused by fixed batch sizes, achieving efficient utilization of network resources. The matching weight formula designed in this invention... This invention incorporates the difference in lag between the delegator and the receiver, as well as the current load of the delegator, into the matching decision. By prioritizing matching in descending order of weights, log distribution tasks are allocated to delegator nodes with lighter loads and higher lag matching rates, effectively avoiding uneven load distribution and resource waste. Furthermore, the invention includes a timeout and retry mechanism during log distribution. For tasks that time out due to network latency or other issues, the leader reassigns the task to the least loaded, idle delegator node, ensuring the timeliness and stability of the distribution process and improving the system's robustness under abnormal conditions.

[0040] Regarding load balancing of the data acquisition engines, this invention adopts a paired deployment approach, configuring each pair of acquisition engines in both load-sharing and disaster recovery modes. During normal operation, the primary and backup engines share the data acquisition tasks, while the backup engine reserves some redundant resources to cope with sudden loads or failures. When the primary acquisition engine fails, the backup engine automatically detects the anomaly through a health check mechanism (such as heartbeat detection) and takes over the task within milliseconds, achieving rapid failover at the acquisition layer. After the task is completed, the backup engine automatically synchronizes the acquired data to the primary engine's storage, ensuring data consistency between the primary and backup engines and supporting rapid system recovery.

[0041] Regarding horizontal scaling of the storage engine, this invention achieves horizontal scaling through a sharding strategy. After a new node is added to the storage cluster, it automatically synchronizes historical log data within the specified range and completes the integration process quickly, solving the problems of insufficient storage capacity and low scaling efficiency in existing technologies. Simultaneously, this invention periodically performs automatic cleanup of historical logs, prioritizing the deletion of archived data in storage to free up storage space. Archived data is stored on low-cost media according to a certain compression ratio and storage strategy, significantly reducing storage costs while ensuring long-term data preservation and query availability.

[0042] This invention achieves 24 / 7 uninterrupted operation in a distributed database monitoring system through the collaborative work of four modules: high availability management of the main service, optimized log distribution, load balancing of the data collection engine, and horizontal scaling of the storage engine. It effectively solves the problems of slow master node switching, high log synchronization latency, and long system recovery time in existing technologies under high load scenarios, significantly improving the system's reliability, scalability, and performance optimization potential. Taking log distribution latency as an example, when a node has a lag of 50 logs and other nodes have a lag of 5 logs, the overall synchronization time of existing technologies is limited by the 50 nodes. This invention, however, addresses this by using a lag threshold (…). When k is 0.55, the threshold is approximately 27.5. Prioritizing the synchronization of nodes with larger lag significantly reduces the overall synchronization time. Combined with dynamic batch windows and load-aware matching mechanisms, log distribution latency can be reduced by 30% to 50% under actual testing. Each module adopts a distributed, dynamically adaptive design, capable of automatically adjusting strategies based on real-time factors such as the number of nodes, network bandwidth, and load conditions. It possesses excellent large-scale cluster scalability and is suitable for increasingly complex database monitoring scenarios. Detailed Implementation

[0043] Example 1

[0044] To improve the synchronization efficiency of distributed log systems and ensure the real-time performance and consistency of log distribution, a distributed log synchronization system based on a window pipeline delegation mechanism is proposed. This system specifically includes the following modules ( Figure 1 ):

[0045] (1) Main service high availability management module

[0046] This module is primarily responsible for monitoring the status of the system's master service node and enabling rapid failover in case of master node failure or anomaly, ensuring the continuity of monitoring tasks. Addressing the issues of low master node election efficiency and high log synchronization latency in existing distributed systems, a dynamic log priority synchronization optimization mechanism based on the RAFT protocol is proposed.

[0047] When the master node fails, an election process is triggered via the RAFT protocol. The new master node must meet the requirement of consistent latest log state, such as the term number and index value of the last log entry of a candidate node being no less than those of other candidate nodes, in order to quickly take over the task. During log synchronization, a dynamic priority synchronization mechanism adjusts the synchronization priority by calculating the degree of log differences among candidate nodes: prioritizing the synchronization of critical nodes and critical log segments, thereby reducing master node switchover time and ensuring rapid task recovery. The synchronization process employs an incremental update strategy, such as using a single log entry as the incremental unit, synchronizing only the log segments missing from the receiver, further improving efficiency.

[0048] (2) Load balancing module for data acquisition engine

[0049] The data acquisition engine load balancing module is primarily responsible for collecting real-time database operating status, performance metrics, and log data, and achieving task balancing through a distributed architecture. To avoid bottlenecks in high-load scenarios, this system adopts a paired deployment approach, configuring each pair of data acquisition engines in both load-sharing and disaster recovery modes.

[0050] During normal operation, the primary and backup engines share the data collection tasks, while the backup engine reserves some redundant resources to cope with sudden loads or failures. When the primary data collection engine fails, the backup engine automatically detects the anomaly through a health check mechanism (such as heartbeat detection) and takes over the task within milliseconds. After the task is completed, the backup engine automatically synchronizes the collected data to the primary engine's storage, ensuring data consistency and supporting rapid recovery.

[0051] (3) Log distribution optimization module

[0052] The log distribution optimization module is responsible for efficiently synchronizing the collected log data to each storage node and ensuring the real-time performance and consistency of the distribution. Existing distributed log distribution suffers from high latency and large network overhead. This system further optimizes the log distribution strategy by proposing a window pipeline-based delegation mechanism to improve distribution performance.

[0053] Specifically, this mechanism dynamically groups nodes into high-priority and low-priority groups, sorted in descending order of latency, with the top 30% in the high-priority group and the remainder in the low-priority group. The grouping strategy is adjusted based on real-time latency. During distribution, a dynamic window control mechanism is introduced to dynamically adjust the batch size of distributed data based on network bandwidth and node processing capacity. A pipeline model is used to achieve segmented log synchronization, effectively reducing global latency in distribution. Furthermore, the system updates the grouping status after each round of distribution to improve the efficiency and fairness of the next round. The optimized algorithm is as follows: Figure 2 As shown.

[0054] 1) Initialize variables

[0055] Define the delegator node group senderVec and the receiver node group receiverVec. senderVec consists of all slave nodes that have completed log synchronization. Define the lag threshold. ,in The `nextIndexSortVec` is a list of lag values, where `nextIndexSortVec` represents an array of node indices sorted in ascending order of `nextIndex`, `nextIndex` represents the next log position that the follower needs to receive, and `k` is an adjustment parameter; the batch distribution window size is defined as... Where R is network bandwidth and C is log entry size. This represents the lag of the receiver node.

[0056] 2) Principal-Recipient Matching

[0057] After applying the grouping strategy for principal and receiver nodes, the matching weight formula is defined as follows:

[0058] in and Let i and j represent the lag amounts for the principal and receiver, respectively. Let `i` be the current load of the delegator node `i`. For all possible... Calculate weights The weights are sorted in descending order for priority matching.

[0059] 3) Log distribution process

[0060] For each receiver node j, find its matching delegator node i. If a match is found, the leader determines the order of events based on the lag of receiver node j and the batch window size. Calculate the log range to be distributed:

[0061]

[0062] ,in and These are the starting and ending indices for the distribution, respectively.

[0063] The leader updates the replication window of the receiver node and sends a log distribution request to the delegator node.

[0064] 4) Timeout handling and retry mechanism

[0065] During log distribution, for tasks that time out due to network latency or other issues, the leader will reassign the task to the idle delegator node with the lowest load to ensure the timeliness and stability of the distribution process.

[0066] 5) Termination conditions

[0067] When the lag of all nodes satisfies When the system reaches consistency, the log distribution process ends.

[0068] (4) Automatic expansion of storage engine

[0069] The storage engine module is responsible for long-term storage and dynamic expansion of collected logs to meet the high-reliability monitoring requirements of the system. To address the issues of insufficient capacity and low expansion efficiency of existing storage engines, this system utilizes a horizontal scaling mechanism and automatic archive cleanup functionality. Specifically, the storage engine rapidly initializes and synchronizes new nodes through a sharding strategy. After a new node is added to the storage cluster, it automatically synchronizes historical log data within its scope and completes integration within a short time. Furthermore, the system periodically performs automatic cleanup of historical logs, prioritizing the deletion of archived data in storage to free up storage space and reduce storage costs. Archived data is stored on low-cost media according to a specific compression ratio and storage strategy to ensure long-term data preservation and retrieval availability.

[0070] Example 2

[0071] This invention primarily reduces log distribution latency through the aforementioned technical means. In traditional distribution mechanisms, log distribution latency is typically affected by factors such as network bandwidth limitations, log synchronization batch size, and node load differences. This invention introduces a window pipeline mechanism, employing dynamically adjusted batch window size and priority synchronization to reduce distribution latency. Specific optimization effects are as follows:

[0072] 1. Dynamic batch window adjustment

[0073] Assuming network bandwidth Log entry size The lag of receiver node 1 The batch window size is .

[0074] The window size is recalculated after each round of distribution. This dynamic adjustment method effectively avoids distribution batches that are too large or too small, ensuring a balance between network bandwidth and latency, and reducing unnecessary waiting and congestion.

[0075] 2. Optimization of hysteresis

[0076] Without optimization mechanisms, differences in lag can lead to longer synchronization times for most receiver nodes. For example, if a node has a large lag (e.g., 50 records) while other nodes have smaller lags (e.g., 5 records), the overall synchronization time will be affected by the slowest node. After optimization, by calculating a lag threshold, nodes are prioritized and synchronized in batches, allowing for the priority synchronization of nodes with larger lags.

[0077] Assume the lag is ,but:

[0078] By dynamically adjusting strategies to optimize nodes with synchronization lag greater than 27.5, the synchronization speed of the entire system can be accelerated.

[0079] 3. Lag and load balancing

[0080] The node load factor was incorporated when calculating the matching weight between the principal and the receiver. Assuming the load of delegator node 1 is... The hysteresis of sender node i The hysteresis of receiver node j The matching weight is then calculated as follows: The matching weight between nodes with large differences in lag will be lower. By selecting the delegator and receiver nodes with higher matching priority through descending weight sorting, the pressure on heavily loaded nodes can be effectively reduced, ensuring the fairness of the task and the rationality of the allocation.

[0081] Based on the specific mathematical derivation and optimization effects in Example 2, the technical advantages of this invention over the prior art are specifically manifested as follows:

[0082] (1) Network utilization advantage brought by dynamic batch window mechanism

[0083] In existing log distribution mechanisms, log distribution latency is typically affected by factors such as network bandwidth limitations, log synchronization batch size, and node load differences. Fixed batch sizes often lead to network congestion (too large a batch) or bandwidth waste (too small a batch), making it difficult to adapt to dynamically changing network environments. To address this technical problem, this invention introduces a dynamic window control mechanism, which adjusts the window size based on network bandwidth R, log entry size C, and receiver node lag. Dynamically calculate the batch distribution window size, i.e. Taking the specific conditions given in Example 2 as an example, assuming the network bandwidth R = 100 Mbps, the log entry size C = 1 kB, and the receiver node lag... If there are 1000 log entries, then the batch window size is 100 × 1000 = 100,000 log entries. This dynamic adjustment method effectively avoids excessively large or small distribution batches, ensures a balance between network bandwidth and latency, and reduces unnecessary waiting and blocking. The fixed batch mechanism of existing technologies cannot achieve this adaptive optimization.

[0084] (2) The overall synchronization efficiency advantage brought about by the lag priority synchronization mechanism

[0085] In existing technologies, under high-load scenarios, when the lag levels of different nodes vary significantly, the overall log synchronization time is limited by the slowest node with the largest lag. For example, if a node has an excessively large lag (e.g., 50 records) while other nodes have smaller lag levels (e.g., 5 records), the overall synchronization time is completely dragged down by this slow node. To address this problem, this invention prioritizes nodes and performs batch synchronization by calculating a lag threshold. Specifically, a lag threshold is defined. Where k is an adjustment parameter. Taking k=0.55 in Example 2 as an example, when the average lag is 50, the threshold is approximately 27.5, and the system automatically prioritizes synchronizing nodes with a lag greater than 27.5. Through this dynamic adjustment strategy, nodes with larger lags obtain higher synchronization priority, thereby accelerating the synchronization speed of the entire system and effectively avoiding the drawback of the overall synchronization efficiency being constrained by a single slow node in the prior art.

[0086] (3) The advantage of fair task allocation brought by the load-aware matching mechanism

[0087] Existing technologies typically employ round-robin or random allocation methods when assigning log distribution tasks, failing to adequately consider the load differences among nodes. This can easily lead to an unbalanced situation where some nodes are overloaded while others are idle. This invention designs a matching weight formula that incorporates node load factors during the matching process between the principal and the receiver. Taking Example 2 as an example, assume the load of principal node 1 is... (Normalized value), the hysteresis of sender node i The hysteresis of receiver node j The matching weight is then calculated as follows: For matching nodes with significant differences in lag, their weights will be reduced accordingly. By selecting delegating and receiving nodes with higher matching priority through descending weight sorting, this invention can effectively reduce the pressure on heavily loaded nodes, ensuring the fairness of tasks and the rationality of allocation. Existing technologies lack this load-aware, fine-grained matching capability.

[0088] By optimizing the delegation mechanism based on window pipelines, and combining it with lag priority synchronization, load balancing, and timeout retry mechanisms, this invention significantly improves log distribution latency, master node switching, and log synchronization success rate. Specifically, the dynamic batch window mechanism solves the network efficiency problem caused by fixed batches, the lag priority synchronization mechanism avoids bottleneck nodes dragging down overall performance, and the load-aware matching mechanism ensures the fairness of task allocation. These optimizations not only improve the system's real-time performance and stability but also effectively reduce resource waste and enhance the system's high availability and fault tolerance. In contrast, existing technologies have obvious efficiency bottlenecks and performance defects in all of the above aspects. This invention surpasses existing technologies through systematic mathematical modeling and dynamic adaptive mechanisms.

[0089] The units, devices, or modules described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. For ease of description, the above devices are described by dividing them into various modules according to their functions. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.

[0090] Those skilled in the art will also know that, besides implementing the controller using purely computer-readable program code, the same functions can be achieved by logically programming the method steps, making the controller function as logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers (PLCs), and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the devices within it used to implement various functions can also be considered structures within that hardware component. Alternatively, the devices used to implement various functions can be considered as both software modules implementing the method and structures within a hardware component.

[0091] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0092] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software and necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0093] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.

[0094] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A distributed log synchronization method based on a window pipeline delegation mechanism, characterized in that, The method includes the following steps: (1) High availability management of the main service: The status of the main service node of the database monitoring system is monitored based on the RAFT protocol. When the main node fails, the election process is triggered. The new main node must meet the consistency of the latest log status. During the log synchronization process, the synchronization priority is dynamically adjusted by calculating the log difference of the candidate nodes, prioritizing the synchronization of key nodes and key log segments, and using an incremental update strategy for log synchronization. (2) Log distribution optimization: A window pipeline-based delegation mechanism is adopted to dynamically divide nodes into high-priority groups and low-priority groups, and the grouping strategy is adjusted according to real-time latency; a dynamic window control mechanism is introduced to dynamically adjust the batch size of distributed data according to network bandwidth and node processing capacity; segmented synchronization of logs is achieved through a pipeline model; and the grouping status is updated after each round of distribution. (3) Load balancing of acquisition engines: Each pair of acquisition engines is configured with dual modes of load sharing and disaster recovery. The primary and backup engines share the acquisition tasks, and the backup engine retains some redundant resources. When the primary acquisition engine fails, the backup engine automatically detects the anomaly through the health check mechanism and takes over the task. After the task is completed, the acquired data is synchronized to the primary engine storage. (4) Storage engine horizontal scaling steps: Add new nodes to the storage cluster through sharding strategy. After the new node is added, the historical log data within the range is automatically synchronized and the access is completed within the preset time. Perform automatic cleanup on historical logs regularly, prioritize the deletion of archived data to free up storage space, and transfer the archived data to low-cost media for long-term storage according to the set compression ratio and storage strategy.

2. The distributed log synchronization method based on a window pipeline delegation mechanism as described in claim 1, characterized in that, The log distribution optimization steps further include: (a) Initialization: Define the delegator node group senderVec and the receiver node group receiverVec. The senderVec consists of all slave nodes that have completed log synchronization; define the lag threshold. ,in The `nextIndexSortVec` is a list of lag values, where `nextIndexSortVec` represents an array of node indices sorted in ascending order of `nextIndex`, `nextIndex` represents the next log position that the follower needs to receive, `k` is an adjustment parameter, `median(ΔL)` is the median lag value of all nodes, and `MAD(ΔL)` is the median absolute deviation of the lag values. The batch distribution window size is defined. R is the network bandwidth, and C is the log entry size. The hysteresis of the receiver node; (b) Principal-Recipient Matching: The matching weight formula is defined as follows ,in , Let be the lags for principal i and receiver j, respectively. Given the current load of the delegator node i; calculate the weights for all possible matching pairs, sort the weights in descending order, and prioritize matching them. (c) Log distribution: For each receiver node j, find its matching delegator node i, and the leader distributes logs based on the lag of receiver node j and the batch window size. Calculate the log range to be distributed: N j ), , and These are the starting and ending indices for distribution; the leader updates the replication window of the receiver node and initiates a log distribution request to the delegator node; (d) Timeout handling: For tasks that time out due to network latency or other issues, (supplement the timeout threshold formula or provide a fixed value), the leader will reassign the task to the idle delegator node with the lowest load; (e) When the lag of all nodes satisfies When the system reaches consistency, the log distribution process ends.

3. The distributed log synchronization method based on a window pipeline delegation mechanism as described in claim 1, characterized in that, In the main service high availability management steps, the dynamic priority synchronization mechanism is as follows: by calculating the log difference degree of candidate nodes, the nodes are sorted from largest to smallest according to their lag, and nodes with a lag greater than the current average lag multiplied by a preset coefficient are synchronized first.

4. The distributed log synchronization method based on a window pipeline delegation mechanism as described in claim 1, characterized in that, In the load balancing step of the acquisition engine, the health check mechanism uses heartbeat detection, and the task takeover time is in the millisecond range.

5. The distributed log synchronization method based on a window pipeline delegation mechanism as described in claim 1, characterized in that, In the horizontal scaling step of the storage engine, new nodes are added to the storage cluster through a sharding strategy. After the new node is added, historical log data within the shard range is automatically synchronized, and the access is completed within a preset time threshold. Regularly clean up historical logs automatically, prioritizing the deletion of archived data in storage to free up storage space; Archived data is compressed according to a set compression ratio and transferred to low-cost media for long-term storage based on storage strategy, ensuring data availability for retrieval.

6. A distributed database monitoring system based on a high availability mechanism, characterized in that, The system consists of the following modules: The main service high availability management module monitors the status of the main service node of the database monitoring system based on the RAFT protocol. When the main node fails, an election process is triggered, and the new main node must meet the requirement of consistency with the latest log status. During the log synchronization process, the synchronization priority is dynamically adjusted by calculating the degree of log difference between candidate nodes, prioritizing the synchronization of key nodes and key log segments, and using an incremental update strategy for log synchronization. Log distribution optimization module: It adopts a window pipeline-based delegation mechanism to dynamically divide nodes into high-priority groups and low-priority groups, and adjusts the grouping strategy according to real-time latency; it introduces a dynamic window control mechanism to dynamically adjust the batch size of distributed data according to network bandwidth and node processing capacity; it realizes segmented synchronization of logs through a pipeline model; and it updates the grouping status after each round of distribution. Data Acquisition Engine Load Balancing Module: Each pair of acquisition engines is configured with dual modes of load sharing and disaster recovery. The primary and backup engines share the acquisition tasks, and the backup engine retains some redundant resources. When the primary acquisition engine fails, the backup engine automatically detects the anomaly through a health check mechanism and takes over the task. After the task is completed, the acquired data is synchronized to the primary engine's storage. Storage engine horizontal scaling module: Adds new nodes to the storage cluster through sharding strategy. After the new node is added, it automatically synchronizes the historical log data within the range and completes the access within a preset time. It also performs automatic cleanup of historical logs on a regular basis, prioritizing the deletion of archived data to free up storage space, and transferring the archived data to low-cost media for long-term storage according to the set compression ratio and storage strategy.

7. A distributed log synchronization system based on a window pipeline delegation mechanism as described in claim 6, characterized in that, The log distribution optimization module completes log distribution optimization using the following steps: (a) Initialization: Define the delegator node group senderVec and the receiver node group receiverVec. The senderVec consists of all slave nodes that have completed log synchronization; define the lag threshold. ,in The `nextIndexSortVec` is a list of lag values, where `nextIndexSortVec` represents an array of node indices sorted in ascending order of `nextIndex`, `nextIndex` represents the next log position that the follower needs to receive, `k` is an adjustment parameter, `median(ΔL)` is the median lag value of all nodes, and `MAD(ΔL)` is the median absolute deviation of the lag values. The batch distribution window size is defined. R is the network bandwidth, and C is the log entry size. The hysteresis of the receiver node; (b) Principal-Recipient Matching: The matching weight formula is defined as follows ,in and Let be the lags for principal i and receiver j, respectively. Given the current load of the delegator node i; calculate the weights for all possible matching pairs, sort the weights in descending order, and prioritize matching them. (c) Log distribution: For each receiver node j, find its matching delegator node i, and the leader distributes logs based on the lag of receiver node j and the batch window size. Calculate the log range to be distributed: N j ), , and These are the starting and ending indices for distribution; the leader updates the replication window of the receiver node and initiates a log distribution request to the delegator node; (d) Timeout handling: For tasks that time out due to network latency or other issues, (supplement the timeout threshold formula or provide a fixed value), the leader will reassign the task to the idle delegator node with the lowest load; (e) When the lag of all nodes satisfies When the system reaches consistency, the log distribution process ends.

8. A distributed log synchronization system based on a window pipeline delegation mechanism as described in claim 6, characterized in that, In the main service high availability management module, the dynamic priority synchronization mechanism is as follows: by calculating the log difference of candidate nodes, the nodes are sorted from largest to smallest according to their lag, and nodes with a lag greater than the current average lag multiplied by a preset coefficient are synchronized first.

9. A distributed log synchronization system based on a window pipeline delegation mechanism as described in claim 6, characterized in that, In the load balancing module of the acquisition engine, the health check mechanism uses heartbeat detection, and the task takeover time is in the millisecond range.

10. A distributed log synchronization system based on a window pipeline delegation mechanism as described in claim 6, characterized in that, The storage engine horizontal scaling module includes the following sub-modules: The new node expansion submodule is used to add new nodes to the storage cluster through sharding strategies. After a new node is added, it automatically synchronizes historical log data within the shard range and completes the access within a preset time threshold. The automatic archive cleanup submodule is used to automatically clean up historical logs on a regular basis, prioritizing the deletion of archived data in storage to free up storage space. The tiered storage submodule is used to compress archived data according to a set compression ratio and transfer it to low-cost media for long-term storage according to the storage strategy, ensuring the availability of data queries.