Data synchronization method and device based on distributed storage
By collecting and analyzing network status data in a distributed storage system and dynamically adjusting synchronization frequency and task plan, the synchronization problem of distributed storage system in complex network environments is solved, data consistency and resource utilization are balanced, and the stability and efficiency of the system are improved.
Patent Information
- Application Number
- CN202510553432.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-18
AI Technical Summary
When facing a complex and dynamic network environment, existing distributed storage systems are difficult to achieve an effective trade-off between synchronization overhead and data timeliness. The static regulation method cannot adapt to real-time changes in network state, resulting in network congestion or data consistency problems.
By collecting network status data of each node, analyzing network delay and bandwidth fluctuations parameters, calculating dynamic synchronization frequency thresholds, adjusting data synchronization plan between nodes, generating a synchronization task list, and executing tasks to achieve real-time synchronization.
It significantly improves the flexibility and robustness of data synchronization in complex network environments, ensures data consistency and real-time, avoids network congestion, improves resource utilization efficiency, and enhances the application stability of the system in key areas.
Smart Images

Figure CN120343041A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a data synchronization method and device based on distributed storage. Background Art
[0002] In modern distributed storage systems, the data synchronization mechanism is an important part of ensuring system consistency and high availability. With the development of cloud computing, edge computing, and hybrid architectures, distributed storage is widely used in key fields such as finance, telecommunications, and manufacturing, which puts forward higher requirements for its synchronization efficiency and the intelligence of synchronization strategies.
[0003] Existing data synchronization methods can be mainly divided into two categories: one is the periodic synchronization method based on fixed time intervals, and the other is the event-driven synchronization method triggered by data changes. These methods can better complete the synchronization tasks in traditional local area network environments or deployment environments with relatively stable network conditions. However, in practical applications with wide-area distribution or multi-region collaboration, due to significant fluctuations in the network environment, such as increased latency, bandwidth jitter, or temporary congestion, fixed or static strategies often cannot balance the real-time nature of data consistency and the efficient utilization of network resources. Specifically, if the synchronization frequency is set too high, in the case of limited bandwidth resources or increased network latency, frequent synchronization may cause network congestion and resource contention, thus affecting the normal operation of the business system; while if the synchronization frequency is set too low, it may lead to lagging data updates, affecting the ultimate consistency and real-time availability of data, and even causing business failures or data consistency errors in some scenarios. In addition, most current synchronization mechanisms lack the ability to perceive the network state in real time, cannot dynamically collect and analyze key network parameters, such as the current end-to-end latency, bandwidth occupancy, etc., and are even less able to adaptively adjust the synchronization strategy according to these parameters. This static regulation method is difficult to achieve an effective balance between synchronization overhead and data timeliness in the face of a complex and dynamic distributed operating environment. Summary of the Invention
[0004] The purpose of the present invention is to provide a data synchronization method and device based on distributed storage, which can regulate the synchronization frequency according to network latency and bandwidth fluctuations to solve the real-time nature and performance balance problems in distributed storage.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A data synchronization method based on distributed storage, the method includes:
[0006] S1. Collect and analyze the network state data of each node to obtain network latency and bandwidth fluctuation parameters;
[0007] S2. Calculate the adaptable dynamic synchronization frequency threshold based on network latency and bandwidth fluctuation parameters, including converting the current network latency and bandwidth fluctuation parameters into a recommended time interval, calculating the current recommended synchronization frequency, which is the dynamic synchronization frequency threshold, and allocating the upper limit of the maximum number of tasks allowed to initiate synchronization per minute according to this rhythm;
[0008] S3. Adjust the data synchronization plan between nodes according to the dynamic synchronization frequency threshold to generate a synchronization task list, including numbering and initially sorting all data blocks to be synchronized. If the synchronization targets are different nodes or different protocols between every two tasks, it is recorded as a task switch, and the total switching overhead is counted. Minimize the total switching overhead and regroup and sort the task list;
[0009] S4. Execute the tasks in the synchronization task list to complete the real-time synchronization of data in the distributed storage system, including assigning a priority weight value to each task, pairing any two tasks, calculating their priority comparison value, and combining multiple tasks to be issued in parallel or sequentially based on the priority comparison value.
[0010] Preferably, S1 includes setting reference values for network latency and bandwidth respectively. Each node periodically collects the current latency and bandwidth, compares them with the reference values respectively to obtain ratios, sets a ratio threshold, and marks the node as a high-pressure node when the ratio is greater than the threshold.
[0011] Preferably, the specific calculation method of obtaining the ratios by comparing the current latency and bandwidth with the reference values in S1 includes setting recommended reference values for two key parameters of network performance, network latency and network bandwidth respectively. Each node regularly collects its own current network status, including the current actual network latency value and the current actual network bandwidth value. Divide the currently measured latency value by the set latency reference value to calculate the balance ratio of latency, and divide the reference bandwidth value by the currently measured bandwidth value to obtain the balance ratio of bandwidth. Analyze the latency ratio and the bandwidth ratio together. If the average value of the two ratios or any one ratio is higher than the preset threshold, mark the node as a high-pressure node.
[0012] Preferably, the specific method of calculating the current recommended synchronization frequency in S2 includes comprehensively analyzing based on the current network latency and bandwidth fluctuation conditions to obtain a recommended time interval, dividing sixty by the above-mentioned recommended time interval to calculate the number of synchronizations that should be performed per minute currently. The calculated number of synchronizations per minute is the dynamic synchronization frequency threshold, and the number of data synchronization tasks allowed to be started at most within each minute is restricted according to this frequency threshold.
[0013] Preferably, the specific method for counting the total switching overhead in S3 includes numbering all the data synchronization tasks to be executed, sorting them according to a preliminary strategy, and then, for each pair of adjacent tasks, determining whether they belong to different synchronization target nodes, protocol types, or network paths. If such differences exist, the relationship between this pair of tasks is regarded as a task switch, and the amount of resources consumed for each task switch is determined, including latency or bandwidth occupancy. Multiply the total number of task switches identified in the previous step by the fixed resource overhead for each switch to obtain the total resource loss caused by switching in the current task arrangement.
[0014] Preferably, the specific method for pairing any two tasks and calculating their priority comparison value in S4 includes setting a priority level value representing the importance or urgency for each data synchronization task to be executed, selecting any two tasks from the task list for pairing, and calculating the numerical difference between them by comparing the corresponding priority level values of these two tasks.
[0015] Preferably, S2 further includes collecting the latency values generated during network communication among multiple nodes in the distributed storage system, calculating the average latency time among all nodes through statistical analysis of the latency data, setting a basic synchronization frequency threshold, comparing the calculated current average latency among nodes with a pre-set reference latency value, dividing the current average latency value by the reference latency value to obtain a scaling factor, and then multiplying this scaling factor by the aforementioned basic synchronization frequency threshold to obtain a new synchronization frequency value. If the calculated dynamic synchronization frequency value is lower than the pre-set minimum allowable frequency threshold, the current frequency value is increased to this minimum threshold.
[0016] Preferably, the specific method for statistically analyzing the latency data and calculating the average latency time among all nodes in S2 includes recording the time required to send data from one node to another during network communication between each pair of nodes, which is the latency for this communication. Uniformly collect and store the latency records generated by all communication operations among nodes to form a data set containing multiple latency data. Perform cleaning operations on the collected latency data, including removing outliers and invalid records. For all the valid latency records after cleaning, sum up each value one by one, count the total number of these summed latency records, and divide the total latency time by the total number of records to obtain the average latency time of communication among nodes in the current system.
[0017] Preferably, the basic synchronization frequency threshold in S2 is set according to the real-time requirements of the service type in the distributed storage system.
[0018] A data synchronization device based on distributed storage, which is used to implement the steps of the data synchronization method based on distributed storage. The device includes:
[0019] A network status collection module, which is used to collect the network status data of each node and analyze to obtain network delay and bandwidth fluctuation parameters;
[0020] A frequency calculation module connected to the network status collection module, which is used to calculate an adapted dynamic synchronization frequency threshold based on the network delay and bandwidth fluctuation parameters;
[0021] A synchronization plan generation module connected to the frequency calculation module, which is used to adjust the data synchronization plan between nodes according to the dynamic synchronization frequency threshold and generate a synchronization task list;
[0022] A task execution module connected to the synchronization plan generation module, which is used to execute the tasks in the synchronization task list to complete the real-time synchronization of data in the distributed storage system.
[0023] As can be seen from the above technical solutions, the present invention has the following beneficial effects:
[0024] The data synchronization method and device based on distributed storage collect and analyze the network status data of each node to obtain network delay and bandwidth fluctuation parameters, calculate an adapted dynamic synchronization frequency threshold based on the network delay and bandwidth fluctuation parameters, adjust the data synchronization plan between nodes according to the dynamic synchronization frequency threshold, generate a synchronization task list, and execute the tasks in the synchronization task list to complete the real-time synchronization of data in the distributed storage system. It can effectively solve the problems of static synchronization strategy, unreasonable frequency setting, and lack of perception of network status in the prior art. By introducing a real-time network status collection and analysis mechanism, the system can dynamically obtain the delay and bandwidth fluctuation parameters between nodes, and further adjust the synchronization frequency threshold accordingly, so as to realize the adaptive optimization of the synchronization strategy, significantly improving the flexibility and robustness of data synchronization in complex network environments. It can increase the synchronization frequency when the network condition is good to ensure data consistency and real-time performance, and automatically reduce the frequency when the network fluctuates severely to avoid congestion and performance degradation caused by excessive transmission, ensuring the efficient use of network resources. At the same time, by introducing dynamic regulation logic, the system has stronger environmental adaptability in scenarios such as multi-region collaboration and cross-regional deployment, and can more stably and efficiently support the implementation of distributed storage systems in key fields such as finance, telecommunications, and manufacturing. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a flowchart of the method of the present invention;
[0026] Figure 2 It is a connection diagram of the device modules of the present invention. Detailed Implementation Manner
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] As Figure 1 shown, the present invention provides a technical solution: a data synchronization method based on distributed storage, and the method includes:
[0029] S1. Collect and analyze the network status data of each node to obtain network delay and bandwidth fluctuation parameters;
[0030] S2. Based on the network delay and bandwidth fluctuation parameters, calculate an adapted dynamic synchronization frequency threshold, including converting the current network delay and bandwidth fluctuation parameters into a recommended time interval, calculating the current recommended synchronization frequency, and this frequency is the dynamic synchronization frequency threshold, and allocate the upper limit of the maximum number of tasks allowed to initiate synchronization per minute according to this rhythm;
[0031] S3. Adjust the data synchronization plan between nodes according to the dynamic synchronization frequency threshold to generate a synchronization task list, including numbering and initially sorting all data blocks to be synchronized, recording a task switch if the synchronization targets are different nodes or different protocols between every two tasks, counting the total switching overhead, minimizing the total switching overhead, and regrouping and sorting the task list;
[0032] S4. Execute the tasks in the synchronization task list to complete the real-time synchronization of data in the distributed storage system, including assigning a priority weight value to each task, pairing any two tasks, calculating their priority comparison value, and combining multiple tasks to be issued in parallel or sequentially based on the priority comparison value.
[0033] Based on the requirements of data consistency and synchronization efficiency in heterogeneous network environments for distributed storage systems, this embodiment designs a dynamically adaptable synchronization scheduling mechanism. The core of this mechanism is to use the network state as an input variable to dynamically calculate the synchronization rhythm that matches the current network conditions, thereby achieving global optimization control of the synchronization process. Specifically, in step S1, the system periodically sends probe packets between nodes to collect the network latency between each node in real time, such as RTT, and bandwidth changes, such as the variance of throughput. The data is filtered and normalized by the edge analysis module to improve the timeliness and robustness of the parameters. The collected parameters are input into the synchronization frequency calculation engine. In step S2, this engine uses a preset mapping function or a trained regression model, such as linear regression or neural network, to convert the current network parameters into the optimal synchronization time interval, and accordingly determines the upper limit of the number of synchronization tasks that can be initiated per minute, forming a dynamic synchronization frequency threshold. In step S3, the system combines information such as the physical distribution of data blocks, synchronization target nodes, and used protocols to preliminarily number and sort the synchronization tasks to be processed. Considering the costs such as handshakes and authentication caused by each target node or protocol switch, the system establishes a switching overhead calculation model, which sums the number of switches between tasks and the unit switching cost weighted. Then, with the minimization of the total overhead as the objective function, through greedy algorithms, dynamic programming, or heuristic rearrangement algorithms, the reorganization and optimization of the synchronization task list are completed. In step S4, to ensure the rationality and concurrency of task scheduling, the system assigns an initial priority weight to each task, which is calculated based on factors such as the urgency of the task, the importance of the data, and the historical synchronization failure rate. During the task execution process, the system pairs tasks two by two, calculates the priority difference between them, and based on this difference, determines whether they are suitable for parallel execution or should be queued hierarchically and executed sequentially. The task scheduler finally generates a scheduling plan to control the task execution queue, ensuring that high-priority tasks are allocated resources first, and low-priority tasks are processed later, achieving dynamic optimal allocation of resources. This method as a whole embodies the distributed scheduling idea of combining multi-dimensional parameter drive, task structure optimization, and priority control, not only ensuring the real-time performance and consistency of data synchronization, but also enhancing the network adaptability and concurrency of the system, and having high engineering operability and system integration compatibility.
[0034] By introducing a dynamic synchronization frequency threshold, this method enables the data synchronization behavior of the distributed storage system to adapt to network environment fluctuations, enhancing the stability and responsiveness of the system. Through the mechanism of minimizing task switching overhead, the synchronization efficiency is improved, especially suitable for heterogeneous environments with multiple nodes and multiple protocols. The priority mechanism combined with the task comparison strategy can achieve reasonable task scheduling and resource utilization, effectively improving the overall synchronization efficiency, reducing system latency, and enhancing the real-time performance and scalability of the system.
[0035] Take a certain video cloud platform as an example. The platform is deployed between data centers in multiple locations, and the video files uploaded by users need to be quickly distributed and backed up globally. The traditional fixed-frequency synchronization mechanism will cause bandwidth congestion between nodes during peak hours, and insufficient resource utilization during low hours. After applying the method of the present invention, the platform can dynamically adjust the synchronization rhythm according to the network conditions of each data center. For example, between the Tokyo node and the Los Angeles node, due to the large fluctuations in the international link, the system automatically identifies frequent bandwidth jitters, reduces the synchronization frequency, and prioritizes batch synchronization during bandwidth stability periods; between nodes in the same city such as New York and Boston, the network stability is high, and the system increases the synchronization frequency to achieve minute-level data consistency updates. This mechanism not only significantly improves the synchronization efficiency, but also reduces unnecessary network load, ensuring the consistency of the user experience when accessing video resources in any area.
[0036] S1 includes setting reference values for network delay and bandwidth respectively. Each node periodically collects the current delay and bandwidth, compares them with the reference values respectively, obtains the ratio, sets the ratio threshold, and when the ratio is greater than the threshold, marks the node as a high-voltage node.
[0037] This implementation method introduces a reference value mechanism to dynamically determine the network status of each node. Specifically, in the initialization phase of the data synchronization system, reference values are first set for network delay and bandwidth, respectively, which can be obtained based on historical network performance statistics or initial test results of the system operation. Each distributed node is equipped with a local monitoring module to collect the current network delay, such as ping value and bandwidth utilization rate, such as data throughput per unit time, according to a predetermined period. The system compares the collected real-time values with the reference values, calculates the ratio of "current value / reference value", and thus quantifies the degree of deviation of the current network status. When the ratio exceeds a preset threshold, for example, 1.2 indicates that the current delay or bandwidth load is 20% higher than the reference state, the system automatically marks the node as a "high-voltage node". The high-voltage node identification will serve as an important basis for subsequent synchronization scheduling decisions, and will be used for the formulation of strategies such as synchronization frequency reduction and task transfer. This mechanism enables the system to identify pressure hotspots in real time, avoid initiating frequent synchronization operations at network congestion nodes, effectively disperse the synchronization load, and maintain the stability and high availability of the overall synchronization process.
[0038] This implementation method provides the system with an efficient and low-cost dynamic network health assessment method through a ratio judgment method. Compared with the traditional method that relies on absolute numerical alarms, this ratio judgment method has stronger adaptability and universality, and can automatically adapt to the network baseline conditions of different nodes. The high-voltage node identification mechanism makes data synchronization scheduling forward-looking, which can avoid synchronization bottlenecks or data delays caused by single-point pressure anomalies, and significantly improve the load balancing capability and data consistency assurance capability of the distributed system.
[0039] In S1, the current delay and bandwidth are respectively compared with the reference values to obtain the ratio. The specific calculation methods include setting recommended reference values for two key parameters of network performance, namely network delay and network bandwidth. Each node regularly collects its own current network status, including the current actual network delay value and the current actual network bandwidth value. Divide the currently measured delay value by the set delay reference value to calculate the balance ratio of the delay. Divide the reference bandwidth value by the currently measured bandwidth value to obtain the balance ratio of the bandwidth. Analyze the delay ratio and the bandwidth ratio together. If the average value of the two ratios or any one of the ratios is higher than the preset threshold, mark this node as a high-pressure node.
[0040] This embodiment details the specific method for calculating the ratio of network delay and bandwidth. The system first sets recommended reference values for the network performance of each node, including the recommended delay reference value, such as 50ms, and the recommended bandwidth reference value, such as 100Mbps. During operation, each node's built-in network monitoring module regularly collects its current actual network status parameters, namely the real-time delay value and the real-time bandwidth value. The delay ratio calculation method is "current delay value ÷ delay reference value", and its result reflects whether the delay of this node at the current moment exceeds the expected baseline; the bandwidth ratio calculation method is "bandwidth reference value ÷ current bandwidth value", which is used to represent the reduction degree of the current available bandwidth relative to the baseline bandwidth. The larger both ratios are, the heavier the network load is. After obtaining the two ratios, execute a joint evaluation strategy: separately extract the delay ratio and the bandwidth ratio, calculate their average value, and compare it with the preset ratio threshold, such as 1.3; or use the threshold for separate judgment. As long as any one ratio exceeds the threshold, it can be determined that this node is in a network pressure state and marked as a "high-pressure node". This mark will be used for optimization processing such as scheduling avoidance or task transfer in subsequent synchronization tasks to ensure the overall stability of the system operation.
[0041] This embodiment provides an accurate and two-dimensional network status judgment mechanism, which can more comprehensively reflect the network bearing capacity of the node. Through the combined analysis of the delay and bandwidth ratios, not only can the bandwidth bottleneck be identified, but also the delay anomaly caused by poor link quality can be captured simultaneously, providing a more targeted load identification basis for the system. The joint evaluation strategy enhances the robustness of anomaly detection and significantly improves the system's response ability to network changes, thus ensuring the real-time and accuracy of resource configuration during the distributed data synchronization process.
[0042] The specific method for calculating the current recommended synchronization frequency in S2 includes comprehensively analyzing the current network delay and bandwidth fluctuation conditions to obtain a recommended time interval, dividing sixty by the above-mentioned recommended time interval to calculate the number of synchronizations that should be performed per minute currently. The calculated number of synchronizations per minute is the dynamic synchronization frequency threshold. Based on this frequency threshold, limit the maximum number of data synchronization tasks allowed to be started within each minute.
[0043] In this embodiment, by introducing a computational synchronization frequency model, a dynamic control mechanism for the triggering frequency of data synchronization tasks is realized. In step S2 of the system, first, a joint analysis is performed on the current network delay values and bandwidth fluctuation parameters collected by each node. When the delay fluctuation is large or the bandwidth drops significantly, the system determines that the network carrying capacity is weakened, and it is necessary to appropriately extend the time interval of the synchronization task. Conversely, the time interval can be shortened to increase the synchronization frequency. The specific calculation process is as follows: The system constructs a time interval recommendation function, with the delay standard deviation and bandwidth variance as the main input variables, and outputs the recommended synchronization interval (in seconds). For example, when the network fluctuation is small and the stability is strong, the recommended interval is 5 seconds, and when the fluctuation is large, it is adjusted to 15 seconds. This recommended time interval reflects the minimum synchronization time window suitable for the current network to carry. Subsequently, 60 (seconds) is divided by the recommended time interval to calculate the maximum number of synchronizations that can be carried per minute theoretically, which is used as the dynamic synchronization frequency threshold. The system uses this threshold as an upper limit restriction mechanism, allowing at most the same amount of data synchronization tasks to be triggered within one minute, preventing task backlog, failed retries, or delay spread caused by deteriorating network conditions, thereby realizing rhythmic and stable scheduling of synchronization loads.
[0044] By dynamically calculating the currently recommended synchronization interval time and deriving the synchronizable frequency therefrom, the present invention greatly enhances the synchronization adaptability of the distributed system in a changing network environment. This method avoids bandwidth congestion and system resource competition caused by too high a synchronization frequency, and also prevents the problem of decreased data timeliness caused by too low a frequency. The introduction of the upper limit of the number of synchronization tasks effectively improves the stability and task success rate of the system, and significantly improves the overall operation efficiency and response ability of the system.
[0045] The specific method for counting the total switching overhead in S3 includes numbering all the data synchronization tasks to be executed and sorting them according to a preliminary strategy. Then, for each pair of adjacent tasks, it is judged whether they belong to different synchronization target nodes, protocol types, or network paths. If there are such differences, the relationship between this pair of tasks is regarded as a task switch, and the amount of resources consumed for each task switch is determined, including delay time or bandwidth occupancy. Multiply the total number of task switches identified in the previous step by the fixed resource overhead for each switch to obtain the total resource loss caused by switching in the current task arrangement.
[0046] This embodiment details a quantitative calculation mechanism for the total task switching overhead, aiming to provide a measurable cost function for subsequent task scheduling optimization. In step S3, the system first numbers all the data synchronization tasks to be executed and performs a preliminary sorting according to criteria such as the importance of the tasks, the amount of data, or the synchronization time window, generating a preliminary task sequence. Then, the system analyzes each pair of adjacent tasks to determine whether they are synchronized to different target nodes, use different data transfer protocols, such as HTTP and FTP, or need to pass through different network paths, such as the backbone network and the edge network segment. If any of the above differences exist, it is determined that there is a "switch" between this pair of tasks. The system sets a set of fixed or adjustable resource cost parameters for each switch, usually including the additional latency caused by task switching, such as handshake latency and connection initialization time, and the bandwidth occupancy, such as the bandwidth waste caused by protocol conversion. The number of all identified task switches is multiplied by the resource overhead of a single switch to obtain the total switching resource loss under the current task arrangement structure. This statistical value will be used as the objective function in the sorting optimization process, and the system can use this function to guide subsequent task rearrangement to achieve the goal of minimizing task switching and improve the coherence of task execution and the utilization efficiency of network resources.
[0047] By constructing a clear and computable task switching cost model, the present invention effectively improves the scientificity and controllability of synchronous task scheduling optimization. Compared with the traditional method of simply sorting based on task content, this method realizes a systematic evaluation of task continuity and resource cost by identifying and evaluating the key switching factors between tasks. The optimized task order can significantly reduce the protocol conversion overhead, network path adjustment time, and connection reconstruction delay, thus accelerating the synchronization process, reducing the failure rate, and saving the overall resource consumption, especially suitable for complex network environments with mixed node types and protocols.
[0048] In S4, any two tasks are paired, and the specific method for calculating their priority comparison value includes setting a priority level value representing the importance or urgency of each data synchronization task to be executed, selecting any two tasks from the task list for pairing, and calculating the numerical difference between them by comparing the priority level values corresponding to these two tasks.
[0049] In this embodiment, by introducing a priority level system, a clear decision-making basis and task combination strategy are provided for the task scheduling process. In step S4, the system first assigns a priority level value to each data synchronization task to be executed. This value can be set based on multiple dimensions of metrics, such as the business impact of the data, the requirements for synchronization timeliness, the user service level agreement, or the historical failure of the previous task, etc., and a comprehensive score is calculated. For example, an integer priority level can be used, ranging from 1 (lowest) to 10 (highest). Subsequently, the scheduler sequentially selects any two tasks from the task list for pairing, extracts their priority level values respectively, and calculates the absolute difference between these two values, which is the "priority comparison value". This comparison value reflects the relative difference in the execution urgency between tasks. The larger the value, the greater the disparity in the priorities of the two tasks. The system uses this priority comparison value to assist in determining whether to combine or hierarchically schedule the two tasks. For example, for tasks with similar priorities, the comparison value is small, and parallel processing can be arranged to improve processing efficiency; for tasks with large priority differences, the high-level tasks are issued first, and the low-level tasks are delayed to ensure that the system resources give priority to guaranteeing critical services. This mechanism realizes task scheduling granularity control and strategy optimization through a digital evaluation method, improving the scheduling intelligence and execution efficiency of the distributed system.
[0050] After introducing the priority comparison mechanism, the task scheduling process becomes more refined and strategic, avoiding the risk that high-value tasks are postponed by the traditional "first come, first served" mode. By making the priority level and its numerical difference explicit, the scheduler can flexibly formulate a task combination plan. On the premise of ensuring that high-priority tasks are completed in a timely manner, the system concurrent resources are reasonably utilized to process secondary tasks, thereby improving the overall resource utilization efficiency and the key business response ability.
[0051] S2 also includes collecting the delay values generated during network communication among multiple nodes in the distributed storage system. By statistically analyzing the delay data, calculating the average delay time among all nodes, setting a basic synchronization frequency threshold, comparing the calculated current average delay among nodes with a pre-set reference delay value, dividing the current average delay value by the reference delay value to obtain a proportionality factor, and then multiplying this proportionality factor by the aforementioned basic synchronization frequency threshold to obtain a new synchronization frequency value. If the calculated dynamic synchronization frequency value is lower than the pre-set minimum allowable frequency threshold, the current frequency value is increased to this minimum threshold.
[0052] In this embodiment, the synchronous frequency adjustment mechanism is further improved through the delay mean modeling method, realizing the overall evaluation and adaptive control of the network performance of the entire system. In step S2 of the system, the network communication delay data between all nodes in the distributed storage system is regularly collected. These delay data are unified in standard through the sampling mechanism, and after removing extreme values, they are averaged to obtain the "average network delay time" within the current system range. To adapt to the overall change of network load, the system sets a "basic synchronous frequency threshold", such as the default of 20 synchronizations per minute, and at the same time presets a "reference delay value", such as 50 ms. Dividing the actually measured "current average delay" by the "reference delay" can obtain the proportional factor of the network pressure contrast. If the proportional factor is greater than 1, it means that the overall system delay is relatively high currently, and the synchronous frequency should be appropriately reduced; if the proportional factor is less than 1, it means that the network load is light and the frequency can be increased. The system multiplies this proportional factor by the basic frequency threshold to calculate the new "dynamic synchronous frequency value". To avoid data delay problems caused by too low frequency, the system also sets a "minimum allowable frequency threshold", such as not less than 5 times per minute. If the dynamic synchronous frequency is lower than this lower limit, the system forcibly increases it to the minimum frequency threshold to ensure data consistency and service quality.
[0053] This method makes an overall evaluation of network delay at the system level, making the frequency adjustment mechanism more comprehensive and robust. By introducing the concept of proportional factor, the system can respond sensitively to network load, automatically tighten and relax the synchronization rhythm, and effectively prevent the frequency setting from deviating from the actual network carrying capacity. At the same time, setting the minimum frequency threshold also ensures that even under high network pressure, the minimum data synchronization can be maintained, preventing node data lag or consistency damage, and improving the service stability and fault tolerance of the system.
[0054] In S2, the delay data is statistically analyzed to calculate the average delay time between all nodes. The specific method includes recording the time required to send data from one node to another when performing network communication between each node. This time is the delay of this communication. All the delay records generated by the communication operations between all nodes are uniformly collected and stored to form a data set containing multiple delay data. The collected delay data is cleaned, including removing outliers and invalid records. For all the valid delay records after cleaning, each value is added up one by one, the total number of these added delay records is counted, and the sum of the delay times is divided by the total number of records to obtain the average delay time of the communication between nodes in the current system.
[0055] This embodiment elaborates in detail the entire process of delayed data statistics and average calculation, aiming to provide accurate and real network performance data support for dynamic synchronization frequency calculation. Specifically, when the system performs data transmission operations between execution nodes, timestamps are marked at the initiation of each communication request and the return of the response. The communication delay for this time is obtained by calculating the difference between these two time points. All communication operations between nodes upload their corresponding delay values in real-time or in batches, and the system aggregates these values uniformly to construct a "delay data set". To improve the accuracy and representativeness of the data analysis results, the system performs a cleaning operation on this set: removing obvious outliers in the data, such as extreme values exceeding the maximum reasonable network delay, as well as invalid records with incorrect formats or missing fields. After cleaning, the system accumulates the values in the remaining valid delay records one by one and records the total number of records. Finally, the "average delay time" of the current inter-node communication of the system is calculated using the method of "total delay time ÷ number of valid records", and this value is used as one of the network performance reference bases for subsequent synchronization frequency adjustment.
[0056] This method provides system-level network status evaluation in a data-driven manner, eliminates the dependence on single or instantaneous node status, and improves the stability and representativeness of the synchronization decision-making basis. The delayed data cleaning mechanism can significantly enhance the system's robustness to network noise and measurement deviations, thereby improving the accuracy of average delay time evaluation. The obtained average value, as one of the core parameters for synchronization frequency adjustment, makes the system adjustment strategy more reasonable and accurate, and enhances the adaptability of the data synchronization process to network load fluctuations.
[0057] The basic synchronization frequency threshold in S2 is set according to the real-time requirements of service types in the distributed storage system. By incorporating the real-time requirements of service types into the basis for setting the synchronization frequency threshold in this embodiment, the basic synchronization rhythm has scenario adaptability and policy flexibility. In step S2, the system will pre-classify the service types served by different data synchronization tasks, such as: online payment services, real-time audio and video processing services, general file backup services, or log archiving, etc. For each type of service, the system defines its corresponding "real-time level", for example, represented by three levels of "high, medium, low" or the maximum allowable delay threshold in milliseconds / second. During the initial system deployment or operation and maintenance process, the corresponding basic synchronization frequency threshold is set according to the actual operation requirements of each type of service. For example, the synchronization of high-real-time financial transactions is set to 60 times per minute, that is, once per second, while the file archiving service with low real-time requirements can be set to 5 times per minute or less. The basic synchronization frequency threshold will be used as a reference starting point for subsequent dynamic synchronization frequency calculation. The system adjusts the upper and lower limits of this frequency based on the real-time collected network status, such as average delay, to achieve a precise and differentiated control logic, ensuring that data synchronization in different service scenarios not only meets performance requirements but also does not cause waste of system resources.
[0058] Linking the basic synchronization frequency with the real-time requirements of the business will help refine the synchronization strategy and optimize resource allocation. Compared with a one-size-fits-all unified frequency strategy, this method can better meet the different synchronization requirements of various business scenarios and improve the overall response efficiency and operational flexibility of the system. At the same time, it can also prevent bandwidth waste and processing redundancy caused by overly high frequency settings in low real-time scenarios, thereby achieving a more energy-efficient system operation mode.
[0059] like Figure 2 As shown, a data synchronization device based on distributed storage is also provided, which is used to implement the steps of the data synchronization method based on distributed storage, and the device includes:
[0060] The network status collection module is used to collect the network status data of each node and analyze the network delay and bandwidth fluctuation parameters;
[0061] A frequency calculation module connected to the network status acquisition module, used to calculate an adaptive dynamic synchronization frequency threshold based on the network delay and bandwidth fluctuation parameters;
[0062] A synchronization plan generation module connected to the frequency calculation module, used to adjust the data synchronization plan between nodes according to the dynamic synchronization frequency threshold and generate a synchronization task list;
[0063] The task execution module connected to the synchronization plan generation module is used to execute the tasks in the synchronization task list to complete the real-time synchronization of data in the distributed storage system.
[0064] Based on the modular design concept, this embodiment constructs a device structure with the ability of full-process data synchronization. The device is divided into four core functional modules, which are connected in sequence and work together to form a closed-loop distributed data synchronization mechanism. First, the network status acquisition module continuously monitors the network connection status of all nodes in the distributed system during the operation of the system. By periodically interacting with each node, it collects the network delay value and the current bandwidth situation, and through local preprocessing and anomaly elimination, generates data packets for calculation, including the mean network delay, fluctuation parameters, and bandwidth change rate. These parameters are important inputs for subsequent synchronization rhythm evaluation. The collected data is received and processed by the frequency calculation module. This module uses a preset or trained model to convert the current network status parameters into an appropriate synchronization frequency recommendation, and calculates the "dynamic synchronization frequency threshold", that is, the maximum upper limit of the reasonable number of synchronization tasks initiated per minute under the current network conditions. The calculation result is updated in real time to provide an input basis for the synchronization plan. Then, the synchronization plan generation module designs the optimal synchronization scheduling scheme based on this frequency threshold. It includes a task analysis sub-module and a sorting and optimization sub-module inside. First, it numbers, determines the target, and differentiates the protocols for all data to be synchronized, identifies potential task switches and resource conflicts; then generates an ordered task list through algorithms such as dynamic orchestration and priority evaluation, and controls the task combination and execution order through constraint conditions to form a complete executable synchronization plan. Finally, the task execution module is responsible for sending the synchronization tasks to each target node, and coordinating the execution process according to scheduling parameters such as priority and resource occupancy. This module supports parallel execution and dynamic scheduling rearrangement mechanisms, and can adaptively adjust tasks in case of network anomalies or node anomalies to ensure the continuity and high availability of the data synchronization process.
[0065] Through modular division and function integration, the device realizes a closed-loop control structure from network status perception to frequency calculation, and then to task scheduling and execution, greatly improving the intelligence and automation of data synchronization in the distributed storage system. Each module can be independently optimized according to the business scenario, improving the scalability and maintainability of the system. Through the dynamic frequency regulation and task optimization mechanism, the device can maintain high-efficiency synchronization performance in complex network environments, significantly reducing resource consumption and delay fluctuations.
[0066] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made in these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A data synchronization method based on distributed storage, characterized in that The method includes: S1. Collect and analyze the network status data of each node to obtain network latency and bandwidth fluctuation parameters; S2. Based on the network latency and bandwidth fluctuation parameters, calculate an appropriate dynamic synchronization frequency threshold, including converting the current network latency and bandwidth fluctuation parameters into a recommended time interval, calculating the current recommended synchronization frequency, which is the dynamic synchronization frequency threshold, and allocating the upper limit of the maximum number of tasks allowed to initiate synchronization per minute according to this rhythm; S3. Adjust the data synchronization plan between nodes according to the dynamic synchronization frequency threshold to generate a synchronization task list, including numbering and initially sorting all data blocks to be synchronized, recording it as a task switch if the synchronization targets of every two tasks are different nodes or different protocols, counting the total switching overhead, minimizing the total switching overhead, and regrouping and sorting the task list; S4. Execute the tasks in the synchronization task list to complete the real-time synchronization of data in the distributed storage system, including assigning a priority weight value to each task, pairing any two tasks, calculating their priority comparison value, and combining multiple tasks based on the priority comparison value for parallel or sequential distribution.
2. The data synchronization method based on distributed storage according to claim 1, characterized in that: The S1 includes setting reference values for network latency and bandwidth respectively. Each node periodically collects the current latency and bandwidth, compares them with the reference values respectively to obtain a ratio, sets a ratio threshold, and marks the node as a high-pressure node when the ratio is greater than the threshold.
3. A data synchronization method based on distributed storage according to claim 1, characterized in that: The specific calculation method of obtaining the ratio by comparing the current latency and bandwidth with the reference values respectively in the S1 includes setting recommended reference values for two key parameters of network performance, network latency and network bandwidth. Each node regularly collects its own current network status, including the current actual network latency value and the current actual network bandwidth value. Divide the currently measured latency value by the set latency reference value to calculate the balance ratio of latency, divide the reference bandwidth value by the currently measured bandwidth value to obtain the balance ratio of bandwidth, and analyze the latency ratio and bandwidth ratio together. If the average value of the two ratios or any one ratio is higher than the preset threshold, mark the node as a high-pressure node.
4. A data synchronization method based on distributed storage according to claim 1, characterized in that: The specific method of calculating the current recommended synchronization frequency in the S2 includes comprehensively analyzing based on the current network latency and bandwidth fluctuation conditions to obtain a recommended time interval, dividing sixty by the above-mentioned recommended time interval to calculate the number of synchronizations that should be performed per minute currently. The calculated number of synchronizations per minute is the dynamic synchronization frequency threshold, and the number of data synchronization tasks allowed to be started at most within each minute is restricted according to this frequency threshold.
5. A data synchronization method based on distributed storage according to claim 1, characterized in that: The specific method of counting the total switching overhead in the S3 includes numbering all data synchronization tasks to be executed and sorting them according to the initial strategy. Then, for each pair of adjacent tasks, judge whether they belong to different synchronization target nodes, protocol types or network paths. If there are such differences, regard the relationship between this pair of tasks as a task switch, determine the amount of resources consumed for each task switch, including latency time or bandwidth occupancy, and multiply the total number of task switches identified in the previous step by the fixed resource overhead of each switch to obtain the total resource loss caused by switching in the current task arrangement.
6. A data synchronization method based on distributed storage according to claim 1, characterized in that: In S4, any two tasks are paired, and the specific method for calculating their priority comparison value includes setting a priority level value representing its importance or urgency for each data synchronization task to be executed. Select any two tasks from the task list for pairing, and by comparing the priority level values corresponding to these two tasks, find the numerical difference between them.
7. A data synchronization method based on distributed storage according to claim 1, characterized in that: S2 also includes collecting the latency values generated during network communication among multiple nodes in the distributed storage system. By statistically analyzing the latency data, calculate the average latency time among all nodes, set a basic synchronization frequency threshold, compare the calculated average latency among the current nodes with a pre-set reference latency value, divide the current average latency value by the reference latency value to obtain a scaling factor. Then, multiply this scaling factor by the aforementioned basic synchronization frequency threshold to obtain a new synchronization frequency value. If the calculated dynamic synchronization frequency value is lower than the pre-set minimum allowable frequency threshold, raise the current frequency value to this minimum threshold.
8. A data synchronization method based on distributed storage according to claim 7, characterized in that: The specific method for statistically analyzing the latency data in S2 to calculate the average latency time among all nodes includes recording the time required to send data from one node to another when network communication is performed between each pair of nodes. This time is the latency of this communication. Uniformly collect and store the latency records generated by all communication operations among nodes to form a data set containing multiple latency data. Perform a cleaning operation on the collected latency data, including removing outliers and invalid records. For all the valid latency records after cleaning, sum up each numerical value one by one, count the total number of these summed latency records, and divide the total latency time by the total number of records to obtain the average latency time of node-to-node communication in the current system.
9. A data synchronization method based on distributed storage according to claim 7, characterized in that: The basic synchronization frequency threshold in S2 is set according to the real-time requirements of the service type in the distributed storage system.
10. A data synchronization device based on distributed storage, which is used to implement the steps of the data synchronization method based on distributed storage according to any one of claims 1-9, characterized in that, The device includes: A network status acquisition module, configured to acquire the network status data of each node and analyze to obtain network latency and bandwidth fluctuation parameters; A frequency calculation module connected to the network status acquisition module, configured to calculate an adapted dynamic synchronization frequency threshold based on the network latency and bandwidth fluctuation parameters; A synchronization plan generation module connected to the frequency calculation module, configured to adjust the data synchronization plan among nodes according to the dynamic synchronization frequency threshold and generate a synchronization task list; A task execution module connected to the synchronization plan generation module, configured to execute the tasks in the synchronization task list to complete the real-time synchronization of data in the distributed storage system.