Data synchronization method, electronic device, and storage medium
Patent Information
- Application Number
- CN202610909405.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-22
AI Technical Summary
[0002]当前分布式存储、跨机房灾备、云端文件同步等业务场景的规模持续扩张,待同步源端数据的体量不断增长,传统数据同步技术大多采用整体传输的工作模式,或是简单对数据均分后直接同步,存在诸多难以规避的技术缺陷,无法满足高精度、高效率的数据同步需求
[0053]This invention calculates data transmission coefficients based on the rate fluctuation slope, enabling it to capture latency fragments with drastic link rate fluctuations. It simultaneously combines fragment data defects and link transmission defects with weighted averages to obtain a synchronization anomaly degree, achieving a comprehensive evaluation of both data quality and network transmission quality. Based on the synchronization anomaly degree, it prioritizes fragment synchronization and generates dedicated synchronization command signals, differentiating the synchronization urgency levels of different fragments. This provides a basis for prioritizing higher-risk fragments for scheduling. The synchronization command signal simultaneously carries fragment management information. By comprehensively considering fragment synchronization priorities and real-time link status, it calculates synchronization execution coefficients and arranges synchronization queues. This ensures that high-risk fragments are synchronized first, while also prioritizing stable links to carry synchronization tasks. It avoids high-priority fragments occupying faulty links for extended periods, causing synchronization blockage. This balances synchronization timeliness and link transmission stability, shortening the overall fragment synchronization time and stably supporting continuous synchronization of large volumes of data.
Smart Images

Figure CN122802515A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data synchronization technology, and more specifically, to data synchronization methods, electronic devices, and storage media. Background Technology
[0002] As the scale of business scenarios such as distributed storage, cross-data center disaster recovery, and cloud file synchronization continues to expand, the volume of source data to be synchronized is constantly increasing. Most traditional data synchronization technologies adopt the overall transmission working mode or simply divide the data equally and then synchronize it directly. There are many technical defects that are difficult to avoid, and they cannot meet the needs of high-precision and high-efficiency data synchronization.
[0003] Traditional synchronization mechanisms do not perform standardized fragmentation of source data. If data corruption or network fluctuations occur during the overall data synchronization process, the entire source data needs to be retransmitted, which not only consumes a large amount of network bandwidth resources but also affects the overall synchronization time. Some traditional solutions that support fragmentation can only rely on a single indicator, the checksum, to determine whether a fragment needs to be retransmitted. This detection dimension is too weak and cannot identify hidden data problems such as incomplete fragment bytes, offset values of key fields, and disordered storage formats. This can easily lead to a mismatch between the data after synchronization at the receiving end and the baseline data at the source end, making it difficult to guarantee the accuracy of data synchronization.
[0004] Existing synchronization scheduling logic only arranges the synchronization execution order based on the fragment generation order, without considering the severity of data defects in the fragments themselves to determine synchronization priorities. Fragments with severe data corruption are synchronized together with fragments without anomalies, and high-risk fragments cannot be prioritized. Furthermore, traditional solutions do not collect real-time status data such as transmission link speed and response time, making it impossible to quantify the synchronization risks caused by link transmission jitter, nor do they incorporate link stability to adjust the synchronization queue order. If high-priority fragments are assigned to a poorly performing link with severe fluctuations, retransmission operations will be frequently triggered, further slowing down the overall synchronization progress. Summary of the Invention
[0005] In view of the shortcomings of the existing technology, the purpose of this invention is to provide a data synchronization method, an electronic device, and a storage medium.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] The data synchronization method includes the following steps:
[0008] The source data is divided into data fragments according to the preset standard fragmentation rules;
[0009] Extract the basic data features of the data shards; wherein, the basic data features include data integrity features and data consistency features; analyze the basic data features based on preset verification rules to obtain the data verification coefficients of the data shards;
[0010] Obtain the transmission status data of each data segment, and evaluate the transmission performance based on the transmission status data to obtain the data transmission coefficient of the data segment;
[0011] The synchronization anomaly degree of data shards is determined based on the data verification coefficient and data transmission coefficient, and the target synchronization shard set is selected and constructed based on the synchronization anomaly degree.
[0012] The synchronization priority and synchronization command signal of each fragment in the target synchronization fragment set are determined based on the synchronization anomaly degree.
[0013] The synchronization execution order of each target synchronization segment is determined by combining the segment synchronization priority and the link status of the data transmission link;
[0014] The control data synchronization receiving end performs data synchronization processing according to the synchronization execution order and synchronization command signals.
[0015] Preferably, the source data includes the total data capacity, data type, and data storage format of the data to be synchronized;
[0016] The data integrity features include the byte length of the data fragment and the checksum value;
[0017] The data consistency features include key field values for data sharding and data format identifiers;
[0018] The transmission status data includes the transmission rate value of data fragments and the transmission response time.
[0019] Preferably, the source data is divided into data fragments according to preset standard fragmentation rules, specifically including the following steps:
[0020] The preset standard sharding rule is the standard sharding capacity;
[0021] Data fragments are obtained by dividing the source data into fragments based on the standard fragment capacity.
[0022] Preferably, the data verification coefficients for data shards are obtained by analyzing the basic data features based on preset verification rules, specifically including the following steps:
[0023] If the difference between the byte length and the preset standard fragment length is greater than the preset length threshold, then the first complete anomaly coefficient is obtained based on the byte length difference and the length threshold.
[0024] Extract the checksum value of the data fragment. If the checksum value is inconsistent with the preset standard checksum, obtain the second complete anomaly coefficient based on the checksum difference value and the preset checksum error threshold.
[0025] The data integrity anomaly coefficient is obtained based on the first integrity anomaly coefficient and the second integrity anomaly coefficient, and the corresponding data fragment is recorded as the abnormal data fragment.
[0026] If the deviation between the key field value and the source baseline field value is greater than the preset field threshold, the first matching anomaly coefficient is obtained based on the field deviation value and the field threshold.
[0027] Identify the data format identifier of the data fragment. If the data format identifier does not match the preset standard format, obtain the second matching anomaly coefficient based on the format difference level and the preset format error threshold.
[0028] The data matching anomaly coefficient is obtained based on the first and second matching anomaly coefficients.
[0029] Set complete anomaly weights and consistent anomaly weights;
[0030] The data verification coefficient for data sharding is obtained based on the integrity anomaly weight, data integrity anomaly coefficient, consistency anomaly weight, and data matching anomaly coefficient.
[0031] Preferably, the data transmission coefficients for data fragmentation are obtained by evaluating transmission performance based on transmission status data, specifically including the following steps:
[0032] The rate fluctuation slope is obtained from the transmission status data;
[0033] If the rate fluctuation slope is greater than the preset rate fluctuation threshold, the data transmission coefficient of the data fragment is obtained based on the rate fluctuation slope and the rate fluctuation threshold.
[0034] Preferably, the synchronization anomaly degree of data shards is determined based on the data verification coefficient and the data transmission coefficient, and the target synchronization shard set is selected and constructed based on the synchronization anomaly degree, specifically including the following steps:
[0035] Set data verification weights and transmission status weights;
[0036] The synchronization anomaly degree of the target synchronization fragment is obtained based on the data verification weight, data verification coefficient, transmission status weight, and data transmission coefficient.
[0037] The abnormal data fragments and delayed data fragments are used to form the target synchronization fragment set of the data to be synchronized.
[0038] Preferably, determining the fragment synchronization priority and synchronization command signal of each fragment in the target synchronization fragment set based on the synchronization anomaly degree specifically includes the following steps:
[0039] The target synchronization fragment is generated based on the synchronization anomaly degree;
[0040] Synchronization command signals for the target synchronization fragment are generated based on the fragment synchronization priority.
[0041] Preferably, the synchronization execution order of each target synchronization segment is determined by combining the segment synchronization priority and the link status of the data transmission link, specifically including the following steps:
[0042] Get the first time point at which the target synchronization fragment initiates a synchronization request, and get the second time point at which the data synchronization receiver responds to the synchronization request;
[0043] The link response duration is obtained based on the first and second time nodes;
[0044] Identify the current network transmission protocol and determine the data transmission rate based on the network transmission protocol;
[0045] Determine the data transmission link between the target synchronization fragment and the data synchronization receiver based on the link response time and data transmission rate.
[0046] Set priority weights and link stability weights;
[0047] The fragmentation priority coefficient of the target synchronization fragment is obtained based on the fragmentation synchronization priority and priority weight.
[0048] The link stability coefficient of the target synchronization fragment is obtained based on the stability level and link stability weight of the data transmission link.
[0049] The synchronization execution coefficient of the target synchronous fragment is obtained based on the fragmentation priority coefficient and the link stability coefficient, and the synchronization execution order of the target synchronous fragment is determined based on the synchronization execution coefficient.
[0050] An electronic device, including a processor, wherein the processor implements a method for synchronizing data when executing the program.
[0051] A storage medium comprising a stored computer program, wherein the computer program is executed by an electronic device to perform a data synchronization method.
[0052] Compared with the prior art, the present invention has the following beneficial effects:
[0053] This invention calculates data transmission coefficients based on the rate fluctuation slope, enabling it to capture latency fragments with drastic link rate fluctuations. It simultaneously combines fragment data defects and link transmission defects with weighted averages to obtain a synchronization anomaly degree, achieving a comprehensive evaluation of both data quality and network transmission quality. Based on the synchronization anomaly degree, it prioritizes fragment synchronization and generates dedicated synchronization command signals, differentiating the synchronization urgency levels of different fragments. This provides a basis for prioritizing higher-risk fragments for scheduling. The synchronization command signal simultaneously carries fragment management information. By comprehensively considering fragment synchronization priorities and real-time link status, it calculates synchronization execution coefficients and arranges synchronization queues. This ensures that high-risk fragments are synchronized first, while also prioritizing stable links to carry synchronization tasks. It avoids high-priority fragments occupying faulty links for extended periods, causing synchronization blockage. This balances synchronization timeliness and link transmission stability, shortening the overall fragment synchronization time and stably supporting continuous synchronization of large volumes of data. Attached Figure Description
[0054] Figure 1 A flowchart illustrating a data synchronization method for embodiments of the present invention;
[0055] Figure 2 This is a schematic diagram illustrating the process of obtaining data fragments in the data synchronization method provided in this embodiment of the invention. Detailed Implementation
[0056] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0057] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0058] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.
[0059] Reference Figures 1-2 As shown.
[0060] The embodiments further illustrate the data synchronization method, electronic device, and storage medium proposed in this invention.
[0061] The data synchronization method includes the following steps:
[0062] The source data is divided into data fragments according to the preset standard fragmentation rules;
[0063] Extract the basic data features of the data shards; these basic data features include data integrity features and data consistency features; analyze the basic data features based on preset verification rules to obtain the data verification coefficients of the data shards;
[0064] Obtain the transmission status data of each data segment, and evaluate the transmission performance based on the transmission status data to obtain the data transmission coefficient of the data segment;
[0065] The synchronization anomaly degree of data shards is determined based on the data verification coefficient and data transmission coefficient, and the target synchronization shard set is selected and constructed based on the synchronization anomaly degree.
[0066] The synchronization priority and synchronization command signal of each fragment in the target synchronization fragment set are determined based on the synchronization anomaly degree.
[0067] The synchronization execution order of each target synchronization segment is determined by combining the segment synchronization priority and the link status of the data transmission link;
[0068] The control data synchronization receiving end performs data synchronization processing according to the synchronization execution order and synchronization command signals.
[0069] The source data includes the total data volume, data type, and data storage format of the data to be synchronized;
[0070] Data integrity features include the byte length of data fragments and checksum values;
[0071] Data consistency characteristics include key field values for data sharding and data format identifiers;
[0072] Transmission status data includes the transmission rate of data fragments and the transmission response time.
[0073] The total data capacity records the overall storage size of all original data to be synchronized. Based on this value and the preset standard shard capacity, the total number of data shards that can be divided from the source data is calculated. Total number of shards = Total data capacity - Remainder of standard shard capacity ÷ Standard shard capacity. If the total source data capacity is 1000 megabytes and the preset standard shard capacity is 100 megabytes, then the total number of shards = 1000 - 0 ÷ 100 = 10 data shards. Data type is used to distinguish whether the source data is text, image, or database form data. The data storage format serves as a baseline standard and is synchronously distributed to the verification stage of each data shard as a comparison reference for data format identification, ensuring that the storage format of all shards remains consistent with the original source data.
[0074] The byte length of a data fragment records the actual storage space occupied by the current fragment. The length difference = actual fragment byte length - standard fragment length. The length difference is compared to a preset length threshold. When the length difference is greater than the threshold, it indicates that the current fragment has a completeness defect due to missing or redundant bytes. The first integrity anomaly coefficient = length difference ÷ length threshold. The higher this coefficient, the more severe the deviation of the fragment byte length from the standard. The checksum value is a checksum calculated for all bytes of data in the current fragment. The source end pre-calculates a standard checksum as a comparison benchmark. The fragment checksum value is compared with the standard checksum. When the two values are inconsistent, the checksum difference = fragment checksum value - standard checksum. The second integrity anomaly coefficient = checksum difference ÷ preset checksum error threshold. The higher this coefficient, the greater the probability that the byte data within the fragment has been tampered with or lost. The data integrity anomaly coefficient = first integrity anomaly coefficient + second integrity anomaly coefficient.
[0075] The key field value records the field values that carry core business information within the shard. It saves the benchmark field value corresponding to the original data from the source as a comparison standard. The field deviation value = shard key field value - source benchmark field value. The field deviation value is compared with the preset field threshold. When the field deviation value is greater than the field threshold, it means that there is an offset error in the core business data within the shard. The first matching anomaly coefficient = field deviation value ÷ preset field threshold. The higher the coefficient value, the deeper the deviation of the shard's core business data from the source benchmark.
[0076] The data format identifier is a set of characters written into the header of the data shard to mark the storage format used by the current shard. A preset global standard format is used as the comparison benchmark. After identifying the data format identifier inherent to the shard, the identifier is matched against the standard format. If they cannot match, a corresponding format difference level is assigned based on the severity of the mismatch. The second matching anomaly coefficient = format difference level ÷ preset format error threshold. The higher this coefficient, the higher the risk of data read failure due to shard storage format errors. The data matching anomaly coefficient = first matching anomaly coefficient + second matching anomaly coefficient, quantifying the overall content matching defect degree of the shard.
[0077] Transmission rate values are continuously collected at fixed time intervals. Multiple sets of time-series transmission rate values are substituted into the calculation to obtain the rate fluctuation slope. The rate fluctuation slope can intuitively reflect the amplitude of the fragment transmission rate oscillation within the transmission cycle. A rate fluctuation threshold is preset. When the rate fluctuation slope is greater than the rate fluctuation threshold, it means that the transmission network of that fragment has severe jitter and is prone to packet loss and delay problems. Data transmission coefficient = rate fluctuation slope ÷ rate fluctuation threshold. The higher the value of this coefficient, the more serious the jitter problem of the fragment transmission link. The transmission response time record shows the total time elapsed from the sending of a fragment synchronization request to the receiving end's response. This value is used in the calculation of subsequent link response times. The first time node when the fragment initiates the synchronization request and the second time node when the receiving end replies are collected. Link response time = second time node - first time node. The link response time, combined with the standard data transmission rate corresponding to the current network transmission protocol, jointly determines the stability level of the transmission link between the fragment and the receiving end. The stability level serves as the basis for calculating the link stability coefficient. Link stability coefficient = link stability level ÷ link level baseline value. Priority weights are configured. Fragment priority coefficient = fragment synchronization priority × priority weight. Synchronization execution coefficient = fragment priority + link stability coefficient. The larger the synchronization execution coefficient value, the earlier the synchronization execution order. This achieves constraint control of the transmission link status on fragment synchronization scheduling.
[0078] The source data is divided into data shards according to the preset standard sharding rules, specifically including the following steps:
[0079] The default standard fragmentation rule is the standard fragmentation capacity;
[0080] Data fragments are obtained by dividing the source data into fragments based on the standard fragment capacity.
[0081] The standard fragment capacity is defined as the only standard fragmentation rule. The standard fragment capacity is a fixed storage capacity value that is configured in advance. This value has a unified capacity measurement unit. The total capacity of the source data is converted to a measurement unit that is completely consistent with the standard fragment capacity in advance. The value of the standard fragment capacity can be flexibly set according to the business scenario. It is set to 100 megabytes for small file synchronization business and 1024 megabytes for large-scale database synchronization business. The unified capacity standard can ensure that all data fragments generated by the split have a similar data size, avoiding fragments with huge size differences from affecting the balance of verification and transmission scheduling. At the same time, the unified capacity standard can also make all fragments in the system use the same set of judgment parameters such as length threshold and verification benchmark.
[0082] The number of complete data shards is calculated as: Total source data capacity ÷ Standard shard capacity. The calculation yields two results: an integer value and a remainder. The integer value represents the number of data shards that can be completely divided into standard-capacity shards, while the remainder represents the residual data that is less than the standard shard capacity after multiple rounds of standard-capacity sharding. Source data is truncated sequentially according to the standard shard capacity, with each truncation operation generating a complete standard-capacity data shard. This continues until the storage capacity of the remaining data is less than the standard shard capacity. Finally, the remaining residual data is used to generate a separate tail shard (not reaching the standard shard capacity). All complete standard shards and tail shards are then combined to form a complete set of data shards to be processed.
[0083] Assuming the standard shard capacity in the business scenario is 200 megabytes, and the total capacity of the original data to be synchronized at the source is 950 megabytes, both are measured in megabytes without additional conversion. The number of complete shards = 950 ÷ 200, which gives an integer number of complete shards of 4, corresponding to 4 standard data shards, each with a storage capacity of 200 megabytes. The remaining data capacity = total data capacity at the source - number of complete shards × standard shard capacity. Substituting the values, the remaining data capacity = 950 - 4 × 200, which gives a result of 150 megabytes. This remaining data generates a tail shard with a capacity of 150 megabytes. In the end, this source data partitioning operation generates 5 data shards.
[0084] The data verification coefficients for data shards are obtained by analyzing the basic data characteristics based on preset verification rules, specifically including the following steps:
[0085] If the difference between the byte length and the preset standard fragment length is greater than the preset length threshold, then the first complete anomaly coefficient is obtained based on the byte length difference and the length threshold.
[0086] Extract the checksum value of the data fragment. If the checksum value is inconsistent with the preset standard checksum, obtain the second complete anomaly coefficient based on the checksum difference value and the preset checksum error threshold.
[0087] The data integrity anomaly coefficient is obtained based on the first integrity anomaly coefficient and the second integrity anomaly coefficient, and the corresponding data fragment is recorded as the abnormal data fragment.
[0088] If the deviation between the key field value and the source baseline field value is greater than the preset field threshold, the first matching anomaly coefficient is obtained based on the field deviation value and the field threshold.
[0089] Identify the data format identifier of the data fragment. If the data format identifier does not match the preset standard format, obtain the second matching anomaly coefficient based on the format difference level and the preset format error threshold.
[0090] The data matching anomaly coefficient is obtained based on the first and second matching anomaly coefficients.
[0091] Set complete anomaly weights and consistent anomaly weights;
[0092] The data verification coefficient for data sharding is obtained based on the integrity anomaly weight, data integrity anomaly coefficient, consistency anomaly weight, and data matching anomaly coefficient.
[0093] The system reads the actual byte length of the current data fragment, retrieves the preset standard fragment length stored in the system, and calculates the length difference as: actual byte length of the fragment - preset standard fragment length. It then compares the length difference with a preset length threshold. If the length difference is greater than the preset length threshold, it indicates that the actual storage volume of the fragment deviates from the standard specification. The first integrity anomaly coefficient is calculated as: length difference ÷ preset length threshold. The system extracts the checksum value inherent to the data fragment and compares it with the preset standard checksum stored in the source. If the fragment checksum value and the preset standard checksum are not equal, it indicates that the binary data stored within the fragment has been tampered with or lost during transmission. The checksum difference value is calculated as: fragment checksum value - preset standard checksum. The second integrity anomaly coefficient is calculated as: checksum difference value ÷ preset checksum error threshold. The higher the value of this coefficient, the more severe the alteration or loss of the original data within the fragment. The data integrity anomaly coefficient = first integrity anomaly coefficient + second integrity anomaly coefficient, thus obtaining a total coefficient that can uniformly characterize the overall integrity defect level of the fragment. At the same time, the current data fragment is marked as an abnormal data fragment, and the marking information is retained for subsequent screening of target synchronization fragment sets.
[0094] The preset standard fragment length is 200 megabytes, the current fragment's actual byte length is 240 megabytes, and the preset length threshold is 30 megabytes. Substituting the length difference (240 - 200), the calculated length difference is 40 megabytes. Since 40 megabytes is greater than 30 megabytes, the first integrity anomaly coefficient is 40 ÷ 30 ≈ 1.33. The fragment checksum value is set to 1260, the preset standard checksum is 1200, and the preset checksum error threshold is 40. The checksum difference is 1260 - 1200 = 60. The two checksum values are not equal, satisfying the calculation condition. The second integrity anomaly coefficient is 60 ÷ 40 = 1.5. The data integrity anomaly coefficient is 1.33 + 1.5 = 2.83, and this fragment is synchronously marked as an abnormal data fragment.
[0095] Read the key field values stored internally in the data shard, retrieve the source baseline field values retained in the original source data, and calculate the field deviation value as: shard key field value - source baseline field value. Compare the field deviation value with a preset field threshold. When the field deviation value is greater than the preset field threshold, it indicates that there is a significant deviation between the core business data carried by the shard and the source baseline data. The first matching anomaly coefficient is calculated as: field deviation value ÷ preset field threshold. The higher the coefficient, the greater the deviation of the shard's core business data from the source baseline.
[0096] The second matching anomaly coefficient = format difference level ÷ preset format error threshold. The higher the coefficient value, the higher the risk of data reading caused by fragmented storage format errors. Data matching anomaly coefficient = first matching anomaly coefficient + second matching anomaly coefficient.
[0097] Assuming the key field value for sharding is 560, the source baseline field value is 500, and the preset field threshold is 40, the field deviation value is 560 - 500 = 60. The first matching anomaly coefficient is 60 ÷ 40 = 1.5. Since the current shard data format identifier cannot match the preset standard format, the format difference level is determined to be 3. The preset format error threshold is 2. Substituting this into the formula, the second matching anomaly coefficient is 3 ÷ 2 = 1.5. The total data matching anomaly coefficient is 1.5 + 1.5 = 3.
[0098] The values for integrity anomaly weight and consistency anomaly weight are pre-set. The integrity anomaly weight adjusts the influence of the data integrity anomaly coefficient on the total verification coefficient, while the consistency anomaly weight adjusts the influence of the data matching anomaly coefficient on the total verification coefficient. These two weights can be flexibly adjusted according to the business scenario. Scenarios emphasizing data integrity detection can increase the integrity anomaly weight, while scenarios emphasizing business content consistency detection can increase the consistency anomaly weight. The data verification coefficient = integrity anomaly weight × data integrity anomaly coefficient + consistency anomaly weight × data matching anomaly coefficient. The higher the final calculated data verification coefficient, the more severe the combined defects of data incompleteness / corruption and business content inconsistency in the current data shard, and the higher the urgency of subsequent synchronization for that shard.
[0099] The data transmission coefficients for data fragmentation are obtained by evaluating transmission performance based on transmission status data, specifically including the following steps:
[0100] The rate fluctuation slope is obtained from the transmission status data;
[0101] If the rate fluctuation slope is greater than the preset rate fluctuation threshold, the data transmission coefficient of the data fragment is obtained based on the rate fluctuation slope and the rate fluctuation threshold.
[0102] Transmission status data includes the transmission rate values and response times of data fragments. Multiple sets of transmission rate values are continuously collected during the fragment transmission process at fixed time intervals. All collected transmission rate values correspond one-to-one with their respective collection time points, forming a time-series data sequence that matches time and transmission rate. A numerical coordinate system is constructed with the collection time as the horizontal axis and the transmission rate value collected at the corresponding time as the vertical axis. The least squares method is used to fit all time-series coordinate points, ultimately calculating the slope of this fitted line, which is the rate fluctuation slope. The magnitude of the rate fluctuation slope directly corresponds to the amplitude of transmission rate oscillations. A larger rate fluctuation slope indicates more drastic fluctuations in the transmission rate within the fragment transmission period, resulting in a higher probability of network link packet loss, transmission delay, and connection interruption. A smaller rate fluctuation slope indicates smoother rate changes during fragment transmission and a more stable network link transmission status.
[0103] The system compares the rate fluctuation slope with a preset rate fluctuation threshold. If the conditions are met, a data transmission coefficient is calculated. The system pre-stores a fixed preset rate fluctuation threshold, which serves as a boundary standard for determining whether there is severe jitter in the network link. The rate fluctuation slope calculated in the first step is compared with the preset rate fluctuation threshold. When the rate fluctuation slope is greater than the preset rate fluctuation threshold, it indicates that the transmission link carrying the data fragment is experiencing rate jitter exceeding the acceptable range. This link is considered a delayed link with a transmission fault. The data transmission coefficient = rate fluctuation slope ÷ preset rate fluctuation threshold. The data transmission coefficient calculated using this formula has a clear numerical meaning. The higher the calculated value of the data transmission coefficient, the greater the rate jitter amplitude of the current fragment transmission link, indicating a higher risk of synchronization failure and data retransmission when the fragment is transmitted on this link. The priority of the fragment's synchronization scheduling is correspondingly increased. If the rate fluctuation slope is less than or equal to the preset rate fluctuation threshold, it means that the transmission rate fluctuation of the current fragment is within the normal range allowed by the system, the link transmission status is stable, no data transmission coefficient needs to be generated, and this fragment will not be classified as a delayed data fragment.
[0104] Six sets of timing transmission rate data for a certain data segment are collected within the transmission cycle. Through fitting calculation, the rate fluctuation slope corresponding to the segment is found to be 7.2. The system preset rate fluctuation threshold is 3, which meets the calculation conditions. The data transmission coefficient is 7.2 ÷ 3 = 2.4.
[0105] The synchronization anomaly degree of data shards is determined based on the data verification coefficient and data transmission coefficient. The target synchronization shard set is then selected and constructed based on the synchronization anomaly degree, specifically including the following steps:
[0106] Set data verification weights and transmission status weights;
[0107] The synchronization anomaly degree of the target synchronization fragment is obtained based on the data verification weight, data verification coefficient, transmission status weight, and data transmission coefficient.
[0108] The abnormal data fragments and delayed data fragments are used to form the target synchronization fragment set of the data to be synchronized.
[0109] Data verification weights are used to control the impact of issues such as incomplete or corrupted data in fragments on the overall synchronization risk. Transmission status weights are matched with data transmission coefficients to control the impact of issues such as network link rate jitter and transmission delay on the overall synchronization risk.
[0110] Synchronization anomaly rate = data verification weight × data verification coefficient + transmission status weight × data transmission coefficient. The data verification coefficient of a certain abnormal data fragment is equal to 2.898. This fragment is also a delayed data fragment, and the corresponding data transmission coefficient is equal to 2.4. For financial synchronization business scenarios, the data verification weight is set to 0.7 and the transmission status weight is set to 0.3. The synchronization anomaly rate = 0.7 × 2.898 + 0.3 × 2.4 = 2.7486. This value directly reflects that the overall synchronization failure risk of the fragment is relatively high.
[0111] The process of calculating the anomaly coefficient for data integrity screening marks abnormal data fragments. These fragments contain inherent data defects such as incomplete bytes, scrambled checksums, offset key fields, or incompatible storage formats. Direct synchronization of these fragments would cause the receiving end to receive incorrect data. The calculation process then identifies these fragments as delayed data fragments. The transmission link rates for these fragments fluctuate drastically, making direct transmission highly susceptible to packet loss and interruptions, preventing the synchronization process from being completed in one go. By consolidating all screened abnormal data fragments and delayed data fragments into a complete target synchronization fragment set, all fragments within this set undergo a synchronization priority allocation and sequence arrangement process. This avoids repeatedly initiating synchronization operations on normal fragments, saving significant network bandwidth and server computing resources.
[0112] The synchronization priority and synchronization command signal of each fragment in the target synchronization fragment set are determined based on the synchronization anomaly degree, specifically including the following steps:
[0113] The target synchronization fragment is generated based on the synchronization anomaly degree;
[0114] Synchronization command signals for the target synchronization fragment are generated based on the fragment synchronization priority.
[0115] Synchronization anomaly is a comprehensive quantitative value that combines the data defects of the shard itself with network transmission defects. The higher the synchronization anomaly value, the higher the overall risk of synchronization failure of the shard, and the higher the priority of synchronization processing. Multiple consecutive numerical intervals are pre-divided, and each numerical interval corresponds to a unique shard synchronization priority level. The synchronization anomaly value of each shard in the target synchronization shard set is read, and the value is compared with the boundary values of each interval preset by the system in turn. After matching the corresponding numerical interval, the corresponding shard synchronization priority is assigned to the shard.
[0116] Three priority levels are preset, each with a corresponding numerical range. A synchronization anomaly score greater than or equal to 2.5 corresponds to priority level 3; a score greater than or equal to 1.5 and less than 2.5 corresponds to priority level 2; and a score less than 1.5 corresponds to priority level 1. The target synchronization fragment set contains two fragments to be processed. The first fragment has a synchronization anomaly score of 2.7486. Comparing this value with the interval boundary, since 2.7486 is greater than 2.5, this fragment is assigned a synchronization priority of 3, representing the highest synchronization urgency. The second fragment has a synchronization anomaly score of 1.82. Since this value is between 1.5 and 2.5, this fragment is assigned a synchronization priority of 2, representing the next highest synchronization urgency.
[0117] The synchronization command signal is a digital signal carrying fragment synchronization control information. Internally, it encapsulates the fragment identifier, fragment synchronization priority level, and fragment synchronization operation type of the current fragment. Upon receiving the synchronization command signal, the receiving end can directly read the priority information encapsulated within the signal to identify the synchronization urgency level of the current fragment. The system pre-stores multiple standardized synchronization command signal templates, each template corresponding one-to-one with a fragment synchronization priority level. The templates have reserved spaces for filling in the fragment identifier and synchronization operation type. The system reads the fragment synchronization priority of a single target fragment, matches it to the corresponding standardized command template, and then fills the template's spaces with the fragment's unique identifier and corresponding retransmission repair synchronization operation information, ultimately generating the synchronization command signal for that fragment.
[0118] Based on the fragment synchronization priority and the link status of the data transmission link, the synchronization execution order of each target synchronization fragment is determined, specifically including the following steps:
[0119] Get the first time point at which the target synchronization fragment initiates a synchronization request, and get the second time point at which the data synchronization receiver responds to the synchronization request;
[0120] The link response duration is obtained based on the first and second time nodes;
[0121] Identify the current network transmission protocol and determine the data transmission rate based on the network transmission protocol;
[0122] Determine the data transmission link between the target synchronization fragment and the data synchronization receiver based on the link response time and data transmission rate.
[0123] Set priority weights and link stability weights;
[0124] The fragmentation priority coefficient of the target synchronization fragment is obtained based on the fragmentation synchronization priority and priority weight.
[0125] The link stability coefficient of the target synchronization fragment is obtained based on the stability level and link stability weight of the data transmission link.
[0126] The synchronization execution coefficient of the target synchronous fragment is obtained based on the fragmentation priority coefficient and the link stability coefficient, and the synchronization execution order of the target synchronous fragment is determined based on the synchronization execution coefficient.
[0127] The target synchronization fragment sends a synchronization request to the data synchronization receiver at the first time node, and simultaneously captures the receiver's reply at the second time node corresponding to this synchronization request. The link response time is calculated as: second time node - first time node. A larger link response time indicates a longer round-trip time between the current fragment and the receiver, resulting in poorer basic link response performance. For example, if the first time node is converted to 1000 milliseconds and the second time node is converted to 1080 milliseconds, then the link response time is 1080 - 1000 = 80 milliseconds.
[0128] The system automatically identifies the network transmission protocol currently running on the fragmented transmission channel. Different network transmission protocols have different theoretical data transmission rates. The system has a pre-stored table of correspondence between network transmission protocols and standard data transmission rates. After identifying the protocol type, the system directly retrieves the matching rate value from the table as the data transmission rate of the current channel. All data transmission rates are uniformly measured in megabytes per second.
[0129] The stability level is determined by comprehensively comparing the link response time and data transmission rate. A shorter response time and a higher data transmission rate indicate excellent data transmission performance, resulting in a higher stability level. Conversely, a longer response time and a lower data transmission rate indicate weaker performance, resulting in a lower stability level. The stability level is indicated by Arabic numerals, with higher numbers representing more stable transmission. For example, a link response time of 80 milliseconds and a matching data transmission rate of 50 megabytes per second would be assigned a stability level of 4.
[0130] Priority weights adjust the influence of fragment synchronization urgency on the final synchronization order, while link stability weights adjust the influence of network link transmission quality on the final synchronization order. Both weights are dimensionless decimals and can be freely adjusted according to the business scenario. If the business scenario prioritizes ensuring the rapid synchronization of abnormal fragments, the priority weight should be increased; if the business scenario prioritizes avoiding congested links and improving overall synchronization throughput, the link stability weight should be increased.
[0131] Fragment synchronization priority is the urgency level of fragments determined earlier based on synchronization anomaly. The fragment priority coefficient = fragment synchronization priority × priority weight. A higher fragment synchronization priority or a larger priority weight results in a higher fragment priority coefficient, indicating a stronger impact of the fragment's synchronization urgency on the sorting result. For example, if the fragment synchronization priority is 3 and the system configuration priority weight is 0.6, substituting these values into the formula (fragment priority coefficient = 3 × 0.6), we get a fragment priority coefficient of 1.8.
[0132] Link stability level represents the network quality level of the current transmission channel. Link stability coefficient = Link stability level × Link stability weight. A higher link stability level or a larger link stability weight results in a higher link stability coefficient, indicating a stronger influence of the link's transmission advantages on the ranking results. For example, if the link stability level is 4 and the system-configured link stability weight is 0.4, the link stability coefficient is 4 × 0.4 = 1.6.
[0133] Synchronization execution coefficient = fragment priority coefficient + link stability coefficient. The higher the synchronization execution coefficient, the higher the urgency of the fragment and the more stable the network link. Such fragments are placed at the front of the synchronization queue for priority synchronization operations, while fragments with lower synchronization execution coefficients are deferred to the back of the queue for processing. Synchronization execution coefficient = 1.8 + 1.6 = 3.4.
[0134] Assuming there are two shards in the target synchronization shard set, the synchronization execution coefficient of the first shard is 3.4 and the synchronization execution coefficient of the second shard is 2.1, the synchronization execution order is determined to be that the first shard performs the synchronization operation first, and the second shard performs the synchronization operation later.
[0135] Electronic devices, including processors, and methods for synchronizing data when the processor executes a program.
[0136] The storage medium includes a stored computer program, wherein the computer program is executed by an electronic device to perform a data synchronization method.
[0137] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0138] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data synchronization method, characterized in that, Includes the following steps: The source data is divided into data fragments according to the preset standard fragmentation rules; Extract the basic data features of the data shards; wherein, the basic data features include data integrity features and data consistency features; analyze the basic data features based on preset verification rules to obtain the data verification coefficients of the data shards; Obtain the transmission status data of each data segment, and evaluate the transmission performance based on the transmission status data to obtain the data transmission coefficient of the data segment; The synchronization anomaly degree of data shards is determined based on the data verification coefficient and data transmission coefficient, and the target synchronization shard set is selected and constructed based on the synchronization anomaly degree. The synchronization priority and synchronization command signal of each fragment in the target synchronization fragment set are determined based on the synchronization anomaly degree. The synchronization execution order of each target synchronization segment is determined by combining the segment synchronization priority and the link status of the data transmission link; The control data synchronization receiving end performs data synchronization processing according to the synchronization execution order and synchronization command signals.
2. The data synchronization method according to claim 1, characterized in that, The source data includes the total data capacity, data type, and data storage format of the data to be synchronized; The data integrity features include the byte length of the data fragment and the checksum value; The data consistency features include key field values for data sharding and data format identifiers; The transmission status data includes the transmission rate value of data fragments and the transmission response time.
3. The data synchronization method according to claim 2, characterized in that, The source data is divided into data shards according to the preset standard sharding rules, specifically including the following steps: The preset standard sharding rule is the standard sharding capacity; Data fragments are obtained by dividing the source data into fragments based on the standard fragment capacity.
4. The data synchronization method according to claim 3, characterized in that, The data verification coefficients for data shards are obtained by analyzing the basic data characteristics based on preset verification rules, specifically including the following steps: If the difference between the byte length and the preset standard fragment length is greater than the preset length threshold, then the first complete anomaly coefficient is obtained based on the byte length difference and the length threshold. Extract the checksum value of the data fragment. If the checksum value is inconsistent with the preset standard checksum, obtain the second complete anomaly coefficient based on the checksum difference value and the preset checksum error threshold. The data integrity anomaly coefficient is obtained based on the first integrity anomaly coefficient and the second integrity anomaly coefficient, and the corresponding data fragment is recorded as the abnormal data fragment. If the deviation between the key field value and the source baseline field value is greater than the preset field threshold, the first matching anomaly coefficient is obtained based on the field deviation value and the field threshold. Identify the data format identifier of the data fragment. If the data format identifier does not match the preset standard format, obtain the second matching anomaly coefficient based on the format difference level and the preset format error threshold. The data matching anomaly coefficient is obtained based on the first and second matching anomaly coefficients. Set complete anomaly weights and consistent anomaly weights; The data verification coefficient for data sharding is obtained based on the integrity anomaly weight, data integrity anomaly coefficient, consistency anomaly weight, and data matching anomaly coefficient.
5. The data synchronization method according to claim 4, characterized in that, The data transmission coefficients for data fragmentation are obtained by evaluating transmission performance based on transmission status data, specifically including the following steps: The rate fluctuation slope is obtained from the transmission status data; If the rate fluctuation slope is greater than the preset rate fluctuation threshold, the data transmission coefficient of the data fragment is obtained based on the rate fluctuation slope and the rate fluctuation threshold.
6. The data synchronization method according to claim 5, characterized in that, The synchronization anomaly degree of data shards is determined based on the data verification coefficient and data transmission coefficient. The target synchronization shard set is then selected and constructed based on the synchronization anomaly degree, specifically including the following steps: Set data verification weights and transmission status weights; The synchronization anomaly degree of the target synchronization fragment is obtained based on the data verification weight, data verification coefficient, transmission status weight, and data transmission coefficient. The abnormal data fragments and delayed data fragments are used to form the target synchronization fragment set of the data to be synchronized.
7. The data synchronization method according to claim 6, characterized in that, The synchronization priority and synchronization command signal of each fragment in the target synchronization fragment set are determined based on the synchronization anomaly degree, specifically including the following steps: The target synchronization fragment is generated based on the synchronization anomaly degree; Synchronization command signals for the target synchronization fragment are generated based on the fragment synchronization priority.
8. The data synchronization method according to claim 7, characterized in that, Based on the fragment synchronization priority and the link status of the data transmission link, the synchronization execution order of each target synchronization fragment is determined, specifically including the following steps: Get the first time point at which the target synchronization fragment initiates a synchronization request, and get the second time point at which the data synchronization receiver responds to the synchronization request; The link response duration is obtained based on the first and second time nodes; Identify the current network transmission protocol and determine the data transmission rate based on the network transmission protocol; Determine the data transmission link between the target synchronization fragment and the data synchronization receiver based on the link response time and data transmission rate. Set priority weights and link stability weights; The fragmentation priority coefficient of the target synchronization fragment is obtained based on the fragmentation synchronization priority and priority weight. The link stability coefficient of the target synchronization fragment is obtained based on the stability level and link stability weight of the data transmission link. The synchronization execution coefficient of the target synchronous fragment is obtained based on the fragmentation priority coefficient and the link stability coefficient, and the synchronization execution order of the target synchronous fragment is determined based on the synchronization execution coefficient.
9. An electronic device, including a processor, characterized in that, When the processor executes the program, it implements the data synchronization method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, The storage medium includes a stored computer program, wherein the computer program is executed by an electronic device to perform the data synchronization method of any one of claims 1 to 8.