Data synchronization method and device, electronic equipment, storage medium and program product
Through batch incremental synchronization and dynamic step adjustment methods, the problem of synchronous deletion of records in big data environments is solved, efficient and synchronous data addition, modification and deletion are achieved, and system performance is optimized.
Patent Information
- Application Number
- CN202510320719.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-08
AI Technical Summary
The existing big data synchronization solution cannot effectively synchronize and delete records, resulting in limited system performance. Especially under the tens of millions of data, when the increase, modify and delete operations are frequent, network congestion and timeout are serious problems.
The batch incremental synchronization method is used to sort the intervals according to the unique identification fields of the source data table and the target data table, and the step size is dynamically adjusted, and the data is added and deleted synchronized. The data creation time and update time are used to determine the increase and change data, and the synchronization performance is optimized through batch processing and dynamic adjustment of the step size.
It realizes efficient synchronous increase, modify and delete data in a big data environment, reduces computing pressure and network load, and improves synchronization efficiency and stability.
Smart Images

Figure CN120277158A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of big data technology, and in particular, to a data synchronization method, apparatus, electronic device, storage medium, and program product. Background Art
[0002] In big data synchronization, a source data table usually contains a large number of fields (for example, more than 100), and the amount of single-line data can reach more than 1KB. As the data volume reaches the tens of millions level, the frequent synchronization requirements for insert, update, and delete operations pose challenges to system performance.
[0003] A common big data synchronization solution is a simple incremental synchronization solution, that is, a timing period is designed, and each time the records in the source data table whose time is greater than the end time of the previous synchronization are written into the target data table; its drawback is that it can only synchronize insert and update records, but cannot synchronize deleted records. Summary of the Invention
[0004] In view of the above problems, the present disclosure is proposed. The present disclosure provides a data synchronization method, apparatus, electronic device, storage medium, and program product.
[0005] According to one aspect of the present disclosure, a data synchronization method is provided. When the source data table is not empty, the method includes:
[0006] Incrementally synchronize the insert and update data in the source data table to the target data table in batches;
[0007] Incrementally synchronize the delete data in the source data table to the target data table in batches according to the divided intervals;
[0008] Wherein, the intervals are divided according to the sorting of the unique identifier fields jointly included in the source data table and the target data table, and during the division process, if the data deletion status of adjacent intervals in the source data table is the same, the step size of the next interval is enlarged, otherwise, the step size of the next interval is reduced.
[0009] In addition, according to the data synchronization method of one aspect of the present disclosure, the incrementally synchronizing the insert and update data in the source data table to the target data table includes:
[0010] Incrementally synchronize the insert and update data in the source data table to the target data table in batches according to a preset single-batch synchronization amount. Wherein, during the process of synchronizing a single batch of the insert and update data to the target data table, if the synchronization performance index is satisfied, the single-batch synchronization amount is increased, otherwise, the single-batch synchronization amount is decreased.
[0011] In addition, according to the data synchronization method of one aspect of the present disclosure, the synchronization performance index includes at least one of the following indexes:
[0012] Whether the data reading of the source data table times out;
[0013] Whether the data writing to the target data table times out;
[0014] Whether the memory usage rate of the server exceeds the memory usage rate threshold;
[0015] Whether the CPU utilization rate of the server exceeds the CPU utilization rate threshold.
[0016] In addition, according to the data synchronization method of one aspect of the present disclosure, the data deletion status includes an undeleted status, a partially deleted status, and a completely deleted status.
[0017] In addition, according to the data synchronization method of one aspect of the present disclosure, dividing the interval includes:
[0018] Dividing the interval step by step according to the step size of the interval and the sorting of the unique identification field;
[0019] Wherein, when dividing the interval step by step, it is judged whether the data deletion statuses of the source data table and the target data table in adjacent intervals are the same. If they are the same, the step size of the next interval is increased; otherwise, the step size of the next interval is decreased.
[0020] In addition, according to the data synchronization method of one aspect of the present disclosure, the fields of the source data table include the data creation time and the data update time. Based on the data creation time and the data update time, the added and modified data of the source data table are determined.
[0021] In addition, according to the data synchronization method of one aspect of the present disclosure, the method includes:
[0022] In the same interval:
[0023] If the data volumes of the source data table and the target data table are the same, it is determined that no data in the source data table has been deleted in this interval;
[0024] If the data volumes of the source data table and the target data table are inconsistent, and the data volume of the source data table is zero, it is determined that all data in the source data table has been deleted in this interval;
[0025] If the data volumes of the source data table and the target data table are inconsistent, and the data volume of the source data table is not zero, it is determined that some data in the source data table has been deleted in this interval, and the unique identification of the deleted data is recursively searched within this interval.
[0026] According to another aspect of the present disclosure, there is provided a data synchronization device, including:
[0027] A first synchronization module, configured to batch incrementally synchronize the added and modified data in the source data table to the target data table when the source data table is not empty;
[0028] A second synchronization module, configured to batch incrementally synchronize the deleted data in the source data table to the target data table in divided intervals when the source data table is not empty;
[0029] Wherein, the intervals are divided according to the sorting of the unique identification fields jointly contained in the source data table and the target data table, and during the division process, if the data deletion status of adjacent intervals in the source data table is the same, the step size of the next interval is enlarged, otherwise, the step size of the next interval is reduced.
[0030] According to another aspect of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of any one of the above methods.
[0031] According to another aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the steps of any one of the above methods are implemented.
[0032] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of any one of the above methods are implemented.
[0033] As will be described in detail below, according to the data synchronization method, apparatus, electronic device, storage medium, and program product of the embodiments of the present disclosure, by batch incrementally synchronizing the added and modified data in the source data table to the target data table, and batch incrementally synchronizing the deleted data in the source data table to the target data table in divided intervals, wherein, during the interval division process, if the data deletion status of adjacent intervals in the source data table is the same, the step size of the next interval is enlarged, otherwise, the step size of the next interval is reduced, dynamic adjustment of the intervals is achieved. The technical solution of the present disclosure uses the method of batch incremental synchronization to synchronize the added, modified, and deleted data, making it more suitable for big data synchronization application scenarios. At the same time, the technical solution of the present disclosure uses the dynamic adjustment of the intervals, so that when synchronizing deleted data, invalid data comparison can be reduced, and the synchronization efficiency of deleted data can be improved.
[0034] It should be understood that both the foregoing general description and the following detailed description are exemplary and are intended to provide further explanation of the claimed technology. Brief Description of the Drawings
[0035] The above and other objects, features, and advantages of the present disclosure will become more apparent by describing the embodiments of the present disclosure in more detail with reference to the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation on the present disclosure. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0036] Figure 1 is a flowchart illustrating a data synchronization method according to an embodiment of the present disclosure.
[0037] Figure 2 is a further flowchart illustrating the method of incrementally synchronizing the added and modified data in the source data table to the target data table in batches in the data synchronization method according to an embodiment of the present disclosure.
[0038] Figure 3 is a further flowchart illustrating the method of incrementally synchronizing the deleted data in the source data table to the target data table in batches in the data synchronization method according to an embodiment of the present disclosure.
[0039] Figure 4 is a block diagram illustrating a data synchronization device according to an embodiment of the present disclosure.
[0040] Figure 5 is a schematic diagram illustrating a computer program product according to an embodiment of the present disclosure.
[0041] Figure 6 is a hardware block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Embodiments
[0042] In order to make the objectives, technical solutions, and advantages of the present disclosure more apparent, exemplary embodiments according to the present disclosure will be described in detail below with reference to the accompanying drawings. It is obvious that the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.
[0043] See Figure 1 , a data synchronization method. When the source data table is not empty, the method includes:
[0044] S101, incrementally synchronizing the added and modified data in the source data table to the target data table in batches.
[0045] In this step, the data to be added or modified can be new data or modified data. The new data or modified data can be separately extracted from the source data table for incremental synchronization to the target data table, or the new data and modified data can be jointly extracted from the source data table and incrementally synchronized to the target data table. The data to be added or modified can be found based on the data creation time create_time and data update time update_time in the table. Generally, indexes are not added to the data creation time create_time and data update time update_time, so when extracting data, a full table scan of the source data table is required to obtain all the data to be added or modified. When the data volume is small, all the new and modified data can be directly extracted at one time and synchronized to the target data table. However, in a big data scenario, both the source data table and the target data table have a magnitude of tens of millions of rows. A full table scan will cause a large computational pressure on the source database, affecting the stability of the database and the customer experience. At the same time, synchronizing all the new and modified data at one time results in a large data volume. For example, if one row of data occupies 1 KByte, 100,000 rows occupy 100 MByte, and the bandwidth between the source database and the target database is only 1 Mbyte, then one synchronization takes 100 * 8 / 1 = 800 seconds, nearly 14 minutes, which may cause problems such as network congestion and timeouts. In the embodiments of the present disclosure, by using the method of batch incremental synchronization, the data to be added or modified in the source data table is gradually synchronized to the target data table in batches, so as to ensure that the data can be updated in a timely manner, while reducing the system load and the computational pressure and bandwidth pressure.
[0046] In an example of step S101, the data to be added or modified in the source data table is batch-incrementally synchronized to the target data table according to a preset single-batch synchronization volume. Among them, during the process of synchronizing the single-batch data to be added or modified to the target data table, if the synchronization performance index is met, the single-batch synchronization volume is increased; otherwise, the single-batch synchronization volume is decreased. In this example, the synchronization performance index is a related index that affects the synchronization performance, which may include at least one of the following indexes: whether the data reading from the source data table times out; whether the data writing to the target data table times out; whether the memory usage rate of the server exceeds the memory usage rate threshold; and whether the CPU utilization rate of the server exceeds the CPU utilization rate threshold. In some cases, the synchronization performance index includes all of the above indexes. Based on the synchronization performance index, the value of the single-batch synchronization volume can be dynamically adjusted, so that when batch-incrementally synchronizing the data to be added or modified, it can be dynamically adapted to the synchronization performance index, and the performance requirements can be met when batch-incrementally synchronizing the data to be added or modified.
[0047] In a specific example, the fields of the source data table include the data creation time and the data update time. Based on the data creation time and the data update time, the added and modified data of the source data table are determined. Specifically, the added and modified data can be determined by using the data creation time, the data update time, and the last synchronization time, without additional data, improving its practicability.
[0048] In a specific example, each row record in the source data table and the target data table has three fields: the primary key id, the data creation time create_time, and the data update time update_time. The id is the unique key, which can be a number or a random string similar to UUID. The data creation time create_time and the data update time update_time can use a time format or timestamp accurate to seconds or milliseconds. The data changes in the source data table can be divided into three categories: addition, modification, and deletion. When adding data, a unique id and the data creation time create_time are generated and written to a new row in the data table. When modifying data, the data update time update_time is updated. When deleting data, the row data is erased. After the source data table and the target data table complete the last synchronization, the maximum data creation time create_time in the target data table is recorded as time_A. After time_A, if new additions, modifications, or deletions occur in the source data table, the newly added data must be in the interval where create_time > time_A, the modified data is in the area where create_time <= time_A and update_time > time_A, and the deleted data is in the interval where create_time <= time_A and update_time <= time_A.
[0049] The data creation time create_time and the data update time update_time can be used to identify newly added and modified data. However, since the data creation time create_time and the data update time update_time generally do not have indexes added, when the amount of data in the source data table is large, if all newly added and modified data is retrieved at once, it will cause a great deal of pressure on the source database. At the same time, too much data may also cause network timeouts. This embodiment proposes a dynamic ramp method to batch process newly added and modified data. At the same time, since common databases all provide merged SQL statements for adding and updating data, the newly added and updated data can be placed in one SQL and imported into the target data table. If there are records with the same primary key in the target data table data, then update them; otherwise, insert new records. For example, the INSERT INTO dst SELECT * FROM src ON DUPLICATE KEY UPDATE statement in MySQL, the MERGE statement in Oracle, the INSERT... ON CONFLICT in PostgreSQL, etc. Therefore, the newly added and modified data in the source data table can be merged into one process for processing.
[0050] In a specific example, refer to Figure 2 , step S101 includes:
[0051] Step S201: Set the initial step size n = 1 and the cumulative step m = 0.
[0052] The step size in this step is the single-batch synchronization volume mentioned above.
[0053] Step S202: Set the source library timeout time timeout_read, the target library timeout time timeout_write, the memory usage threshold mem_max, and the CPU utilization threshold cpu_max.
[0054] The source library timeout time is the timeout time for reading data from the source data table mentioned above. When reading data from the source data table exceeds this timeout time, it represents a read timeout; the target library timeout time is the timeout time for writing data to the target data table. When writing data to the target data table exceeds this timeout time, it represents a write timeout. The memory usage threshold is the memory usage threshold of the server. The CPU utilization threshold is the CPU utilization threshold of the server.
[0055] Step S203: Obtain the maximum value time_begin of the data creation time from the target data table.
[0056] Read all the data creation times create_time in the data from the target data table and take the maximum value of the data creation times create_time.
[0057] Step S204: Obtain the total amount M of the added and modified data from the source data table.
[0058] Reference SQL: select count(id) where create_time >= time_begin or update_time >= time_begin.
[0059] Step S205: Enter the loop while(n>0 && m<M). If so, enter Step S206. If not, end.
[0060] Step S206: Obtain the first n newly added and modified data from the source data table.
[0061] Reference SQL: select id...(the full fields of the source table are omitted here, the same below) from src where create_time >= time_begin limit n offset m.
[0062] Step S207: Determine whether the reading of the source data table times out. If so, enter Step S214. If not, enter Step S208.
[0063] Step S208: Write the data into the target data table.
[0064] Step S209: Determine whether the writing of the data into the target data table times out. If so, enter Step S214. If not, enter Step S210.
[0065] Step S210: Obtain the current CPU utilization rate cpu_cur and memory utilization rate mem_cur of the server;
[0066] Step S211: Determine whether the CPU utilization rate is less than the (CPU utilization rate) threshold and the memory utilization rate is less than the (memory usage rate) threshold. If so, enter Step S212. If not, enter Step S214.
[0067] That is, determine whether cpu_cur < cpu_max && mem_cur < mem_max.
[0068] Step S212: Coordinate progression, m = m + n.
[0069] Step S213: Increase the step size, n = n * 2, and return to the 5th step for loop.
[0070] Step S214: Decrease the step size, n = int(n / 2), and return to the 5th step for loop.
[0071] S102: In batches and incrementally synchronize the deleted data in the source data table to the target data table according to the divided intervals. The intervals are divided based on the sorting of the unique identifier field common to both the source data table and the target data table. Moreover, during the division process, if the data deletion statuses of adjacent intervals in the source data table are the same, then increase the step size of the next interval; otherwise, decrease the step size of the next interval.
[0072] In this step, after completing the synchronization of the added and modified data, the deleted data in the source data table is incrementally synchronized to the target data table in batches according to the divided intervals to ensure that the data can be updated in a timely manner while reducing the computational pressure. The intervals are divided based on the sorting of the unique identifier field common to both the source data table and the target data table (such as ascending order), enabling the data deletion status of the same interval to be judged based on the number of data within the interval, reducing the difficulty of judgment and enhancing the judgment speed. For example, if the number of data records is the same in the same interval of the source data table and the target data table, it can be judged that no data has been deleted.
[0073] When increasing the step size of the next interval or decreasing the step size of the next interval, the amplitude of the increase or decrease can be set according to requirements. For example, the increased step size is twice the current step size, and the decreased step size is 1 / 2 of the current step size and rounded down.
[0074] In one example, assuming that when dividing intervals by the unique identifier field, the divided intervals are interval Q1, interval Q2, interval Q3, ……, interval QN, then interval Q1 and interval Q2 are adjacent intervals, interval Q2 and interval Q3 are adjacent intervals, and so on; if the data deletion status of interval Q1 and interval Q2 in the source data table is the same, then expand interval Q3; similarly, if the data deletion status of interval Q1 and interval Q2 in the source data table is different, then shrink interval Q3. It should be understood that a minimum interval threshold and a maximum interval threshold can be set, so that when expanding or shrinking, if the corresponding threshold is exceeded, the expansion or shrinkage can be stopped. The data deletion status refers to the deletion status of the data in the source data table relative to the data in the target data table, which can generally be divided into an undeleted status, a partially deleted status, and a completely deleted status. For example, if the data in interval Q1 of the source data table is the same as the data in interval Q1 of the target data table (for example, when the number of records in interval Q1 of the source data table is the same as the number of records in interval Q1 of the target data table, it can be considered that the data in interval Q1 of the source data table is the same as the data in interval Q1 of the target data table), it indicates that the data deletion status of the source data table in interval Q1 is the undeleted status; if there is no data in interval Q1 of the source data table and there is data in interval Q1 of the target data table (for example, when the number of records in interval Q1 of the source data table is 0 and the number of records in interval Q1 of the target data table is greater than 0, it can be considered that the data deletion status of the source data table in interval Q1 is the completely deleted status), it indicates that the data deletion status of the source data table in interval Q1 is the completely deleted status; if there is data in interval Q1 of the source data table and the amount of data is less than the data in interval Q1 of the target data table (for example, when the number of records in interval Q1 of the source data table is not 0 and is less than the number of records in interval Q1 of the target data table), it indicates that the data deletion status of the source data table in interval Q1 is the partially deleted status.
[0075] In one example, dividing intervals includes: gradually dividing intervals according to the step size of the intervals and the sorting of the unique identifier field;
[0076] Among them, when gradually dividing intervals, it is judged whether the data deletion status of the source data table and the target data table in adjacent intervals is the same. If it is the same, the step size of the next interval is expanded; otherwise, the step size of the next interval is shrunk.
[0077] It can be known that when gradually dividing intervals according to the step size of the intervals and the sorting of the unique identifier field. The unique identifier field can be sorted, and the unique identifier field can be divided according to the step size of the intervals to obtain the divided intervals. Among them, in order to dynamically adjust the step size of the intervals, therefore, when gradually dividing intervals, according to the data deletion status of adjacent regions, it is determined whether to expand or shrink the step size of the intervals.
[0078] It should be understood that determining whether the data deletion statuses of the source data table and the target data table in adjacent intervals are the same is judged in the case where there are adjacent regions. It can be known that before gradually dividing the first interval and the second interval, there are no two adjacent intervals, so the step size of the interval can be maintained; when dividing the third and subsequent intervals, there are two adjacent intervals, so it can be judged whether the data deletion statuses of the source data table and the target data table in the adjacent intervals are the same. If they are the same, the step size of the next interval is increased; otherwise, the step size of the next interval is decreased. For example, if the data deletion statuses of the first interval and the second interval are the same, then the step size of the third interval is increased. Among them, the way of increasing the step size can be set according to requirements. For example, it can be doubled. Similarly, the way of decreasing the step size can be set according to requirements. For example, it is halved.
[0079] In one example, when batch-incrementally synchronizing the deleted data in the source data table to the target data table according to the divided intervals, first determine the data deletion status of the source data table in the interval, and determine the deleted data in the source data table based on the data deletion status. In the same interval, if the data volumes of the source data table and the target data table are the same, it is determined that no data has been deleted in the source data table in this interval; if the data volumes of the source data table and the target data table are different, and the data volume of the source data table is zero, it is determined that all data has been deleted in the source data table in this interval; if the data volumes of the source data table and the target data table are different, and the data volume of the source data table is not zero, it is determined that some data has been deleted in the source data table in this interval, and recursively search for the unique identifier of the deleted data within this interval. When recursively searching for the unique identifier of the deleted data within this interval, smaller sub-intervals can continue to be divided within this interval until all sub-intervals are either undeleted data or all-deleted data, then the unique identifier of the deleted data in this interval can be determined according to the unique identifiers of the sub-intervals corresponding to the all-deleted data.
[0080] Taking the unique identifier field as id as an example, in a specific example, refer to Figure 3 , step S102, includes:
[0081] Step S301, obtain the minimum value id_min_dst and the maximum value id_max_dst of the id in the target data table.
[0082] Since id is the primary key, the speed is very fast and it will not affect the performance of the target database.
[0083] Step S302: Enter the recursive function sync_delete(id_min, id_max), and use id_min_dst and id_max_dst as parameters.
[0084] Step S303: Obtain the minimum value id_min_src and the maximum value id_max_src of the source data table id within [id_min_dst, id_max_dst].
[0085] Step S304: Delete the records in the target data table within the intervals [id_min, id_min_src) and (id_max_src, id_max].
[0086] Deleting the records in the target data table within the intervals [id_min, id_min_src) and (id_max_src, id_max] means deleting the redundant records at the beginning and end of the target data table.
[0087] Step S305: Count the total number of records cnt_total_src of the source data table id within [id_min, id_max], and count the total number of records cnt_total_dst of the target data table id within [id_min, id_max].
[0088] The specific SQL can refer to select count(id) from src / dst where id >= id_min and id <= id_max. Since the id field has an index, the calculation speed is very fast and it will not cause pressure on the database.
[0089] Step S306: Determine whether cnt_total_src == cnt_total_dst. If so, it means that the records in the source data table and the target data table are consistent within the range of id_min and id_max at this time, and the process ends; otherwise, go to Step S307.
[0090] Step S307: Initialize the progression coefficient n = 0, the data deletion status status = null, and the id coordinates id_first = id_last = id_min.
[0091] Among them, the progression coefficient n is used to calculate the step size of each progression of the coordinate id_last. The data deletion status status has three values: green, red, and null. Green means that the number of records in the source data table and the target data table was the same within a certain id interval during the previous comparison. Red means that all the records in the source data table within a certain id region were deleted during the previous comparison. Null means that some data in the source data table within a certain id region was deleted during the previous comparison. Here, green, red, and null are just values taken for convenience of description. In fact, they can also be replaced with 1, 2, 3, or other values without affecting the results of this algorithm.
[0092] Step S308: Enter the loop while (id_first != null). If id_first != null, enter Step S309; otherwise, the process ends.
[0093] Step S309: Count the total number of records cnt_src in the source data table between id_first and id_last; count the total number of records cnt_dst in the target data table with id between id_first and id_last.
[0094] The specific SQL can refer to select count(id) from src / dst where id >= id_first and id <= id_last. Since the id field has an index, the calculation speed is very fast and will not cause pressure on the database.
[0095] Step S310: Determine whether cnt_src != cnt_dst and cnt_src != 0. If so, it means that some records in the source data table between id_first and id_last have been deleted but not completely. At this time, the coordinates cannot be advanced further, and Step S311 needs to be entered; otherwise, either no data has been deleted in the source data table between id_first and id_last, or all data in the source data table has been deleted. In these two cases, the coordinates can be advanced, and Step S312 is entered.
[0096] Step S311: Set the data deletion status to null, status = null, and use id_first and id_last as parameters of the sync_delete function to recursively find the deleted ids in [id_first, id_last].
[0097] Since only some records have been deleted in the source data table between id_first and id_last, the data deletion status can be set to null, status = null, and use id_first and id_last as parameters of the sync_delete function to recursively find the deleted ids in [id_first, id_last].
[0098] Step S312: Advance the id_first coordinate. The advancement position is the next id after the current id_last in the target data table. The SQL can refer to: select id from dst where id > id_last and id <= id_max limit 1 offset 0.
[0099] Step S313: Increment the id_last coordinate. Since the computer can use left shift to accelerate the 2-fold calculation, the increment position is generally set to the next power of 2 of the current id_last in the target data table. ^ The next n - 1 ids, where n is the increment coefficient set in Step 7. However, it can also be set to increment positions of 3-fold, 4-fold, or any multiple according to requirements, without affecting the processing results of this disclosure. For reference, the SQL statement can be: select id from dst where id>id_last and id<=id_max limit1offset 2^n-1.
[0100] Step S314: Conditional judgment. If cnt_src == cnt_dst, it means that no data has been deleted from the source data table in the range [id_first, id_last], and proceed to Step S315. If cnt_src == 0, it means that all data in the source data table in the range [id_first, id_last] has been deleted, and proceed to Step S318.
[0101] Step S315: Check if status == green. If so, proceed to Step S316; otherwise, proceed to Step S317.
[0102] Step S316: Increment the increment coefficient by 1, i.e., n = n + 1, and return to Step S308 to enter the while loop.
[0103] For the previous comparison interval and the current comparison interval, the number of records in the source data table and the target data table is the same. To speed up the comparison, the step size of id_last needs to be increased. Therefore, increment the increment coefficient by 1, i.e., n = n + 1, and return to Step S308 to enter the while loop.
[0104] Step S317: Set the increment coefficient to 0, and set the data deletion status to green, i.e., flip the status to status == green, and return to Step S308 to enter the while loop.
[0105] For the previous comparison interval, the number of records in the source data table and the target data table is different, while for the current interval, the number of records in the source data table and the target data table is the same. The comparison results in adjacent intervals are different. Generally, it is likely that the user has performed a series of business operations here, increasing the possibility of data changes. At this time, the step size of id_last needs to be shrunk. Therefore, set the increment coefficient to 0, and flip the status to status == green, and return to Step S308 to enter the while loop.
[0106] Step S318: Delete the data in the target data table in the range [id_first, id_last] to be consistent with the source data table.
[0107] Step S319: Determine whether status == red. If yes, go to step 320; otherwise, go to step 321.
[0108] Step 320: Increment the progression coefficient by 1, i.e., n = n + 1, and return to step S308 to enter the while loop.
[0109] For the previous comparison interval and the current comparison interval, all data in the source data table are deleted. To speed up the comparison, the step size of id_last needs to be increased. Therefore, increment the progression coefficient by 1, i.e., n = n + 1, and return to step S308 to enter the while loop.
[0110] Step 321: Set the progression coefficient to 0, set the data deletion status to green red, i.e., flip the status to status = red, and return to step S308 to enter the while loop.
[0111] For the previous comparison interval, not all data in the source data table were deleted, while for the current interval, all data in the source data table are deleted. The comparison results in adjacent intervals are different. At this time, the step size of id_last needs to be shrunk. Therefore, set the progression coefficient to 0, and flip the status to status = red, and return to step S308 to enter the while loop.
[0112] In a specific example, it is described as follows according to the undeleted status, fully deleted status, and partially deleted status:
[0113] 1. Undeleted status of the source data table:
[0114] Referring to Table 1, no data in the source data table is deleted in a certain interval. For the first comparison, the step size is 1, the id interval is [1, 1], the number of records is the same, and the status is green. For the second comparison, the step size is still 1, the id interval is [2, 2], the number of records is the same, and the status is also green. For the third comparison, since the statuses of the previous two times are the same, the step size doubles to 2, the id interval becomes [3, 4], the number of records is the same, and the status is still green. For the fourth comparison, the step size doubles again to 4, the id interval is [5, 8], the number of records is the same, and the status is still green, and so on. Since the degree of change of the step size is a multiple of 2, the comparison speed is very fast. Therefore, in the case where no data in the source data table is deleted, the unchanged area can be quickly determined.
[0115] Table 1
[0116] Primary key of the source data table Primary key of the target data table 1 1 2 2 3 3 4 4 5 5 6 6 7 7 8 8
[0117] 2. Fully deleted status:
[0118] Refer to Table 2. The source data table is completely deleted within a certain range (batch deletion). During the first comparison, the step size is 1, the id range is [1, 1], the record in the source data table is 0, and the status is red. The records in the target data table within this range are deleted. During the second comparison, the step size is still 1, the id range is [2, 2], the record in the source data table is also 0, and the status is also red. The records in the target data table within this range are deleted. During the third comparison, since the statuses of the previous two times are the same, the step size doubles to 2, the id range becomes [3, 4], the record in the source data table is still 0, and the status is still red. The records in the target data table within this range are deleted. During the fourth comparison, the step size doubles again to 4, the id range is [5, 8], the record in the source data table is still 0, and the status is still red. The records in the target data table within this range are deleted, and so on. Since the degree of change of the step size is a multiple of 2, the comparison speed is very fast. Therefore, in the case of batch deletion of data in the source data table, the data in the same area of the target data table can be quickly deleted.
[0119] Table 2
[0120]
[0121]
[0122] 3. Random deletion of the source data table:
[0123] Refer to Table 3. The source data table is randomly deleted within a certain range, such as the processing process under deleting every other row. During the first comparison, the step size is 1, the id range is [1, 1], the number of records is the same, and the status is green. During the second comparison, the step size is still 1, the id range is [2, 2], the record in the source data table is 0, and the status becomes red. The records in the target data table within this range are deleted. During the third comparison, since the statuses of the previous two times are different, the step size remains unchanged at 1, the id range is [3, 3], the number of records is the same, and the status is green. During the fourth comparison, since the statuses of the previous two times are different, the step size remains unchanged at 1, the id range is [4, 4], the record in the source data table is 0, and the status becomes red. The records in the target data table within this range are deleted, and so on. It can be found that there are no redundant comparisons in this algorithm. If the traditional binary method is used, first count the total number of data in the whole table, then divide the id into two parts, and count the number of data in the two ranges. Finally, 15 comparisons are required to complete the data comparison and synchronization. Therefore, in the case of batch deletion of data in the source data table, this algorithm can also handle it efficiently.
[0124] Table 3
[0125] Primary key of the source data table Primary key of the target data table 1 1 2 (Deleted after the second comparison) 3 3 4 (Deleted after the fourth comparison) 5 5 6 (Deleted after the sixth comparison) 7 7 8 (Deleted after the eighth comparison)
[0126] Only a small number of variables are utilized, without the need to synchronize a large amount of data from the source data table and the target data table, and only the primary key id is operated on, making full use of the characteristics of the primary key index, greatly reducing the computing, memory, and network pressure on the source database, target database, and the server where the synchronization program is located. During the process of flipping and progressing, the step size of the coordinates is dynamically adjusted through an algorithm, and the incremental synchronization of the deleted data can be efficiently completed.
[0127] It can be known that steps S101 and S201 are executed when the source data table is not empty. If the source data table is empty, the target data table can be directly cleared, so that both the source data table and the target data table are empty.
[0128] See Figure 4 , the embodiments of the present disclosure also provide a data synchronization device, including:
[0129] The first synchronization module 401 is used to incrementally synchronize the added and modified data in the source data table to the target data table in batches when the source data table is not empty;
[0130] The second synchronization module 402 is used to incrementally synchronize the deleted data in the source data table to the target data table in batches according to the divided intervals when the source data table is not empty;
[0131] Among them, the intervals are divided according to the sorting of the unique identifier fields jointly contained in the source data table and the target data table. And during the division process, if the data deletion status in adjacent intervals in the source data table is the same, the step size of the next interval is enlarged; otherwise, the step size of the next interval is reduced.
[0132] In an example, when the first synchronization module 401 is used to incrementally synchronize the added and modified data of the source data table to the target data table in batches, it is specifically used for:
[0133] Incrementally synchronize the added and modified data in the source data table to the target data table in batches according to the preset single-batch synchronization amount. Among them, during the process of synchronizing the single-batch added and modified data to the target data table, if the synchronization performance index is met, the single-batch synchronization amount is increased; otherwise, the single-batch synchronization amount is decreased.
[0134] In an example, the synchronization performance index includes at least one of the following indexes:
[0135] Whether the data reading of the source data table times out;
[0136] Whether the data writing to the target data table times out;
[0137] Whether the memory usage rate of the server exceeds the memory usage rate threshold;
[0138] Whether the CPU utilization rate of the server exceeds the CPU utilization rate threshold.
[0139] In one example, the data deletion status includes an undeleted status, a partially deleted status, and a completely deleted status.
[0140] In one example, the second synchronization module includes a partitioning module, and the partitioning module is configured to:
[0141] Gradually partition an interval according to the step size of the interval and the sorting of the unique identification fields;
[0142] When gradually partitioning the interval, it is determined whether the data deletion statuses of the source data table and the target data table in adjacent intervals are the same. If they are the same, the step size of the next interval is increased; otherwise, the step size of the next interval is decreased.
[0143] In one example, the fields of the source data table include the data creation time and the data update time, and the device further includes a data determination module, which is configured to determine the added and modified data of the source data table based on the data creation time and the data update time.
[0144] In one example, the device further includes a determination module, which is configured to, in the same interval:
[0145] If the data volumes of the source data table and the target data table are the same, it is determined that the source data table has no deleted data in this interval;
[0146] If the data volumes of the source data table and the target data table are different, and the data volume of the source data table is zero, it is determined that the source data table has been completely deleted in this interval;
[0147] If the data volumes of the source data table and the target data table are different, and the data volume of the source data table is not zero, it is determined that the source data table has been partially deleted in this interval, and the unique identification of the deleted data is recursively searched within this interval.
[0148] An exemplary embodiment of the present disclosure further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program that can be executed by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiments of the present disclosure.
[0149] An exemplary embodiment of the present disclosure further provides a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiments of the present disclosure.
[0150] Reference Figure 5 , an exemplary embodiment of the present disclosure further provides a computer program product 500, including a computer program 501, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiments of the present disclosure.
[0151] Reference Figure 6 Now, the structural block diagram of the electronic device 600 that can be used as the server or client of the present disclosure will be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0152] The electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 602 or the computer program loaded from the storage unit 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for device operation can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0153] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. The input unit 606 can be any type of device that can input information into the electronic device 600. The input unit 606 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 607 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 608 can include but is not limited to magnetic disks, optical disks. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0154] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above. For example, in some embodiments, the methods of the embodiments of the present disclosure can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 600 via the ROM 602 and / or the communication unit 609. In some embodiments, the computing unit 601 can be configured to execute the methods of the embodiments of the present disclosure in any other suitable manner (e.g., by means of firmware).
[0155] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for the purposes of illustration and facilitating understanding, rather than limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0156] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.
[0157] In addition, as used herein, the "or" used in the listing of items starting with "at least one" indicates a disjunctive listing, so that for example, the listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the term "exemplary" does not mean that the described examples are preferred or better than other examples.
[0158] It should also be noted that in the systems and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.
[0159] Various changes, substitutions, and alterations to the technology herein can be made without departing from the technology of the teachings defined by the appended claims. Additionally, the scope of the claims of the present disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Processes, machines, manufactures, compositions of events, means, methods, or acts that currently exist or will later be developed that perform substantially the same function or achieve substantially the same result as the corresponding aspects herein can be utilized. Accordingly, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.
[0160] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0161] The above description has been presented for purposes of illustration and description. Additionally, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub - combinations thereof.
Claims
1. A data synchronization method, characterized in that, When the source data table is not empty, the method includes: Incrementally synchronize the added and modified data in the source data table to the target data table in batches; Incrementally synchronize the deleted data in the source data table to the target data table in batches according to the divided intervals; Wherein, the intervals are divided according to the sorting of the unique identification fields jointly contained in the source data table and the target data table, and during the division process, if the data deletion status of the adjacent intervals in the source data table is the same, the step size of the next interval is enlarged, otherwise, the step size of the next interval is reduced.
2. The method according to claim 1, wherein The incrementally synchronizing the added and modified data in the source data table to the target data table includes: Incrementally synchronize the added and modified data in the source data table to the target data table in batches according to a preset single-batch synchronization amount. Wherein, during the process of synchronizing a single batch of the added and modified data to the target data table, if the synchronization performance index is met, the single-batch synchronization amount is increased, otherwise, the single-batch synchronization amount is reduced.
3. The method according to claim 2, wherein The synchronization performance index includes at least one of the following indexes: Whether the data reading from the source data table times out; Whether the data writing to the target data table times out; Whether the memory usage rate of the server exceeds the memory usage rate threshold; Whether the CPU utilization rate of the server exceeds the CPU utilization rate threshold.
4. The method according to claim 1, characterized in that The data deletion status includes an undeleted status, a partially deleted status, and a fully deleted status.
5. The method according to claim 1, wherein Dividing the intervals includes: Gradually divide the intervals according to the step size of the intervals and the sorting of the unique identification fields; Wherein, when gradually dividing the intervals, it is judged whether the data deletion status of the source data table and the target data table in the adjacent intervals is the same. If it is the same, the step size of the next interval is enlarged, otherwise, the step size of the next interval is reduced.
6. The method according to claim 1, characterized in that, The fields of the source data table include the data creation time and the data update time. Based on the data creation time and the data update time, determine the added and modified data in the source data table; And / or In the same interval: If the data amounts of the source data table and the target data table are the same, it is determined that the source data table has no deleted data in this interval; If the data amounts of the source data table and the target data table are inconsistent, and the data amount of the source data table is zero, it is determined that the source data table has all been deleted in this interval; If the data amounts of the source data table and the target data table are inconsistent, and the data amount of the source data table is not zero, it is determined that the source data table has been partially deleted in this interval, and recursively search for the unique identifier of the deleted data within this interval.
7. A data synchronization device, characterized in that, Includes: A first synchronization module, configured to incrementally synchronize the added and modified data in the source data table to the target data table in batches when the source data table is not empty; A second synchronization module, configured to incrementally synchronize the deleted data in the source data table to the target data table in batches according to the divided intervals when the source data table is not empty; The intervals are divided according to the order of the unique identification fields commonly contained in the source data table and the target data table, and, during the division process, if the data deletion status of adjacent intervals in the source data table is consistent, the step size of the next interval is enlarged, otherwise, the step size of the next interval is reduced.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of any one of claims 1 to 6.
9. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instruction is executed by a processor, the steps of any method described in claims 1 to 6 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by a processor, the steps of any method described in claims 1 to 6 are implemented.