Incremental data processing system
Through the incremental data processing system, the complexity and resource waste of incremental data processing in the prior art are solved, efficient and accurate incremental data processing is achieved, and data resource utilization and processing efficiency are optimized.
Patent Information
- Application Number
- CN202510761962.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-05
AI Technical Summary
The prior art is difficult to efficiently process incremental data during data analysis, especially in the case of real deletion, which requires full processing, resulting in wasted resources and time, and the incremental calculation logic is complex, making it difficult to achieve accuracy and efficiency.
It provides an incremental data processing system, including an incremental data processing subsystem, an incremental data entry subsystem, an incremental data aggregation subsystem and an incremental data association subsystem. By monitoring data source changes, pseudo-deletion processing, unified flag bit configuration, incremental data integration programs and algorithm libraries, it uses Lean-join and idempotent mode for aggregation and association, and combines partition pruning and optimization association operations.
It realizes the accuracy and efficiency of incremental data processing, ensures the integrity and consistency of data processing, optimizes the utilization of resources and time, and improves data processing efficiency.
Smart Images

Figure CN120596493A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and in particular relates to a system for processing incremental data. Background Art
[0002] In data analysis, especially big data analysis, using incremental data analysis is an important means of accelerating data processing. For example, if the full data set contains three years, or approximately 1,000 days of data, and the daily incremental data can be processed daily, the amount of data processed will only account for approximately 0.1% of the full data set, significantly improving processing speed and reducing computing resource consumption.
[0003] However, in the actual data analysis process, there are still some situations that require data analysis to process the entire data every day, resulting in a large waste of resources and time. These situations are as follows:
[0004] (1) The data source and data processing process include true deletions: Whether the data in the table is truly deleted during data source analysis or data analysis, the subsequent processing process cannot know the truly deleted data except for reading the upstream table in full, and has to perform full processing;
[0005] (2) Complex incremental calculation logic: The algorithms for incremental aggregation and incremental connection are complex and completely different from each other. When considering true deletion, the problem becomes even more complicated and difficult to use. Summary of the Invention
[0006] In view of the above-mentioned deficiencies in the prior art, the present invention provides a system for processing incremental data, which solves the problem of difficulty in efficiently processing incremental data.
[0007] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:
[0008] In one aspect, the present invention provides a system for processing incremental data, comprising the following steps:
[0009] The incremental data processing subsystem monitors data changes in the data source, processes and marks truly deleted data as pseudo-deleted, uniformly configures the pseudo-deletion and true-deletion flags in the data source, and provides an incremental data integration program and an incremental data processing library embedded with incremental data processing algorithms.
[0010] The incremental data ingestion subsystem is used to convert incremental data from the data source into the data lake format and then write it into the data lake.
[0011] Incremental data aggregation subsystem, used to perform incremental data aggregation based on the lean-join incremental aggregation mode or the idempotent incremental aggregation result merge mode;
[0012] The incremental data association subsystem is used to perform inner joins, left joins, or anti-joins when both tables have incremental data, and upsert the corresponding incremental join results into the historical aggregation result table by updating or inserting them.
[0013] The beneficial effects of the present invention are as follows: the present invention provides an incremental data processing system, which ensures the accuracy and efficiency of data processing by using the incremental data processing subsystem to monitor data source changes, process true deletion data and unify flags, and provide incremental data integration programs and algorithm libraries; uses the incremental data entry subsystem to convert incremental data into a data lake format and write it to achieve unified storage and management of data; uses the incremental data aggregation subsystem to perform aggregation based on two modes to ensure the accuracy and consistency of data aggregation results; uses the incremental data association subsystem to achieve effective association between incremental data through multiple association methods and corresponding processing logic, and writes the association results into the historical aggregation result table in an update or insert manner to ensure the integrity and accuracy of data association; at the same time, the present invention uses the partition pruning related module to optimize the association operation through partitioning and pre-scanning screening, thereby improving data processing efficiency; the present invention realizes efficient processing of incremental data when data in the data source changes, effectively saving data processing resources and time.
[0014] Furthermore, the incremental data processing subsystem includes:
[0015] The incremental data integration module is used to monitor data changes in the data source, perform pseudo-deletion processing on truly deleted data and mark it, and uniformly configure the pseudo-deletion flag and true-deletion flag in the data source;
[0016] The incremental data processing module is used to provide an incremental data integration program and an incremental data processing library embedded with an incremental data processing algorithm.
[0017] Furthermore, the incremental data integration module includes:
[0018] The data change monitoring submodule is used to monitor data changes in the data source through the change data capture CDC method and send data change information to the message queue;
[0019] The data pseudo-deletion submodule is used to convert the truly deleted data into pseudo-deleted data and mark the pseudo-deleted data with a unified pseudo-deletion field;
[0020] The flag bit unification submodule is used to convert the truly deleted data into pseudo-deleted data when monitoring the data change information in the message queue, and configure the pseudo-deleted flag bit and the truly deleted flag bit in the data source to the same flag bit.
[0021] Furthermore, the incremental data aggregation subsystem includes:
[0022] A first incremental aggregation module, configured to perform incremental data aggregation based on a Lean-join incremental aggregation mode;
[0023] The second incremental aggregation module is used to perform incremental data aggregation based on an idempotent incremental aggregation result merging mode.
[0024] Furthermore, the method for the first incremental aggregation module to perform incremental data aggregation based on the Lean-join incremental aggregation mode includes the following steps:
[0025] Output the historical data that has been aggregated according to the aggregation key to the historical aggregation result table;
[0026] If there is incremental data in the table to be aggregated, perform a lean join on the incremental data and the historical data in the historical aggregation result table to obtain the first lean join result;
[0027] Aggregate the first Lean-join result again based on the aggregation key to obtain the result to be updated or inserted;
[0028] Upsert the results to be updated or inserted into the historical aggregation result table through update or insert operations.
[0029] Furthermore, the method for the second incremental aggregation module to perform incremental data aggregation based on the idempotent incremental aggregation result merging mode includes the following steps:
[0030] Aggregate the incremental data according to the aggregation key to obtain the incremental data aggregation table;
[0031] If the historical aggregation result table has the same aggregation key as the incremental data aggregation table, the incremental data aggregation table and the historical aggregation result table will be merged by updating or inserting the aggregation key, and the data entries with the same aggregation key will be merged at the same time.
[0032] If the aggregation objects when aggregating the incremental data aggregation table and the historical aggregation result table include numerical accumulation fields, the corresponding numerical accumulation fields are accumulated;
[0033] During the incremental aggregation payload writing phase, a flag field is added to the historical aggregation result table. The flag field indicates the latest write time of the incremental aggregation data.
[0034] Perform a Lean-join on the historical aggregation result table and the incremental data aggregation table to obtain the second Lean-join aggregation result;
[0035] The flag field is used to determine whether the second lean-join result meets the idempotence condition. If not, full aggregation is performed again based on the aggregation key until the aggregation result is reached.
[0036] Add the delete flag to the aggregation key by adding a field.
[0037] Furthermore, the incremental data association subsystem includes:
[0038] The inner join module is used to perform inner joins when both tables have incremental data, and upsert the incremental results of the inner join into the historical aggregation result table by updating or inserting them.
[0039] The calculation expression of the Inner-join incremental result is as follows:
[0040] ,
[0041] in, Indicates the inner association incremental result, represents the first table with incremental data, Indicates the incremental data in the first table, Indicates internal association, Represents the incremental data in the second table, Indicates parallel connection, Represents the second table with incremental data;
[0042] For inner joins, if any data with the delete flag exists in the first table with incremental data or the second table with incremental data, it will be marked as deleted in the historical aggregation result table;
[0043] The left join module is used to perform a left join when both tables have incremental data, and upsert the left join incremental results into the historical aggregation result table by updating or inserting them.
[0044] The calculation expression of the left association result is as follows:
[0045] ,
[0046] in, Represents a left-associative incremental result, Indicates left association;
[0047] For left joins, if any data with the delete flag exists in the first table with incremental data or the second table with incremental data, it is marked as deleted in the historical aggregation result table;
[0048] The anti-association module is used to perform anti-association when both tables have incremental data, and upsert the anti-association incremental results into the historical aggregation result table by updating or inserting them.
[0049] The calculation expression of the anti-correlation increment result is as follows:
[0050] ,
[0051] in, Indicates the anti-correlation incremental result, Indicates anti-association;
[0052] For anti-association, if there is any data with a delete flag in the first table with incremental data, it will be marked as deleted in the historical aggregation result table. If there is any data with a delete flag in the second table with incremental data, as long as there is at least one data in the aggregated data without a delete flag, the flag in the second table with incremental data will not be processed. Otherwise, the anti-association incremental result will be added. .
[0053] Furthermore, the incremental data association subsystem further includes:
[0054] The partition pruning prerequisite module is used to partition the first and second tables according to the join key, or the partition key ensures that the same join key is in the same partition;
[0055] The incremental partition pruning module is used to pre-scan the distribution of incremental data in the first table and the incremental data in the second table before performing inner joins, left joins, or anti-joins. When the join contains existing historical data, it pre-screens the partition data corresponding only to the incremental data in the first table or the incremental data in the second table.
[0056] Other advantages of the present invention will be analyzed in more detail in subsequent embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0058] Figure 1 This is a block diagram of a system for processing incremental data in an embodiment of the present invention. DETAILED DESCRIPTION
[0059] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0060] Change Data Capture (CDC) is a method for capturing and tracking data changes in a database. It plays a key role in data synchronization and integration and can be used to achieve incremental data synchronization.
[0061] like Figure 1 As shown, in one embodiment of the present invention, the present invention provides an incremental data processing system, comprising:
[0062] The incremental data processing subsystem monitors data changes in the data source, processes and marks truly deleted data as pseudo-deleted, uniformly configures the pseudo-deletion and true-deletion flags in the data source, and provides an incremental data integration program and an incremental data processing library embedded with incremental data processing algorithms.
[0063] The incremental data processing subsystem includes:
[0064] The incremental data integration module is used to monitor data changes in the data source, perform pseudo-deletion processing on truly deleted data and mark it, and uniformly configure the pseudo-deletion flag and true-deletion flag in the data source;
[0065] The incremental data integration module includes:
[0066] The data change monitoring submodule is used to monitor data changes in the data source through the change data capture CDC method and send data change information to the message queue;
[0067] The data change monitoring software using the change data capture (CDC) method in this solution can ensure that data change information can still be sent even after the data is actually deleted.
[0068] The data pseudo-deletion submodule is used to convert the truly deleted data into pseudo-deleted data and mark the pseudo-deleted data with a unified pseudo-deletion field;
[0069] In this solution, true deletions are treated as pseudo deletions, so that the truly deleted data can be applied to subsequent data processing processes.
[0070] The flag bit unification submodule is used to convert the truly deleted data into pseudo-deleted data when monitoring the data change information in the message queue, and configure the pseudo-deleted flag bit and the truly deleted flag bit in the data source to the same flag bit.
[0071] In this solution, by configuring the pseudo-deletion flag and the true-deletion flag to be the same flag, unified processing can be achieved during data processing.
[0072] The incremental data processing module is used to provide an incremental data integration program and an incremental data processing library embedded with an incremental data processing algorithm.
[0073] In this embodiment, the incremental data integration program usually uses data change monitoring software to monitor data changes in the data source. This type of software can ensure that data change information can be sent even if the data is actually deleted.
[0074] The incremental data ingestion subsystem is used to convert incremental data from the data source into the data lake format and then write it into the data lake.
[0075] The incremental change data is incremental data of data change transmitted in the data source.
[0076] In this embodiment, microservices and Spark programs are used to monitor message queues and convert data formats according to the data lake format. For example, the CDC change data format is converted into structured / semi-structured data accepted by the data lake, and then written to the data lake. In addition, modifications to key fields used in downstream data lake calculations, including join keys, data record keys, aggregation keys, etc., require special processing: if changes to these keys are detected, the original data record needs to be converted into two: the data for the old key needs to be marked as truly deleted, and the data for the new key needs to be marked as newly added. Before performing these special key conversions, it is necessary to pre-configure a lookup table for use during the conversion.
[0077] Incremental data aggregation subsystem, used to perform incremental data aggregation based on the lean-join incremental aggregation mode or the idempotent incremental aggregation result merge mode;
[0078] The incremental data aggregation subsystem includes:
[0079] A first incremental aggregation module, configured to perform incremental data aggregation based on a Lean-join incremental aggregation mode;
[0080] The method for the first incremental aggregation module to perform incremental data aggregation based on the Lean-join incremental aggregation mode includes the following steps:
[0081] Output the historical data that has been aggregated according to the aggregation key to the historical aggregation result table;
[0082] If there is incremental data in the table to be aggregated, perform a lean join on the incremental data and the historical data in the historical aggregation result table to obtain the first lean join result;
[0083] In this solution, Lean-join means selecting only the required fields in the aggregation result table during join, reducing the amount of data read and write.
[0084] Aggregate the first Lean-join result again based on the aggregation key to obtain the result to be updated or inserted;
[0085] Upsert the results to be updated or inserted into the historical aggregation result table through update or insert operations.
[0086] The second incremental aggregation module is used to perform incremental data aggregation based on an idempotent incremental aggregation result merging mode.
[0087] The method for the second incremental aggregation module to perform incremental data aggregation based on the idempotent incremental aggregation result merging mode includes the following steps:
[0088] Aggregate the incremental data according to the aggregation key to obtain the incremental data aggregation table;
[0089] If the historical aggregation result table has the same aggregation key as the incremental data aggregation table, the incremental data aggregation table and the historical aggregation result table will be merged by updating or inserting the aggregation key, and the data entries with the same aggregation key will be merged at the same time.
[0090] If the aggregation objects when aggregating the incremental data aggregation table and the historical aggregation result table include numerical accumulation fields, the corresponding numerical accumulation fields are accumulated;
[0091] Some current data lake technologies, such as the Payload function, are not idempotent when merging data. Running incremental aggregation calculations multiple times can result in erroneous data. For example, if the program is unexpectedly interrupted during the Payload writing phase, rerunning the incremental aggregation will result in erroneous aggregation results. This solution improves the incremental algorithm to make the data merging operation idempotent:
[0092] During the incremental aggregation payload writing phase, a flag field is added to the historical aggregation result table. The flag field indicates the latest write time of the incremental aggregation data.
[0093] Perform a Lean-join on the historical aggregation result table and the incremental data aggregation table to obtain the second Lean-join aggregation result;
[0094] The flag field is used to determine whether the second lean-join result meets the idempotence condition. If not, full aggregation is performed again based on the aggregation key until the aggregation result is reached. To ensure the idempotence of incremental aggregation results.
[0095] In this embodiment, let the latest write time of the aggregate data corresponding to aggregation key A in the historical aggregation data table be t1. If the incremental data also includes write data from t2 to t3, then when the incremental aggregation result table is used to upsert to the historical aggregation result table through update or insert, the aggregation result of the aggregate data corresponding to aggregation key A in the historical aggregation result table is marked as t3, to indicate that the aggregation data corresponding to aggregation key A includes the aggregation results of data up to t3. In this case, if the earliest write time of the new incremental data is earlier than t3, the merge operation should be rejected during the update or insert, and the full aggregation of aggregation key A should be re-executed. This is because the incremental data may contain duplicate, already aggregated data, which damages idempotence and requires the full aggregation of aggregation key A to be re-executed to correct it.
[0096] Add the delete flag to the aggregation key by adding a field.
[0097] In this solution, a delete flag is added to the aggregation key by adding a field to ensure that downstream computing tasks can incrementally read the deleted data, thereby ensuring that deleted data is handled correctly. The method of adding fields is similar to the need to correctly handle the delete flag when obtaining data from the data source. The delete flag also needs to be correctly handled throughout the entire chain and process of the entire data lake calculation process. For example, during data calculation, if a piece of data in the historical aggregation result table is truly deleted, then in the downstream data calculation, the truly deleted data cannot be detected through incremental queries, which will lead to incorrect downstream data calculation results. Therefore, all true deletions in the calculation process should be processed as pseudo-deletions according to the correct semantics. The difference from data extraction is that during the data processing process of the data lake, changes to the aggregation key and association key do not require special processing. Data changes should be processed as true deletions and additions during the data entry process.
[0098] The incremental data association subsystem is used to perform inner joins, left joins, or anti-joins when both tables have incremental data, and upsert the corresponding incremental join results into the historical aggregation result table by updating or inserting them.
[0099] The incremental data association subsystem includes:
[0100] The inner join module is used to perform inner joins when both tables have incremental data, and upsert the incremental results of the inner join into the historical aggregation result table by updating or inserting them.
[0101] The calculation expression of the Inner-join incremental result is as follows:
[0102] ,
[0103] in, Indicates the inner association incremental result, represents the first table with incremental data, Indicates the incremental data in the first table, Indicates internal association, Represents the incremental data in the second table, Indicates parallel connection, Represents the second table with incremental data;
[0104] For inner joins, if any data with the delete flag exists in the first table with incremental data or the second table with incremental data, it will be marked as deleted in the historical aggregation result table;
[0105] This solution significantly reduces computational complexity through Inner-join incremental joins. For example, using the same computing resources, calculating the full join of 5 billion and 7.5 billion data points takes 10 hours. However, using the Inner-join incremental join provided by this solution, the time can be reduced to 2 hours, a five-fold increase in speed.
[0106] The left join module is used to perform a left join when both tables have incremental data, and upsert the left join incremental results into the historical aggregation result table by updating or inserting them.
[0107] The calculation expression of the left association result is as follows:
[0108] ,
[0109] in, Represents a left-associative incremental result, Indicates left association;
[0110] For left joins, if any data with the delete flag exists in the first table with incremental data or the second table with incremental data, it is marked as deleted in the historical aggregation result table;
[0111] The anti-association module is used to perform anti-association when both tables have incremental data, and upsert the anti-association incremental results into the historical aggregation result table by updating or inserting them.
[0112] The calculation expression of the anti-correlation increment result is as follows:
[0113] ,
[0114] in, Indicates the anti-correlation incremental result, Indicates anti-association;
[0115] For anti-association, if there is any data with a delete flag in the first table with incremental data, it will be marked as deleted in the historical aggregation result table. If there is any data with a delete flag in the second table with incremental data, as long as there is at least one data in the aggregated data without a delete flag, the flag in the second table with incremental data will not be processed. Otherwise, the anti-association incremental result will be added. .
[0116] The incremental data association subsystem further includes:
[0117] The partition pruning prerequisite module is used to partition the first and second tables according to the join key, or the partition key ensures that the same join key is in the same partition;
[0118] In this solution, this goal can be achieved in many cases by properly adjusting the partition key. For example, if the transaction ID needs to be used to join the transaction header table and transaction detail table, the transaction time (year / month / day) can be used as the partition key. This is because the transaction header data and transaction detail data with the same transaction ID are always in the same partition (because the transaction time of the transaction header and transaction detail data with the same transaction ID is always the same). Another partitioning method is to directly use the modulo of the hash value of the transaction ID as the partition key. This approach also ensures that the transaction header data and transaction detail data with the same transaction ID are always in the same partition.
[0119] The incremental partition pruning module is used to pre-scan the distribution of incremental data in the first table and the incremental data in the second table before performing inner joins, left joins, or anti-joins. When the join contains existing historical data, it pre-screens the partition data corresponding only to the incremental data in the first table or the incremental data in the second table.
[0120] In this embodiment, the table containing inventory historical data is generally a large table, for example, a table containing three years of historical data, while the table containing one day of inventory historical data is a small table. The amount of data that meets the standard is 1095 times that of the small table, where 1095=365×3.
[0121] In this plan, implementation When the first table with incremental data meets the requirements, its data volume is M, and the data volume of the incremental data of the first table and the incremental data of the second table is N, and N is much smaller than M. Then the time complexity of the inner join algorithm is Since N is much smaller than M, the impact of N on the algorithm complexity can be ignored. The main amount of calculation is the data shuffling during Join. If the join key of the incremental data of the second table is limited to a limited number of partitions, and the two tables can ensure that the same Join key is in the same partition, the use of partition pruning technology will greatly improve the efficiency of the calculation. For example, the first table or the second table stores two years of data and is partitioned by month. When the corresponding incremental data of the first table or the incremental data of the second table is in 1-2 partitions, and no more than 2, then after partition pruning, the algorithm complexity is reduced to When the daily data increment is large, for example, the daily incremental data exceeds 60MB, we can partition by day, then the algorithm complexity after partition pruning will be further reduced to The larger the data increment, the more obvious the effect of incremental Join partition pruning. If the incremental data exceeds 60MB per hour and is partitioned by hour, the algorithm complexity will be Furthermore, data I / O is significantly reduced to 1 / 12, 1 / 365, and 1 / 8760 of the previous values, respectively. Considering that data I / O is a crucial component of big data analysis time, using the partition pruning method proposed in this solution in Join operations can significantly optimize data processing performance.
[0122] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention.
Claims
1. A system for processing incremental data, characterized in that: include: The incremental data processing subsystem monitors data changes in the data source, processes and marks truly deleted data as pseudo-deleted, uniformly configures the pseudo-deletion and true-deletion flags in the data source, and provides an incremental data integration program and an incremental data processing library embedded with incremental data processing algorithms. The incremental data ingestion subsystem is used to convert incremental data from the data source into the data lake format and then write it into the data lake. Incremental data aggregation subsystem, used to perform incremental data aggregation based on the lean-join incremental aggregation mode or the idempotent incremental aggregation result merge mode; The incremental data association subsystem is used to perform inner joins, left joins, or anti-joins when both tables have incremental data, and upsert the corresponding incremental join results into the historical aggregation result table by updating or inserting them.
2. The incremental data processing system according to claim 1, characterized in that: The incremental data processing subsystem includes: The incremental data integration module is used to monitor data changes in the data source, perform pseudo-deletion processing on truly deleted data and mark it, and uniformly configure the pseudo-deletion flag and true-deletion flag in the data source; The incremental data processing module is used to provide an incremental data integration program and an incremental data processing library embedded with an incremental data processing algorithm.
3. The incremental data processing system according to claim 2, characterized in that: The incremental data integration module includes: The data change monitoring submodule is used to monitor data changes in the data source through the change data capture CDC method and send data change information to the message queue; The data pseudo-deletion submodule is used to convert the truly deleted data into pseudo-deleted data and mark the pseudo-deleted data with a unified pseudo-deletion field; The flag bit unification submodule is used to convert the truly deleted data into pseudo-deleted data when monitoring the data change information in the message queue, and configure the pseudo-deleted flag bit and the truly deleted flag bit in the data source to the same flag bit.
4. The incremental data processing system according to claim 3, characterized in that: The incremental data aggregation subsystem includes: A first incremental aggregation module, configured to perform incremental data aggregation based on a Lean-join incremental aggregation mode; The second incremental aggregation module is used to perform incremental data aggregation based on an idempotent incremental aggregation result merging mode.
5. The incremental data processing system according to claim 4, characterized in that: The method for the first incremental aggregation module to perform incremental data aggregation based on the Lean-join incremental aggregation mode includes the following steps: Output the historical data that has been aggregated according to the aggregation key to the historical aggregation result table; If there is incremental data in the table to be aggregated, perform a lean join on the incremental data and the historical data in the historical aggregation result table to obtain the first lean join result; Aggregate the first Lean-join result again based on the aggregation key to obtain the result to be updated or inserted; Upsert the results to be updated or inserted into the historical aggregation result table through update or insert operations.
6. The incremental data processing system according to claim 4, characterized in that: The method for the second incremental aggregation module to perform incremental data aggregation based on the idempotent incremental aggregation result merging mode includes the following steps: Aggregate the incremental data according to the aggregation key to obtain the incremental data aggregation table; If the historical aggregation result table has the same aggregation key as the incremental data aggregation table, the incremental data aggregation table and the historical aggregation result table will be merged by updating or inserting the aggregation key, and the data entries with the same aggregation key will be merged at the same time. If the aggregation objects when aggregating the incremental data aggregation table and the historical aggregation result table include numerical accumulation fields, the corresponding numerical accumulation fields are accumulated; During the incremental aggregation payload writing phase, a flag field is added to the historical aggregation result table. The flag field indicates the latest write time of the incremental aggregation data. Perform a Lean-join on the historical aggregation result table and the incremental data aggregation table to obtain the second Lean-join aggregation result; The flag field is used to determine whether the second lean-join result meets the idempotence condition. If not, full aggregation is performed again based on the aggregation key until the aggregation result is reached. Add the delete flag to the aggregation key by adding a field.
7. The incremental data processing system according to claim 4, characterized in that: The incremental data association subsystem includes: The inner join module is used to perform inner joins when both tables have incremental data, and upsert the incremental results of the inner join into the historical aggregation result table by updating or inserting them. The calculation expression of the Inner-join incremental result is as follows: , in, Indicates the inner association incremental result, represents the first table with incremental data, Indicates the incremental data in the first table, Indicates internal association, Represents the incremental data in the second table, Indicates parallel connection, Represents the second table with incremental data; For inner joins, if any data with the delete flag exists in the first table with incremental data or the second table with incremental data, it will be marked as deleted in the historical aggregation result table; The left join module is used to perform a left join when both tables have incremental data, and upsert the left join incremental results into the historical aggregation result table by updating or inserting them. The calculation expression of the left association result is as follows: ; in, Represents a left-associative incremental result, Indicates left association; For left joins, if any data with the delete flag exists in the first table with incremental data or the second table with incremental data, it is marked as deleted in the historical aggregation result table; The anti-association module is used to perform anti-association when both tables have incremental data, and upsert the anti-association incremental results into the historical aggregation result table by updating or inserting them. The calculation expression of the anti-correlation increment result is as follows: , in, Indicates the anti-correlation incremental result, Indicates anti-association; For anti-association, if there is any data with a delete flag in the first table with incremental data, it will be marked as deleted in the historical aggregation result table. If there is any data with a delete flag in the second table with incremental data, as long as there is at least one data in the aggregated data without a delete flag, the flag in the second table with incremental data will not be processed. Otherwise, the anti-association incremental result will be added. .
8. The incremental data processing system according to claim 7, characterized in that: The incremental data association subsystem further includes: The partition pruning prerequisite module is used to partition the first and second tables according to the join key, or the partition key ensures that the same join key is in the same partition; The incremental partition pruning module is used to pre-scan the distribution of incremental data in the first table and the incremental data in the second table before performing inner joins, left joins, or anti-joins. When the join contains existing historical data, it pre-screens the partition data corresponding only to the incremental data in the first table or the incremental data in the second table.
Citation Information
Patent Citations
Automatic system and method for data to enter lake
CN116842023A
Database idempotent data increment synchronization method and system of statement-level log
CN117290449A
Geographic space vector data increment updating method based on information feature code
CN119271754A
Data migration method and device and computer equipment
CN119645962A