General incremental calculation method based on intermediate state
By introducing the concept of intermediate state, the logical execution plan is rewritten and matched, and the incremental execution plan is generated, which solves the problem of repeated calculations in the existing technology and achieves efficient incremental calculation effect.
Patent Information
- Application Number
- CN202510741668.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing incremental computing technology does not fully consider state storage when generating execution plans, resulting in repeated calculations, inefficient efficiency, and the advantages of incremental computing are not fully utilized.
By introducing the concept of intermediate state, the logical execution plan is rewritten, the target operator for persistence is selected, and the intermediate state is persisted storage and matching is performed to generate incremental execution plans, and the intermediate state is monitored and updated in real time.
Repeated calculations are avoided, and the efficiency and performance of incremental calculations are significantly improved, calculation time is shortened, resource consumption is reduced, and data processing speed and accuracy is improved.
Smart Images

Figure CN120256469A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and particularly to a general incremental calculation method based on an intermediate state. Background Art
[0002] General incremental calculation: In a big data system, the incremental calculation technology is an efficient method for processing data changes. It only calculates the newly added or modified parts of the data instead of processing the full volume of data each time, which can significantly improve the processing efficiency and reduce resource consumption. For example, in an e-commerce system, a large amount of new order data is generated every day. By using incremental calculation, only the newly generated orders of the day need to be processed, rather than recalculating all the order data in the past. The general incremental calculation technology is for general scenarios and unifies the current stream, batch, and interactive modes through a set of incremental calculation logics. It is different from the incremental calculation of stream computing.
[0003] SQL query: SQL query is an important means for users to interact with a big data system. It allows users to extract the required data from one or more tables according to specific conditions and requirements. By writing SQL query statements, users can specify the data columns to be retrieved, filtering conditions, sorting rules, and the grouping method of the data, etc.
[0004] SQL query optimizer: The optimizer is a crucial component in a big data system, mainly responsible for generating an efficient execution plan for a query, thereby significantly improving the query performance.
[0005] Execution plan: The execution plan describes a series of operation steps and sequences that a data system takes to complete a specific SQL query. It is the optimal execution plan generated by the optimizer based on the query statement, metadata (such as table structure, index information, etc.), and statistical information (such as data distribution, number of rows, etc.). It will be allocated and placed on a real physical machine for execution, and finally the calculation result will be obtained.
[0006] Intermediate state: The intermediate state is a temporary result set in the incremental calculation process. It needs to be persisted on a storage medium. Each incremental task can rely on the intermediate state refreshed in the previous incremental update and continue with the incremental update to accelerate the execution efficiency of some tasks.
[0007] Scan operator: The scan operator is a basic operator in the query execution plan, mainly used to read data from physical storage (such as disks, memory) and provide the original data input for subsequent data processing steps.
[0008] Aggregate Operator: The aggregate operator is a key component in query processing and big data processing frameworks for summarizing and statistical operations on data. Its main function is to perform aggregate calculations on a set of data according to specific rules, converting multiple input rows into one or more output rows.
[0009] Window Operator: The window operator is a key component in the big data framework that aggregates and sorts data by partition and then performs corresponding calculations. Join Operator: The join operator is an important operation in query and data processing frameworks for merging data from two or more data sources.
[0010] Sink Operator: The sink operator is a key component in the data processing flow and plays an important role in both stream processing and batch processing systems. It is mainly responsible for outputting the processed data to an external storage system or other target locations and is the "exit" of the data processing pipeline.
[0011] In the prior art, generating a general incremental calculation execution plan can be divided into the following steps, as Figure 2 shown: 1. Query Parsing The big data system first performs lexical analysis on the input SQL query statement, splitting the statement into individual lexical units (Tokens), such as keywords (e.g., SELECT, FROM, WHERE), identifiers (table names, column names), operators (e.g., +, -, *, / ), and constants (numbers, strings), etc. According to the lexical units obtained from lexical analysis, a syntax tree (SyntaxTree) is constructed according to the grammar rules of the SQL language. The syntax tree is a tree structure that reflects the syntax structure and logical relationship of the query statement. Semantic checking is performed on the syntax tree to ensure that the query statement is semantically correct. This includes checking whether table names and column names exist, whether data types match, and whether permissions are sufficient, etc.
[0012] 2. Generate Logical Execution Plan Convert the verified AST into a logical query plan (Logical Query Plan). The logical query plan is a relational algebra-based representation that describes the logical operation steps of the query without involving specific physical implementation details. Common logical operations include selection, projection, join, aggregation, etc.
[0013] 3. Query Optimization The query optimizer performs a series of equivalent transformations on the logical execution plan, converting it into a form that is easier to optimize and execute. For example, it converts subqueries into join operations or simplifies complex expressions. The system uses two types of optimizers to optimize the execution plan: Rule-based Optimizer (RBO) and Cost-based Optimizer (CBO).
[0014] RBO first rewrites the execution plan. This process is based on a series of optimization rules that can perform equivalent transformations on the logical plan to obtain a better logical plan. Common logical optimizations include Predicate Pushdown, Projection Pruning, and join order adjustment. For example, the Predicate Pushdown rule advances the filtering conditions as early as possible to the data source for execution, reducing the amount of data for subsequent operations.
[0015] After completing the basic optimization operations, CBO begins to convert the optimized logical plan into a Physical Query Plan. The Physical Query Plan describes the specific execution method of the query, including the algorithms used, data access methods, parallel execution strategies, etc. At this stage, the optimizer considers factors such as the characteristics of the data source and system resources and selects appropriate physical operators to implement the logical operations. For example, for a join operation, it selects an appropriate join algorithm (such as nested loop join, hash join, etc.).
[0016] At the same time, during the CBO stage, it also attempts to perform incremental rewriting of the execution plan. For example, it replaces the data source (Scan) with a mode that reads increments and gradually replaces each operator upwards with an incremental calculation form (different operators have different incremental algorithms) until the generation of the incremental execution plan is completed. Figure 3 The following shows an example of the incremental execution plan rewriting. The Scan only needs to be simply replaced with a mode that reads incremental data. Subsequently, the Filter naturally has the ability to consume any streaming data without any rewriting. The final Sink operator needs to be converted from a mode that covers all data to a mode that appends incremental data. In this way, a simple incremental task rewriting is completed.
[0017] 4. Execution Plan The final physical query plan is handed over to the specific execution engine for execution. The execution engine reads data from the data source according to the description of the physical plan, processes it according to the specified operation steps, and returns the query result.
[0018] It can be seen that the generation of an incremental execution plan is completed in the CBO engine. In this complex process, different operators generate corresponding sub-execution plans according to their specific algorithmic logics, aiming to accurately output the incremental data corresponding to the operator. This process seems orderly, but actually hides efficiency hazards. An ordinary incremental computing engine has a significant shortcoming, that is, the key concept of state storage is not introduced. State storage is like a "memory bank" of data, which can record intermediate results and state information during the calculation process, providing a fast way to obtain historical results for subsequent calculations.
[0019] Reference Figure 4 , the figure shows the incremental rewriting algorithms for the Scan operator and the Aggregate operator: Scan only needs to simply change to the mode of reading incremental data; however, Aggregate needs to scan historical data, recalculate the previous results, then read the incremental data and merge it with the historical data, calculate the current full-scale result, and then cancel each other out to obtain the current incremental result. This algorithm is very inefficient, equivalent to calculating a large amount of data twice, not only wasting computing resources but also increasing the computing time. The core goal of incremental computing is to only consume incremental data and obtain the latest computing results at the lowest cost. But the existing algorithms obviously do not achieve this goal, making the advantages of incremental computing unable to be fully exerted.
[0020] In view of this, there is an urgent need for a brand-new computing framework to improve the overall performance of incremental computing. This new framework should fully consider the importance of state storage and avoid unnecessary repeated calculations by reasonably recording and utilizing intermediate states. Summary of the Invention
[0021] The purpose of the present invention is to provide a general incremental computing method based on intermediate states to overcome the deficiencies in the prior art.
[0022] To achieve the above object, the present invention provides the following technical solutions: The present application discloses a general incremental computing method based on intermediate states, which rewrites a logical execution plan to obtain its corresponding incremental execution plan, including the following steps: S1: Obtain the user's SQL and parse it to generate a logical execution plan; S2: Screen and judge the operators, and select the target operators for persistent processing; S3: Perform persistent storage on the target operators for persistent processing; S4: Match the states of the logical execution plan of the intermediate state and the current execution plan through an algorithm, so that all operators find their corresponding equivalent operators; S5: Read the aggregation result of the previous incremental execution plan, merge it with the incremental aggregation result calculated from the current incremental data, and obtain the final execution plan; S6: Monitor, update, and maintain the final execution plan in real time.
[0023] The S2 includes the following sub-steps: S21: Determine whether the operator has a state, where the aggregation operator, window operator, and join operator are operators with a state; S22: Determine whether the state of the operator is itself. If not, rewrite it; S23: If the state required by the current operator is its input node, select the most worthy position to be saved.
[0024] The following is included in the S22: The aggregation operator makes a judgment based on the type of aggregation function: If it is of the SUM or COUNT type, it is itself the state; if it is of the AVG type, rewrite it to the SUM and COUNT types; if it is of the MIN or MAX type, when its input data is append-only, its output result is regarded as the state, otherwise its input data is the state node.
[0025] The following is included in the S22: S221: Determine the type of the input data of the operator itself, including append-only, delete-only, and both; S222: Determine whether the intermediate state has explicitly appeared in the initial logical execution plan. If not, rewrite it to make it a state worthy of being saved.
[0026] The following is included in the S3: For the target operator of the persistence process, after selecting the creation position of the intermediate state, add an additional sink to the current execution plan, persistently store the content of the current node in the storage medium, and record relevant information to trace several created intermediate states.
[0027] The following is included in the S4: Implement the matching of two logical execution plans through an intermediate state matching algorithm. Match the operators of the logical execution plan of the intermediate state and the logical execution plan generated by the query from bottom to top until the root node. After the matching is completed, replace the logical execution plan of the entire view with a scan operation on the view; At this time, a tree consisting of several compensation operators followed by a scan operator is generated, and this tree is equivalent to a certain node in the logical execution plan generated by the query.
[0028] The matching of the two logical execution plans in S4 includes the following: The matching results include: non-matching, complete matching, and partial matching; in the case of partial matching, a compensation operator is calculated and generated, and this compensation operator is continuously pulled up until it pulls through a series of subsequent nodes that need to be matched; if the compensation operator cannot be calculated or pulled up, the view matching fails.
[0029] S6 includes the following: S61: Monitor the update of the intermediate state in real time to ensure data consistency; S62: Automatically recycle the state that fails to match multiple times and release the occupied storage space; S63: Adaptively update the state storage when the SQL changes.
[0030] This application also discloses a general incremental computing device based on the intermediate state, including a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the above-mentioned general incremental computing method based on the intermediate state.
[0031] This application also discloses a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the above-mentioned general incremental computing method based on the intermediate state.
[0032] Advantages of the present invention: (1) By introducing the concept of state storage into the general incremental computing framework in this solution, the complex calculation result set is cached in the persistent storage medium, avoiding repeated calculation of the same historical data, greatly shortening the time and resources required for calculation. When dealing with large-scale data, it can obtain results at a faster speed, providing more timely and accurate support for the decision-making of enterprises and organizations, showing excellent advantages in improving the execution performance of incremental calculation, and is expected to play an important role in the field of big data processing and promote the development and progress of the industry.
[0033] (2) This solution completely proposes the specific maintenance behaviors of each life cycle link of the intermediate state from creation to deletion, enabling it to play a role adaptively in complex systems and a large number of query jobs, significantly reducing the work burden of data developers and data practitioners, enabling them to focus more on the development of business systems rather than optimizing each massive detail of the existing link one by one, greatly liberating the productivity of enterprises and individuals.
[0034] The features and advantages of the present invention will be described in detail through embodiments in conjunction with the accompanying drawings. Description of the Drawings
[0035] Figure 1It is the flowchart of the steps of a general incremental calculation method based on intermediate states in the present invention; Figure 2 It is the schematic diagram of generating a general incremental calculation execution plan in the prior art of the present invention; Figure 3 It is the schematic diagram of rewriting the incremental execution plan in the prior art of the present invention; Figure 4 It is the incremental rewriting algorithm for Scan operator and Aggregate operator in the prior art of the present invention; Figure 5 It is the schematic diagram of rewriting the logical execution plan in the present invention; Figure 6 It is the schematic diagram of matching the logical execution plan of the intermediate state with the current execution plan in the present invention; Figure 7 It is the schematic diagram of the final execution plan in the present invention; Figure 8 It is the schematic diagram of the device in the present invention; Figure 9 It is the schematic diagram of the judgment process of the state operator in the present invention. Specific Embodiments
[0036] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below through the accompanying drawings and embodiments. However, it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessarily confusing the concepts of the present invention.
[0037] Refer to Figure 1 , the embodiment of the present invention provides a general incremental calculation framework based on intermediate states. By caching a complex calculation result set on persistent storage, the execution performance of incremental calculation is greatly improved. The main functions are as follows: With this efficient caching mechanism, many advanced and efficient incremental calculation algorithms can be introduced into this framework, enabling it to use many efficient incremental calculation algorithms to complete query calculations. By overloading the state cached in the previous calculation in each incremental calculation task, the huge overhead caused by repeated calculations can be avoided, thus greatly improving the execution efficiency of incremental calculation.
[0038] At the same time, this framework needs to systematically solve various problems related to intermediate states: First is the problem of selecting intermediate states. In the face of massive data and complex calculation processes, how to accurately select those intermediate states that truly have a key impact on the calculation results is a very challenging task.
[0039] Secondly, how to achieve a complete match for the intermediate state is also a crucial step. During the calculation process, the intermediate state may change dynamically as the data changes and the calculation progresses. It is necessary to ensure that the corresponding intermediate state can be accurately found during each query.
[0040] Furthermore, the update of the intermediate state cannot be ignored. As new data continuously pours in and the calculation continues, the intermediate state needs to be updated in a timely manner to ensure its consistency with the latest data and calculation results.
[0041] Finally, maintaining the validity of the intermediate state is the cornerstone for the stable operation of the entire framework. In practical applications, the intermediate state may be affected by various factors, such as data errors, system failures, etc., which pose a threat to its validity. It is necessary to establish a complete set of monitoring and maintenance mechanisms to monitor the validity of the intermediate state in real time and be able to repair and adjust it in a timely manner once problems are found.
[0042] At the same time, it is also necessary to periodically clean up and optimize the intermediate state, removing useless or expired information to improve the operation efficiency and performance of the entire framework.
[0043] It includes the following content: 1. Determine the location of the intermediate state First of all, it needs to be clear that not all operators require an intermediate state, and having an intermediate state does not necessarily make the calculation execute faster. Therefore, it is necessary to determine which operator results are worth being persistently stored and ensure that these results can be used in future calculations to achieve the purpose of acceleration. The previous text Figure 3 and Figure 4 gave the incremental execution plans for two different operators. Figure 3 The calculation shown mainly passes through the results of the input data, so there is no need to trace back any historical data; however, Figure 4 is an aggregation operation that requires calculating historical data together with the newly added input data to complete.
[0044] Figure 3 The partial example code shown includes: INSERT OVERWRITE TABLE res SELECT* FROM T WHERE date>'0101' Therefore, it is necessary to first make a distinction between operator types: stateful operators and stateless operators. Stateful operators mean that any calculation depends on a historical calculation result and cannot solely rely on incremental data to complete the current calculation; conversely, stateless operators can complete any calculation solely by themselves without having to look back at any historical results. There are mainly three categories of stateful operators: Aggregate operators, Window operators, and Join operators. The intermediate states that need to be generated mainly depend on these three types of operators.
[0045] Secondly, whether an operator can use intermediate states is also strongly correlated with the type of the input data itself (whether the input data is append-only / delete-only or both), and different properties will result in completely different incremental algorithms being applied. For example, for the aggregation function MAX, if there are deletion operations in the input data and the result of the deletion happens to be the value of MAX itself, then it is impossible to rely on incremental data to complete the calculation of the new result anyway.
[0046] At the same time, there is another factor that affects intermediate states, that is, whether the intermediate state has explicitly appeared in the initial logical execution plan. Usually, some operators need to be equivalently rewritten to a certain extent before they can become states worth saving. For example, for a common aggregation function AVG, it needs to be rewritten into the form of SUM and COUNT. Saving AVG directly will not have any effect, so an additional step is required to first rewrite AVG to calculate SUM and COUNT, and then obtain the original AVG result through a scalar calculation of SUM / COUNT.
[0047] Based on the above, the following steps for screening intermediate states can be obtained: Is the current operator a stateful calculation (Aggregate / Window / Join are regarded as stateful calculations)? If "yes": a. Is the state of the current operator itself (for example, for Aggregate, it needs to be judged according to the type of the aggregation function. SUM / COUNT are themselves states; AVG needs to be rewritten; MIN / MAX can regard themselves as states only when the input data is append-only, otherwise its input data is the state node). b. If the state required by the current operator is its input node, then the most worthy position to be saved needs to be found. When specifically judging whether a node is the most valuable position to be saved currently, the following indicators can be referred to: 1. The number of data rows of the current node. The number of data rows is an evaluation metric that can most directly reflect the storage space occupied by this intermediate state. Usually in a big data system, the number of input data rows may increase or decrease depending on the type of calculation. For example, a filtering operation usually results in fewer rows, while a multi-table join operation leads to more data. Storing a node with a relatively small amount of data can bring benefits in state storage and reuse; 2. The computational complexity of the current node. When storing the input node of the current operator, traverse downward along the current logical execution plan. Each time an operator is traversed downward, it means that the operators that have been traversed need to be recalculated during state reuse. For example, given the execution plan A <- B <- C <-..., if the state of node B is directly saved for node A, then when calculating A next time, the state of node B can be directly read to avoid repeated calculations. At this time, C can be regarded as a state that can be saved. However, for A, after reading the state of C each time, the calculation of node B still needs to be executed again. In summary, a balance can be found between the number of data rows and the computational complexity of each node, that is, when the cost of reading the data of lower-level operators plus the overhead of recalculating the higher-level nodes with unsaved states is less than the cost of directly reading the data of higher-level nodes, a more suitable storage location is found; 3. Whether the current node contains columns that can perform pruning operations. During the calculation of incremental tasks, a lot of data does not need to be read again. For example, for aggregation operations, all affected results in a single incremental calculation are selected by the aggregation key (Group Key) of the incremental data, and the rest of the data does not participate in the calculation. Therefore, the intermediate state can be pruned according to the aggregation key. However, if the operator used as the state node does not have these columns, then the full amount of data needs to be scanned forcedly, and the acceleration effect will be greatly reduced.
[0048] Figure 9 The sample code of the content shown includes: With cte AS ( SELECT*FROM( SELECT* FROM A WHERE date ='0101' )LATERAL VIEW EXPLODE(attributes) exploded_attr AS single_attr ) Select * FROM cte LEFT JOIN B on cte.id = B.id WHERE B.COI2 IS NOT NULL; As Figure 9 shown in a specific example, for the Join calculation, both the left and right sides of it are states that need to be saved. For the left side, there are three options: 1. Save Explode; 2. Save Filter; 3. Save Scan. Among these three options, the judgment can be made according to the previous rules. First, when encountering the state operator Explode, since it is an operation that expands one row of data into multiple rows, it will bring a multiple-level data expansion to the state node, which is not a relatively optimal option. Exploring down along Explode, the operator Filter is encountered. It is a filtering operator that filters the input data and only retains partial results. Therefore, what is worth considering is that if the filtering is very good, that is, it can filter the input data to only retain a small part of the results, even if Explode is recalculated once, it will show a positive gain. However, if the filtering is very poor, then continue to explore other operators. Finally, the operator Scan is encountered. It is a read operation, that is, the data to be read has been persisted, so there is no need to save it again. Therefore, Scan will not be regarded as a state node. To sum up, saving the result of Filter may be the best choice to a certain extent.
[0049] c. If neither of the above two conditions is met, then no state is saved. Most common calculation logics will meet the judgment criteria of a or b, that is, the judgment rules cover the vast majority of actual scenarios.
[0050] If "No": The operator itself does not need to create additional state storage. At this time, it is necessary to judge whether the current node needs to be saved according to the subsequent nodes.
[0051] The more detailed discrimination rules are recorded in Table 1.
[0052] Table 1 Once the position for creating intermediate states is selected, an additional Sink will be attached to the current execution plan to persist the content of the current node to the storage medium. The persisted data is also stored as an ordinary table (readable and writable, no different from the table directly created by the user), and the relevant information is recorded in the Meta metadata of the currently refreshed target table to track all the intermediate states created by different tasks.
[0053] Among them, the information to be recorded mainly includes: 1. The name of the table after the current intermediate state is persisted in the form of a table; 2. The logical execution plan of the current intermediate state, which is used for node matching during reuse; 3. The version numbers of the various source tables on which the current intermediate state depends, that is, which data it has consumed. Based on these three types of metadata information, the intermediate state can be completely restored in the subsequent state reuse phase and embedded into the incremental execution plan.
[0054] In addition, for stateful operators (Join / Agg / Window), the states they depend on usually have a strong correlation with several columns of the incremental data (JoinKey / GroupKey / PartitionKey). Additional constraints can be automatically established on these associated columns for the intermediate state to accelerate its reuse performance. Through these features, the process of state loading can be made very fast, and most state files that do not need to be loaded at all will be directly filtered out, further improving the execution performance of queries on the basis of saving repeated calculations. The specific details are recorded in Table 2.
[0055] Table 2 The specific generation process is as Figure 5 shown. First, the SQL passed in by the user is compiled into a logical execution plan. Subsequently, according to the above judgment rules, the AVG in the Aggregate operator is split into COUNT and SUM parts, and an external Sink operator is connected for persistent storage. It should be noted that Join is also a stateful operator, but since the left side of it is a Scan operator that already expresses the semantics of persistence and the right side Aggregate has already been saved, Join does not need to store any additional state.
[0056] Figure 5 The partial example code shown includes: INSERT OVERWRITE TABLE res SELECT*FROM (SELECT *,AVG(c1) as _avg FROM T2) t2 INNER JOIN (SELECT* FROM T1) t1 on t1.join_key = t2.join_key 2. Matching the intermediate state After selecting the intermediate state and completing the persistent storage of the intermediate table, the next issue to consider is how to apply the existing intermediate state to the current query optimization. There are the following two subsequent solutions to complete the reuse of the intermediate state: a. Directly break the original execution plan from the intermediate state, regard it as two upper and lower execution sub-fragments, and directly solidify these sub-fragments in the Meta metadata. When executing next time, take them out and execute them directly in sequence (first execute the part from the data source to the intermediate state, and then execute the part from the intermediate state to the output). Once the intermediate state is created, all execution plans will be executed step by step, and it will be ensured that the intermediate state will definitely be used.
[0057] b. Keep the original execution plan unchanged. Each time the intermediate state is regarded as a node that can be equivalently replaced and is selectively placed in the original execution plan. The optimizer freely chooses whether to use the previously saved intermediate state according to the cost model.
[0058] After repeated evaluation, it is decided to use the second solution, which can decouple the intermediate state from the user's original query. That is, the state is only a tool for accelerating the query that can be dynamically used, and is not an essential part of the original query. Whenever the user decides to change the SQL of the query, these decoupled states can be adaptively upgraded / eliminated / regenerated independently. This gives the system and the user more freedom.
[0059] Thus, a brand-new intermediate state matching algorithm (View-based Optimization, hereinafter referred to as Vbo) is derived. The problem of intermediate state matching is regarded as a sub-problem of view rewriting (that is, how to rewrite the user's arbitrary query Query with a defined View), and it is solved through techniques related to the optimizer. From the perspective of the optimizer, view matching can be abstracted into the matching problem of two logical execution plans. That is, the intermediate state has a logical execution plan P1, and the logical execution plan generated by the query to be accelerated is P2, and both of them are composed of several operators. When these two logical execution plans are input into Vbo, the framework will start matching operators one by one from the bottom up, that is, starting from the bottom Scan root node and matching layer by layer.
[0060] There are several possibilities for each layer of matching: No match: At this time, the materialized view matching fails; Complete match: Then continue to match upward; Partial match: That is, the rows, columns, etc. required by the query are less than those of the view. At this time, a compensatory filter or project (or other operators) needs to be calculated. This compensation will be pulled up step by step through a series of subsequent nodes to be matched. If the compensation cannot be calculated or pulled up, it also means that the view matching fails. If in the end, the view...Figure 1 If the root node is directly matched, the logical execution plan of the entire view can be replaced with a scan operation on the view.
[0061] At this time, a logical execution plan will be formed, consisting of several compensation operators connected in series, with a data reading operation on the view at the root node. That is, after a series of transformations on the previously saved state, it can be equivalently recombined with a certain position in the currently matched execution plan through short calculations. This enables great flexibility in saving intermediate states. Even if the system has iterated for some time, resulting in some differences between the current execution plan and the execution plan used for state saving at the earliest, the old state nodes can still be used in the current execution plan through a series of computational transformations.
[0062] The specific process is as Figure 6 shown. The previously saved intermediate state will put the logical execution plan that generated it and the current logical execution plan into the Vbo for state matching. First, the Scan operator is compared. After successful matching, the remaining operators will be matched upward. Since Sink is a special persistent operator that only stores data in the storage medium without changing any logical semantics, it does not participate in state matching. State matching is completed when Aggregate is matched, and all operators have found their corresponding equivalent operators. So the subsequent steps are as Figure 7 shown. A Scan(state) operator will be directly placed in the incremental logical execution plan, which can directly read the previous aggregation result and merge it with the incremental aggregation result calculated from the current incremental data to obtain the final incremental calculation expression.
[0063] 3. Update and maintain intermediate states After state matching is completed, the intermediate state will be placed in the execution plan in the form of Scan(state), and a corresponding Sink(state) will be additionally generated to be responsible for writing out the update situation of the state in this refresh task. Any state will have a Sink(state) to ensure data consistency.
[0064] In addition, if a state fails to match in multiple different refresh tasks consecutively (for example, due to SQL changes / large version code changes, resulting in relatively large changes in the execution plan and no longer being compatible with the old version), it will be automatically recycled by the system to release the occupied storage space. Its corresponding metadata information will also be automatically recycled.
[0065] In particular, when SQL changes, the state storage will also be adaptively updated.
[0066] For example, the original SQL is “... SELECT COUNT(*) FROM SubQuery GROUP BY C1, C2;” The state at this time is equivalent to the SQL: “... SELECT COUNT(*) FROM SubQuery GROUP BY C1, C2;” The changed SQL is “... SELECT COUNT(*) FROM SubQuery GROUP BY C1;” The new state is equivalent to the SQL: “... SELECT COUNT(*) FROM SubQuery GROUP BY C1;” It can be easily adaptively updated from the first state to the second state, that is new_state = SELECT SUM(COUNT(*)) FROM old_state GROUP BY C1; The upgrade of the link can be completed with greatly reduced overhead. It should be noted that this equivalent upgrade will only be executed once. After the new state is successfully created, the old state will be automatically recycled to release the space it occupies. The upgrade process is also completed by the Vbo matching algorithm, that is, an additional aggregation operator needs to be added to the old state for equivalent compensation, and it is considered that the result after re-aggregation based on the View is equivalent to the result of the current operator.
[0067] An embodiment of a general incremental computing device based on an intermediate state according to the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by a processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory and running. From the hardware level, as Figure 8 shown, it is a hardware structure diagram of any device with data processing capabilities where a general incremental computing device based on an intermediate state according to the present invention is located. Except for Figure 8In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities where the device in the embodiment is located may generally include other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here. The specific implementation processes of the functions and roles of each unit in the above device can be found in the implementation processes of the corresponding steps in the above method, which will not be elaborated here.
[0068] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0069] The embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a general incremental computing device based on an intermediate state in the above embodiment.
[0070] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store the data that has been output or will be output.
[0071] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, or improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A general incremental calculation method based on intermediate states, characterized in that: It includes the following steps: S1: Obtain the user's SQL and parse it to generate a logical execution plan; S2: Filter and judge the operators, and select the target operators for persistent processing; S3: Equivalently rewrite the target operators for persistent processing to generate an intermediate state, and persistently store the intermediate state in a storage medium; S4: Match the state of the logical execution plan of the intermediate state with the current execution plan through an algorithm, so that all operators find their corresponding equivalent operators; S5: Read the persisted intermediate state and incremental data, perform combined calculations, and output the final result; S6: Monitor, update, and maintain the final execution plan in real time.
2. The general incremental calculation method based on an intermediate state according to claim 1, wherein: The following sub-steps are included in S2: S21: Judge whether the operator has a state. Among them, aggregation operators, window operators, and join operators are stateful operators; S22: Judge whether the state of the operator is itself. If not, rewrite it; S23: If the state required by the current operator is its input node, select the most worthy position to be saved.
3. The general incremental calculation method based on an intermediate state according to claim 2, wherein: The following content is included in S22: The aggregation operator makes a judgment according to the type of aggregation function: if it is of the SUM or COUNT type, it is itself the state; if it is of the AVG type, it is rewritten as the SUM and COUNT types; if it is of the MIN or MAX type, when its input data is only appendable, its output result is regarded as the state, otherwise its input data is the state node.
4. A general incremental calculation method based on an intermediate state according to claim 2, characterized in that: The following content is included in S22: S221: Judge the type of the input data of the operator itself, including only appendable, only deletable, and both; S222: Judge whether the intermediate state has explicitly appeared in the initial logical execution plan. If not, rewrite it to make it a state worthy of being saved.
5. The general incremental calculation method based on the intermediate state according to claim 1, wherein: The following content is included in S3: For the target operators for persistent processing, after selecting the creation position of the intermediate state, add a sink to the current execution plan additionally, persistently store the content of the current node in the storage medium, and record relevant information to track several created intermediate states.
6. The general incremental calculation method based on the intermediate state according to claim 1, wherein: The following content is included in S4: Implement the matching of two logical execution plans through an intermediate state matching algorithm. Match the operators of the logical execution plan of the intermediate state and the logical execution plan generated by the query from bottom to top until the root node. After the matching is completed, replace the entire logical execution plan of the view with a scan operation on the view; At this time, a tree composed of several compensation operators followed by a scan operator is generated, and this tree is equivalent to a certain node in the logical execution plan generated by the query.
7. The general incremental calculation method based on an intermediate state according to claim 6, characterized in that: The matching of the two logical execution plans in S4 includes the following content: The matching results include: unmatched, completely matched, and partially matched; in the case of partial matching, calculate and generate a compensation operator, and this compensation operator is continuously pulled up until it pulls over a series of subsequent nodes to be matched; if the compensation operator cannot be calculated or pulled up, the view matching fails.
8. A general incremental calculation method based on an intermediate state according to claim 1, characterized in that: The following content is included in S6: S61: Monitor the update situation of the intermediate state in real time to ensure data consistency; S62: Automatically recycle the states that have failed to match multiple times and release the occupied storage space; S63: Adaptive update the status storage when the SQL changes.
9. A general incremental computing device based on an intermediate state, characterized in that: It includes a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement a general incremental calculation method based on the intermediate state according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: A program is stored thereon. When the program is executed by a processor, it implements a general incremental calculation method based on the intermediate state according to any one of claims 1 to 8.
Citation Information
Patent Citations
Database query processing method and cloud computing platform and device
CN116680284A
Scheduling processing method and system for pre-aggregation calculation of time sequence database
CN116821209A
Query acceleration method and device based on materialized view, electronic equipment and medium
CN117688032A
Inquiry processing optimization device in relational data base system
JP2001222452A
Re-costing for on-line optimization of parameterized queries with guarantees
US20180329955A1