Method for improving visual modeling performance optimization

By analyzing and optimizing the operation process in the SQL model, moving row operations, column operations and table removal, the performance degradation and resource waste caused by redundancy and operation sequence of the computing process in the prior art is solved, and more efficient computing resource utilization and storage resource management are achieved.

CN120144557APending Publication Date: 2025-06-13FIBERHOME TELECOMMUNICATION TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510218599.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In existing visual modeling technologies, performance degradation and waste of storage resources caused by redundancy in computing process, as well as computational performance degradation caused by the order of computing process operation.

Method used

By analyzing the operation process in the SQL model, moving row operations, column operations and table removal, the model achieves the optimal calculation process. Specific steps include model SQL structure tree analysis, determining whether it meets the simplest form, filtering and aggregation operations to the underlying layer translation, non-aggregation column operations translation, short-circuit attribute detection and no-operation intermediate table removal.

Benefits of technology

It effectively solves the performance degradation and waste of storage resources caused by redundancy in computing process, as well as the performance degradation caused by the operation sequence of computing process, and optimizes the utilization of computing resource and storage resource management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144557A_ABST
    Figure CN120144557A_ABST
Patent Text Reader

Abstract

The invention discloses a method for improving visual modeling performance optimization, and belongs to the field of big data visual modeling. Analyzing a model SQL structure tree; judging whether the simplest form of the SQL structure tree is met or not; sequentially translating to the bottom layer through filtering operation and polymerization operation; performing non-aggregated column operation translation; short circuit attribute detection; removing an intermediate table without operation; by analyzing the operation process in the SQL model, moving row operation, column operation and table removal, the final model achieves the optimal calculation process, and through moving aggregation, filtering, column calculation and other operations in the SQL model, the depth of the whole SQL model is reduced, the calculation sequence is more reasonable, and meanwhile it is completely guaranteed that the calculation result is consistent with the original result; the problems of calculation performance reduction and storage resource waste caused by calculation process redundancy and calculation performance reduction caused by a calculation process operation sequence are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of big data visualization modeling, and particularly relates to a method for optimizing and improving visualization modeling performance. Background Art

[0002] Existing visualization modeling technologies mainly implement the calculation logic of SQL business models through visualization operators and configuration items, thereby reducing the modeling threshold for non-data mining professionals. However, due to the lack of consideration by general business personnel for the utilization of computing performance and storage resources, the SQL models generated by visualization modeling suffer from serious waste of computing resources and storage resources.

[0003] Problems existing in the prior art: The decline in computing performance and waste of storage resources caused by redundant computing processes, and the decline in computing performance caused by the operation sequence of computing processes. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for optimizing and improving visualization modeling performance in view of the deficiencies of the background art, effectively solving the decline in computing performance and waste of storage resources caused by redundant computing processes, and the decline in computing performance caused by the operation sequence of computing processes.

[0005] The present invention adopts the following technical solutions to solve the above technical problems:

[0006] A method for optimizing and improving visualization modeling performance, by analyzing the operation process in the SQL model and moving row operations, column operations, and table removals, so that the final model achieves an optimal computing process, specifically including the following steps;

[0007] Step 1, analyzing the SQL structure tree of the model;

[0008] Step 2, determining whether the simplest form of the SQL structure tree is satisfied;

[0009] Step 3, sequentially translating the filtering operation and the aggregation operation to the bottom layer;

[0010] Step 4, translating non-aggregation column operations;

[0011] Step 5, detecting short-circuit attributes;

[0012] Step 6, removing intermediate tables with no operations.

[0013] As a further preferred solution of the method for optimizing the visualization modeling performance of the present invention, in step 2, the simplest form of SQL is the result summarized based on general SQL optimization experience. By means of filtering operations and aggregation operations, the amount of data for subsequent processing is reduced. Through join operations, hash pairing or loop traversal is performed on two tables, and the operation cost is relatively high. Reducing the amount can significantly optimize the performance. It is the basis for judging whether the final SQL model reaches the optimal state. If the optimal state is not reached, translation transformation and removal operations are performed.

[0014] As a further preferred solution of the method for optimizing the visualization modeling performance of the present invention, in step 3, through filtering operations and aggregation operations, they are translated downward layer by layer in sequence, specifically including the following;

[0015] Translation conditions for row operations:

[0016] In the case of non-order row operations and when the filtered field exists, the filtering operation can unconditionally move the execution order in the data processing flow;

[0017] The aggregation operation can be translated across the filtering of group key values or when performing column operations other than field selection;

[0018] Translation conditions for column operations:

[0019] When all column-related operations meet the calculation consistency, the aggregation column calculation is disassembled equivalently;

[0020] When meeting the Cartesian product operation consistency, the column calculation operation can penetrate the Join operation for translation;

[0021] The column calculation of a general table can be unconditionally translated under the dependence of row operations and table operations.

[0022] As a further preferred solution of the method for optimizing the visualization modeling performance of the present invention, in step 5, short-circuit attribute detection: refers to the optimization of filtering conditions. For example, if there is a condition of count > 1 in the first step and a condition of count > 10 in the second step, then the condition of count > 1 is invalid; and unused fields in the subsequent process are removed in advance.

[0023] As a further preferred solution of the method for optimizing the visualization modeling performance of the present invention, in step 6, removal of invalid operations is as follows:

[0024] If a field is not used in the subsequent process, the field can be removed;

[0025] If the same field is continuously aggregated, one of them can be removed;

[0026] If a table does not contain any aggregation, association, sorting, or column calculation functions, then the table can be removed.

[0027] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects:

[0028] A method for optimizing the performance of visual modeling according to the present invention analyzes the operation process in the SQL model and moves row operations, column operations, and table removal, so that the final model reaches the optimal calculation process; in the SQL model, operations such as aggregation, filtering, and column calculation can be moved (because the positions of the operations are relative, so there is no need to move the operation order of JOIN anymore), reducing the depth of the entire SQL model and making the calculation order more reasonable, while fully ensuring that the calculation result is the same as the original; effectively solving the problems of reduced computing performance and wasted storage resources caused by redundant calculation processes and reduced computing performance caused by the operation order of the calculation process. Description of the Drawings

[0029] Figure 1 is the flowchart of the overall model optimization of the present invention;

[0030] Figure 2 is the process of moving column calculation in B-SQL to A-SQL according to the present invention;

[0031] Figure 3 is the process of merging A-SQL and B-SQL into C-SQL according to the present invention. Detailed Description of the Invention

[0032] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings:

[0033] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. The present invention will be described in detail below according to the drawings and preferred embodiments, and the purpose and effect of the present invention will become more apparent. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0034] SQL operations are mainly divided into row operations and column operations. Row operations include filtering (where condition), aggregation (group, distinct), association (join), and sorting (order). Column operations include field renaming, system functions, and custom functions, etc. In the SQL model, both row operations and column operations can move forward to the previous operation process under certain conditions. The order of calculation will affect the amount of calculated data, thereby optimizing computing resources. If an intermediate table has no operations, then this intermediate table can be removed to achieve the purpose of optimizing storage.

[0035] Translation conditions for row operations:

[0036] In the case of non-order row operations and when the filtered field exists, the filtering operation can unconditionally move the execution order in the data processing flow.

[0037] The aggregation operation can be translated across the filtering of group key values or when performing column operations other than field selection.

[0038] Translation conditions for column operations:

[0039] When all column-related operations meet the calculation consistency, the decomposition of aggregated column calculations is equivalent.

[0040] When meeting the Cartesian product operation consistency, the column calculation operation can be translated through the Join operation.

[0041] The column calculation of a general table can be unconditionally translated under the dependence of row operations and table operations.

[0042] Removal of invalid operations:

[0043] If a field is not used in the subsequent process, then this field can be removed;

[0044] If continuously aggregating the same field, then one of them can be removed;

[0045] If a table does not contain any aggregation, association, sorting, or column calculation functions, then this table can be removed.

[0046] By analyzing the operation process in the SQL model and moving row operations, column operations, and table removal, the final model reaches the optimal calculation process.

[0047] As Figure 1 shown, the flow chart of the overall model optimization, where:

[0048] Simplest form judgment: The simplest form of SQL is summarized based on general SQL optimization experience. Filter operations (filter operations) and aggregation operations (Group) can significantly reduce the amount of data to be processed subsequently. When the calculation result remains unchanged and conditions permit, it is better to perform them earlier. Since the join operation requires hash pairing or loop traversal of two tables, with relatively large operation overhead, reducing the quantity in advance can greatly optimize performance. This is the basis for us to judge whether the final SQL model has reached the optimal state. If it has not reached the optimal state, operations such as translation transformation and removal are performed.

[0049] Short-circuit attribute detection: mainly refers to filter condition optimization. For example, if there is a condition of count > 1 in the first step and a condition of count > 10 in the second step, then the condition of count > 1 is invalid. And for unused fields subsequently, remove them in advance.

[0050] Removal of intermediate tables with no operations: directly use the previous table without any operations as the next table. Such intermediate tables will waste read / write resources and storage space, and removal is used to optimize performance. After operation translation, there will be such intermediate processes without operations, which is also the key reason why the optimization method is effective.

[0051] The following figure illustrates the specific process of moving column calculation optimization through two simple SQLs.

[0052] As Figure 2 shown, it shows the process of moving column calculations in B-SQL to A-SQL:

[0053] Columns D1 and D2 in table T2 are not calculated and do not need to be changed.

[0054] The calculation F2 of column D3 in table T2 is moved forward to C3 in table T1 and becomes F2(C3).

[0055] The calculation F4 of column D4 in table T2 is moved forward. Since there is already a calculation in C4 of table T1, it becomes F4(F3(C4)).

[0056] After moving, B-SQL has no operations on table T2, so B-SQL meets the removal condition.

[0057] As Figure 3 shown, it shows the process of merging A-SQL and B-SQL into C-SQL. Since there are no operations in B-SQL after translation, it can be directly removed. To ensure the normal progress of subsequent operations, we need to ensure that the input columns and output classes remain unchanged. That is, select the input fields of the original input table as the input, and the output of the final result table as the output, and retain the intermediate column calculations. The figure shows a simple process of calculation translation and table removal.

[0058] The removal methods of GROUP (aggregation) and FILTER (where or having conditions) are not shown in the figure. The specific translation methods of conditions and aggregation operations need to meet their respective conditions.

[0059] Translation of filtering conditions:

[0060] Prerequisites for condition translation: When translating the conditions of B-SQL to A-SQL, it is necessary to ensure that the filtered fields of B-SQL have been calculated or can be calculated in A-SQL.

[0061] Steps for condition translation:

[0062] If A-SQL does not contain a group operation: Remove the filtering conditions of B-SQL and add the corresponding filtering conditions in A-SQL. If the filtered fields do not exist in A-SQL, replace them with the corresponding calculation process.

[0063] If A-SQL contains a group operation: Remove the filtering conditions of B-SQL and add the corresponding filtering conditions in A-SQL. If the filtered fields do not exist in A-SQL, replace them with the corresponding calculation process, and modify them to having filtering conditions.

[0064] If B-SQL contains a JOIN operation: Check which table among the multiple tables in B-SQL the filtering conditions are applied to, and the conditions can be moved only to the corresponding table.

[0065] If A-SQL contains a union operation: The filtering conditions of B-SQL need to be moved to the SQL statements of each table in the union.

[0066] Translation of aggregation operations:

[0067] Prerequisites for aggregation translation: When translating the aggregation operations of B-SQL to A-SQL, it is required that A itself does not contain calculations related to aggregation. Except for group, distinct, max, and row_number window functions, etc. are all operations related to aggregation, and A-SQL does not contain a union operation.

[0068] If B-SQL contains a join operation: It is necessary to check the joined fields. If one of the fields of the two joined tables satisfies the unique value, the group operation can be moved forward, and at the same time, the columns for aggregation calculation are also moved forward.

[0069] If B-SQL does not contain a join operation: The aggregation operation can be directly moved forward, and the corresponding column calculations are also moved forward. If B-SQL contains a having filter, it needs to be changed to a where filter or moved forward at the same time.

[0070] As shown above, in the SQL model, operations such as moving aggregation, filtering, and column calculation can be performed (since the positions of the operations are relative, there is no need to move the JOIN operation sequence anymore), reducing the depth of the entire SQL model, making the calculation order more reasonable, and at the same time fully ensuring that the calculation results are the same as the original ones.

[0071] Those of ordinary skill in the art can understand that the above are only preferred examples of the invention and are not used to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, for those skilled in the art, they can still modify the technical solutions described in the foregoing examples, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, etc. made within the spirit and principle of the invention shall be included within the protection scope of the invention. All technical features in this embodiment can be freely combined according to actual needs.

[0072] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for improving visual modeling performance optimization, characterized in that: By analyzing the operation process in the SQL model, moving row operations, column operations, and table removal, the final model achieves the optimal calculation process, which specifically includes the following steps; Step 1: Model SQL structure tree analysis; Step 2, determine whether the simplest form of the SQL structure tree is satisfied; Step 3: Move to the bottom layer in sequence through filtering and aggregation operations; Step 4: non-aggregate column operations are translated; Step 5, short circuit property detection; Step 6: Remove the intermediate table without any operation.

2. The method for improving the performance optimization of visual modeling according to claim 1, characterized in that: In step 2, the simplest form of SQL is the result summarized based on general SQL optimization experience. The amount of data to be processed later is reduced through filtering and aggregation operations. The two tables are hashed or traversed through association operations. The operation overhead is large, and reducing the number can greatly optimize performance. It is the basis for judging whether the final SQL model has reached the optimal state. If it has not reached the optimal state, translation transformation and removal operations are performed.

3. The method for improving the performance optimization of visual modeling according to claim 1, characterized in that: In step 3, the filtering and aggregation operations are sequentially moved to the bottom layer, which specifically includes the following: The translation conditions for row operations: In the case of non-order row operations, if the filtered field exists, the filtering operation can unconditionally move the execution order in the data processing flow; Aggregation operations can be translated across filtering of group key values ​​or when there is no column operation other than field selection; Translation conditions for column operations: When all column-related operations satisfy computational consistency, the computation and decomposition of aggregate columns are equivalent; When the Cartesian product operation consistency is met, the column calculation operation can be translated through the Join operation; Generally, column calculations of a table can be unconditionally shifted as long as the dependencies of row and table operations are met.

4. The method for improving the performance optimization of visual modeling according to claim 1, characterized in that: In step 5, short-circuit attribute detection refers to the optimization of filtering conditions. For example, if there is a condition of count>1 in the first step and a condition of count>10 in the second step, the condition of count>1 is invalid; and unused fields are removed in advance.

5. The method for improving the performance optimization of visual modeling according to claim 1, characterized in that: In step 6, invalid operations are removed as follows: If a field is not used in subsequent processes, it can be removed; If you aggregate the same field continuously, you can remove one of them; If a table does not contain any aggregation, association, sorting, or column calculation functions, the table can be removed.