Data processing method and device, equipment, storage medium and product
By performing operator chain splitting and compiling analysis on the query plan, and adaptively selecting the execution mode, the problem of inefficiency of the existing data query engine is solved, and more efficient query processing and analysis performance is achieved.
Patent Information
- Application Number
- CN202410039737.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-11
AI Technical Summary
Existing data query engines cannot adaptively select the appropriate execution engine based on the query plan, resulting in inefficient query processing.
By splitting the query plan into multiple operator chains, and compiling and analyzing each operator chain, obtaining compilation income and memory indication information, adaptively determine the execution mode, fusing vectorized execution and compilation execution modes, and optimizing the execution strategy of the query plan.
It significantly improves query processing efficiency and analysis performance, and can find the optimal execution strategy in different scenarios, reducing query response delay.
Smart Images

Figure CN120296039A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a data processing method, apparatus, device, storage medium, and product. Background Art
[0002] With the advent of the big data era, the analysis of massive data requires the assistance of some database processing technologies to achieve more efficient processing. Currently, various execution engines can be used to execute query plans, thereby executing the corresponding calculation logic to query data from the database. The execution engines here include, but are not limited to: vectorized execution engines, compiled execution engines, etc. In current data query engines, the corresponding execution engines are generally fixed and cannot adaptively select a suitable execution engine to execute the query plan according to the query plan. During the process of executing the query plan through the fixed execution engines, the queries executed by these execution engines are limited in terms of efficiency and performance, resulting in low query processing efficiency. Summary of the Invention
[0003] Embodiments of this application provide a data processing method, apparatus, device, storage medium, and product that can adaptively fuse one or more execution modes to complete a query plan.
[0004] On the one hand, embodiments of this application provide a data processing method, including:
[0005] Obtain a query plan to be processed, and perform splitting processing on the query plan to obtain multiple operator chains;
[0006] Perform compilation analysis processing on each operator chain in the multiple operator chains to obtain compilation benefit indication information for each operator chain and compilation memory indication information for each operator chain; the compilation memory indication information is used to indicate the amount of memory space required for the code generated after the corresponding operator chain is compiled;
[0007] Determine the execution mode of each operator chain according to the compilation benefit indication information of each operator chain and the compilation memory indication information of each operator chain;
[0008] Execute the corresponding operator chain according to the determined execution mode of each operator chain to complete the query plan.
[0009] On the other hand, embodiments of this application provide a data processing apparatus, including:
[0010] An acquisition unit, configured to obtain a query plan to be processed, and perform splitting processing on the query plan to obtain multiple operator chains;
[0011] A processing unit for performing compilation analysis processing on each operator chain in a plurality of operator chains to obtain compilation benefit indication information of each operator chain and compilation memory indication information of each operator chain; the compilation memory indication information is used to indicate the amount of memory space required for the code generated after the corresponding operator chain is compiled.
[0012] The processing unit is further configured to determine an execution mode of each operator chain according to the compilation benefit indication information of each operator chain and the compilation memory indication information of each operator chain.
[0013] The processing unit is further configured to execute the corresponding operator chain according to the determined execution mode of each operator chain to complete the query plan.
[0014] On the other hand, an embodiment of the present application provides a computer device, which includes an input interface and an output interface, and the computer device further includes: a processor and a computer storage medium;
[0015] Wherein, the processor is adapted to implement one or more instructions, and the computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by the processor to perform the data processing method mentioned above.
[0016] On the other hand, an embodiment of the present application provides a computer storage medium, in which one or more instructions are stored, and the one or more instructions are adapted to be loaded and executed by the processor to perform the data processing method mentioned above.
[0017] On the other hand, an embodiment of the present application provides a computer program product, which includes a computer program; when the computer program is executed by a processor, the data processing method mentioned above is implemented.
[0018] In the embodiment of the present application, a query plan to be processed can be obtained and split to obtain a plurality of operator chains, and then each operator chain is subjected to compilation analysis processing to obtain compilation benefit indication information and compilation memory indication information of each operator chain. Based on the compilation benefit indication information and compilation memory indication information of each operator chain, the execution mode of each operator chain can be determined. In this process, for a complete query plan, taking the split operator chains as units, indication information corresponding to the compilation time dimension and the storage space dimension can be obtained based on the compilation analysis of the operator chains, and based on the indication information in these two dimensions, the execution mode of the operator chains can be adaptively determined. The determined execution mode is more reasonable and reliable, and can significantly improve the query execution efficiency. Finally, the corresponding operator chain can be executed according to the determined execution mode of each operator chain to complete the query plan. In this way, for a query plan, it supports flexibly integrating one or more execution modes, thereby improving the query processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0020] Figure 1 is an architecture diagram of a data processing system provided by an embodiment of the present application;
[0021] Figure 2 is a flowchart of a data processing method provided by an embodiment of the present application;
[0022] Figure 3 is a schematic diagram of splitting a query plan into multiple operator chains provided by an embodiment of the present application;
[0023] Figure 4 is a flowchart of another data processing method provided by an embodiment of the present application;
[0024] Figure 5 is a schematic diagram of a model training process provided by an embodiment of the present application;
[0025] Figure 6 A flowchart of another data processing method provided by an embodiment of the present application;
[0026] Figure 7 is a schematic diagram of the structure of a data processing device provided by an embodiment of the application;
[0027] Figure 8 is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0029] The present application proposes a data processing method, which is a technical solution that supports the integration of one or more execution modes. This technical solution splits the query plan to be processed into multiple operator pipelines, and makes decisions on the execution strategy for each operator pipeline. Specifically, it first performs compilation analysis on each operator pipeline to obtain the compilation benefit indication information and compilation memory indication information for each operator pipeline, and determines the execution mode for each operator pipeline based on the compilation benefit indication information and compilation memory indication information for each operator pipeline. Then, it executes the operator pipelines according to the execution modes of each operator pipeline to complete the execution of the query plan. This technical solution can analyze the benefits generated by the compilation of each operator pipeline and the memory consumption after compilation, and combine the analyzed benefits and memory consumption to reasonably determine the required execution mode for each operator pipeline, complete the decision-making of the execution strategy of the query plan, and thus execute the operator pipelines according to the determined execution mode, which can improve the query processing efficiency and further improve the query analysis performance. Further, if the technical solution of the present application is applied to the data analysis scenario, it can also improve the data analysis efficiency in the data analysis scenario.
[0030] A query plan refers to the logical process for organizing a query. This query plan can also be referred to as an execution plan or a physical query plan, and can include a set of steps for completing the query. The set of steps can include one or more execution steps, and these steps describe the database operations (also referred to as query operations) required to create the query result. Such database operations are, for example: SELECT, JOIN, TableScan, etc. Specifically, a query plan is the execution plan generated by the database before executing the query, and the query plan includes all the operations involved in the query, the operation order, the connection method, the index usage, and other information. It can help the database optimize the query performance, for example, by selecting the best index or connection order. A query plan can correspond to a query statement, and this query statement can be an SQL (Structured Query Language, a structured query language) statement, which can be used to define, manage, and query data in a relational database-like system.
[0031] An operator pipeline (pipeline) is the result of splitting a query plan. An operator pipeline can include multiple query operators (abbreviated as operators for short), and one query operator corresponds to one query operation (or sub-operation) in the query plan. A complete query plan consists of multiple independent query operations, and each query operation can be represented as a query operator in the operator pipeline. For example, a selection operation can be represented by a "SELECT" operator, and a join operation can be represented by a "JOIN" operator.
[0032] For any operator chain, the execution mode of an operator chain can be the vectorized execution mode or the compilation execution mode. Then, by adopting this solution, it can be determined which operator chains among the split operator chains need to be compiled and executed, and which operator chains need to be vectorized and executed, so as to execute the query plan more efficiently and improve the data retrieval efficiency.
[0033] Vectorized execution is the core technology for MPP (Massively Parallel Processing) databases to achieve high-performance query analysis. An MPP database refers to a database adopting the MPP architecture. Based on MPP, tasks can be parallelly distributed to all cluster nodes. After the calculation is completed on each node, the results of each part are aggregated together to obtain the final result. During the vectorized execution process, the basic operations in the query statement can be abstracted into basic vectorized operators. When processing a query, a batch of data is fetched each time and organized in columnar form in memory. Each operator takes one or more columns as input and generates a new intermediate temporary column as the result for return. After the final result is obtained, the intermediate temporary column will be discarded. The engine used to implement vectorized execution is called the vectorized execution engine (or vectorized execution engine). The vectorized execution engine is essentially still an interpreted execution engine, that is, at runtime, according to the SQL query statement input by the user, it is necessary to dynamically judge the input situation when executing each operator and jump to the corresponding code segment for execution. The vectorized execution engine is a special type of compilation execution engine. Compared with the traditional row-based execution engine, such as the compilation execution engine that processes one row of data each time, the vectorized execution engine processes a batch of data each time and uses vectorized operations to accelerate query execution. The vectorized operation here refers to performing the same operation on multiple data elements simultaneously instead of one by one. Specifically, batch operations are performed with a column of the same type as the basic unit. This method can reduce the number of loops and virtual function (a member function of a class) calls, and can significantly improve the CPU (Central Processing Unit) utilization rate and cache hit rate, thereby improving query performance. Currently, databases adopting the vectorized execution engine include open-source databases such as ClickHouse, Byconity, and Hologres. Generally speaking, the vectorized execution engine processes a batch of data each time and performs batch operations with a column of the same type as the basic unit, which can significantly improve the CPU utilization rate and cache hit rate. However, in some complex expression scenarios, the frequent memory application, reading and writing, and destruction of intermediate temporary columns also bring non-negligible additional overheads, significantly slowing down the overall query processing efficiency.
[0034] Compiled execution, also known as query compilation (or query compilation execution), is the process of converting a query statement (such as an SQL statement) entered by the user into executable machine code. The compiled execution engine checks the syntax and semantics of the query statement and converts it into assembly language or other intermediate representation forms, and finally compiles it into machine code for immediate execution. This process can help the database optimize query performance and improve CPU utilization, for example, eliminating unnecessary branches, virtual function calls, etc. Different from vectorized execution, query compilation can eliminate all unnecessary branches at runtime and generate code specialized for the dynamic query statement entered by the user, and its performance can be comparable to handwritten code for a specific query statement. The compiled execution engine is a component in the database responsible for executing queries. It receives the SQL query statement entered by the user, scans the data according to the content of the query statement, executes the corresponding calculation logic, and retrieves and converts data from the database. The compiled execution engine can be used to implement query compilation (or query compilation execution, compiled execution). Query compilation technology can dynamically fuse multiple query operators into a single operator at runtime, avoiding the generation and reading / writing of intermediate temporary columns. However, the additional costs it brings cannot be ignored: usually, a single compilation takes more than 15 ms, and the more complex the query, the longer the compilation time. Generally speaking, the compiled execution engine can directly compile a query statement (such as an SQL statement) into compact and efficient machine code, eliminate all unnecessary branches and conditional jumps, and can generate code friendly to modern CPU architectures, with a speed comparable to that of a handwritten query execution plan. Currently, execution engines based on query compilation are, for example, HyPerDB. However, the duration consumed by a single query compilation will increase with the increase in query compilation complexity, the compilation time overhead is high, and since the machine code requires additional memory for storage, the storage space overhead is also large.
[0035] Based on the above method, an adaptive fusion execution engine can be designed to execute the data processing method provided in the embodiments of the present application. The adaptive fusion execution engine supports multiple execution modes, including but not limited to the vectorized execution mode and the compiled execution mode. When the execution mode of each operator chain is the vectorized execution mode, then the adaptive fusion execution engine is equivalent to a vectorized execution engine. When the execution mode of each operator chain is the compiled execution mode, then the adaptive fusion execution engine is equivalent to a traditional compiled execution engine. When the execution mode of some operator chains is the vectorized execution mode and the execution mode of some operator chains is the compiled execution mode, then the adaptive fusion execution engine is a brand-new fusion execution engine. It can adaptively determine which operator chains need to enable query compilation according to the specific situation of the query, and the query operators it supports for compilation are not limited to specific operators, but any operator is applicable under corresponding conditions, thus being able to improve the acceleration effect.
[0036] At present, some fusion execution engines introduce partial query compilation capabilities on the basis of vectorized execution, and can accelerate some simple expressions compilation and certain operations in operators through query compilation. In this application, if the execution modes of different operator chains involve at least two modes, then the execution engine implementing this solution is equivalent to a completely new fusion execution engine. Compared with the fusion execution engine that controls whether to enable query compilation through a global switch, the operator chains that can be compiled by the fusion execution engine in this solution are adaptively determined based on the indication information of the corresponding dimension obtained from the compilation analysis of the operator chains. The basic unit of query compilation is the operator chain, and the execution mode is determined based on the compilation benefit indication information and the compilation memory indication information, balancing the processing duration consumed and the storage space occupied.
[0037] The technical solution provided by the embodiments of this application can be applied to various data processing scenarios, including but not limited to: AP (On-Line Analytical Processing, OLAP; hereinafter referred to as online analytical processing) scenarios and other business scenarios. Among them, the AP scenario is mainly oriented to the analysis scenario, with many complex queries, long running times, and a focus on response latency. Based on the technical solution provided by this application, for queries in different scenarios, an optimal execution strategy can be found for the query plan, thereby improving the query execution efficiency and significantly reducing the query response latency. This application takes the operator chain as the unit, and the execution strategy is composed of the execution modes of each operator chain. In the formulation of the execution strategy, based on the compilation analysis of the operator chain, efforts can be made to improve the overall query efficiency and make more scientific decisions on the execution strategy. In addition, this application can integrate the advantages of different execution modes (such as vectorized execution mode and compilation execution mode), adaptively select a relatively faster execution mode for each operator chain, so as to achieve reasonable allocation of computing resources and optimize the execution plan, significantly improve the query analysis performance, and improve the query efficiency.
[0038] Next, the architecture of the data processing system provided by the embodiments of this application will be introduced in conjunction with the accompanying drawings.
[0039] Please refer to Figure 1 , which is an architecture diagram of a data processing system provided by the embodiments of this application. As Figure 1As shown in the figure, the data processing system includes a database 100 and a computer device 101; a communication connection can be established between the database 100 and the computer device 101 in a wired or wireless manner. Among them, the computer device 101 is used to execute the data processing process; the database 100 is used to provide data support for the data processing process of the computer device 101. For example, the database 100 can be used to store query plans, statistical information of query statements, and a target compilation analysis model for compiling and analyzing operator chains. The computer device mentioned above can include any one or both of a terminal and a server. The terminal includes, but is not limited to: smart phones, tablets, smart wearable devices, smart voice interaction devices, smart home appliances, personal computers, in-vehicle terminals, smart cameras, virtual reality devices, etc., and this application does not make any restrictions on this. Regarding the number of terminals, this application does not make any restrictions. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, but is not limited thereto. Regarding the number of servers, this application does not make any restrictions.
[0040] The general process for the computer device 101 to execute the data processing method may include:
[0041] ① Obtain a query plan. In an implementable manner, a query statement can be obtained, and the query statement can be parsed and optimized to obtain the query plan corresponding to the query statement. Optionally, a query optimizer and a fusion execution engine can be deployed in the computer device. The data processing method provided in this application can be called by the computer device to execute the fusion execution engine. For the process of obtaining the query plan, it can include: the computer device obtains the query statement, calls the query optimizer to parse and optimize the query statement to obtain the query plan corresponding to the query statement. Then the query optimizer sends the query plan to the fusion execution engine, or the computer device calls the fusion execution engine to obtain the query plan from the query optimizer so that the fusion execution engine can split the obtained query plan.
[0042] ② Split the query plan. In one embodiment, the fusion execution engine can be called to split the query plan to obtain multiple operator chains. An operator chain can be composed of multiple query operators, and each query operator corresponds to a query operation in the query plan. The query operations represented by the operators between different operator chains can be the same. Based on this, there is also a dependency relationship between multiple operator chains. Through the operator chains, the execution order of each query operation in the query plan can be reorganized to improve the execution efficiency of the query plan. For example, two operator chains without a dependency relationship can be executed in parallel.
[0043] ③ Compile and analyze each operator chain. By compiling and analyzing each operator chain, the compilation benefit indication information and compilation memory indication information of each operator chain can be obtained. The compilation benefit indication information can be used to indicate the compilation benefits that can be generated when the corresponding operator chain is executed in different execution modes. The compilation benefit can specifically be the duration benefit. The duration benefit indicated by the compilation benefit indication information can be positive or negative. For example, if the duration taken for a certain operator chain to execute query compilation is x seconds less than the duration taken for vectorized execution, then the duration benefit indicated by the compilation benefit indication information of the operator chain can be x seconds. Conversely, if it is y seconds more, then the duration benefit indicated by the compilation benefit indication information can be -y seconds. The compilation memory indication information can be used to indicate the amount of memory space occupied by the code generated after the corresponding operator chain is compiled. Generally speaking, it is used to indicate the memory consumption after the corresponding operator chain is compiled. For example, if the code volume generated by a certain operator chain during query compilation is 0.5 KB (Kilobyte), that is, 0.5 KB of memory space is required to store the compiled code, then the amount of memory space occupied indicated by the compilation memory indication information is 0.5 KB.
[0044] In one embodiment, step ③ above can be executed by calling the target compilation analysis model. The training of the target compilation analysis model can be completed in the computer device 101 or in other computer devices. The trained target compilation analysis model can be stored in the database. Any computer device can obtain it from the database and deploy it in the fusion execution engine when it needs to use the target compilation analysis model. Alternatively, after the target compilation analysis model is trained, it can be directly deployed in the fusion execution engine for the fusion execution engine to use directly. Optionally, an execution module corresponding to the execution mode (such as a vectorized execution module and a compilation execution module) can also be deployed in the fusion execution engine to support calling the corresponding execution module to execute the operator chain according to the determined execution mode after the corresponding execution mode is determined for each operator chain.
[0045] ④ Determine the execution mode of each operator chain. For any operator chain, the execution mode of the operator chain can be determined based on the indication information in two dimensions: the compilation benefit indication information and the compilation memory indication information of the operator chain. This is beneficial to balance the duration required for the actual execution of the operator chain and the memory consumption required during the actual execution process, so as to significantly improve the query execution efficiency under limited computing resources.
[0046] ⑤ Execute the operator chain according to the determined execution mode. For each operator chain obtained by splitting, the corresponding execution mode will be determined. Thus, each operator chain can be executed according to the corresponding execution mode. And the execution of the operator chain means that the content in the query plan is executed. After all the split operator chains are executed, the query plan can be completed.
[0047] It can be seen that in the above data processing process, by splitting the query plan to be processed into multiple operator chains, first performing compilation analysis processing on each operator chain to obtain the compilation benefit indication information and the compilation memory indication information of each operator chain, and determining the execution mode of each operator chain based on the compilation benefit indication information and the compilation memory indication information of each operator chain, and then executing the operator chain according to the execution mode of each operator chain to complete the execution of the query plan. This technical solution can estimate the compilation benefit generated by the compilation of each operator chain and the memory consumption after compilation through compilation analysis, and combine the estimated compilation benefit and memory consumption to reasonably determine the required execution mode for each operator chain, complete the decision-making of the execution strategy of the query plan, so as to execute the operator chain according to the determined execution mode, which can improve the query processing efficiency and further improve the query analysis performance.
[0048] The data processing method provided by the embodiments of the present application will be elaborated in detail below.
[0049] Please refer to Figure 2 , which is a schematic flowchart of a data processing method provided by the embodiments of the present application. This data processing method can be executed by a computer device (such as Figure 1 the computer device 101 in the data processing system shown) and can include the content described in the following S201 - S204.
[0050] S201, Obtain a query plan and perform a splitting process on the query plan to obtain multiple operator chains.
[0051] In a specific implementation, the logic for obtaining a query plan can be as follows: First, obtain the query statement, then perform parsing and optimization processing on the query statement to obtain a physical execution plan, and then determine this physical execution plan as the query plan. Optionally, a query parser and a query optimizer can be deployed in the computer device. In the implementation of the parsing and optimization processing, the query parser can be called to parse the query statement to obtain a parse tree, which is an initial execution plan. To ensure more efficient processing, further, the query optimizer can be called to perform optimization processing on this initial execution plan to obtain an optimized initial execution plan, which is a better physical execution plan and can be used as the query plan processed in the embodiments of this application.
[0052] It should be noted that since the query optimizer is unaware of the splitting of the query plan and the determination of the execution strategy (such as vectorization / compilation / fused execution), a fused execution engine can also be deployed in the computer device. The query optimizer can send the query plan to the fused execution engine. After receiving the query plan, the fused execution engine needs to perform low-level execution layer optimization before actually executing the query plan, and the execution layer optimization includes: splitting of the query plan and determination of the execution strategy. Among them, the splitting of the query plan refers to: splitting the query plan into multiple operator chains, and the determination of the execution strategy refers to: determining the execution mode of each operator chain based on the respective indication information obtained from the compilation analysis of the operator chain. In this way, the optimal execution strategy can be determined, and then the corresponding query can be executed.
[0053] Each operator chain obtained by splitting the query plan includes multiple query operators, and one query operator is used to represent a query operation in the query plan. Each operator chain can accept one input and has one output. The input of the starting operator (i.e., the first operator) of an operator chain can come from a read local file or an external data source, or can be the output data from other operator chains it depends on, or the calculation result of the end (i.e., the last operator) of other operator chains. The end operator of an operator chain can output the calculation result to a disk or an external data source, or can send the calculation result to a downstream operator chain that depends on itself. There can be the same query operators between different operator chains to represent the inclusion of the same query operation, and there is a dependency relationship between such operator chains. Optionally, the query plan can be a query tree, and each operator chain can be regarded as a subtree of this query tree. Exemplarily, as Figure 3 the schematic diagram of splitting the query plan into multiple operator chains shown. Figure 3 The query plan shown in Figure 3 is the physical execution plan corresponding to TPCH-Q5 (a query statement), and is divided into 8 operator chains P1 - P8 ( Figure 3 the content in each dashed border in Figure 3In the shown operator chain, there are also upstream and downstream data dependencies between Pipelines, specifically including the following: P2 depends on P1; P3 depends on P2; P6 depends on P3, P4, and P5; P7 depends on P6; P8 depends on P7. In one implementation, there is a dependency relationship between two operator chains, and this dependency relationship can be used to determine the order of scheduling operator chains to run in a thread pool. During the actual execution of the operator chain, based on the dependency relationships among multiple operator chains, multiple operator chains can be alternately scheduled to run in the thread pool, so that the operator chains that can be executed in parallel are executed in parallel, and the operator chains with dependency relationships are executed serially, improving the execution efficiency of the operator chain and thus the execution efficiency of the query plan.
[0054] S202. Compile and analyze each operator chain in multiple operator chains to obtain the compilation benefit indication information of each operator chain and the compilation memory indication information of each operator chain.
[0055] Among them, the compilation benefit indication information is used to indicate the duration benefit that can be generated when the corresponding operator chain is executed in different execution modes, and the compilation memory indication information is used to indicate the amount of memory space required for the code generated after the corresponding operator chain is compiled. It should be noted that the content indicated by the compilation benefit indication information and the compilation memory indication information are both estimated contents obtained through compilation analysis and are used to guide the determination of the execution mode. The compilation memory indication information is the indication information obtained by compiling and estimating the target execution mode, where the target execution mode is a mode used to compile the operator chain into a code segment and then compile it, such as the compilation execution mode. Among them, the code segment is a string of instruction codes executed by a computer, which can be called machine code. It is a string of numbers composed of binary digits and is a form of the computer's underlying language. After the operator chain is compiled in the target execution mode, it will generate a certain volume of code (or called code segment), and a certain volume of code needs to occupy the corresponding storage space for storage. Therefore, both the code volume and the amount of memory space can be measured in bytes (Byte, a basic storage unit of computer data). Exemplarily, the compilation memory indication information indicates that the amount of memory space required for the operator chain after being compiled is 0.5KB (Kilobyte), which means that the code generated by the operator chain during compilation in the compilation execution mode needs to occupy 0.5KB of memory space for storage.
[0056] In one implementation, the duration benefit indicated by the compilation benefit indication information refers to the relative difference between the predicted durations required to execute the operator chain under two execution modes. The positive or negative value of the duration benefit can be used to represent the time saved or extra time spent by one execution mode compared to another. For example, the execution modes include the vectorized execution mode and the compilation execution mode, and the compilation benefit indication information is used to indicate the compilation benefit when the operator chain is executed in the compilation execution mode compared to the vectorized execution mode. This compilation benefit can be denoted as y. If the compilation benefit y of a certain operator chain is 100, it means that the operator chain can save 100 seconds when executed in the compilation execution mode compared to the vectorized execution mode, that is, the compilation execution mode is 100 seconds faster than the vectorized execution mode. Another example, if y = -10, it means that the operator chain will spend 10 more seconds when executed in the compilation execution mode compared to the vectorized execution mode, that is, the compilation execution mode is 10 seconds slower than the vectorized execution mode.
[0057] For any one of the multiple operator chains, the logic of the compilation analysis process can be generally as follows: First, perform feature extraction processing on the operator chain to obtain the feature vector of the operator chain. This feature vector is used to represent some statistical features of the operator chain, including but not limited to: the number of input rows and output rows of the operator chain, the number of input bits and output bits, the number of each type of query operator included in the operator chain, the number of each type of expression or each function included in the operator chain, and so on. Then, based on the relationship between the feature vector of the operator chain and the compilation benefit, predict the compilation benefit indication information for each operator chain, and based on the relationship between the feature vector and the memory occupancy, predict the compilation memory indication information for each operator chain. In a feasible way, the relationship between the feature vector and the compilation benefit includes: the larger the corresponding eigenvalue in the feature vector, the smaller the compilation benefit and the smaller the duration benefit that can be brought; the smaller the eigenvalue, the larger the compilation benefit and the larger the duration benefit that can be brought. The relationship between the feature vector and the memory occupancy includes: the feature vector is positively correlated with the memory occupancy. The larger each eigenvalue in the feature vector, the larger the memory occupancy; the lower each eigenvalue in the feature vector, the smaller the memory occupancy.
[0058] S203. Determine the execution mode of each operator chain according to the compilation benefit indication information of each operator chain and the compilation memory indication information of each operator chain.
[0059] In a specific implementation, based on the duration gain indicated by the compilation gain indication information and the amount of memory space required indicated by the compilation memory indication information, the execution mode can be determined for the operator chain from the dimensions of the estimated duration gain and the memory occupancy, so as to balance the duration consumption and the memory consumption, and improve the execution efficiency as much as possible under limited resources. For a query plan, the process of determining the execution mode of each operator chain can be understood as a decision-making process of an execution strategy. The execution strategy of the query plan is used to indicate the execution mode of each operator chain obtained by splitting. The decision-making goal of this execution strategy is to make the overall execution efficiency of the query plan relatively optimal. Therefore, the determination logic for the execution mode of each operator chain can generally include the following content:
[0060] In one embodiment, the execution mode includes a vectorized execution mode and a compilation execution mode. The execution mode of an operator chain can be one of the vectorized execution mode and the compilation execution mode. In the above manner, there can be the following situations for the execution modes of multiple operator chains obtained by splitting: (1) The execution mode of each operator chain in the multiple operator chains is the compilation execution mode; (2) The execution mode of each operator chain in the multiple operator chains is the vectorized execution mode; (3) The execution mode of some operator chains in the multiple operator chains is the compilation execution mode, and the execution mode of the remaining operator chains is the vectorized execution mode.
[0061] S204, Execute the corresponding operator chain according to the determined execution mode of each operator chain to complete the query plan.
[0062] In one embodiment, the execution mode of an operator chain is the vectorized execution mode or the compilation execution mode. Based on this, the determination of the execution mode of each operator chain is also equivalent to the determination of whether to trigger compilation execution for each operator chain. For example, if the execution mode of a certain operator chain is the compilation execution mode, it means that sufficient benefits can be obtained by performing query compilation on this operator chain, so the execution mode can be determined to be the compilation execution mode, and this operator chain will trigger query compilation during the actual execution process. Therefore, for the execution of a query plan, there are the following three situations: ① All operator chains trigger compilation execution, and the execution process can be equivalent to the pure compilation execution implemented by the compilation execution engine; ② None of the operator chains trigger compilation execution. In other words, all operator chains trigger vectorized execution, and the execution process can be equivalent to the pure vectorized execution implemented by the vectorized execution engine; ③ Some operator chains trigger compilation execution, and other operator chains adopt vectorized execution, and the execution process is a new type of fusion execution implemented by the fusion execution engine.
[0063] For any operator chain, if the execution mode is the compilation execution mode, then query compilation processing can be performed on the query operators in the operator chain to obtain the code segment corresponding to the operator chain, and execute the code segment corresponding to the corresponding operator chain. By compiling the operator into fully specialized machine code, all unnecessary conditional jumps and virtual function calls can be eliminated. The compiled operator chain (Pipeline) is a compiled function. Given an input chunk, a chunk can be returned as the output. A chunk is a small part of all the data. Specifically, it can be a data block obtained by splitting the data. For example, adjacent data rows belong to the same chunk, or data rows with the same hash value belong to the same chunk. It can be seen that in the compilation execution mode, each operator can be compiled into a code segment. Based on the execution of the code segment, the execution of the query operation corresponding to the operator is completed, and after the code segments corresponding to all operator chains are executed, the query plan execution is completed, and thus the query result can be obtained. If the execution mode of a certain operator chain is the vectorized execution mode, then the operator chain can be executed according to the vectorized execution mode, and both the input and output of the operator chain are also a chunk. Here, a chunk is also a small part of all the data and is organized in columnar storage form. It should be noted that when executing the operator chain in the vectorized execution mode, the operator chain does not need to be converted into a code segment but can be directly executed, and a new intermediate temporary column is generated as the result and discarded after the final result is obtained.
[0064] The data processing method provided by the embodiments of the present application can obtain the query plan to be processed and perform splitting processing on the query plan to obtain multiple operator chains, and then perform compilation analysis processing on each operator chain to obtain the compilation benefit indication information and compilation memory indication information of each operator chain. Based on the compilation benefit indication information and compilation memory indication information of each operator chain, the execution mode of each operator chain can be determined. In this process, taking the split operator chain as a unit, the indication information corresponding to the compilation time dimension and the storage space dimension can be obtained based on the compilation analysis of the operator chain, and based on the indication information in these two dimensions, the execution mode of the operator chain can be adaptively determined. The determined execution mode is more reasonable and reliable, and can maximize the query execution efficiency. Finally, the corresponding operator chain can be executed according to the execution mode of each operator chain determined to complete the query plan. In this way, for a query plan, it supports flexibly integrating one or more execution modes, thereby improving the query processing efficiency.
[0065] Please refer to Figure 4 , which is a schematic flowchart of a data processing method provided by the embodiments of the present application. This data processing method can be performed by a computer device (such as Figure 1be executed by the computer device 101) in the data processing system shown, and the data processing method may include the content described in the following S401 - S405.
[0066] S401. Obtain a query plan, and perform a splitting process on the query plan to obtain a plurality of operator chains.
[0067] In one embodiment, the query plan is a query tree, the query tree includes a plurality of nodes, and each node is used to represent a query operation in the query plan. Among the plurality of nodes, there is at least one leaf node, and the leaf node refers to the terminal node in the query tree, that is, the node that does not branch out new child nodes. During the splitting process of the query plan, the complete query plan can be split into a plurality of operator chains (pipelines) with the boundary of whether the operator is blocked. Specifically, when the computer device performs a splitting process on the query plan to obtain a plurality of operator chains, it can be implemented according to the following steps S11 - S14.
[0068] S11. Start traversing from the leaf node of the query tree, and perform a blocking analysis on the query operation represented by the currently traversed node.
[0069] In a specific implementation, the number of leaf nodes of the query tree is one or more. During traversal, one or more leaf nodes can be traversed in parallel. For the sake of description, taking the currently traversed node as an example for exemplary illustration, for each traversed node, a similar process can be executed. Since each node can represent a query operation, in order to determine whether the current node can be used as the splitting boundary of the operator chain, a blocking analysis needs to be performed on the query operation represented by the current node. Here, the blocking analysis includes: determining whether the query operation represented by the current node needs to wait until all data is received before outputting the result.
[0070] S12. If it is analyzed that the query operation represented by the current node is blocked, then split out an operator chain according to the nodes included in the path between the leaf node and the current node.
[0071] Based on the logic of the blocking analysis, the query operation represented by the current node being blocked means that: in order to calculate the correct result, the current node needs to wait until all the data to be processed is received before performing the processing and outputting. For example, if the query operation is a sort operation, then all data needs to be received before the correct sorting result can be output. Therefore, the sort operation will be blocked and can be used as the splitting boundary of an operator chain. Specifically, the query operations represented by all the nodes included in the path between the leaf node and the current node can be used as the query operators in the operator chain. One query operation corresponds to one query operator, this path corresponds to one operator chain, and the split operator chain can be regarded as a subtree of the query tree. For example, Figure 3For the query plan shown, each leaf node is a TableScan node. Starting from each leaf node and traversing, when the current node is a HashJoin node, blocking analysis shows that the HashJoin operation needs to wait until data from other tables is received before it can be executed. Thus, the TableScan node and the HashJoin node can be regarded as operators in the operator chain, obtaining the operator chain P4.
[0072] S13. If blocking occurs in the query operation represented by the next node of the current node, then split out an operator chain according to the current node and the next node of the current node.
[0073] The query operation represented by the current node not having blocking means that: when the current node receives the data to be processed, it can process and output the calculation result. If it is analyzed that the query operation represented by the current node does not have blocking, then the traversal can continue, and the blocking analysis of the query operation represented by the next node of the current node can continue. The next node of the current node refers to: the node traversed after the current node and without other nodes intervening between it and the current node. If it is analyzed that the query operation represented by the current node has blocking, the traversal can also continue, and the blocking analysis of the query operation represented by the next node of the current node can continue. In the case where the current node has blocking and the next node of the current node also has blocking, the query operations represented by the current node and the next node of the current node can also be regarded as query operators in the operator chain, that is, the path formed by the current node and its next node corresponds to an operator chain. Exemplarily, as Figure 3 shown in the split operator chain, the HashAggregation node is a node with blocking in the operator chain P6, and the TopN node also has blocking. Thus, the operator chain P7 can be split according to the HashAggregation node and the TopN node. The same applies to the operator chain P8.
[0074] S14. After traversing the query tree, multiple operator chains are obtained.
[0075] Among the multiple nodes in the query tree, there is also a root node. The root node refers to the top - most node of the query tree, that is, the node without a parent node. For example Figure 3 the Projection node shown. When traversing to the root node of the query tree and performing blocking analysis on the query operation represented by the root node, after obtaining the analysis result, it indicates the end of the traversal of the query tree. Since blocking analysis is performed for each traversed node to determine whether an operator chain can be split out, after traversing all the nodes in the query tree, multiple operator chains can be obtained. Subsequently, based on the compilation and analysis processing of the multiple operator chains, the execution strategy of the query plan can be determined.
[0076] S402. Perform compilation analysis processing on each of the multiple operator chains to obtain the compilation benefit indication information and the compilation memory indication information for each operator chain.
[0077] In one embodiment, the specific implementation of the above S402 may include the following steps (1)-(2).
[0078] (1) Invoke the target compilation analysis model to perform compilation analysis processing on each of the multiple operator chains to obtain the compilation benefit value and the compilation code metric information for each operator chain.
[0079] To support the decision of the execution strategy, a target compilation analysis model can be trained. The target compilation analysis model is a machine learning model for estimating the compilation benefit and the amount of memory space occupied, and can specifically be used to estimate the compilation benefit that any split operator chain can obtain and the code volume generated after the operator chain is compiled.
[0080] In a feasible implementation, the target compilation analysis model can be invoked to perform feature extraction processing on each operator chain to obtain the feature vector of each operator chain, and then the compilation benefit value and the compilation code metric information of each operator chain can be predicted based on the feature vector of each operator chain.
[0081] Among them, the target compilation analysis model is a compilation analysis model trained based on multiple training samples. Each training sample includes a sample operator chain and training supervision information; the target compilation analysis model can be trained through the indication information and the training supervision information obtained by invoking the compilation analysis model to perform compilation analysis processing on the sample operator chain. The sample operator chain refers to the operator chain obtained by splitting the sample query plan. The sample query plan is a query plan generated according to the sample query statement for training. Any sample query statement can be a query statement generated according to a preset query template or a historical query statement provided by the user. This application does not make any restrictions on this. The training supervision information includes the compilation benefit supervision information and the compilation memory supervision information of the sample operator chain. In a specific implementation, considering that XGBoost performs well in solving regression problems and can complete fast training and inference only relying on the CPU, XGBoost can be selected as the underlying machine learning model, that is, the target compilation analysis model in this application can adopt XGBoost. In addition, the target compilation analysis model can also adopt other machine learning models, such as LightGBM (a boosting integrated model, which has better performance than XGBoost). This application does not make any restrictions on this.
[0082] The compilation benefit value refers to the time difference between the execution duration required for the corresponding operator chain in the vectorized execution mode and the execution duration required in the compilation execution mode; the compilation code metric information includes: the code volume that needs to be generated after the corresponding operator chain is compiled and executed in the compilation execution mode. This code volume matches the amount of memory space required.
[0083] (2) Determine the compilation benefit value of each operator chain as the compilation benefit indication information of the corresponding operator chain, and determine the compilation code metric information of each operator chain as the compilation memory indication information of the corresponding operator chain.
[0084] In this way, the compilation benefit indication information of each operator chain is the compilation benefit value of the corresponding operator chain, and the compilation memory information of each operator chain may include the code volume generated after the corresponding operator chain is compiled.
[0085] It can be seen that in the above method, each operator chain is compiled and analyzed by calling the target compilation analysis model. Since the target compilation analysis model is an artificial intelligence model trained to have a sufficiently high accuracy, it is possible to achieve a fast and accurate analysis of the operator chain, improve the determination efficiency of the compilation benefit indication information in the time dimension and the compilation memory indication information in the memory dimension, which is conducive to improving the fast and accurate determination of the execution mode, and thus can improve the overall processing efficiency.
[0086] The training process of the target compilation analysis model will be introduced in detail below, specifically including the following steps (1.1)-(1.4).
[0087] (1.1) Obtain the target training sample.
[0088] In one implementation, the computer device can obtain multiple training samples, and each training sample includes a sample operator chain and training supervision information. Then, a training sample is obtained from the multiple training samples as the target training sample. Since similar content will be executed during the training process for each training sample, for the sake of illustration, taking one training sample as an example, it can be understood that during the specific training process, the model can train the model once based on a batch of training samples, or train the model once based on multiple batches of training samples. This application does not limit this. The target training sample includes a sample operator chain and training supervision information, and the training supervision information includes the compilation benefit supervision information of the sample operator chain and the compilation memory supervision information of the sample operator chain.
[0089] The compilation benefit supervision information of the sample operator chain is used to indicate the compilation benefits obtained by the sample operator chain under different execution modes. The compilation benefit can be specifically defined as the time difference of the execution duration spent under different execution modes. The compilation memory supervision information of the sample operator chain is used to indicate the code volume generated after the sample operator chain is compiled and executed, and can specifically be used to reflect the amount of memory space occupied. These supervision information (including compilation benefit supervision information and compilation memory supervision information) are all data obtained after the actual vectorization execution and compilation execution of the sample operator chain, and can be used as the data referenced during model training. That is, the closer the indication information output during the model training process is to the corresponding supervision information, the better the model prediction ability, and it also represents that the trained model has better compilation analysis ability, can more accurately predict the compilation benefits and code volume of other operator chains, and thus obtain more accurate indication information.
[0090] For the construction process of each training sample, the specific implementation may include the following steps ① - ③.
[0091] Step ①: Obtain the sample query statement, and generate multiple sample operator chains corresponding to the sample query statement according to the sample query statement.
[0092] In a specific implementation, the query set can be obtained first. The query set includes at least one query statement, and then each query statement in the query set is determined as the sample query statement. In this way, at least one sample query statement can be obtained.
[0093] In one implementation manner, the sample query statement can be generated based on the query statement template and query parameters. The query statement template is a preset set of query templates, and the query parameters are the parameters to be filled into the query statement template, which can be input by the user or take default values. For example, for the parameters of the query range, if it is specified to query a certain table, then the query parameter can include the name of the table. Optionally, each query statement generated based on the query statement template and query parameters can include any one of the following operations: TableScan, Project, Filter, Aggregate, HashJoin, Sort, Union, Intersect, Minus. In another implementation manner, the obtained query set can include a historical query set provided by the user himself / herself, and each historical query statement in the historical query set can be used as the sample query statement. Thus, the sample query statement includes the historical query statements actually used by the user. In yet another implementation manner, the multiple sample query statements can include both the sample query statements generated based on the query statement template and query parameters and the historical query statements provided by the user.
[0094] After obtaining the sample query statement, the sample query statement can be first converted into a sample query plan, that is: the sample query statement is parsed and optimized to obtain the sample query plan. Then, the sample query plan is split to obtain multiple sample operator chains, and these sample operator chains correspond to the sample query statement. Similar to the operator chain, any sample operator chain includes multiple sample operators, and each sample operator corresponds to a query operation in the corresponding sample query plan. It can be understood that, for the sake of convenience of explanation, the above content is described by taking the processing of one sample query statement as an example. For each sample query statement, the corresponding sample operator chain can be generated in the above manner, so as to construct training samples based on multiple sample operator chains and the supervision information obtained by actually executing each sample operator chain.
[0095] Step ②: Execute each sample operator chain according to the vectorized execution mode and the compilation execution mode respectively, so as to obtain the compilation benefit supervision information of each sample operator chain and the compilation memory supervision information of each sample operator chain.
[0096] In a specific implementation, for a sample operator chain, it can be executed twice according to the vectorized execution mode and the compilation execution mode, so as to obtain the corresponding compilation benefit supervision information and compilation memory supervision information. The compilation benefit supervision information is used to indicate the time difference between the execution duration already spent by the corresponding sample operator chain in the vectorized execution mode and the compilation execution mode, and the compilation memory supervision information is used to indicate the amount of memory space occupied by the code generated after the corresponding operator chain is compiled in the compilation execution mode. In one implementation manner, the specific implementation of step ② may include the following steps a-step b.
[0097] Step a: Execute each sample operator chain according to the vectorized execution mode to obtain the first compilation statistic information of each sample operator chain; and execute each sample operator chain according to the compilation execution mode to obtain the second compilation statistic information and the reference code metric information of each sample operator chain.
[0098] For any sample operator chain, executing the sample operator chain in the vectorized execution mode can collect the first compilation statistic information, which can be used to indicate the duration spent on the vectorized execution of the sample operator chain. For example, the CPU time when the sample operator chain is vectorized is denoted as tv, and the unit is seconds (s). Executing the sample operator chain in the compilation execution mode can collect the second compilation statistic information, which can be used to indicate the duration spent on the compilation execution of the sample operator chain, including: ① The CPU time when the sample operator chain is executed after compilation is denoted as tc, and the unit is seconds (s); ② The CPU time consumed by the compilation of the sample operator chain is denoted as e, and the unit is seconds (s). The reference code metric information is used to indicate the code volume generated after the sample operator chain is compiled, specifically including: the code volume generated after the sample operator chain is compiled and executed, denoted as y2, and the unit is bytes (byte).
[0099] Step b: Determine the compilation benefit supervision information for each sample operator chain according to the first compilation statistic information and the second compilation statistic information of each sample operator chain; and determine the reference code metric information of each sample operator chain as the compilation memory supervision information for the corresponding sample operator chain.
[0100] In a specific implementation, the difference between the duration indicated by the first compilation statistic information and the duration indicated by the second compilation statistic information can be determined as the compilation benefit supervision information. The specific expression of the compilation benefit can be as follows: y1 = (tv - e - tc). In addition, the reference code metric information of the sample operator chain can be directly determined as the compilation memory supervision information, that is, the code volume indicated by the compilation memory supervision information is y2.
[0101] Generally speaking, if a query set is obtained, all the query statements in the query set can be run. During the process of running the query statements, it includes converting the query statements in the query set into sample query plans, splitting the sample query plans into sample operator chains, and then executing each sample operator chain according to the vectorized execution mode and the compilation execution mode mentioned above. In this way, the statistic information of each sample operator chain after the splitting of each query statement can be collected, and then the supervision information can be obtained based on the statistic information.
[0102] Step ③: Combine each sample operator chain, the compilation benefit supervision information of each sample operator chain, and the compilation memory supervision information of each sample operator chain to obtain the training sample corresponding to each sample operator chain.
[0103] In a specific implementation, each sample operator chain, the compilation benefit supervision information of each sample operator chain, and the compilation memory supervision information can be directly used as the training samples corresponding to the respective sample operator chains. Thus, the training samples corresponding to a sample operator chain can include: a sample operator chain, the compilation benefit supervision information of the sample operator chain, and the compilation memory supervision information. In this way, each row of data in the training samples can include a sample operator chain and two inputs (the compilation benefit supervision information y1 and the compilation memory supervision information y2 respectively). In another specific implementation, the feature vector of each sample operator chain can be obtained according to each sample operator chain, and then the feature vector of each sample operator chain, the compilation benefit supervision information of each sample operator chain, and the compilation memory supervision information are combined to obtain the training samples corresponding to the respective sample operator chains. Thus, the training samples corresponding to a sample operator chain can include: the feature vector of a sample operator chain, the compilation benefit supervision information of the sample operator chain, and the compilation memory supervision information. In this way, each row of data in the training samples can include a feature vector and two inputs (the compilation benefit supervision information y1 and the compilation memory supervision information y2 respectively). Based on the one-to-one correspondence between the sample operator chains and the training samples, the number of sample operator chains is equal to the number of training samples.
[0104] The construction process of the training samples shown in the above steps ① - ③ obtains the data for supervising the model training and obtains the complete training samples by actually executing the sample query statements, specifically based on the actual processing of the sample operator chains in the vectorized execution mode and the compilation execution mode, and benefiting from the above statistical information, so as to construct a training set including multiple training samples. Then, the training set can be used to train a machine learning model for subsequent compilation benefit estimation and code volume estimation, so as to optimize the execution strategy based on the estimated data. And during the model training process, the data during actual processing can be referred to for supervising the model training, so as to ensure the accuracy of the model training.
[0105] (1.2) Invoke the compilation analysis model to perform compilation analysis processing on the sample operator chain included in the target training sample, and obtain the compilation benefit indication information and the compilation memory indication information of the target training sample.
[0106] In this application, the estimation of compilation benefits and the size of the compiled code can be abstracted as a regression problem: given a feature vector x, predict the benefit value y1 that can be obtained by compilation execution compared to vectorized execution and the code size y2 generated by compilation execution. The benefit value y1 can be defined as the computational time saved by compilation execution compared to vectorized execution, in seconds. For example, y1 = 100 means that compilation execution can save 100 s compared to vectorized execution, and y1 = -10 means that compilation execution is 10 s slower than vectorized execution. y2 represents the memory size occupied by the code segment generated by compilation execution, in bytes. Therefore, the training samples constructed in the embodiments of this application may include the feature vector x of the sample operator chain, and the feature vector x may be obtained through pre-statistics or processed based on the feature extraction ability of the model.
[0107] Specifically, if the target training sample includes the feature vector of the sample operator chain, then the feature vector can be obtained by performing feature extraction processing on the sample operator chain during the construction of the training sample. In another embodiment, if the target training sample does not include the feature vector of the sample operator chain, then during the process of calling the compilation analysis model to perform compilation analysis processing on the sample operator chain, the compilation analysis model can be called to first perform feature extraction processing on the sample operator chain to obtain the feature vector of the sample operator chain, and then perform compilation analysis processing based on the feature vector to obtain the compilation benefit indication information of the sample operator chain and the compilation memory indication information of the sample operator chain. Among them, based on the process of model training, the compilation analysis model can be an initialized compilation analysis model or a compilation analysis model after adjusting the model parameters one or more times. This application does not limit this.
[0108] For the feature extraction processing process of the sample operator chain, the following takes the sample operator chain in a training sample as an example for detailed introduction. Specifically, it may include the following steps 1-step 2. Among them, the training sample to be processed can be any training sample constructed, such as the target training sample.
[0109] Step 1: Perform feature extraction processing on the sample operator chain in the training sample in at least one dimension to obtain the sub-features of the sample operator chain in the training sample in each dimension.
[0110] In the specific processing process, the features of the sample operator chain in the training samples can be statistically analyzed, and a feature vector can be constructed by statistically analyzing the sub-features of each dimension. In a feasible implementation, the sample operator chain in the training samples includes an input and an output. The input of the sample operator chain in the training samples can come from an external data source or a local file read, or from the output of the dependent sample operator chain. After the input is sequentially calculated through the query operations corresponding to each operator in the sample operator chain in the training samples, a calculation result can be finally output, simply referred to as the output.
[0111] The dimensions involved in feature extraction include at least one of the following: input dimension, output dimension, and content dimension. Based on these dimensions, the implementation methods corresponding to step 1 include at least one of the following:
[0112] ① If the dimension includes the input dimension, perform feature extraction processing on the sample operator chain in the training samples under the input dimension to obtain the input sub-features of the sample operator chain in the training samples under the input dimension. The input sub-features include at least one of the following: the number of input rows, the number of input bytes, and the number of columns of each data type included in the input. Optionally, both the number of input rows and the number of input bytes can be estimated by calling the CBO (Cost-Based Optimizer) module of the query optimizer. Each data type included in the input can be a data type supported by the database, such as int and double types. The sample operator chain can contain multiple columns of data of different types. For example, it contains one column of int type data and two columns of double type data, so that the number of columns of the two data types can be statistically obtained as 1 and 2 respectively, and these are all used as input sub-features.
[0113] ② If the dimension includes the output dimension, perform feature extraction processing on the sample operator chain in the training samples under the output dimension to obtain the output sub-features of the sample operator chain in the training samples under the output dimension; the output sub-features include at least one of the following: the number of output rows, the number of output bytes, and the number of columns of each data type included in the output. Similar to the feature extraction type under the input dimension, the number of output rows and the number of output bytes can also be estimated through the CBO module in the query optimizer under the output dimension. Each data type included in the output is a data type supported by the database, such as int and double types. Thus, the number of columns of each data type included in the output can all be used as output sub-features. Assuming that the database supports two types, int and double, and the eigenvalue of the feature vector corresponding to the output sub-features is [1, 2], it means that the input of the sample operator chain contains one column of int and two columns of double.
[0114] ③ If the dimension includes the content dimension, feature extraction processing is performed on the sample operator chain in the training sample under the content dimension to obtain the content sub-features of the sample operator chain in the training sample under the content dimension. The content sub-features include at least one of the following: the number of each query operation included in the sample operator chain, the number of each expression included in the sample operator chain, and the number of each function included in the sample operator chain. The execution layer of the database supports multiple query operations (or called query plan nodes), such as TableScan, Project, Filter, etc., and each query operation will generate corresponding feature values. This is because in a sample operator chain, a query operation may appear once or multiple times, so the number of occurrences of the query operation representing the corresponding type can be counted in the sample operator chain to obtain the number of each query operation. In addition, the sample operator chain may also include different types of expressions and functions, etc. For example, if the database supports three expressions of addition, subtraction, and multiplication, and the sample operator chain contains these three expressions, and the numbers of these three expressions in the sample operator chain are 2, 1, and 3 respectively, then some feature values of the content sub-features can be [2, 1, 3]. Another example is that if the database supports the split function and the sample operator chain contains this split function, then the number of this function can be counted and added to the content sub-features.
[0115] Step 2: Perform fusion processing on the sub-features of the sample operator chain in the training sample under each dimension to obtain the feature vector of the sample operator chain in the training sample.
[0116] In specific implementation, if the feature extraction obtains the sub-features under one dimension, then the sub-features under this dimension can be directly determined as the feature vector of the sample operator chain in the training sample. If the feature extraction obtains the sub-features under two or more dimensions, then the sub-features under multiple dimensions can be fused. Specifically, the sub-features under each dimension can be concatenated to obtain the feature vector of the sample operator chain in the training sample. Exemplarily, if the dimension includes: input dimension, output dimension, and content dimension, then the input sub-features, output sub-features, and content sub-features can be concatenated to obtain the feature vector of the sample operator chain in the training sample. Then the feature vector of the sample operator chain can be composed of the following sub-features: ① The number of input rows and output rows of the sample operator chain. ② The number of input bytes and output bytes of the sample operator chain. Optionally, both ① and ② can be realized by predicting through calling the CBO module in the query optimizer. ③ The number of columns of each data type included in the input of the sample operator chain and the number of columns of each data type included in the output. ④ The number of each query operation, each expression, and each function included in the sample operator chain.
[0117] It can be seen that the feature vectors obtained by the above feature extraction method are obtained by fusing sub-features in one or more dimensions. This can more accurately describe the sample operator chain and is conducive to exploring the relationship between the feature vectors of the sample operator chain and the output data, so as to more accurately predict the duration consumption and memory consumption in the actual compilation process, and provide more accurate guidance for determining the execution mode.
[0118] Further, if the feature vectors of the sample operator chain are included in the target training sample, then the feature vectors can be compiled and analyzed. Through encoding analysis, the compilation benefit indication information and compilation memory indication information of the target training sample can be obtained. The indication information in these two dimensions is actually the prediction information corresponding to the sample operator chain included in the target training sample, and can be used for differential calculation with the supervision information in the corresponding dimensions included in the target training sample. Thus, based on the calculated difference information, the model parameters of the compilation analysis model are adjusted to make the prediction information continuously approach the supervision information during the iterative training process, and further continuously adjust the accuracy of the target compilation analysis model.
[0119] (1.3) Determine the benefit difference corresponding to the target training sample according to the compilation benefit indication information of the target training sample and the compilation benefit supervision information included in the target training sample, and determine the code volume difference corresponding to the target training sample according to the compilation memory indication information of the target training sample and the compilation memory supervision information included in the target training sample.
[0120] In a specific implementation, both the compilation benefit indication information and the compilation memory indication information of the target training sample are prediction data, while the compilation benefit supervision information and the compilation memory supervision information are both supervision data. Through the differential calculation between the prediction data and the supervision data, the difference information in the corresponding dimension can be obtained, and the model is trained based on this difference information. In a specific implementation, the difference between the compilation benefit indicated by the compilation benefit indication information of the target training sample and the compilation benefit indicated by the compilation benefit supervision information can be determined as the benefit difference, and the difference between the code volume indicated by the compilation memory indication information of the target training sample and the code volume indicated by the compilation benefit supervision information can be determined as the code volume difference. Both the code volume difference and the benefit difference can be used to reflect the processing accuracy of the model. The smaller the code volume difference and the benefit difference, the higher the processing accuracy of the model, and the larger the code volume difference and the benefit difference, the lower the processing accuracy of the model.
[0121] (1.4) Train the compilation analysis model according to the benefit difference and the code volume difference corresponding to the target training sample to obtain the target compilation analysis model.
[0122] In a specific implementation, a loss can be constructed based on the revenue difference and code size difference corresponding to the target training samples, and then the parameters of the compilation analysis model can be adjusted in the direction of reducing the loss until the compilation analysis model converges to obtain the target compilation analysis model. The convergence mentioned here can include any one or more of the following: the loss of the compilation analysis model during training reaches the minimum; the loss of the compilation analysis model during training reaches stability, and this loss no longer changes or the change range is less than a preset threshold as the number of training times increases; the training duration of the compilation analysis model reaches the preset training duration; the number of iterative training times of the translation analysis model reaches the preset number of training times; and so on.
[0123] It can be understood that this loss can be used as the backpropagation parameter of the model to adjust the model parameters of the compilation analysis model. One adjustment of the model parameters of the compilation analysis model can be regarded as one training. The process of training the compilation analysis model is an iterative process. Before the compilation analysis model converges, the steps described in (1.1)-(1.4) or (1.2)-(1.4) above can be repeatedly executed, and the training is repeated multiple times. During the training process, the parameters of the compilation analysis model are continuously optimized until the compilation analysis model converges to obtain the target compilation analysis model.
[0124] Based on the above description, a schematic diagram of the model training process as shown in Figure 5 can be provided. As shown in Figure 5 , K (K is a positive integer) training samples are provided. Each training sample corresponds to a sample operator chain. The sample operator chains corresponding to different training samples are different, and each training sample includes the feature vector x of the corresponding sample operator chain, the compilation revenue supervision information y1, and the compilation memory supervision information y2. The K training samples are input into the compilation analysis model. The compilation analysis model can process each feature vector xj (j ∈ [1, K]) to obtain the compilation revenue indication information y1' and the compilation memory indication information y2' of the corresponding sample operator chain. Thus, the difference information can be determined based on the obtained indication information and supervision information, and this difference information is backpropagated to adjust the model parameters of the compilation analysis model. Thus, through continuous iterative adjustment, the target compilation analysis model can finally be obtained.
[0125] The target compilation analysis model can be a compilation analysis model trained using a training set and having better compilation analysis capabilities. With sufficient data provided by the training set, it is beneficial to improve the overall performance of the target compilation analysis model. In the training process shown in the above steps (1.1) - (1.4), the compilation analysis model can be trained based on the difference between the indication information obtained through the compilation analysis process of the sample operator chain and the supervision information corresponding to the sample operator chain. By using the supervision information to supervise and guide the model training, the obtained target compilation analysis model can be made more accurate. After training the target compilation analysis model, the indication information can be predicted using the target compilation analysis model and used for optimizing the execution policy guidance.
[0126] Generally speaking, in an implementable manner, after receiving the query plan sent by the query optimizer, the hybrid execution engine can split the complete query plan into multiple operator chains, and then call the trained target compilation analysis model to estimate the compilation benefit and compilation code size of each operator chain. Among them, since the compilation benefit is relative to the vectorized execution mode in the compilation execution mode, this compilation benefit can also be called the acceleration benefit. A value greater than 0 indicates that the compilation execution is faster than the vectorized execution, a value less than 0 indicates that the compilation execution is slower than the vectorized execution, and a value of 0 indicates that the compilation execution is equal to the vectorized execution. The compilation code size is the code size generated after the predicted compilation execution. Thus, based on the compilation benefit and compilation code size of the operator chain, with the goal of optimizing the execution efficiency of the query plan, the execution mode of each operator chain can be determined as the vectorized execution mode or the compilation execution mode, and the execution mode of the query plan can be optimized.
[0127] The solution provided in the embodiments of this application can be applied to the Starrocks extreme full-scenario MPP analysis platform, and the function of the hybrid execution engine built based on the technical solution of this application is default closed and can be enabled through the following corresponding system parameters.
[0128] / / Enable support for the hybrid execution engine
[0129] set enable_hybrid_pipeline_engine=true;
[0130] In addition, the hybrid execution engine needs to first collect sufficient statistical information and train the underlying machine learning model before it can generate a better execution policy. Therefore, users can collect sufficient statistical information first through the following command:
[0131] / / Collect statistical information of the specified table
[0132] analyze table[table_name];
[0133] / / Enable the collection of query information on the specified table
[0134] alter table [table_name] PROPERTIES ("enable_hybrid_pipeline_profile" = "true");
[0135] After collecting sufficient statistical information, the collection of statistical information can be turned off and the machine model can be trained using the following command:
[0136] / / Disable the collection of query information on the specified table
[0137] alter table [table_name] PROPERTIES ("enable_hybrid_pipeline_profile" = "false");
[0138] / / Train the model
[0139] train hybrid_pipeline_engine_model [model_name] on [table_name];
[0140] After that, the queries passed in by the user can automatically trigger the execution policy optimization. The fusion execution engine can select an appropriate fusion execution policy to execute the user's query and improve the query execution efficiency. For the optimization process of the execution policy, it is specifically the process of determining the execution mode of each operator chain. However, for different query plans, the optimal execution methods may be completely different, and there are many factors affecting the query execution efficiency. How to select the optimal execution policy to maximize the query execution efficiency is an important issue. This mainly involves two difficult problems. One is the selection of the execution policy, that is, how to select any one of vectorized execution, query compilation execution, and fusion execution to make the query efficiency relatively optimal. The other is for the fusion execution policy, how to select the execution mode of each operator chain to make the overall efficiency relatively optimal. To solve this difficult problem, this application can, after obtaining the compilation benefit indication information of each operator chain and the compilation memory indication information of each operator chain, determine the execution mode of each operator chain according to the memory space quota, the compilation benefit indication information of each operator chain, and the compilation memory indication information of each operator chain. Thus, under the given memory limit, a reasonable execution mode for each operator chain can be formulated, and the execution policy of the query plan can be obtained, thereby solving the two difficult problems mentioned above. Executing the query plan according to this execution policy can integrate the advantages of different execution modes, adaptively accelerate each query, and achieve relatively excellent query execution efficiency with the least memory consumption.
[0141] S403. Obtain the memory space quota.
[0142] Among them, the content space quota refers to the maximum memory capacity configured for storing the code generated during the execution of the query plan. That is, it is the maximum amount of memory space allocated for the query plan, or it can be understood as the upper limit value of the storage space for storing the code generated by the operator chain during compilation and execution. Exemplarily, the memory space quota is 10GB, that is, 10GB of memory space can be allocated to store the code generated after the operator chain executes query compilation. This memory space quota can be used to limit the amount of memory space occupied when executing the query plan in units of operator chains, and can avoid low query efficiency caused by occupying too much memory space.
[0143] S404. Divide multiple operator chains into a first operator chain set and a second operator chain set according to the memory space quota, the compilation benefit indication information of each operator chain, and the compilation memory indication information of each operator chain, where the execution mode associated with the first operator chain set is the compilation execution mode, and the execution mode associated with the second operator chain set is the vectorization execution mode.
[0144] In a specific implementation, the compilation benefit indication information is relative to the vectorization execution mode for the compilation execution mode, the compilation memory indication information is obtained in the compilation execution mode. The execution mode associated with the first operator chain set is the compilation execution mode, which is equivalent to that each operator chain in the first operator chain set can adopt the compilation execution mode. The execution mode associated with the second operator chain set is the vectorization execution mode, which is equivalent to that the operator chains in the second operator chain set can adopt the vectorization execution mode. One of the first operator chain set and the second operator chain set may be empty. If the first operator chain set is empty, all the split operator chains are divided into the second operator chain set, so that each operator chain adopts the vectorization execution mode. If the second operator chain set is empty, all the split operator chains are divided into the first operator chain set, so that each operator chain adopts the compilation execution mode. If both the first operator chain and the second operator chain are non-empty sets, then the first operator chain set may include some operator chains, and the second operator chain set may include some operator chains. Then the execution strategy of the query plan is a fusion execution strategy, that is, the operator chains in the first operator chain set adopt the compilation execution mode, and the operator chains in the second operator chain set adopt the vectorization execution mode.
[0145] Suppose the number of operator chains obtained after query plan splitting is N. The compilation benefits indicated by the compilation benefit indication information of each operator chain can be denoted as t[i], and the volume of the compiled code indicated by the compilation memory indication information is v[i], where i ∈ [1, N]. To control the memory usage, a memory space quota is given here, requiring that the volume of the code generated after compilation during the execution of the query plan is less than or equal to this memory space quota V. The optimization goal of the execution strategy is to maximize the overall compilation benefit of this query plan. Based on this, the problem of determining the execution mode of operator chains can be abstracted into a 0-1 knapsack problem. The Knapsack Problem is a classic combinatorial optimization problem. Its goal is to select the optimal combination of items under the given weights and values of a set of items and the limited capacity of the knapsack, so that the total value of the items in the knapsack is maximized. The knapsack problem has the following characteristics:
[0146] Input: A set of items, each item having a corresponding weight and value; the upper limit of the capacity of the knapsack.
[0147] Output: The combination of items selected to maximize the total value of the items in the knapsack.
[0148] Constraints: The sum of the weights of the selected items cannot exceed the upper limit of the capacity of the knapsack. Each item is selected one and cannot be divided.
[0149] For this application, each operator chain (Pipeline) can be regarded as an item, and their code volume (memory occupancy) and compilation benefits correspond to the weight and value of the knapsack respectively. The goal is to select some operator chains and put them into the knapsack under the given memory limit V to maximize the sum of their compilation benefits. That is:
[0150] Input: A set of Pipelines, each Pipeline having a corresponding volume v[i] and value t[i]; the upper limit of the memory limit V.
[0151] Output: The combination of Pipelines selected to maximize the total value.
[0152] Constraints: The sum of the volumes of the selected Pipelines cannot exceed the preset upper limit of the memory capacity. Each Pipeline is selected for compilation or not.
[0153] For the solution to the above problem, a dynamic programming (DP) algorithm can be used to generate a feasible solution to finally determine whether each operator chain performs query compilation, so that after the operator chains to be compiled are compiled into machine code, each operator chain can be scheduled to execute the corresponding machine code concurrently according to the dependency relationship. Corresponding to step S404 above, since code will be generated only after the operator chain is compiled and executed, there is a corresponding code volume. Then: The operator chains put into the backpack (i.e., the feasible solutions) are all the operator chains in the first operator chain set, so each operator chain in the first operator chain set adopts the compilation execution mode. For the operator chains not put into the backpack, they are all the operator chains in the second operator chain set, and each operator chain in the second operator chain set adopts the vectorization execution mode.
[0154] In one implementation, the computer device can execute the above S404 specifically according to the following steps s21 - step s24.
[0155] Step s21: Based on the memory space quota and the compilation memory indication information of each operator chain, generate at least one operator chain set from multiple operator chains.
[0156] Each operator chain set includes at least one operator chain. If the number of operator chain sets includes two or more, the same operator chain may exist between different operator chain sets. For example, there may be one same operator chain between operator chain set A and operator chain set B. The memory space occupancy corresponding to the sub-chain set is less than the memory space quota, and the memory space occupancy is determined based on the memory space amount required to be occupied as indicated by the compilation memory indication information of the operator chains included in the corresponding operator chain set. This is because each operator chain corresponds to compilation memory indication information, and this compilation memory indication information is used to indicate the memory space amount required for the code generated after the corresponding operator chain is compiled in the compilation execution mode. Thus, for any operator chain set, by summing up the memory space amounts required to be occupied as indicated by the compilation memory indication information of the operator chains included in the operator chain set, the memory space occupancy corresponding to this operator chain set can be obtained. For example, operator chain set S1 includes 4 operator chains, and the memory space amounts required to be occupied as indicated by the corresponding compilation memory indication information are {p1, p2, p3, p4}, then the memory space occupancy corresponding to this operator chain set = p1 + p2 + p3 + p4.
[0157] Step s22: Determine the total compilation revenue value of each operator chain set according to the compilation revenue indication information of the operator chains included in each operator chain set.
[0158] Based on the compilation benefits indicated by the compilation benefit indication information for each operator chain, that is, the compilation benefits obtained in the vectorized execution mode relative to the compilation execution mode, for a set of operator chains, by summing up the compilation benefits indicated by the compilation benefit indication information of each operator chain included in the set of operator chains, the total compilation benefit value of the set of operator chains can be obtained. The total compilation benefit value can be used to reflect the total difference between the total duration of each operator chain in the set of operator chains in the compilation execution mode and the total duration in the vectorized execution mode, and can specifically be used to determine the execution mode adopted by the operator chains in the set of operator chains.
[0159] Step S23: According to the total compilation benefit value of each determined set of operator chains, select the target set of operator chains with the largest total compilation benefit value from at least one set of operator chains.
[0160] In one implementation manner, if the number of sets of operator chains includes multiple, the computer device can sort each set of operator chains based on the total compilation benefit value of each set of operator chains. Based on the sorted result, the set of operator chains with the largest total compilation benefit value can be selected relatively quickly as the target set of operator chains. Specifically, if sorted in descending order, the first sorted set of operator chains can be selected as the target set of operator chains; if sorted in ascending order, the last sorted set of operator chains can be selected as the target set of operator chains. If the number of sets of operator chains only includes one, then there is only one set of operator chains in at least one set of operator chains, and this set of operator chains can be directly determined as the target set of operator chains.
[0161] Step S24: Determine the target set of operator chains as the first set of operator chains, and obtain the second set of operator chains according to the operator chains among the multiple operator chains except the operator chains in the target set of operator chains.
[0162] In one embodiment, the target set of operator chains can be directly determined as the first set of operator chains, and the operator chains among the multiple operator chains except the operator chains in the target set of operator chains can form the second set of operator chains.
[0163] In another embodiment, the total compilation benefit may be positive or negative. If the total compilation benefit of the set of target operator chains is positive, it indicates that the total execution duration of each operator chain in the set of operator chains with the maximum compilation benefit in the compilation execution mode is shorter than that in the vectorization execution mode. Each operator chain in the set of operator chains can save more time in the compilation execution mode compared to the vectorization execution mode, and the memory space occupancy corresponding to the set of target operator chains can also be controlled within the preset memory space quota. Thus, step S24 can be triggered to obtain the first set of operator chains and the second set of operator chains. Conversely, if the total compilation benefit of the set of target operator chains is negative, it indicates that the total execution duration of each operator chain in the set of operator chains with the maximum compilation benefit in the compilation execution mode is longer than that in the vectorization execution mode. Since the maximum total compilation benefit is negative, it means that the total compilation benefits of other sets of operator chains are also negative. Thus, each operator chain obtained by splitting can be directly determined as an operator chain in the second set of operator chains.
[0164] In another embodiment, the specific implementation of S404 may also include the following steps S25 - S26.
[0165] Step S25: If the total compilation benefits of all sets of operator chains are negative, or the amount of memory space required indicated by the compilation memory indication information of each operator chain is greater than the memory space quota, then each operator chain is determined as an operator chain in the second set of operator chains.
[0166] That is to say, ① if the total compilation benefits of the set of operator chains are all negative, it indicates that the total execution duration spent by each operator chain in the set of operator chains in the vectorization execution mode is shorter than that in the compilation execution mode. That is, for the set of operator chains, a certain amount of time can be saved as a whole. Thus, each operator chain can be determined as an operator chain in the second set of operator chains. ② If the amount of memory space required indicated by the compilation memory indication information of each operator chain is greater than the memory space quota, it indicates that no matter how these operator chains are combined, they will be greater than the memory space quota. Thus, these operator chains can be directly determined as operator chains in the second set of operator chains.
[0167] In a feasible embodiment, the compilation benefit indicated by the compilation benefit indication information of any sample operator chain may be positive or negative. When it is positive, it indicates that the compilation execution mode saves time compared to the vectorization execution mode. When it is negative, it indicates that the vectorization execution mode saves time compared to the compilation execution mode. If the compilation benefits indicated by the compilation benefit indication information of each operator chain are all negative, then no matter how the operator chains are combined, the total compilation benefit value of the combined operator chain set is also negative, which also indicates that each operator chain can save time when executed in the vectorization execution mode compared to the compilation execution mode. Thus, each operator chain can be determined as an operator chain in the second operator chain set and associated with the vectorization execution mode. If the compilation benefit indicated by the compilation benefit indication information of an operator chain is negative, the operator chain can also be directly determined as an operator chain in the second operator chain set.
[0168] Step S26: If the compilation benefit values indicated by the compilation benefit indication information of each operator chain are all positive, and the sum value of the memory space amounts required to be occupied indicated by the compilation memory indication information of each operator chain does not exceed the memory space quota, then each operator chain is determined as an operator chain in the first operator chain set.
[0169] Specifically, when the compilation benefit values indicated by the compilation benefit indication information of each operator chain are all positive, it indicates that the compilation execution mode takes less time than the vectorization execution mode. However, due to the limitation of the memory space quota, it is also necessary to determine the total memory space occupation according to the memory space amounts required to be occupied indicated by the compilation memory indication information of each operator chain. When the sum value of the content space amounts occupied by the code generated after all operator chains are compiled does not exceed the memory space quota, it indicates that no matter how the operator chains are combined, the memory space occupation amount corresponding to the operator chain set can be controlled within the memory space quota. Thus, when the sum value of this memory space occupation is less than or equal to the memory space quota, each operator chain can be determined as an operator chain in the first operator chain set, that is, the first operator chain set is directly obtained from multiple operator chains.
[0170] Based on the above division of the first operator chain set and the second operator chain set, and the execution modes associated with each operator chain set, the execution mode of each operator chain can be determined, and thus step S405 can be executed.
[0171] S405: Execute the corresponding operator chain according to the determined execution mode of each operator chain to complete the query plan.
[0172] Based on the above, the execution mode of an operator chain is either the vectorized execution mode or the compilation execution mode. Therefore, when the computer device executes S405, it may specifically include any one of the following situations: ① If the execution mode of each determined operator chain is the compilation execution mode, query compilation processing is performed on each operator chain according to the compilation execution mode; ② If the execution mode of each determined operator chain is the vectorized execution mode, vectorized execution processing is performed on each operator chain according to the vectorized execution mode; ③ If the execution mode of at least one first operator chain among multiple operator chains is the compilation execution mode and the execution mode of at least one second operator chain is the vectorized execution mode, query compilation processing is performed on each first operator chain according to the compilation execution mode, and vectorized execution processing is performed on each second operator chain according to the vectorized execution mode; where the first operator chain refers to the operator chain from the first operator chain set, and the second operator chain refers to the operator chain from the second operator chain set.
[0173] Based on the data processing method introduced in the above embodiments, an exemplary schematic diagram of a model training process and a policy decision process can be provided as Figure 6 shown. As Figure 6 shown, in the model training stage, statistical information can be collected, and a compilation analysis model can be trained to obtain a trained compilation analysis model, that is, the target compilation analysis model. After the model is trained, when the execution engine receives a query plan, it can first split the query plan into multiple operator chains (pipelines) to obtain a list of operator chains, and these operator chains in the list can be defaulted to execute in the vectorized mode first. Then, the target compilation analysis model can be called to perform compilation analysis processing on each operator chain, estimate the compilation benefits that each operator chain can generate when executed in the compilation execution mode relative to the vectorized mode, and estimate the amount of memory space that the code generated after compilation of each operator chain in the compilation execution mode needs to occupy. Furthermore, in combination with the compilation benefits and memory consumption, it can be determined which operator chains adopt the compilation execution mode, and the remaining operator chains adopt the default vectorized execution mode.
[0174] The data processing method provided in this embodiment can split a query plan into an operator chain, determine the execution mode based on the operator chain as a unit, and thus determine an execution strategy that is beneficial to improving the query efficiency, realizing an efficient query. Specifically: ① For each operator chain, by predicting the compilation benefit of the operator chain, the speed of vectorized execution and compilation execution can be compared, providing reasonable and quantifiable data for reference in determining the execution mode. ② Predict the amount of memory space required for the code generated after compiling each operator chain. In this way, under the given memory limit, based on the compilation benefit and the volume of the compiled code, an execution mode that makes the execution efficiency of the overall query plan higher can be determined for each operator chain, balancing the compilation time overhead and the storage space overhead, and further improving the execution efficiency of the query. ③ Not limited to a specific operator chain, for each operator chain obtained by splitting the query plan, an execution mode can be determined, so that the execution of the query plan can flexibly integrate multiple execution modes, rather than being limited to only one execution mode. This is an adaptive fusion execution strategy determined according to the specific situation, which can achieve a higher query execution efficiency. ④ An adaptive fusion execution engine based on vectorized execution and query compilation can be designed. By constructing an AI model for estimating the compilation benefit in the vectorized engine, the actual benefit of any subtree in the compiled query tree can be estimated, and a corresponding optimization algorithm is designed to adaptively select the query execution strategy that maximizes the overall benefit, and finally generate a pure vectorized, pure query compilation, or fusion execution plan to maximize the query analysis performance. In this way, the advantages of both vectorized execution and query compilation technologies can be taken into account, and the optimal execution method can be adaptively selected to maximize the query analysis performance.
[0175] In one implementation, the parallel scheduling process of the operator chain can also be implemented with the help of a thread pool to improve the execution efficiency of the operator chain and manage the ordered execution of the operator chain. Specifically, the following steps can be executed: First, according to the dependency relationship between multiple operator chains, at least one operator chain among the multiple operator chains is scheduled into the thread pool to execute the operator chain in the thread pool according to the execution mode of the corresponding operator chain; during the execution of the operator chain, if it is determined that an exception occurs in the target operator chain, the target operator chain is removed from the thread pool.
[0176] In a specific implementation, based on the dependency relationships among multiple operator chains, operator chains that support parallel execution and operator chains that support serial execution can be selected from multiple operator chains. Specifically, if there is no dependency relationship between any two operator chains, then these two operator chains can be determined as operator chains that support parallel execution. Conversely, if one operator chain needs to wait for the output result of another operator chain that has been executed before it can execute depending on the output result of the executed operator chain, then these two operator chains can be determined as operator chains that support serial execution. Subsequently, the operator chains that support parallel execution can be scheduled into the thread pool first, and the operator chains scheduled into the thread pool can be executed in the thread pool according to the determined execution mode. Exemplarily, as Figure 3 shown, operator chain P4 and operator chain P2 support parallel execution. Thus, these two operator chains can be scheduled into the thread pool first, and operator chain P4 performs query compilation according to the compilation execution mode, while operator chain P2 performs vectorization execution according to the vectorization execution mode.
[0177] During the execution of the operator chains, there may be one or more operator chains that encounter exceptions for some reasons. The target operator chain refers to one of the at least one operator chains that have been scheduled into the thread pool. An exception occurring in the target operator chain means that: the operator chain on which the target operator chain depends has not been executed, or the execution duration of the target operator chain exceeds the duration threshold. On the one hand, if the target operator chain depends on other operator chains and the operator chains it depends on have not been executed, then the target operator chain will be blocked due to the lack of upstream output, resulting in an exception in the target operator chain. On the other hand, a duration threshold can be set corresponding to each operator chain, so as to monitor whether the target operator chain is executed normally based on the duration threshold. If the execution duration of the target operator chain is greater than or equal to the duration threshold, that is, the execution of a single operator chain exceeds the predetermined duration, then it is also considered that an exception has occurred in the target operator chain.
[0178] To prevent the target operator chain with an exception from occupying resources in the thread pool but not being processed smoothly, in the case of an exception, the target operator chain with the exception can be temporarily removed from the thread pool, so that other operator chains can be scheduled into the thread pool and run in the thread pool. Further, in the case where the target operator chain is removed from the thread pool due to an exception, the target operator chain can also be rescheduled into the thread pool for execution under corresponding conditions. Based on the type of exception situation, it can specifically include the content described in (1) and (2) below.
[0179] (1) When an exception occurs in the target operator chain, which means that the operator chain on which the target operator chain depends has not been executed, after waiting for the operator chain in the thread pool that has a dependency relationship with the corresponding operator chain to be executed, the target operator chain is scheduled to the thread pool. In this way, the target operator chain is removed from the thread pool because the target operator chain lacks the output of the dependent operator chain. Therefore, after the dependent operator chain is executed, the upstream output required by the target operator chain can be obtained, and then the target operator chain can be scheduled to the thread pool, and this upstream output can be used as the input of the target operator chain to execute the target operator chain.
[0180] (2) When an exception occurs in the target operator chain, which means that the execution duration of the target operator chain exceeds the duration threshold, when the waiting duration meets the waiting condition, the target operator chain is scheduled to the thread pool. In this way, the target operator chain is removed from the thread pool because the execution duration of the target operator chain exceeds the duration threshold. Therefore, the waiting duration after the target operator chain is removed from the thread pool can be counted and compared with the set waiting duration threshold. Then, the waiting duration meeting the waiting condition can include: the waiting duration reaches the waiting duration threshold. If the waiting duration threshold is equal to the duration threshold, then when the waiting duration reaches the duration threshold, an operator chain in the thread pool may also be executed or removed from the thread pool due to some abnormal reasons, so there are sufficient thread resources, and thus the target operator chain can be scheduled to the thread pool and the target operator chain can be re-executed in the thread pool.
[0181] In the above embodiments, based on the dependency relationship between the operator chains (Pipelines) obtained by splitting, multiple operator chains can be scheduled in parallel. Specifically, each operator chain (including the operator chain for vectorized execution or the operator chain for compilation execution) can be alternately scheduled to the thread pool to run. When a single Pipeline exceeds the predetermined time or is blocked due to lack of upstream output, the Pipeline is temporarily removed from the thread pool. After other Pipelines that need to be executed are completed, this Pipeline is continued to be scheduled to run.
[0182] Based on the data processing method provided in the embodiments of the present application, a fusion execution engine can be provided, which has the following advantages: ① Accelerate the execution of query statements and improve the efficiency of complex query analysis. Through practice, by testing on the TPCH standard test set, the fusion execution engine applying the solution provided in the present application performs excellently. Compared with a pure vectorized execution engine, it can improve the execution efficiency by up to 2 to 7 times at most. For AP scenarios such as decision support, it can significantly speed up and provide users with a faster data query and analysis experience. ② Efficiently utilize resources and reduce hardware costs. Since the fusion execution engine reasonably allocates computing resources and optimizes the execution plan, this can effectively improve the hardware utilization rate and reduce hardware costs. And this enables the business side to use fewer resources when purchasing servers and storage devices to achieve higher computing performance, helping the business side achieve cost reduction and efficiency improvement. ③ Real-time data analysis and accelerate business decisions. The fusion execution engine supports real-time data analysis, can quickly respond to business requirements, and provide key data support for relevant business sides. This will help the business side make business decisions faster and improve the overall operation efficiency.
[0183] In the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0184] Please refer to Figure 7 , Figure 7 FIG. is a schematic structural diagram of a data processing device provided in the embodiments of the present application. This data processing device can be set in the computer device provided in the embodiments of the present application. Figure 7 The data processing device shown can be a computer program (including program code) running in a computer device. This data processing device can be used to execute Figure 2 or Figure 4 Some or all of the steps in the method embodiments shown. Please refer to Figure 7 , and this data processing device can include the following units:
[0185] An acquisition unit 701, configured to acquire a query plan to be processed, and perform splitting processing on the query plan to obtain a plurality of operator chains;
[0186] The processing unit 702 is configured to perform compilation analysis processing on each of the multiple operator chains to obtain compilation benefit indication information for each operator chain and compilation memory indication information for each operator chain; the compilation memory indication information is used to indicate the amount of memory space required for the code generated after the corresponding operator chain is compiled.
[0187] The processing unit 702 is further configured to determine the execution mode of each operator chain according to the compilation benefit indication information for each operator chain and the compilation memory indication information for each operator chain.
[0188] The processing unit 702 is further configured to execute the corresponding operator chain according to the determined execution mode of each operator chain to complete the query plan.
[0189] In one embodiment, the execution modes include a vectorized execution mode and a compilation execution mode; when determining the execution mode of each operator chain according to the compilation benefit indication information for each operator chain and the compilation memory indication information for each operator chain, the processing unit 702 is specifically configured to: obtain a memory space quota; the memory space quota refers to the memory capacity configured for storing the code generated during the execution of the query plan; divide the multiple operator chains into a first operator chain set and a second operator chain set according to the memory space quota, the compilation benefit indication information for each operator chain, and the compilation memory indication information for each operator chain; wherein, the execution mode associated with the first operator chain set is the compilation execution mode, and the execution mode associated with the second operator chain set is the vectorized execution mode.
[0190] In one embodiment, when dividing the multiple operator chains into a first operator chain set and a second operator chain set according to the memory space quota, the compilation benefit indication information for each operator chain, and the compilation memory indication information for each operator chain, the processing unit 702 is specifically configured to: generate at least one operator chain set according to the memory space quota and the compilation memory indication information for each operator chain; wherein, each operator chain set includes at least one of the multiple operator chains, and the memory space occupancy corresponding to each operator chain set is less than the memory space quota, and the memory space occupancy is determined based on the memory space amount indicated by the compilation memory indication information of the operator chains included in the corresponding operator chain set; determine the total compilation benefit of each operator chain set according to the compilation benefit indication information of the operator chains included in each operator chain set; select the target operator chain set with the largest total compilation benefit from the at least one operator chain set according to the determined total compilation benefit of each operator chain set; determine the target operator chain set as the first operator chain set, and obtain the second operator chain set according to the operator chains among the multiple operator chains other than the operator chains in the target operator chain set.
[0191] In one embodiment, the processing unit 702 is further configured to: if the total compilation benefit value of each operator chain set is negative, or the amount of memory space required indicated by the compilation memory indication information of each operator chain is greater than the memory space quota, then determine each operator chain as an operator chain in the second operator chain set; if the compilation benefit values indicated by the compilation benefit indication information of each operator chain are all positive, and the sum of the amounts of memory space required indicated by the compilation memory indication information of each operator chain does not exceed the memory space quota, then determine each operator chain as an operator chain in the first operator chain set.
[0192] In one embodiment, when the processing unit 702 performs compilation analysis processing on each operator chain in a plurality of operator chains to obtain the compilation benefit indication information of each operator chain and the compilation memory indication information of each operator chain, it is specifically configured to: call a target compilation analysis model to perform compilation analysis processing on each operator chain in the plurality of operator chains to obtain the compilation benefit value of each operator chain and the compilation code metric information of each operator chain; wherein, the compilation benefit value refers to the difference between the execution duration required for the corresponding operator chain in the vectorized execution mode and the execution duration required in the compilation execution mode; the compilation code metric information includes: the code volume to be generated after the corresponding operator chain is compiled and executed in the compilation execution mode; determine the compilation benefit value of each operator chain as the compilation benefit indication information of the corresponding operator chain, and determine the compilation code metric information of each operator chain as the compilation memory indication information of the corresponding operator chain.
[0193] In one embodiment, the processing unit 702 is further configured to: obtain a target training sample; the target training sample includes a sample operator chain and training supervision information; the training supervision information includes: the compilation benefit supervision information of the sample operator chain and the compilation memory supervision information of the sample operator chain; call the compilation analysis model to perform compilation analysis processing on the sample operator chain included in the target training sample to obtain the compilation benefit indication information and the compilation memory indication information of the target training sample; determine the benefit difference corresponding to the target training sample according to the compilation benefit indication information of the target training sample and the compilation benefit supervision information included in the target training sample; determine the code volume difference corresponding to the target training sample according to the compilation memory indication information of the target training sample and the compilation memory supervision information included in the target training sample; train the compilation analysis model according to the benefit difference and the code volume difference corresponding to the target training sample to obtain the target compilation analysis model.
[0194] In one embodiment, the obtaining unit 701 is further configured to: obtain a sample query statement, and generate a plurality of sample operator chains corresponding to the sample query statement according to the sample query statement;
[0195] The processing unit 702 is further configured to: execute each sample operator chain according to the vectorized execution mode and the compilation execution mode respectively, so as to obtain the compilation benefit supervision information of each sample operator chain and the compilation memory supervision information of each sample operator chain; and combine each sample operator chain, the compilation benefit supervision information of each sample operator chain and the compilation memory supervision information of each sample operator chain to obtain a training sample corresponding to each sample operator chain.
[0196] In one embodiment, when the processing unit 702 executes each sample operator chain according to the vectorized execution mode and the compilation execution mode respectively to obtain the compilation benefit supervision information of each sample operator chain and the compilation memory supervision information of each sample operator chain, it is specifically configured to: execute each sample operator chain according to the vectorized execution mode to obtain the first compilation statistic information of each sample operator chain; and execute each sample operator chain according to the compilation execution mode to obtain the second compilation statistic information of each sample operator chain and the reference code metric information; the reference code metric information includes: the code volume to be generated after the corresponding sample operator chain is compiled and executed in the compilation execution mode; determine the compilation benefit supervision information of each sample operator chain according to the first compilation statistic information and the second compilation statistic information of each sample operator chain; and determine the reference code metric information of each sample operator chain as the compilation memory supervision information of the corresponding sample operator chain.
[0197] In one embodiment, the processing unit 702 is further configured to: perform feature extraction processing on the sample operator chains in the training sample in at least one dimension to obtain sub-features of the sample operator chains in the training sample in each dimension; and perform fusion processing on the sub-features of the sample operator chains in the training sample in each dimension to obtain a feature vector of the sample operator chains in the training sample, where the feature vector is used for the target compilation analysis model to perform compilation analysis processing on the sample operator chains in the training sample.
[0198] In one embodiment, the sample operator chain in the training sample includes an input and an output; the dimensions include at least one of the following: input dimension, output dimension, and content dimension. The processing unit 702, when performing feature extraction processing on the sample operator chain in the training sample under at least one dimension to obtain sub-features of the sample operator chain in the training sample under each dimension, is specifically used for at least one of the following: If the dimension includes the input dimension, perform feature extraction processing on the sample operator chain in the training sample under the input dimension to obtain the input sub-feature of the sample operator chain in the training sample under the input dimension; the input sub-feature includes any one or more of: the number of input rows, the number of input bytes, and the number of columns of each data type included in the input; or, if the dimension includes the output dimension, perform feature extraction processing on the sample operator chain in the training sample under the output dimension to obtain the output sub-feature of the sample operator chain in the training sample under the output dimension; the output sub-feature includes any one or more of: the number of output rows, the number of output bytes, and the number of columns of each data type included in the output; or, if the dimension includes the content dimension, perform feature extraction processing on the sample operator chain in the training sample under the content dimension to obtain the content sub-feature of the sample operator chain in the training sample under the content dimension, and the content sub-feature includes any one or more of: the number of each query operation included in the sample operator chain, the number of each expression included in the sample operator chain, and the number of each function included in the sample operator chain.
[0199] In one embodiment, the processing unit 702 is further configured to: according to the dependency relationship between multiple operator chains, schedule at least one operator chain among the multiple operator chains to a thread pool to execute the operator chain in the thread pool according to the execution mode of the corresponding operator chain; during the execution of the operator chain, if it is determined that an exception occurs in the target operator chain, remove the target operator chain from the thread pool; wherein, the target operator chain refers to one operator chain among the multiple operator chains, and an exception occurs in the target operator chain means that: the operator chain on which the target operator chain depends has not been executed, or the execution duration of the target operator chain exceeds the duration threshold.
[0200] In one embodiment, the processing unit 702 is further configured to: in the case where an exception occurs in the target operator chain means that the operator chain on which the target operator chain depends has not been executed, after waiting for the operator chain having a dependency relationship with the target operator chain in the thread pool to be executed, schedule the target operator chain to the thread pool; or, in the case where an exception occurs in the target operator chain means that the execution duration of the target operator chain exceeds the duration threshold, when the waiting duration meets the waiting condition, schedule the target operator chain to the thread pool.
[0201] In one embodiment, the query plan is a query tree, the query tree includes multiple nodes, and each node is used to represent a query operation in the query plan; when the processing unit 702 performs splitting processing on the query plan to obtain multiple operator chains, it is specifically used for: starting from the leaf nodes of the query tree, traversing, and performing blocking analysis on the query operation represented by the currently traversed node; if it is analyzed that the query operation represented by the currently traversed node is blocked, then split out an operator chain according to the nodes included in the path between the leaf node and the current node; if the query operation represented by the next node of the currently traversed node is blocked, then split out an operator chain according to the current node and the next node of the current node; after traversing the query tree, obtain multiple operator chains.
[0202] It can be understood that the specific functions of the various units of the data processing device described in the embodiments of the present application can be specifically implemented according to the methods in the above method embodiments, and the specific implementation process can refer to the relevant descriptions of the above method embodiments, which will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either.
[0203] Next, the computer device provided in the embodiments of the present application will be elaborated.
[0204] The embodiments of the present application also provide a structural schematic diagram of a computer device, and the structural schematic diagram of the computer device can be seen Figure 8 ; the computer device may include: a processor 801, an input interface 802, an output interface 803, and a memory 804. The above-mentioned processor 801, input interface 802, output interface 803, and memory 804 are connected through a bus. The memory 804 is used to store a computer storage medium, and a computer program is stored in the computer storage medium. The computer program includes program instructions, and the processor 801 is used to execute the program instructions stored in the memory 804.
[0205] In one embodiment, the processor 801 performs the following operations by running the computer program in the memory 804: obtaining a query plan to be processed, and performing splitting processing on the query plan to obtain multiple operator chains; performing compilation analysis processing on each operator chain in the multiple operator chains to obtain compilation benefit indication information for each operator chain and compilation memory indication information for each operator chain; the compilation memory indication information is used to indicate the amount of memory space required for the code generated after the corresponding operator chain is compiled; determining the execution mode of each operator chain according to the compilation benefit indication information of each operator chain and the compilation memory indication information of each operator chain; executing the corresponding operator chain according to the determined execution mode of each operator chain to complete the query plan.
[0206] In one embodiment, the execution modes include a vectorized execution mode and a compiled execution mode; when determining the execution mode of each operator chain according to the compilation benefit indication information of each operator chain and the compilation memory indication information of each operator chain, the processor 801 is specifically configured to: obtain a memory space quota; the memory space quota refers to the memory capacity configured for storing the code generated during the execution of the query plan; divide the multiple operator chains into a first operator chain set and a second operator chain set according to the memory space quota, the compilation benefit indication information of each operator chain, and the compilation memory indication information of each operator chain; wherein, the execution mode associated with the first operator chain set is the compiled execution mode, and the execution mode associated with the second operator chain set is the vectorized execution mode.
[0207] In one embodiment, when dividing the multiple operator chains into a first operator chain set and a second operator chain set according to the memory space quota, the compilation benefit indication information of each operator chain, and the compilation memory indication information of each operator chain, the processor 801 is specifically configured to: generate at least one operator chain set according to the memory space quota and the compilation memory indication information of each operator chain; wherein, each operator chain set includes at least one operator chain among the multiple operator chains, and the memory space occupancy corresponding to each operator chain set is less than the memory space quota, and the memory space occupancy is determined based on the memory space amount indicated by the compilation memory indication information of the operator chains included in the corresponding operator chain set; determine the total compilation benefit value of each operator chain set according to the compilation benefit indication information of the operator chains included in each operator chain set; select, from the at least one operator chain set, a target operator chain set with the largest total compilation benefit value according to the determined total compilation benefit value of each operator chain set; determine the target operator chain set as the first operator chain set, and obtain the second operator chain set according to the operator chains other than the operator chains in the target operator chain set among the multiple operator chains.
[0208] In one embodiment, the processor 801 is further configured to: if the total compilation benefit values of all operator chain sets are negative, or the memory space amounts required to be occupied indicated by the compilation memory indication information of each operator chain are all greater than the memory space quota, then determine all operator chains as the operator chains in the second operator chain set; if the compilation benefit values indicated by the compilation benefit indication information of all operator chains are positive, and the sum of the memory space amounts required to be occupied indicated by the compilation memory indication information of all operator chains does not exceed the memory space quota, then determine all operator chains as the operator chains in the first operator chain set.
[0209] In one embodiment, when the processor 801 performs compilation analysis processing on each operator chain in a plurality of operator chains to obtain the compilation benefit indication information of each operator chain and the compilation memory indication information of each operator chain, it is specifically configured to: call a target compilation analysis model to perform compilation analysis processing on each operator chain in the plurality of operator chains to obtain the compilation benefit value of each operator chain and the compilation code metric information of each operator chain; wherein, the compilation benefit value refers to the difference between the execution duration required for the corresponding operator chain in the vectorized execution mode and the execution duration required in the compilation execution mode; the compilation code metric information includes: the code volume to be generated after the corresponding operator chain is compiled and executed in the compilation execution mode; determine the compilation benefit value of each operator chain as the compilation benefit indication information of the corresponding operator chain, and determine the compilation code metric information of each operator chain as the compilation memory indication information of the corresponding operator chain.
[0210] In one embodiment, the processor 801 is further configured to: obtain a target training sample; the target training sample includes a sample operator chain and training supervision information; the training supervision information includes: the compilation benefit supervision information of the sample operator chain and the compilation memory supervision information of the sample operator chain; call the compilation analysis model to perform compilation analysis processing on the sample operator chain included in the target training sample to obtain the compilation benefit indication information and the compilation memory indication information of the target training sample; determine the benefit difference corresponding to the target training sample according to the compilation benefit indication information of the target training sample and the compilation benefit supervision information included in the target training sample; determine the code volume difference corresponding to the target training sample according to the compilation memory indication information of the target training sample and the compilation memory supervision information included in the target training sample; train the compilation analysis model according to the benefit difference and the code volume difference corresponding to the target training sample to obtain the target compilation analysis model.
[0211] In one embodiment, the processor 801 is further configured to: obtain a sample query statement, and generate a plurality of sample operator chains corresponding to the sample query statement according to the sample query statement; execute each sample operator chain respectively in the vectorized execution mode and the compilation execution mode to obtain the compilation benefit supervision information of each sample operator chain and the compilation memory supervision information of each sample operator chain; combine each sample operator chain, the compilation benefit supervision information of each sample operator chain and the compilation memory supervision information of each sample operator chain to obtain a training sample corresponding to each sample operator chain.
[0212] In one embodiment, when the processor 801 executes each sample operator chain in the vectorized execution mode and the compilation execution mode respectively to obtain the compilation benefit supervision information of each sample operator chain and the compilation memory supervision information of each sample operator chain, it is specifically configured to: execute each sample operator chain in the vectorized execution mode to obtain the first compilation statistic information of each sample operator chain; and execute each sample operator chain in the compilation execution mode to obtain the second compilation statistic information of each sample operator chain and the reference code metric information; the reference code metric information includes: the code volume that needs to be generated after the corresponding sample operator chain is compiled and executed in the compilation execution mode; determine the compilation benefit supervision information of each sample operator chain according to the first compilation statistic information and the second compilation statistic information of each sample operator chain; and determine the reference code metric information of each sample operator chain as the compilation memory supervision information of the corresponding sample operator chain.
[0213] In one embodiment, the processor 801 is further configured to: perform feature extraction processing on the sample operator chains in the training samples in at least one dimension to obtain sub-features of the sample operator chains in the training samples in each dimension; perform fusion processing on the sub-features of the sample operator chains in the training samples in each dimension to obtain a feature vector of the sample operator chains in the training samples, and the feature vector is used for the target compilation analysis model to perform compilation analysis processing on the sample operator chains in the training samples.
[0214] In one embodiment, the sample operator chains in the training samples include inputs and outputs; the dimensions include at least one of the following: input dimension, output dimension, and content dimension. When the processor 801 performs feature extraction processing on the sample operator chains in the training samples in at least one dimension to obtain sub-features of the sample operator chains in the training samples in each dimension, it is specifically configured to perform at least one of the following: if the dimension includes the input dimension, perform feature extraction processing on the sample operator chains in the training samples in the input dimension to obtain input sub-features of the sample operator chains in the training samples in the input dimension; the input sub-features include any one or more of: the number of input rows, the number of input bytes, and the number of columns of each data type included in the input; or, if the dimension includes the output dimension, perform feature extraction processing on the sample operator chains in the training samples in the output dimension to obtain output sub-features of the sample operator chains in the training samples in the output dimension; the output sub-features include any one or more of: the number of output rows, the number of output bytes, and the number of columns of each data type included in the output; or, if the dimension includes the content dimension, perform feature extraction processing on the sample operator chains in the training samples in the content dimension to obtain content sub-features of the sample operator chains in the training samples in the content dimension, and the content sub-features include any one or more of: the number of each query operation included in the sample operator chain, the number of each expression included in the sample operator chain, and the number of each function included in the sample operator chain.
[0215] In one embodiment, the processor 801 is further configured to: according to the dependency relationships among multiple operator chains, schedule at least one of the multiple operator chains to a thread pool to execute the operator chain in the thread pool according to the execution mode of the corresponding operator chain; during the execution of the operator chain, if it is determined that an exception occurs in the target operator chain, remove the target operator chain from the thread pool; wherein, the target operator chain refers to one of the multiple operator chains, and an exception occurring in the target operator chain means that: the operator chain on which the target operator chain depends has not been executed, or the execution duration of the target operator chain exceeds the duration threshold.
[0216] In one embodiment, the processor 801 is further configured to: in the case where an exception occurring in the target operator chain means that the operator chain on which the target operator chain depends has not been executed, after waiting for the operator chain having a dependency relationship with the target operator chain in the thread pool to be executed, schedule the target operator chain to the thread pool; or, in the case where an exception occurring in the target operator chain means that the execution duration of the target operator chain exceeds the duration threshold, when the waiting duration meets the waiting condition, schedule the target operator chain to the thread pool.
[0217] In one embodiment, the query plan is a query tree, the query tree includes multiple nodes, and each node is used to represent a query operation in the query plan; when the processor 801 performs a splitting process on the query plan to obtain multiple operator chains, it is specifically configured to: start traversing from the leaf nodes of the query tree, and perform blocking analysis on the query operation represented by the currently traversed node; if it is analyzed that the query operation represented by the currently traversed node is blocked, split out an operator chain according to the nodes included in the path from the leaf node to the current node; if the query operation represented by the next node of the currently traversed node is blocked, split out an operator chain according to the current node and the next node of the current node; after traversing the query tree, obtain multiple operator chains.
[0218] It should be understood that the computer device described in the embodiments of the present application can execute the descriptions of the data processing method in the corresponding previous embodiments, and can also execute the descriptions of the data processing device in the corresponding previous embodiments, which will not be elaborated here. In addition, the descriptions of the beneficial effects of adopting the same method will not be elaborated either.
[0219] In addition, it should be pointed out here that: the embodiments of the present application further provide a computer-readable storage medium, and a computer program is stored in the computer-readable storage medium, and the computer program includes program instructions. When the processor executes the above program instructions, it can execute the methods in the corresponding previous Figure 2 and Figure 4 embodiments, and therefore, it will not be elaborated here.
[0220] According to one aspect of the present application, there is provided a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device can execute the method in the foregoing Figure 2 and Figure 4 corresponding embodiments. Therefore, it will not be repeated here.
[0221] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.
[0222] The foregoing disclosure is only a preferred embodiment of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A data processing method, characterized in that The method includes: Obtain a query plan to be processed, and perform splitting processing on the query plan to obtain a plurality of operator chains; Perform compilation analysis processing on each operator chain in the plurality of operator chains to obtain compilation benefit indication information for each operator chain and compilation memory indication information for each operator chain; the compilation memory indication information is used to indicate the amount of memory space required for the code generated after the corresponding operator chain is compiled; Determine the execution mode of each operator chain according to the compilation benefit indication information of each operator chain and the compilation memory indication information of each operator chain; Execute the corresponding operator chain according to the determined execution mode of each operator chain to complete the query plan.
2. The method according to claim 1, wherein The execution mode includes a vectorized execution mode and a compilation execution mode; the determining the execution mode of each operator chain according to the compilation benefit indication information of each operator chain and the compilation memory indication information of each operator chain includes: Obtain a memory space quota; the memory space quota refers to the memory capacity configured for storing the code generated during the execution of the query plan; Divide the plurality of operator chains into a first operator chain set and a second operator chain set according to the memory space quota, the compilation benefit indication information of each operator chain, and the compilation memory indication information of each operator chain; Among them, the execution mode associated with the first operator chain set is the compilation execution mode, and the execution mode associated with the second operator chain set is the vectorized execution mode.
3. The method according to claim 2, wherein The dividing the plurality of operator chains into a first operator chain set and a second operator chain set according to the memory space quota, the compilation benefit indication information of each operator chain, and the compilation memory indication information of each operator chain includes: Generate at least one operator chain set according to the memory space quota and the compilation memory indication information of each operator chain; wherein, each operator chain set includes at least one operator chain in the plurality of operator chains, and the memory space occupancy corresponding to each operator chain set is less than the memory space quota, and the memory space occupancy is determined based on the memory space amount indicated by the compilation memory indication information of the operator chains included in the corresponding operator chain set; Determine the total compilation benefit value of each operator chain set according to the compilation benefit indication information of the operator chains included in each operator chain set; Select the target operator chain set with the largest total compilation benefit value from at least one operator chain set according to the determined total compilation benefit value of each operator chain set; Determine the target operator chain set as the first operator chain set, and obtain the second operator chain set according to the operator chains in the plurality of operator chains other than the operator chains in the target operator chain set.
4. The method according to claim 3, wherein The method further includes: If the total compilation benefit value of each operator chain set is negative, or the amount of memory space required indicated by the compilation memory indication information of each operator chain is greater than the memory space quota, then determine each operator chain as an operator chain in the second operator chain set; If the compilation benefit values indicated by the compilation benefit indication information of each operator chain are all positive, and the sum of the memory space amounts required to be occupied indicated by the compilation memory indication information of each operator chain does not exceed the memory space quota, then each operator chain is determined as an operator chain in the first operator chain set.
5. The method according to claim 1, wherein The performing compilation analysis processing on each operator chain among the multiple operator chains to obtain the compilation benefit indication information of each operator chain and the compilation memory indication information of each operator chain includes: Invoking a target compilation analysis model to perform compilation analysis processing on each operator chain among the multiple operator chains to obtain the compilation benefit value of each operator chain and the compilation code metric information of each operator chain; wherein, the compilation benefit value refers to the difference between the execution duration required for the corresponding operator chain in the vectorized execution mode and the execution duration required in the compilation execution mode; the compilation code metric information includes: the code volume that needs to be generated after the corresponding operator chain is compiled and executed in the compilation execution mode. Determining the compilation benefit value of each operator chain as the compilation benefit indication information of the corresponding operator chain, and determining the compilation code metric information of each operator chain as the compilation memory indication information of the corresponding operator chain.
6. The method according to claim 5, wherein The method further includes: Obtaining a target training sample; the target training sample includes a sample operator chain and training supervision information; the training supervision information includes: the compilation benefit supervision information of the sample operator chain and the compilation memory supervision information of the sample operator chain. Invoking a compilation analysis model to perform compilation analysis processing on the sample operator chain included in the target training sample to obtain the compilation benefit indication information and the compilation memory indication information of the target training sample. Determining the benefit difference corresponding to the target training sample according to the compilation benefit indication information of the target training sample and the compilation benefit supervision information included in the target training sample. Determining the code volume difference corresponding to the target training sample according to the compilation memory indication information of the target training sample and the compilation memory supervision information included in the target training sample. Training the compilation analysis model according to the benefit difference and the code volume difference corresponding to the target training sample to obtain the target compilation analysis model.
7. The method according to claim 6, characterized in that, The method further includes: Obtaining a sample query statement, and generating multiple sample operator chains corresponding to the sample query statement according to the sample query statement. Executing each sample operator chain respectively in the vectorized execution mode and the compilation execution mode to obtain the compilation benefit supervision information of each sample operator chain and the compilation memory supervision information of each sample operator chain. Combining each sample operator chain, the compilation benefit supervision information of each sample operator chain, and the compilation memory supervision information of each sample operator chain to obtain a training sample corresponding to each sample operator chain.
8. The method according to claim 7, wherein The executing each sample operator chain respectively in the vectorized execution mode and the compilation execution mode to obtain the compilation benefit supervision information of each sample operator chain and the compilation memory supervision information of each sample operator chain includes: Execute each sample operator chain according to the vectorized execution mode to obtain the first compilation statistic information of each sample operator chain; and, execute each sample operator chain according to the compilation execution mode to obtain the second compilation statistic information and reference code metric information of each sample operator chain; the reference code metric information includes: the code volume that needs to be generated after the corresponding sample operator chain is compiled and executed in the compilation execution mode. Determine the compilation benefit supervision information of each sample operator chain according to the first compilation statistic information and the second compilation statistic information of each sample operator chain; and, determine the reference code metric information of each sample operator chain as the compilation memory supervision information of the corresponding sample operator chain.
9. The method according to any one of claims 6 - 8, characterized in that, The method further includes: Perform feature extraction processing on the sample operator chains in the training sample in at least one dimension to obtain sub-features of the sample operator chains in the training sample in each dimension; Perform fusion processing on the sub-features of the sample operator chains in the training sample in each dimension to obtain a feature vector of the sample operator chains in the training sample, and the feature vector is used for the target compilation analysis model to perform compilation analysis processing on the sample operator chains in the training sample.
10. The method according to claim 9, characterized in that, The sample operator chains in the training sample include inputs and outputs; the dimensions include at least one of the following: input dimension, output dimension, and content dimension. The performing feature extraction processing on the sample operator chains in the training sample in at least one dimension to obtain sub-features of the sample operator chains in the training sample in each dimension includes at least one of the following: If the dimension includes the input dimension, perform feature extraction processing on the sample operator chains in the training sample in the input dimension to obtain input sub-features of the sample operator chains in the training sample in the input dimension; the input sub-features include: any one or more of the number of input rows, the number of input bytes, and the number of columns of each data type included in the input; or, If the dimension includes the output dimension, perform feature extraction processing on the sample operator chains in the training sample in the output dimension to obtain output sub-features of the sample operator chains in the training sample in the output dimension; the output sub-features include: any one or more of the number of output rows, the number of output bytes, and the number of columns of each data type included in the output; or, If the dimension includes the content dimension, perform feature extraction processing on the sample operator chains in the training sample in the content dimension to obtain content sub-features of the sample operator chains in the training sample in the content dimension, and the content sub-features include: any one or more of the number of each query operation included in the sample operator chain, the number of each expression included in the sample operator chain, and the number of each function included in the sample operator chain.
11. The method according to claim 1, characterized in that, The query plan is a query tree, and the query tree includes multiple nodes, and each node is used to represent a query operation in the query plan; The splitting the query plan to obtain multiple operator chains includes: Start traversing from the leaf nodes of the query tree, and perform blocking analysis on the query operation represented by the currently traversed node; If it is analyzed that the query operation represented by the current node is blocked, split out an operator chain according to the nodes included in the path between the leaf node and the current node; If the query operation represented by the next node of the currently traversed node is blocked, split out an operator chain according to the current node and the next node of the current node; After traversing the query tree, multiple operator chains are obtained.
12. The method according to claim 1, wherein The method further includes: According to the dependency relationships among the multiple operator chains, schedule at least one of the multiple operator chains to a thread pool to execute the operator chains in the thread pool according to the execution modes of the corresponding operator chains; During the execution of the operator chain, if it is determined that an exception occurs in the target operator chain, remove the target operator chain from the thread pool; Wherein, the target operator chain refers to one of the multiple operator chains, and the occurrence of an exception in the target operator chain means that the operator chain on which the target operator chain depends has not been executed, or the execution duration of the target operator chain exceeds the duration threshold.
13. The method according to claim 12, wherein The method further includes: In the case where the occurrence of an exception in the target operator chain means that the operator chain on which the target operator chain depends has not been executed, after waiting for the operator chain that has a dependency relationship with the target operator chain in the thread pool to be executed, schedule the target operator chain to the thread pool; or, In the case where the occurrence of an exception in the target operator chain means that the execution duration of the target operator chain exceeds the duration threshold, when the waiting duration meets the waiting condition, schedule the target operator chain to the thread pool.
14. A data processing device, characterized in that, The device includes: An acquisition unit, configured to acquire a query plan to be processed, and perform splitting processing on the query plan to obtain multiple operator chains; A processing unit, configured to perform compilation analysis processing on each of the multiple operator chains to obtain compilation benefit indication information of each operator chain and compilation memory indication information of each operator chain; the compilation memory indication information is used to indicate the amount of memory space required for the code generated after the corresponding operator chain is compiled; The processing unit is further configured to determine the execution mode of each operator chain according to the compilation benefit indication information of each operator chain and the compilation memory indication information of each operator chain; The processing unit is further configured to execute the corresponding operator chain according to the determined execution mode of each operator chain to complete the query plan.
15. A computer device includes an input interface and an output interface, characterized in that, It further includes: A processor and a computer storage medium; Wherein, the processor is adapted to implement one or more instructions, and the computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by the processor to perform the data processing method according to any one of claims 1-13.
16. A computer storage medium, characterized in that, One or more instructions are stored in the computer storage medium, and the one or more instructions are adapted to be loaded and executed by the processor to perform the data processing method according to any one of claims 1-13.
17. A computer program product, characterized in that, The computer program product includes a computer program or computer instructions, and the computer program or computer instructions are executed by a processor to implement the data processing method according to any one of claims 1-13.