Data processing method and database system
By converting query plans into intermediate representations (IR), the problem of difficulty in sharing and collaborating query plans between heterogeneous database systems is solved, reducing computational resources waste and improving query plans execution efficiency.
Patent Information
- Application Number
- PCT/CN2024/124968
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-30
- Filing Date
- 2024-10-15
- Publication Date
- 2025-05-22
AI Technical Summary
Between heterogeneous database systems, it is difficult for prior art to effectively share and collaborate query plans, resulting in waste of computing resources, especially when repeated compilation and optimization of query plans are required.
By converting query plans to intermediate representations (IR), which are described in languages supported by all relevant database systems, thus sharing and execution among database systems, avoiding duplicate compilation and optimization.
It reduces the waste of computing power overhead due to repeated compilation and optimization, and improves query plan sharing and collaboration efficiency between database systems.
Smart Images

Figure CN2024124968_22052025_PF_FP_ABST
Abstract
Description
A data processing method and database system
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 14, 2023, with application number 202311525323.8 and application name “A request processing method and related equipment”, and claims priority to the Chinese patent application filed with the State Intellectual Property Office on April 30, 2024, with application number 202410559375.5 and application name “A data processing method and database system”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of database technology, and in particular to a data processing method and a database system. Background Art
[0003] A distributed database is a logically unified database formed by connecting multiple physically dispersed database units using a computer network. Each connected database unit is called a site or node.
[0004] In a distributed database system, the base tables of a subquery and the base tables of a parent query are distributed across multiple data nodes (DNs). In existing technologies, when executing a subquery, all base table data is scanned, passed to the upper-level operator, and filtered using predicate conditions. Finally, the results of the subquery and parent query are returned to the CN, which then returns the aggregated results to the user.
[0005] The query plans obtained by compiling and optimizing heterogeneous database systems are often not interoperable. In other words, the query plans generated by the database systems cannot be accurately parsed by the heterogeneous database systems. When implementing query plan sharing and collaboration between heterogeneous database systems, for example, database system A determines through analysis (for example, from the perspective of performance, cost, and other factors) that the query plan (partial or complete) is more optimal to execute in database system B, database system A needs to pass the corresponding information to database system B. Since database system B cannot parse the query plan already obtained by database system A, the existing technology passes the original user query to database system B. After receiving the query, database system B needs to repeatedly compile and optimize to obtain the query plan, which will lead to a waste of computing resources.
[0006] Summary of the Invention
[0007] In a first aspect, the present application provides a data processing method, which is applied to a fused database system, wherein the fused database system includes a first database system and a second database system, and the second database system and the first database system have different language descriptions of query plans. The method includes: the first database system obtains a first query plan for executing the first query based on a first query input by a user, and the first query plan is described in a language supported for parsing by the first database system; when the first database system determines, based on the first query plan, that the second database system is more suitable for executing the first query plan than the first database system, the first database system converts the first query plan into an intermediate representation IR of the first query plan, and the IR of the first query plan is described in a language supported for parsing by both the first database system and the second database system; the first database system sends the IR of the first query plan to the second database system; the second database system converts the IR of the first query plan into a second query plan, and the second query plan is described in a language supported for parsing by the second database system; the second database system executes the second query plan to obtain a first query result.
[0008] The query plans obtained by heterogeneous database systems (such as the first database system and the second database system in the embodiment of the present application) through compilation and optimization are often not interoperable. That is, the query plan generated by the first database system cannot be accurately parsed by the second database system (or can be described as not supporting parsing). When implementing query plan sharing and collaboration between heterogeneous database systems, for example, the first database system determines through analysis (for example, from the perspective of performance, cost, etc.) that the query plan (partial or complete) is more optimal to execute in the second database system (that is, more suitable for execution in the second database system). The first database system needs to pass the corresponding information to the second database system. Since the second database system cannot parse the query plan already obtained by the first database system, the existing technology passes the original user query to the second database system. After receiving the query, the second database system needs to repeatedly compile and optimize to obtain the query plan, which will result in a waste of computing resources. In an embodiment of the present application, the first database system can convert the query plan into a language description that can be parsed by both the first database system and the second database system (that is, the intermediate representation of the first query plan in the embodiment of the present application). When it is determined that the query plan is more suitable for execution on other database systems (such as the second database system), the corresponding intermediate representation can be passed to the second database system. Since the intermediate representation is a language description that can be parsed by the second database system, the second database system does not need to recompile and optimize, but can directly convert the intermediate representation back to the query plan and execute it (or, if it can be determined that there is another database system that is more suitable for execution, the intermediate representation can be passed to the other database system). In this way, the waste of computing power caused by repeated compilation and optimization can be reduced.
[0009] In one possible implementation, the first database system may determine that the query plan (or part of it) is more suitable for execution on the second database system (for example, when the execution engine of the first database is highly loaded and the engine of the second database is idle, the query shards may be forwarded to the engine of the second database system for execution). Therefore, an intermediate representation that is more suitable for execution on the second database system needs to be passed to the second database system. Since the first database system and the second database system are heterogeneous database systems, the first database system needs to convert the first query plan into a language that the second database system supports for parsing (that is, the intermediate representation IR of the first query plan).
[0010] For example, if a query submitted by a user to database system A is more suitable for execution on database system B, the query plan generated by database system A can be passed to database system B for execution through Substrait.
[0011] This approach avoids the cost of repeated parsing optimization: in the optimal database selection scenario, the optimization cost of traditional Double SQL parsing is avoided.
[0012] In a possible implementation, the method further includes: the first database system receiving the first query result from the second database system and returning the result to the user; or the second database system returning the first query result to the user.
[0013] In one possible implementation, the method also includes: the first database system obtains a third query plan for executing the second user query based on the received second user query, and the third query plan is described in a language supported by the first database system for parsing; the first database system divides the third query plan into at least two plan shards, and the at least two plan shards include a first plan shard and a second plan shard; when the first database system determines based on the first plan shard that the second database system is more suitable for executing the first plan shard than the first database system, the first database system converts the first plan shard into an IR of the first plan shard, and the IR of the first plan shard is described in a language supported by both the first database system and the second database system for parsing; the first database system sends the IR of the first plan shard to the second database system; the second database system converts the IR of the first plan shard into a fourth query plan, and the fourth query plan is described in a language supported by the second database system for parsing; the second database system executes the fourth query plan to obtain a second query result; when the first database system determines based on the second plan shard that the first database system is more suitable for executing the second plan shard than the second database system, the first database system executes the second plan shard to obtain a third query result.
[0014] Through the above method, when it is determined that part of the query plan (that is, the first plan shard) is more suitable for execution on the second database system, and part of the query plan (that is, the second plan shard) is more suitable for execution on itself (that is, the first database system), the plan suitable for execution on the second database system can be converted into a language supported by the second database system for parsing (that is, the IR of the first plan shard), and sent to the second database system, and the plan suitable for execution on itself (that is, the second plan shard) can be executed.
[0015] In one possible implementation, the fusion database system further includes a third database system, and the third database system and the first database system have different language descriptions of the query plan. The method further includes: the first database system obtains a third query plan for executing the second user query based on the received second user query, and the third query plan is described in a language supported by the first database system for parsing; the first database system divides the third query plan into at least two plan shards, and the at least two plan shards include a first plan shard and a second plan shard; when the first database system determines based on the first plan shard that the second database system is more suitable for executing the first plan shard than the first database system, the first database system converts the first plan shard into an IR of the first plan shard, and the IR of the first plan shard is described in a language supported by both the first database system and the second database system for parsing; the first database system converts the IR of the first plan shard Send to the second database system; the second database system converts the IR of the first plan shard into a fourth query plan, and the fourth query plan is described in a language supported by the second database system for parsing; the second database system executes the fourth query plan to obtain a second query result; when the first database system determines that the third database system is more suitable for executing the second plan shard than the first database system based on the second plan shard, the first database system converts the second plan shard into the IR of the second plan shard, and the IR of the second plan shard is described in a language supported by both the first database system and the third database system for parsing; the first database system sends the IR of the second plan shard to the third database system; the third database system converts the IR of the second plan shard into a fifth query plan, and the fifth query plan is described in a language supported by the third database system for parsing; the third database system executes the fifth query plan to obtain a third query result.
[0016] In the above manner, when it is determined that part of the query plan (i.e., the first plan shard) is more suitable for execution on the second database system, and part of the query plan (i.e., the second plan shard) is more suitable for execution on the third database system, the plan suitable for execution on the second database system can be converted into a language supported by the second database system for parsing (i.e., the IR of the first plan shard) and sent to the second database system, and the plan suitable for execution on the third database system can be converted into a language supported by the third database system for parsing (i.e., the IR of the second plan shard) and sent to the third database system.
[0017] In one possible implementation, the third query plan includes multiple sub-query plans, and one plan shard among the multiple plan shards corresponds to at least one sub-query plan among the multiple sub-query plans; or, the query plan includes multiple operators, and one plan shard among the multiple plan shards corresponds to at least one operator among the multiple operators.
[0018] A query can be split into multiple shards. For example, coarse-grained partitioning can be based on subquery plans, meaning the query can be split into multiple subquery plans, and each plan shard can correspond to one or more subquery plans. Another example is fine-grained partitioning based on operators, meaning the query can be split into multiple operators, and each plan shard can include one or more operators.
[0019] The first database system can reasonably split queries and allocate database systems based on the performance or operating costs of other database systems, thereby improving overall performance.
[0020] In one possible implementation, the CN can perform a unified compilation and optimization to obtain a query plan and directly send it to multiple DNs. In this way, only one compilation and optimization is required, and the SQL is forwarded to the DN for execution. This not only omits the query compilation and query optimization overhead of the DN, but also ensures that each DN executes the same query plan.
[0021] In one possible implementation, a field for describing database-specific information can be added to the intermediate representation to facilitate parsing of heterogeneous databases. In one possible implementation, the intermediate representation includes fields and associated identifiers, where the identifiers are used to indicate that the fields are data of a unique type in the first database. The second database includes a parser, which is used to parse the intermediate representation. The parser is configured to recognize the identifiers and parse the unique type of data.
[0022] In one possible implementation, language is a substrait.
[0023] In one possible implementation, the second database system is more suitable for executing the first query plan than the first database system, including: the execution performance of the first query plan on the second database system is higher than that on the first database system; or the execution cost of the first query plan on the second database system is lower than that on the first database system. For example, the first database system is overloaded while the second database system is idle, or the engine of the second database system is more efficient in parsing and executing the first query plan than the engine of the first database system.
[0024] In a second aspect, the present application provides a fusion database system, which includes a first database system and a second database system. The second database system and the first database system have different language descriptions of query plans:
[0025] The first database system is configured to obtain a first query plan for executing the first query based on a first query input by a user, wherein the first query plan is described in a language that the first database system supports parsing;
[0026] The first database system is configured to convert the first query plan into an intermediate representation (IR) of the first query plan when the first database system determines that the second database system is more suitable for executing the first query plan than the first database system based on the first query plan, wherein the IR of the first query plan is described in a language that is parseable by both the first database system and the second database system;
[0027] The first database system is configured to send the IR of the first query plan to the second database system;
[0028] A second database system is configured to convert the IR of the first query plan into a second query plan, where the second query plan is described in a language that the second database system supports parsing;
[0029] The second database system is used to execute the second query plan to obtain the first query result.
[0030] In a possible implementation, the first database system is further configured to receive the first query result from the second database system and return it to the user; or,
[0031] The second database system is further configured to return the first query result to the user.
[0032] In one possible implementation, the first database system is further configured to obtain, based on the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that the first database system supports parsing.
[0033] The first database system is further configured to divide the third query plan into at least two plan shards, where the at least two plan shards include a first plan shard and a second plan shard;
[0034] The first database system is further configured to, when the first database system determines, based on the first planned sharding, that the second database system is more suitable for executing the first planned sharding than the first database system, convert the first planned sharding into an IR of the first planned sharding, where the IR of the first planned sharding is described in a language that both the first database system and the second database system can parse;
[0035] The first database system is further configured to send the IR of the first plan shard to the second database system;
[0036] The second database system is further configured to convert the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system supports parsing.
[0037] The second database system is further configured to execute the fourth query plan to obtain a second query result;
[0038] The first database system is further configured to execute the second planned sharding to obtain a third query result when the first database system determines, based on the second planned sharding, that the first database system is more suitable for executing the second planned sharding than the second database system.
[0039] In a possible implementation, the third query plan includes multiple sub-query plans, and one plan slice among the multiple plan slices corresponds to at least one sub-query plan among the multiple sub-query plans; or,
[0040] The query plan includes multiple operators, and one plan shard among the multiple plan shards corresponds to at least one operator among the multiple operators.
[0041] In a possible implementation, the fusion database system further includes a third database system, and the third database system and the first database system have different language descriptions of the query plan;
[0042] The first database system is further configured to obtain, based on the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that can be parsed by the first database system;
[0043] The first database system is further configured to divide the third query plan into at least two plan shards, where the at least two plan shards include a first plan shard and a second plan shard;
[0044] The first database system is further configured to, when the first database system determines, based on the first planned sharding, that the second database system is more suitable for executing the first planned sharding than the first database system, convert the first planned sharding into an IR of the first planned sharding, where the IR of the first planned sharding is described in a language that both the first database system and the second database system can parse;
[0045] The first database system is further configured to send the IR of the first plan shard to the second database system;
[0046] The second database system is further configured to convert the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system supports parsing.
[0047] The second database system is further configured to execute the fourth query plan to obtain a second query result;
[0048] The first database system is further configured to, if the first database system determines, based on the second plan sharding, that the third database system is more suitable for executing the second plan sharding than the first database system, convert the second plan sharding into an IR of the second plan sharding, where the IR of the second plan sharding is described in a language that both the first database system and the third database system can parse;
[0049] The first database system is further configured to send the IR of the second plan shard to the third database system;
[0050] The third database system is further configured to convert the IR of the second plan shard into a fifth query plan, where the fifth query plan is described in a language that the third database system supports parsing.
[0051] The third database system executes the fifth query plan to obtain a third query result.
[0052] In one possible implementation, the intermediate representation includes a field and an associated identifier, where the identifier is used to indicate that the field is data of a unique type in the first database system. The second database system includes a parser, which is used to parse the intermediate representation. The parser is configured to recognize the identifier and parse the unique type of data.
[0053] In one possible implementation, the language of the IR is substrait.
[0054] In a possible implementation, the idleness of the load of the second database system is more suitable for executing the first query plan than the idleness of the load of the first database system; or,
[0055] The execution performance of the second database system is more suitable for executing the first query plan than the execution performance of the first database system.
[0056] In a third aspect, the present application provides a computing device cluster. The computing device cluster includes at least one computing device, each of which includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is configured to execute instructions stored in the at least one memory, causing the computing device or computing device cluster to perform the method according to the first aspect or any implementation of the first aspect.
[0057] In a fourth aspect, the present application provides a computer-readable storage medium storing instructions, the instructions instructing a computing device or a computing device cluster to execute the method executed by the database system of the above-mentioned first aspect or any implementation of the first aspect.
[0058] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when run on a computing device or a computing device cluster, enables the computing device or the computing device cluster to execute the method executed by the database system of the above-mentioned first aspect or any implementation of the first aspect.
[0059] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figures 1A to 1D are schematic diagrams of the application framework of this application;
[0061] Figures 2A and 2B are schematic diagrams of the application framework of this application;
[0062] FIG3 is a flowchart of a data processing method according to an embodiment of the present application;
[0063] FIG4 is a schematic diagram of the process of processing a query in an embodiment of the present application;
[0064] FIG5 is a schematic diagram of an intermediate representation of an embodiment of the present application;
[0065] FIG6 is a schematic diagram of query sharding according to an embodiment of the present application;
[0066] FIG7 is a schematic diagram of a data processing method according to an embodiment of the present application;
[0067] FIG8 is a schematic diagram of a database according to an embodiment of the present application;
[0068] FIG9 is a schematic diagram of a data processing method according to an embodiment of the present application;
[0069] 10 to 13 are schematic diagrams of the structures of the data processing devices according to the embodiments of the present application. DETAILED DESCRIPTION
[0070] The following describes the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present application, rather than all the embodiments. Those skilled in the art will appreciate that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0071] The terms "first," "second," and the like in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.
[0072] The method provided in the embodiment of the present application can be applied to a database system. FIG1A shows a typical logical architecture of a database system. According to FIG1A , a database system 100 includes a database 110 and a database management system (DBMS) 130 .
[0073] Among them, the database 110 is an organized data set stored in the data storage 120, that is, a related data set organized, stored and used according to a specific data model. According to the different data models used to organize data, data can be divided into multiple types, such as relational data, graph data, time series data, etc. Relational data is data modeled using a relational model, usually represented as a table, and the rows in the table represent a set of related values of an object or entity. Graph data, referred to as "graph", is used to represent the relationship between objects or entities, such as social relationships. Time series data, referred to as time series data, is a data column recorded and indexed in chronological order, used to describe the state change information of an object in the time dimension.
[0074] The database management system 130 is the core of the database system and is the system software used to organize, store, and maintain data. Clients 200 can access the database 110 through the database management system 130, and database administrators also use the database management system to perform database maintenance. The database management system 130 provides various functions for clients 200, which can be applications or user devices, to create, modify, and query databases. The functions provided by the database management system 130 may include but are not limited to the following: (1) Data definition function. The database management system 130 provides a data definition language (DDL) to define the structure of the database 110. DDL is used to describe the database framework and can be saved in the data dictionary; (2) Data access function. The database management system 130 provides a data manipulation language (DML) to implement basic access operations on the database 110, such as retrieval, insertion, modification and deletion; (3) Database operation management function. The database management system 130 provides a data control function to effectively control and manage the operation of the database 110 to ensure that the data is correct and valid; (4) Database establishment and maintenance function, including loading the initial data of the database, dumping, restoring, and reorganizing the database, system performance monitoring, analysis and other functions; (5) Database transmission. The database management system provides transmission of processed data to realize communication between the client and the database management system, which is usually coordinated with the operating system.
[0075] The data storage 120 includes, but is not limited to, solid-state drives (SSDs), disk arrays, cloud storage, or other types of non-transitory computer-readable storage media. Those skilled in the art will appreciate that a database system may include fewer or more components than those shown in FIG1A , or may include components different from those shown in FIG1A . FIG1A merely illustrates components that are more relevant to the implementation disclosed in the embodiments of the present invention.
[0076] The database system provided in the embodiments of the present application may be a distributed database system (DDBS). During transaction processing, a DDBS typically employs a global transaction manager (GTM) to manage transactions in order to achieve concurrency control between transactions. The following describes a DDBS with reference to Figures 1B and 1C.
[0077] Figure 1B is a schematic diagram of a distributed database system using a shared-storage architecture, including one or more coordinator nodes (CN), multiple data nodes (DN), and one or more GTMs (such as the first and second GTMs in Figure 1B). The first GTM serves as the master GTM, and the second GTM is used to back up the data of the first GTM and take over the work of the first GTM when the first GTM fails, thus ensuring the high reliability of the DDBS. The CN and DN communicate through a network channel. In one embodiment, the network channel can be composed of network devices such as switches, routers, and gateways. The CN, DN, and GTM jointly implement the functions of the database management system, providing clients with database retrieval, insertion, modification, and deletion services. In one embodiment, a database management system is deployed on each CN, DN, and GTM. The shared data storage stores data that can be shared by multiple DNs, and the DNs can perform read and write operations on the data in the data storage through the network channel. The shared data storage can be a shared disk array. The CN, DN, first GTM or second GTM in the distributed database system can be a physical machine, such as a database server, or a virtual machine (VM) or container running on abstract hardware resources. In one embodiment, the CN, DN, first GTM or second GTM is a virtual machine or container, and the network channel is a virtual switching network, which includes a virtual switch. The database management system deployed in the CN, DN, first GTM or second GTM is a DBMS instance, which can be a process or a thread. These DBMSs work together to complete the functions of the database relational system. In another embodiment, the CN, DN, first GTM or second GTM is a physical machine, and the network channel includes one or more switches, which are storage area network (SAN) switches, Ethernet switches, fiber switches or other physical switching devices.
[0078] Figure 1C is a schematic diagram of a distributed database system using a shared-nothing architecture. Each DN has its own dedicated hardware resources (such as data storage), operating system, and database. The CN, DN, and first or second GTM communicate via a network channel, which can be understood by referring to the corresponding description of Figure 1B above. In this system, data is distributed to each DN based on the database model and application characteristics. Query tasks are divided into several parts by the CN and executed in parallel on all DNs, collaborating with each other to provide database services as a whole. All communication functions are implemented on a high-bandwidth network interconnection system. Similar to the distributed database system with a shared-storage architecture described in Figure 1B, the CN, DN, first or second GTM can be either physical machines or virtual machines.
[0079] In all embodiments of the present application, the data storage of the database system includes but is not limited to solid state drives (SSDs), disk arrays, or other types of non-transient computer-readable media. Although the database is not shown in Figures 1B-1C, it should be understood that the database is stored in the data storage. Those skilled in the art will understand that a database system may include fewer or more components than those shown in Figures 1A-1C, or include components different from those shown in Figures 1A-1C, and Figures 1A-1C only illustrate components that are more relevant to the implementation methods disclosed in the embodiments of the present application. However, those skilled in the art will understand that a distributed database system may include any number of CNs and DNs. The database management system functions of each CN and DN may be implemented by an appropriate combination of software, hardware, and / or firmware running on each CN and DN, respectively.
[0080] The distributed database system described in FIG. 1B and FIG. 1C includes multiple DNs and multiple CNs, wherein the function of each DN is substantially the same, and the function of each CN is also substantially the same.
[0081] FIG1D shows an application data management system provided by an embodiment of the present application. Specifically, as shown, a data node cluster may include multiple data node DNs. Application developers may deploy the relevant data of the developed application in the corresponding data node DNs. The application may send a service request to the application server, and the application server may convert the service request into a data operation request; the application server may send the data operation request to the distributed database system. Then, one or more data node DNs in the distributed database system may perform the operation corresponding to the operation request on the data in the database memory. Specifically, the application server may send the data operation request to the coordination node in the coordination node cluster, and the coordination node may forward the operation request to the relevant DN, and the relevant DN may perform the data operation corresponding to the operation request.
[0082] Refer to Figure 2A, which is a system architecture diagram of an embodiment of the present application: wherein, the present application can be applied to the architecture of a heterogeneous database system. As shown in Figure 2A, the heterogeneous database includes a database system A and a database system B. Database system A and database system B are heterogeneous databases. The so-called heterogeneous databases can be understood as databases with different processing tasks, structures or characteristics.
[0083] For example, Gauss (DWS) is a powerful distributed database suitable for large-scale data processing and analysis; SparkSQL is a database system built on Apache Spark, focusing on big data processing; Clickhouse is a columnar database suitable for high-speed data query and analysis; and Velox is an emerging scalable database designed to process large-scale graph data.
[0084] There are many differences and characteristics between different database systems. Based on their data storage models, they can be categorized as relational databases (such as MySQL and PostgreSQL) and non-relational databases (such as MongoDB and Cassandra). Furthermore, databases can be categorized based on their primary purpose into transactional databases, designed to support complex data transactions, and analytical databases, designed for high-performance data analysis and report generation.
[0085] As shown in Figure 2B, in a typical system containing heterogeneous databases, if subquery plans are to be transferred between databases, there will be a double parsing optimization problem and unnecessary computational overhead. For example, in Figure 2B, a query sent by a user to database system B (①) is only found to be more suitable for execution in database system A after query compilation and optimization. The query is then forwarded to database system A in the form of raw SQL (②), and query compilation and optimization are performed on database system A. In database system A, the CN forwards the query task in the form of a query plan or SQL to a specific DN for execution (③). Compared to the ideal situation where the user directly sends the query to database system A (④), this architecture has the overhead of repeated query compilation and query optimization. In addition, this architecture cannot support more fine-grained optimal database selection, such as subqueries or operations.
[0086] In order to solve the above problems, an embodiment of the present application provides a data processing method. Referring to FIG3 , FIG3 is a flow chart of a data processing method provided in an embodiment of the present application, including:
[0087] 301. A first database system obtains a first query plan for executing the first query according to a first query input by a user. The first query plan is described in a language that the first database system supports parsing.
[0088] In one possible implementation, query compilation and query optimization may be performed on a first query (e.g., an SQL query) input by a user to obtain a first query plan. For example, query compilation and query optimization may be performed on the first query input by the user by a CN node in a database system to obtain the first query plan.
[0089] 302. When the first database system determines, based on the first query plan, that the second database system is more suitable for executing the first query plan than the first database system, the first database system converts the first query plan into an intermediate representation IR of the first query plan, where the IR of the first query plan is described in a language that both the first database system and the second database system support parsing.
[0090] The query plans obtained by heterogeneous database systems (such as the first database system and the second database system in the embodiment of the present application) through compilation and optimization are often not interoperable. That is, the query plan generated by the first database system cannot be accurately parsed by the second database system (or can be described as not supporting parsing). When implementing query plan sharing and collaboration between heterogeneous database systems, for example, the first database system determines through analysis (for example, from the perspective of performance, cost, etc.) that the query plan (partial or complete) is more optimal to execute in the second database system (that is, more suitable for execution in the second database system). The first database system needs to pass the corresponding information to the second database system. Since the second database system cannot parse the query plan already obtained by the first database system, the existing technology passes the original user query to the second database system. After receiving the query, the second database system needs to repeatedly compile and optimize to obtain the query plan, which will result in a waste of computing resources. In an embodiment of the present application, the first database system can convert the query plan into a language description that can be parsed by both the first database system and the second database system (that is, the intermediate representation of the first query plan in the embodiment of the present application). When it is determined that the query plan is more suitable for execution on other database systems (such as the second database system), the corresponding intermediate representation can be passed to the second database system. Since the intermediate representation is a language description that can be parsed by the second database system, the second database system does not need to recompile and optimize, but can directly convert the intermediate representation back to the query plan and execute it (or, if it can be determined that there is another database system that is more suitable for execution, the intermediate representation can be passed to the other database system). In this way, the waste of computing power caused by repeated compilation and optimization can be reduced.
[0091] For example, referring to FIG8 , the two heterogeneous database systems (DB1 and DB2) in FIG8 may each include a module for generating a query plan, a module for converting into an intermediate representation, a parser, and an optimizer. The database systems may transfer the intermediate representation to collaborate in query execution.
[0092] The language used to describe the intermediate representation may be, but is not limited to, substrait or other common languages of heterogeneous database systems.
[0093] Taking substrait as an example, its universal query plan representation allows users and applications to define and describe queries in a unified manner, regardless of the differences in the underlying databases. This not only simplifies the query and data operation process but also promotes collaboration and interoperability between different database systems. It provides users with greater flexibility, allowing them to leverage the respective strengths of different database systems without having to worry too much about the underlying technical details. This helps improve data consistency and collaboration in multi-database environments, enabling users to better meet the needs of different application scenarios and improving overall system performance and efficiency.
[0094] Optionally, embodiments of the present application can use substrait to uniformly express query plans generated by multiple databases. Referring to Figure 4 , Figure 4 illustrates the SQL to Substrait conversion process. For a user-entered query (SQL), after query compilation and query optimization in the database, a query plan is generated, which can then be converted to Substrait.
[0095] The substrait shown in Figure 4 describes a query plan through four types of objects. Their names and functions are as follows:
[0096] Type: Used to describe the data type. Common type definitions such as i8 and i32 clarify the data type and word length to avoid data type incompatibilities between different databases.
[0097] Relation: Used to describe operation types and their connections. An operation node in a database query plan generally corresponds to a Relation node in Substrait. It records information such as the operation type and the columns it operates on. It also records the input relations of the operation. A tree-like query plan is represented in Substrait as a tree of relations.
[0098] Expression: defines an expression. Some operation expressions in the query plan are recorded through Expression.
[0099] Function: defines the functions used and their parameters and return value types.
[0100] In addition to the objects defined above, fields for describing database-specific information can be added to the intermediate representation to facilitate parsing of heterogeneous databases. In one possible implementation, the intermediate representation includes fields and associated identifiers, where the identifiers are used to indicate that the fields are data of a unique type in the first database. The second database includes a parser, which is used to parse the intermediate representation. The parser is configured to recognize the identifiers and parse the unique type of data.
[0101] For example, the associated identifier could be "Node." See Figure 5, which illustrates an intermediate representation. The Node field is added to Substrait to support recording some database-specific information, improving support for unified cross-database representation. Databases can design scalable parsers for the unified IR representation to parse and deparse unique information, thereby supporting efficient execution across heterogeneous databases.
[0102] 303. The first database system sends the IR of the first query plan to the second database system.
[0103] In the diverse database ecosystem, choosing the best database system for a specific scenario is crucial. Specifically, when a query is received on one database, it might be more suitable for execution on another database system. This decision (i.e., which database system to choose) requires a comprehensive consideration of multiple factors, two of which are performance and cost.
[0104] First, performance is a key factor in database selection. Different database systems exhibit different performance characteristics across various computing tasks. Some database systems may excel at transaction processing, while others are more advantageous for large-scale data analysis. Furthermore, database performance can vary significantly under varying load conditions. Therefore, it is crucial to select the most appropriate database system based on the application's performance requirements and workload.
[0105] Secondly, cost is another factor that requires careful consideration. Different database systems require significantly different computing resources (e.g., CPU, memory) and storage resources, which directly impacts operational costs. Some database systems may require high-performance hardware and large amounts of memory, resulting in significant hardware costs. Others, on the other hand, may operate well with relatively low hardware configurations, thus reducing overall costs. Choosing the most cost-effective database system can effectively control the project budget.
[0106] Considering both performance and cost, you need to find a balance to ensure that your database choice can meet your performance requirements while operating within a reasonable cost range. This may require in-depth performance testing and cost analysis to determine the best database choice.
[0107] Existing technologies have provided some useful modeling tools for database selection. These modeling methods can evaluate database systems based on multiple dimensions, including cloud environment billing models, performance estimation models, and resource estimation models. These models can help decision makers better understand the performance and cost characteristics of different database systems, thereby making more informed choices about the database system that best suits their application scenarios.
[0108] In one possible implementation, the first database may determine that the query plan (or part of it) is better executed on the second database (for example, when the execution engine of the first database is highly loaded and the engine of the second database is idle, the query shards may be forwarded to the engine of the second database for execution), and therefore an intermediate representation that is more suitable for execution on the second database needs to be passed to the second database.
[0109] In a possible implementation, when the execution performance of the first query plan on the second database system is higher than the execution performance of the first query plan on the first database system, it can be considered that the second database system is more suitable for executing the first query plan than the first database system.
[0110] In a possible implementation, when the execution cost of executing the first query plan on the second database system is lower than the execution cost of executing the first query plan on the first database system, it can be considered that the second database system is more suitable for executing the first query plan than the first database system.
[0111] For example, the first database system is in an overloaded state, while the second database system is in an idle state. For another example, the efficiency of the engine of the second database system in parsing and executing the first query plan is higher than that of the engine of the first database system.
[0112] For example, if a query submitted by a user to database system A is more suitable for execution on database system B, the query plan generated by database system A can be passed to database system B for execution through Substrait. This approach avoids the cost of repeated parsing optimization: in the optimal database selection scenario, the optimization cost of traditional Double SQL parsing is avoided.
[0113] In one possible implementation, the method also includes: the first database system obtains a third query plan for executing the second user query based on the received second user query, and the third query plan is described in a language supported by the first database system for parsing; the first database system divides the third query plan into at least two plan shards, and the at least two plan shards include a first plan shard and a second plan shard; when the first database system determines based on the first plan shard that the second database system is more suitable for executing the first plan shard than the first database system, the first database system converts the first plan shard into an IR of the first plan shard, and the IR of the first plan shard is described in a language supported by both the first database system and the second database system for parsing; the first database system sends the IR of the first plan shard to the second database system; the second database system converts the IR of the first plan shard into a fourth query plan, and the fourth query plan is described in a language supported by the second database system for parsing; the second database system executes the fourth query plan to obtain a second query result; when the first database system determines based on the second plan shard that the first database system is more suitable for executing the second plan shard than the second database system, the first database system executes the second plan shard to obtain a third query result.
[0114] Through the above method, when it is determined that part of the query plan (that is, the first plan shard) is more suitable for execution on the second database system, and part of the query plan (that is, the second plan shard) is more suitable for execution on itself (that is, the first database system), the plan suitable for execution on the second database system can be converted into a language supported by the second database system for parsing (that is, the IR of the first plan shard), and sent to the second database system, and the plan suitable for execution on itself (that is, the second plan shard) can be executed.
[0115] In one possible implementation, the fusion database system further includes a third database system, and the third database system and the first database system have different language descriptions of the query plan. The method further includes: the first database system obtains a third query plan for executing the second user query based on the received second user query, and the third query plan is described in a language supported by the first database system for parsing; the first database system divides the third query plan into at least two plan shards, and the at least two plan shards include a first plan shard and a second plan shard; when the first database system determines based on the first plan shard that the second database system is more suitable for executing the first plan shard than the first database system, the first database system converts the first plan shard into an IR of the first plan shard, and the IR of the first plan shard is described in a language supported by both the first database system and the second database system for parsing; the first database system converts the IR of the first plan shard Send to the second database system; the second database system converts the IR of the first plan shard into a fourth query plan, and the fourth query plan is described in a language supported by the second database system for parsing; the second database system executes the fourth query plan to obtain a second query result; when the first database system determines that the third database system is more suitable for executing the second plan shard than the first database system based on the second plan shard, the first database system converts the second plan shard into the IR of the second plan shard, and the IR of the second plan shard is described in a language supported by both the first database system and the third database system for parsing; the first database system sends the IR of the second plan shard to the third database system; the third database system converts the IR of the second plan shard into a fifth query plan, and the fifth query plan is described in a language supported by the third database system for parsing; the third database system executes the fifth query plan to obtain a third query result.
[0116] In the above manner, when it is determined that part of the query plan (i.e., the first plan shard) is more suitable for execution on the second database system, and part of the query plan (i.e., the second plan shard) is more suitable for execution on the third database system, the plan suitable for execution on the second database system can be converted into a language supported by the second database system for parsing (i.e., the IR of the first plan shard) and sent to the second database system, and the plan suitable for execution on the third database system can be converted into a language supported by the third database system for parsing (i.e., the IR of the second plan shard) and sent to the third database system.
[0117] In one possible implementation, the third query plan includes multiple sub-query plans, and one plan shard among the multiple plan shards corresponds to at least one sub-query plan among the multiple sub-query plans; or, the query plan includes multiple operators, and one plan shard among the multiple plan shards corresponds to at least one operator among the multiple operators.
[0118] A query can be split into multiple shards. For example, coarse-grained partitioning can be based on subquery plans, meaning the query can be split into multiple subquery plans, and each plan shard can correspond to one or more subquery plans. Another example is fine-grained partitioning based on operators, meaning the query can be split into multiple operators, and each plan shard can include one or more operators.
[0119] The first database system can reasonably split queries and allocate database systems based on the performance or operating costs of other database systems, thereby improving overall performance.
[0120] In one possible implementation, the CN can perform a unified compilation and optimization to obtain a query plan and directly send it to multiple DNs. In this way, only one compilation and optimization is required, and the SQL is forwarded to the DN for execution. This not only omits the query compilation and query optimization overhead of the DN, but also ensures that each DN executes the same query plan.
[0121] Taking Substrait as an example, the query plan sharding technology in the embodiments of this application can support more fine-grained optimal database execution selection for subqueries or operations. By performing more fine-grained query plan sharding on Substrait's unified IR, the most appropriate database execution engine can be selected for each sub-plan and even each operation.
[0122] For example, referring to FIG. 6 , FIG. 6 is a schematic diagram of a query sharding, which includes execution engines of three databases. The same query can be executed by the execution engines of the three databases.
[0123] In a possible implementation, the coordination node CN of the first database converts the query plan into an intermediate representation (IR), and the data node DN of the first database executes the query plan corresponding to the second sub-representation to obtain a second query result.
[0124] For example, as shown in Figure 7, the CN needs to distribute the generated query plan to multiple DNs for execution. The client submits the SQL to the CN, which obtains the intermediate representation and then sends it to the corresponding DN for execution. After receiving the intermediate representation, the DN converts it into a query plan for execution. Compared to the traditional method of directly forwarding the SQL to the DN for execution, this method not only eliminates the query compilation and optimization overhead of the DN, but also ensures that each DN executes the same query plan.
[0125] For example, referring to FIG9 , FIG9 shows a specific conversion and execution process, wherein the query plan can be converted into an intermediate representation described by the substrait through “to substrait”, and the intermediate representation described by the substrait can be converted into a query plan through “from substrait”.
[0126] 304. The second database system converts the IR of the first query plan into a second query plan, where the second query plan is described in a language that the second database system supports parsing.
[0127] 305. The second database system executes the second query plan to obtain the first query result.
[0128] In a possible implementation, the first database system may receive the first query result from the second database system and return it to the user; or the second database system returns the first query result to the user.
[0129] Referring to Figure 10 , an embodiment of the present application further provides a fusion database system 1000 that can be used to implement the data processing method provided by any possible implementation of the method embodiments described above. Fusion database system 1000 includes a first database system 1001 and a second database system 1002 . The second database system 1002 and the first database system 1001 have different language descriptions for query plans:
[0130] The first database system 1001 is configured to obtain a first query plan for executing the first query based on a first query input by a user, wherein the first query plan is described in a language that the first database system 1001 supports parsing.
[0131] The first database system 1001 is configured to convert the first query plan into an intermediate representation (IR) of the first query plan when the first database system 1001 determines that the second database system 1002 is more suitable for executing the first query plan than the first database system 1001 based on the first query plan. The IR of the first query plan is described in a language that is parseable by both the first database system 1001 and the second database system 1002.
[0132] The first database system 1001 is configured to send the IR of the first query plan to the second database system 1002;
[0133] The second database system 1002 is configured to convert the IR of the first query plan into a second query plan, where the second query plan is described in a language that the second database system 1002 supports parsing.
[0134] The second database system 1002 is configured to execute the second query plan to obtain the first query result.
[0135] In a possible implementation, the first database system 1001 is further configured to receive the first query result from the second database system 1002 and return the result to the user; or,
[0136] The second database system 1002 is further configured to return the first query result to the user.
[0137] In one possible implementation, the first database system 1001 is further configured to obtain, based on the received second user query, a third query plan for executing the second user query, where the third query plan is described in a language that the first database system 1001 supports parsing.
[0138] The first database system 1001 is further configured to divide the third query plan into at least two plan shards, where the at least two plan shards include a first plan shard and a second plan shard;
[0139] The first database system 1001 is further configured to convert the first planned sharding into an IR of the first planned sharding when the first database system 1001 determines that the second database system 1002 is more suitable for executing the first planned sharding than the first database system 1001 based on the first planned sharding. The IR of the first planned sharding is described in a language that is parseable by both the first database system 1001 and the second database system 1002.
[0140] The first database system 1001 is further configured to send the IR of the first plan shard to the second database system 1002;
[0141] The second database system 1002 is further configured to convert the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system 1002 supports parsing.
[0142] The second database system 1002 is further configured to execute the fourth query plan to obtain a second query result;
[0143] The first database system 1001 is further configured to execute the second planned sharding to obtain a third query result when the first database system 1001 determines, based on the second planned sharding, that the first database system 1001 is more suitable for executing the second planned sharding than the second database system 1002.
[0144] In a possible implementation, the third query plan includes multiple sub-query plans, and one plan slice among the multiple plan slices corresponds to at least one sub-query plan among the multiple sub-query plans; or,
[0145] The query plan includes multiple operators, and one plan shard among the multiple plan shards corresponds to at least one operator among the multiple operators.
[0146] In a possible implementation, the fusion database system further includes a third database system, and the third database system and the first database system 1001 have different language descriptions of the query plan;
[0147] The first database system 1001 is further configured to obtain, based on the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that the first database system 1001 supports parsing.
[0148] The first database system 1001 is further configured to divide the third query plan into at least two plan shards, where the at least two plan shards include a first plan shard and a second plan shard;
[0149] The first database system 1001 is further configured to convert the first planned sharding into an IR of the first planned sharding when the first database system 1001 determines that the second database system 1002 is more suitable for executing the first planned sharding than the first database system 1001 based on the first planned sharding. The IR of the first planned sharding is described in a language that is parseable by both the first database system 1001 and the second database system 1002.
[0150] The first database system 1001 is further configured to send the IR of the first plan shard to the second database system 1002;
[0151] The second database system 1002 is further configured to convert the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system 1002 supports parsing.
[0152] The second database system 1002 is further configured to execute the fourth query plan to obtain a second query result;
[0153] The first database system 1001 is further configured to convert the second planned sharding into an IR of the second planned sharding when the first database system 1001 determines, based on the second planned sharding, that the third database system is more suitable for executing the second planned sharding than the first database system 1001, where the IR of the second planned sharding is described in a language that both the first database system 1001 and the third database system can parse;
[0154] The first database system 1001 is further configured to send the IR of the second plan shard to the third database system;
[0155] The third database system is further configured to convert the IR of the second plan shard into a fifth query plan, where the fifth query plan is described in a language that the third database system supports parsing.
[0156] The third database system executes the fifth query plan to obtain a third query result.
[0157] In one possible implementation, the intermediate representation includes a field and an associated identifier, where the identifier is used to indicate that the field is data of a unique type in the first database system 1001. The second database system 1002 includes a parser, which is used to parse the intermediate representation. The parser is configured to recognize the identifier and parse the unique type of data.
[0158] In one possible implementation, the language of the IR is substrait.
[0159] In a possible implementation, the idleness of the load of the second database system 1002 is more suitable for executing the first query plan than the idleness of the load of the first database system 1001; or,
[0160] The execution performance of the second database system 1002 is more suitable for executing the first query plan than the execution performance of the first database system 1001 .
[0161] Specifically, the specific implementation of the various operations of the data processing method performed by the fusion database system 1000 can refer to the description of the relevant content in the above method embodiment, which will not be repeated here.
[0162] In the embodiment of the present application, the first database system and the second database system can be implemented by software or hardware. The following takes the first database system as an example to introduce the implementation of the first database system, and the implementation of the second database system can be used as a reference.
[0163] As an example of a software functional unit, the first database system may include code running on a compute instance. A compute instance may be at least one of a physical host (computing device), a virtual machine, a container, or other computing devices. Furthermore, the computing device may be one or more. For example, a node may include code running on multiple hosts, virtual machines, or containers. It should be noted that the multiple hosts, virtual machines, or containers running the code may be distributed in the same AZ or different AZs, with each AZ comprising a data center or multiple geographically close data centers. The multiple hosts, virtual machines, or containers running the code may be distributed in the same region or different regions. Typically, a region may include multiple AZs, and a VPC is set up within a region. Cross-region communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0164] Similarly, multiple hosts, virtual machines, or containers running the code can be distributed in the same VPC or in multiple VPCs. Typically, a region can include multiple AZs.
[0165] As an example of a hardware functional unit, the first database system may include at least one computing device, such as a server. Alternatively, the module may be implemented using an ASIC or a PLD. The PLD may be implemented using a CPLD, FPGA, GAL, or any combination thereof.
[0166] The multiple computing devices included in the first database system can be distributed in the same AZ or in different AZs. The multiple computing devices included in the module can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the module can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs. The present application also provides a computing device 1000. As shown in Figure 11, the computing device 1000 includes: a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate with each other via the bus 1002. The computing device 1000 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 1000.
[0167] Bus 1002 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG11 shows only one line, but this does not imply a single bus or type of bus. Bus 1002 may include a path for transmitting information between various components of computing device 1000 (e.g., memory 1006, processor 1004, and communication interface 1008).
[0168] The processor 1004 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0169] The memory 1006 may include a volatile memory, such as a random access memory (RAM). The memory 1006 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD). The memory 1006 stores executable program code, and the processor 1004 executes the executable program code to implement the method executed by the aforementioned database system. Specifically, the memory 1006 stores instructions for the method executed by the database system (e.g., the first database system and the second database system).
[0170] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 and other devices or a communication network.
[0171] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0172] 12 , the computing device cluster includes at least one computing device 1000. Instructions for executing the method executed by the database system (eg, the first database system and the second database system) may be stored in the memory 1006 of one or more computing devices 1000 in the computing device cluster.
[0173] In some possible implementations, one or more computing devices 1000 in the computing device cluster may also be used to execute some instructions of the method executed by the database system (e.g., the first database system and the second database system). In other words, the combination of one or more computing devices 1000 can jointly execute the instructions of the method executed by the database system (e.g., the first database system and the second database system).
[0174] It should be noted that the memories 1006 in different computing devices 1000 in the computing device cluster may store different instructions for executing partial functions of the heterogeneous database system 100 .
[0175] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), etc. FIG13 illustrates a possible implementation. As shown in FIG13 , two computing devices 120A and 120B are connected via a network. Specifically, the connection to the network is achieved through a communication interface in each computing device. In this type of possible implementation, the memory 116 in the computing device 120A may store instructions for executing the functions of the second database system. Simultaneously, the memory 116 in the computing device 120B may store instructions for executing the functions of the first database system. Alternatively, the memory 116 in the computing device 120A may store instructions for executing part of the functions of the first database system. Simultaneously, the memory 116 in the computing device 120B may store instructions for executing another part of the functions of the first database system.
[0176] It should be understood that the functionality of the computing device 120A shown in Figure 13 may also be implemented by multiple computing devices. Similarly, the functionality of the computing device 120B may also be implemented by multiple computing devices.
[0177] The present application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection relationship between the computing device clusters in Figures 12 and 13. However, the memory 116 in one or more computing devices in this computing device cluster can store the same instructions for executing the template generation method.
[0178] In some possible implementations, the memory 116 of one or more computing devices in the computing device cluster may also store partial instructions for executing the template generation method. In other words, the combination of one or more computing devices can jointly execute instructions for executing the template generation method.
[0179] It should be noted that the memory 116 in different computing devices in the computing device cluster can store different instructions for executing portions of the template generation method. In other words, the instructions stored in the memory 116 in different computing devices can implement the functions of one or more modules in the second database system and the first database system.
[0180] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned method, training sample generation method, and model training method applied to the optimization problem solving system 100 for executing the database system execution.
[0181] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the aforementioned database system execution method, training sample generation method, and model training method.
[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A data processing method, characterized in that: Applied to a fusion database system, the fusion database system includes a first database system and a second database system, the second database system and the first database system have different language descriptions of query plans, the method includes: The first database system obtains a first query plan for executing the first query according to a first query input by a user, wherein the first query plan is described in a language that the first database system supports parsing; When the first database system determines, based on the first query plan, that the second database system is more suitable for executing the first query plan than the first database system, the first database system converts the first query plan into an intermediate representation IR of the first query plan, where the IR of the first query plan is described in a language that both the first database system and the second database system can parse; The first database system sends the IR of the first query plan to the second database system; The second database system converts the IR of the first query plan into a second query plan, where the second query plan is described in a language that the second database system supports parsing; The second database system executes the second query plan to obtain the first query result.
2. The method according to claim 1, characterized in that The method further comprises: The first database system receives the first query result from the second database system and returns it to the user; or, The second database system returns the first query result to the user.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: The first database system obtains, according to the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that the first database system supports parsing; The first database system divides the third query plan into at least two plan slices, where the at least two plan slices include a first plan slice and a second plan slice; When the first database system determines that the second database system is more suitable for executing the first plan sharding than the first database system based on the first plan sharding, the first database system converts the first plan sharding into an IR of the first plan sharding, where the IR of the first plan sharding is described in a language that both the first database system and the second database system support parsing; The first database system sends the IR of the first plan shard to the second database system; The second database system converts the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system supports parsing; The second database system executes the fourth query plan to obtain a second query result; When the first database system determines, based on the second plan sharding, that the first database system is more suitable for executing the second plan sharding than the second database system, the first database system executes the second plan sharding to obtain a third query result.
4. The method according to claim 3, characterized in that The third query plan includes a plurality of sub-query plans, and one plan slice of the plurality of plan slices corresponds to at least one sub-query plan of the plurality of sub-query plans; or, The query plan includes multiple operators, and one plan slice among the multiple plan slices corresponds to at least one operator among the multiple operators.
5. The method according to claim 1 or 2, characterized in that: The fusion database system further includes a third database system, and the third database system and the first database system have different language descriptions of query plans. The method further includes: The first database system obtains, according to the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that the first database system supports parsing; The first database system divides the third query plan into at least two plan slices, where the at least two plan slices include a first plan slice and a second plan slice; In the case where the first database system determines, based on the first planned sharding, that the second database system is more suitable for executing the first planned sharding than the first database system, the first database system converts the first planned sharding into the first planned sharding. IR of a slice, where the IR of the first planned slice is described in a language that both the first database system and the second database system can parse; The first database system sends the IR of the first plan shard to the second database system; The second database system converts the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system supports parsing; The second database system executes the fourth query plan to obtain a second query result; When the first database system determines that the third database system is more suitable for executing the second plan sharding than the first database system based on the second plan sharding, the first database system converts the second plan sharding into an IR of the second plan sharding, where the IR of the second plan sharding is described in a language that both the first database system and the third database system support parsing; The first database system sends the IR of the second plan shard to the third database system; The third database system converts the IR of the second plan shard into a fifth query plan, where the fifth query plan is described in a language that the third database system supports parsing; The third database system executes the fifth query plan to obtain a third query result.
6. The method according to any one of claims 1 to 5, characterized in that: The intermediate representation includes a field and an associated identifier, wherein the identifier is used to indicate that the field is a unique type of data in the first database system. The second database system includes a parser, which is used to parse the intermediate representation, and the parser is configured to recognize the identifier and parse the unique type of data.
7. The method according to any one of claims 1 to 6, characterized in that: The language of the IR is substrait.
8. The method according to any one of claims 1 to 7, characterized in that: The second database system is more suitable for executing the first query plan than the first database system, including: The execution performance of the first query plan executed on the second database system is higher than the execution performance of the first query plan executed on the first database system; or, An execution cost of executing the first query plan on the second database system is lower than an execution cost of executing the first query plan on the first database system.
9. A fusion database system, characterized in that: The fusion database system includes a first database system and a second database system, wherein the second database system and the first database system have different language descriptions of query plans: The first database system is used to obtain a first query plan for executing the first query according to a first query input by a user, wherein the first query plan is described in a language that the first database system supports parsing; the first database system being configured to convert the first query plan into an intermediate representation IR of the first query plan when the first database system determines that the second database system is more suitable for executing the first query plan than the first database system based on the first query plan, wherein the IR of the first query plan is described in a language that both the first database system and the second database system can parse; The first database system is used to send the IR of the first query plan to the second database system; The second database system is used to convert the IR of the first query plan into a second query plan, where the second query plan is described in a language that the second database system supports parsing; The second database system is used to execute the second query plan to obtain the first query result.
10. The system according to claim 9, characterized in that The first database system is further configured to receive the first query result from the second database system and return the result to the user; or, The second database system is further used to return the first query result to the user.
11. The system according to claim 9 or 10, characterized in that: The first database system is further configured to obtain, according to the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that the first database system supports parsing; The first database system is further used to divide the third query plan into at least two plan slices, wherein the at least two plan slices include a first plan slice and a second plan slice; The first database system is further configured to convert the first plan shard into an IR of the first plan shard if the first database system determines that the second database system is more suitable for executing the first plan shard than the first database system based on the first plan shard, wherein the IR of the first plan shard is described in a language that both the first database system and the second database system support parsing; The first database system is further used to send the IR of the first plan shard to the second database system; The second database system is further used to convert the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system supports parsing; The second database system is further used to execute the fourth query plan to obtain a second query result; The first database system is further configured to execute the second plan sharding to obtain a third query result if the first database system determines, based on the second plan sharding, that the first database system is more suitable for executing the second plan sharding than the second database system.
12. The system according to claim 11, characterized in that The third query plan includes a plurality of sub-query plans, and one plan slice of the plurality of plan slices corresponds to at least one sub-query plan of the plurality of sub-query plans; or, The query plan includes multiple operators, and one plan slice among the multiple plan slices corresponds to at least one operator among the multiple operators.
13. The system according to claim 9 or 10, characterized in that The fusion database system further includes a third database system, wherein the third database system and the first database system have different language descriptions of query plans; The first database system is further configured to obtain, according to the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that the first database system supports parsing; The first database system is further used to divide the third query plan into at least two plan slices, wherein the at least two plan slices include a first plan slice and a second plan slice; The first database system is further configured to convert the first plan shard into an IR of the first plan shard if the first database system determines that the second database system is more suitable for executing the first plan shard than the first database system based on the first plan shard, wherein the IR of the first plan shard is described in a language that both the first database system and the second database system support parsing; The first database system is further used to send the IR of the first plan shard to the second database system; The second database system is further used to convert the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system supports parsing; The second database system is further used to execute the fourth query plan to obtain a second query result; The first database system is further configured to, when the first database system determines based on the second plan sharding that the third database system is more suitable for executing the second plan sharding than the first database system, convert the second plan sharding into an IR of the second plan sharding, wherein the IR of the second plan sharding is described in a language that both the first database system and the third database system support parsing; The first database system is further used to send the IR of the second plan shard to the third database system; The third database system is further used to convert the IR of the second plan shard into a fifth query plan, where the fifth query plan is described in a language that the third database system supports parsing; The third database system executes the fifth query plan to obtain a third query result.
14. The system according to any one of claims 9 to 13, characterized in that: The intermediate representation includes a field and an associated identifier, wherein the identifier is used to indicate that the field is a unique type of data in the first database system. The second database system includes a parser, which is used to parse the intermediate representation, and the parser is configured to recognize the identifier and parse the unique type of data.
15. The system according to any one of claims 9 to 14, characterized in that: The language of the IR is substrait.
16. The system according to any one of claims 9 to 15, characterized in that: The execution performance of executing the first query plan on the second database system is higher than the execution performance of executing the first query plan on the first database system; or, An execution cost of executing the first query plan on the second database system is lower than an execution cost of executing the first query plan on the first database system.
17. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 8.
18. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 8.
19. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Statement conversion method and system between databases and, terminal
CN109857757A
Data query platform, method, and equipment and storage medium
CN110222072A
Language conversion method and device of database, electronic equipment and storage medium
CN111061757A
Cross-domain data query method and device
CN111190924A
Distributed memory data query optimization method and equipment
CN113568930A