Data processing method and database system
By converting query plans into intermediate representations (IR), the problem of difficulty in sharing and collaboration between heterogeneous database systems is solved, reducing waste of computing resources and improving efficiency.
Patent Information
- Application Number
- CN202410559375.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-14
- Filing Date
- 2024-04-30
- Publication Date
- 2025-05-16
AI Technical Summary
Between heterogeneous database systems, it is difficult for the prior art to effectively share and collaborate in query plans, resulting in wasted computing resources.
By converting query plans into intermediate representations (IRs), it communicates between multiple database systems, avoiding duplicate compilation and optimization.
It reduces the waste of computing power overhead due to repeated compilation and optimization, and improves query plan sharing and collaboration efficiency between database systems.
Smart Images

Figure CN120011394A_ABST
Abstract
Description
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 14, 2023, with application number 202311525323.8 and application name “A request processing method and related equipment”, all contents of which are incorporated by reference in this application. Technical Field
[0002] The present application relates to the field of database technology, and in particular to a data processing method and a database system. Background Art
[0003] A distributed database is a logically unified database formed by connecting multiple physically dispersed database units using a computer network. Each connected database unit is called a site or node.
[0004] In a distributed database system, the base tables of the subquery and the parent query are distributed on multiple data nodes DN. In the existing technology, for a subquery, all base table data is scanned during execution, passed to the upper-level operator, and predicate conditions are filtered. Finally, the results of the subquery and the parent query are returned to CN, and CN returns the aggregated results to the user.
[0005] The query plans obtained by compilation and optimization of heterogeneous database systems are often not interoperable. In other words, the query plans generated by the database systems cannot be accurately parsed by the heterogeneous database systems. When query plans are shared and collaborated between heterogeneous database systems, for example, database system A determines through analysis (for example, analysis from factors such as performance and cost) that it is better to execute the query plan (partial or complete) in database system B. Database system A needs to transmit the corresponding information to database system B. Since database system B cannot parse the query plan that database system A has obtained, the original user query is transmitted to database system B in the prior art. After receiving the query, database system B needs to perform repeated compilation and optimization to obtain the query plan, which will lead to a waste of computing resources. Summary of the invention
[0006] In a first aspect, the present application provides a data processing method, which is applied to a fused database system, wherein the fused database system includes a first database system and a second database system, and the second database system and the first database system have different language descriptions of query plans. The method includes: the first database system obtains a first query plan for executing the first query based on a first query input by a user, and the first query plan is described in a language supported for parsing by the first database system; when the first database system determines that the second database system is more suitable for executing the first query plan than the first database system based on the first query plan, the first database system converts the first query plan into an intermediate representation IR of the first query plan, and the IR of the first query plan is described in a language supported for parsing by both the first database system and the second database system; the first database system sends the IR of the first query plan to the second database system; the second database system converts the IR of the first query plan into a second query plan, and the second query plan is described in a language supported for parsing by the second database system; the second database system executes the second query plan to obtain a first query result.
[0007] The query plans obtained by compilation and optimization of heterogeneous database systems (for example, the first database system and the second database system in the embodiment of the present application) are often not interoperable. That is, the query plan generated by the first database system cannot be accurately parsed by the second database system (or can be described as not supporting parsing). When sharing and collaborating query plans between heterogeneous database systems, for example, the first database system determines through analysis (for example, analysis from factors such as performance and cost) that it is better to execute the query plan (partially or completely) in the second database system (that is, it is more suitable for execution in the second database system), and the first database system needs to pass the corresponding information to the second database system. Since the second database system cannot parse the query plan already obtained by the first database system, the original user query is passed to the second database system in the prior art. After receiving the query, the second database system needs to perform repeated compilation and optimization to obtain the query plan, which will result in a waste of computing resources. In the embodiment of the present application, the first database system can convert the query plan into a language description that can be parsed by both the first database system and the second database system (that is, the intermediate representation of the first query plan in the embodiment of the present application). When it is determined that the query plan is more suitable for execution on other database systems (such as the second database system), the corresponding intermediate representation can be passed to the second database system. Since the intermediate representation is a language description that can be parsed by the second database system, the second database system does not need to recompile and optimize, but can directly convert the intermediate representation back to the query plan and execute it (or, if it can be determined that there are other database systems that are more suitable for execution, the intermediate representation can be passed to other database systems). In the above manner, the waste of computing power overhead caused by repeated compilation and optimization can be reduced.
[0008] In one possible implementation, the first database system may determine that the query plan (or part of it) is more suitable for execution on the second database system (for example, when the execution engine of the first database is under high load and the engine of the second database is idle, the query shards may be forwarded to the engine of the second database system for execution), and therefore an intermediate representation that is more suitable for execution on the second database system needs to be delivered to the second database system. Since the first database system and the second database system are heterogeneous database systems, the first database system needs to convert the first query plan into a language that the second database system supports parsing for description (that is, the intermediate representation IR of the first query plan).
[0009] For example, if a query submitted by a user to database system A is more suitable for execution on database system B, the query plan generated by database system A can be passed to database system B for execution through Substrait.
[0010] Through the above method, the optimization cost of repeated analysis can be avoided: in the scenario of selecting the optimal database, the optimization cost of traditional Double SQL analysis is avoided.
[0011] In a possible implementation, the method further includes: the first database system receives the first query result from the second database system and returns it to the user; or the second database system returns the first query result to the user.
[0012] In a possible implementation, the method further includes: the first database system obtains a third query plan for executing the second user query according to the received second user query, and the third query plan is described in a language supported by the first database system for parsing; the first database system divides the third query plan into at least two plan shards, and the at least two plan shards include a first plan shard and a second plan shard; when the first database system determines based on the first plan shard that the second database system is more suitable for executing the first plan shard than the first database system, the first database system converts the first plan shard into an IR of the first plan shard, and the IR of the first plan shard is described in a language supported by both the first database system and the second database system for parsing; the first database system sends the IR of the first plan shard to the second database system; the second database system converts the IR of the first plan shard into a fourth query plan, and the fourth query plan is described in a language supported by the second database system for parsing; the second database system executes the fourth query plan to obtain a second query result; when the first database system determines based on the second plan shard that the first database system is more suitable for executing the second plan shard than the second database system, the first database system executes the second plan shard to obtain a third query result.
[0013] Through the above method, when it is determined that part of the query plan (that is, the first plan shard) is more suitable for execution on the second database system, and part of the query plan (that is, the second plan shard) is more suitable for execution on itself (that is, the first database system), the plan suitable for execution on the second database system can be converted into a language supported by the second database system for parsing (that is, the IR of the first plan shard), and sent to the second database system, and the plan suitable for execution on itself (that is, the second plan shard) is executed.
[0014] In a possible implementation, the fusion database system also includes a third database system, and the third database system and the first database system have different language descriptions of the query plan. The method also includes: the first database system obtains a third query plan for executing the second user query based on the received second user query, and the third query plan is described in a language that the first database system supports parsing; the first database system divides the third query plan into at least two plan slices, and the at least two plan slices include a first plan slice and a second plan slice; when the first database system determines that the second database system is more suitable for executing the first plan slice than the first database system based on the first plan slice, the first database system converts the first plan slice into an IR of the first plan slice, and the IR of the first plan slice is described in a language that both the first database system and the second database system support parsing; the first database system converts the IR of the first plan slice Send to the second database system; the second database system converts the IR of the first plan shard into a fourth query plan, and the fourth query plan is described in a language that the second database system supports for parsing; the second database system executes the fourth query plan to obtain a second query result; when the first database system determines that the third database system is more suitable for executing the second plan shard than the first database system based on the second plan shard, the first database system converts the second plan shard into the IR of the second plan shard, and the IR of the second plan shard is described in a language that both the first database system and the third database system support for parsing; the first database system sends the IR of the second plan shard to the third database system; the third database system converts the IR of the second plan shard into a fifth query plan, and the fifth query plan is described in a language that the third database system supports for parsing; the third database system executes the fifth query plan to obtain a third query result.
[0015] Through the above method, when it is determined that part of the query plan (that is, the first plan shard) is more suitable for execution on the second database system, and part of the query plan (that is, the second plan shard) is more suitable for execution on the third database system, the plan suitable for execution on the second database system can be converted into a language supported by the second database system for parsing (that is, the IR of the first plan shard) and sent to the second database system, and the plan suitable for execution on the third database system can be converted into a language supported by the third database system for parsing (that is, the IR of the second plan shard) and sent to the third database system.
[0016] In one possible implementation, the third query plan includes multiple sub-query plans, and one plan slice among the multiple plan slices corresponds to at least one sub-query plan among the multiple sub-query plans; or, the query plan includes multiple operators, and one plan slice among the multiple plan slices corresponds to at least one operator among the multiple operators.
[0017] A query can be split into multiple shards. For example, if the coarse-grained division is based on the subquery plan, the query can be split into multiple subquery plans, and each plan shard can correspond to one or more subquery plans. For another example, the fine-grained division can be based on the operator, that is, the query can be split into multiple operators, and each plan shard can include one or more operators.
[0018] The first database system can reasonably split the query and allocate the database systems based on the performance or operation cost of other database systems, so as to improve the overall performance.
[0019] In a possible implementation, CN can perform a unified compilation and optimization to obtain a query plan and directly send it to multiple DNs. In this way, only one compilation and optimization is required to forward the SQL to DN for execution. This not only omits the query compilation and query optimization overhead of DN, but also ensures that each DN executes the same query plan.
[0020] In one possible implementation, a field for describing information unique to the database may be added to the intermediate representation to facilitate parsing of heterogeneous databases. In one possible implementation, the intermediate representation includes fields and associated identifiers, and the identifiers are used to indicate that the fields are data of a unique type in the first database. The second database includes a parser, and the parser is used to parse the intermediate representation, and the parser is configured to recognize the identifiers and parse the unique type of data.
[0021] In one possible implementation, language is a substrait.
[0022] In a possible implementation, the second database system is more suitable for executing the first query plan than the first database system, including: the execution performance of the first query plan on the second database system is higher than the execution performance of the first query plan on the first database system; or the execution cost of the first query plan on the second database system is lower than the execution cost of the first query plan on the first database system. For example, the first database system is in an overloaded state, while the second database system is in an idle state, or for another example, the efficiency of the engine of the second database system in parsing and executing the first query plan is higher than that of the engine of the first database system.
[0023] In a second aspect, the present application provides a fusion database system, which includes a first database system and a second database system, wherein the second database system and the first database system have different language descriptions of query plans:
[0024] A first database system is used to obtain a first query plan for executing the first query according to a first query input by a user, wherein the first query plan is described in a language that the first database system supports parsing;
[0025] The first database system is configured to convert the first query plan into an intermediate representation IR of the first query plan when the first database system determines that the second database system is more suitable for executing the first query plan than the first database system based on the first query plan, wherein the IR of the first query plan is described in a language that both the first database system and the second database system can parse;
[0026] The first database system is used to send the IR of the first query plan to the second database system;
[0027] A second database system is used to convert the IR of the first query plan into a second query plan, where the second query plan is described in a language that the second database system supports parsing;
[0028] The second database system is used to execute the second query plan to obtain the first query result.
[0029] In a possible implementation, the first database system is further configured to receive the first query result from the second database system and return the result to the user; or,
[0030] The second database system is also used to return the first query result to the user.
[0031] In a possible implementation, the first database system is further configured to obtain, according to the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that the first database system supports parsing;
[0032] The first database system is further used to divide the third query plan into at least two plan slices, where the at least two plan slices include a first plan slice and a second plan slice;
[0033] The first database system is further configured to convert the first plan sharding into an IR of the first plan sharding when the first database system determines that the second database system is more suitable for executing the first plan sharding than the first database system based on the first plan sharding, wherein the IR of the first plan sharding is described in a language that both the first database system and the second database system support parsing;
[0034] The first database system is further used to send the IR of the first plan shard to the second database system;
[0035] The second database system is further used to convert the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system supports parsing;
[0036] The second database system is further used to execute the fourth query plan to obtain the second query result;
[0037] The first database system is further configured to execute the second plan sharding to obtain a third query result when the first database system determines, based on the second plan sharding, that the first database system is more suitable for executing the second plan sharding than the second database system.
[0038] In a possible implementation, the third query plan includes multiple sub-query plans, and one plan slice among the multiple plan slices corresponds to at least one sub-query plan among the multiple sub-query plans; or,
[0039] The query plan includes multiple operators, and one plan slice among the multiple plan slices corresponds to at least one operator among the multiple operators.
[0040] In a possible implementation, the fusion database system further includes a third database system, and the third database system and the first database system have different language descriptions of the query plan;
[0041] The first database system is further configured to obtain, according to the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that the first database system supports parsing;
[0042] The first database system is further used to divide the third query plan into at least two plan slices, where the at least two plan slices include a first plan slice and a second plan slice;
[0043] The first database system is further configured to convert the first plan sharding into an IR of the first plan sharding when the first database system determines that the second database system is more suitable for executing the first plan sharding than the first database system based on the first plan sharding, wherein the IR of the first plan sharding is described in a language that both the first database system and the second database system support parsing;
[0044] The first database system is further used to send the IR of the first plan shard to the second database system;
[0045] The second database system is further used to convert the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system supports parsing;
[0046] The second database system is further used to execute the fourth query plan to obtain the second query result;
[0047] The first database system is further configured to convert the second plan shard into an IR of the second plan shard when the first database system determines that the third database system is more suitable for executing the second plan shard than the first database system based on the second plan shard, wherein the IR of the second plan shard is described in a language that both the first database system and the third database system can parse;
[0048] The first database system is further used to send the IR of the second plan shard to the third database system;
[0049] The third database system is further used to convert the IR of the second plan shard into a fifth query plan, where the fifth query plan is described in a language that the third database system supports parsing;
[0050] The third database system executes the fifth query plan to obtain a third query result.
[0051] In one possible implementation, the intermediate representation includes a field and an associated identifier, where the identifier is used to indicate that the field is a data of a unique type in the first database system. The second database system includes a parser, which is used to parse the intermediate representation, and the parser is configured to recognize the identifier and parse the data of the unique type.
[0052] In one possible implementation, the language of the IR is substrait.
[0053] In a possible implementation, the idleness of the load of the second database system is more suitable for executing the first query plan than the idleness of the load of the first database system; or,
[0054] The execution performance of the second database system is more suitable for executing the first query plan than the execution performance of the first database system.
[0055] In a third aspect, the present application provides a computing device cluster. The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is used to execute instructions stored in at least one memory, so that the computing device or the computing device cluster performs the method of the first aspect or any implementation of the first aspect.
[0056] In a fourth aspect, the present application provides a computer-readable storage medium storing instructions, wherein the instructions instruct a computing device or a computing device cluster to execute a method executed by a database system of the above-mentioned first aspect or any implementation of the first aspect.
[0057] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computing device or a computing device cluster, enables the computing device or the computing device cluster to execute the method executed by the database system of the above-mentioned first aspect or any implementation of the first aspect.
[0058] Based on the implementations provided in the above aspects, this application can also be further combined to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figures 1A to 1D This is an illustration of the application framework of this application;
[0060] Figure 2A and Figure 2B This is an illustration of the application framework of this application;
[0061] Figure 3 The following is a flowchart of a data processing method according to an embodiment of the present application;
[0062] Figure 4 The following is a flowchart of query processing in an embodiment of the present application;
[0063] Figure 5 It is a schematic diagram of an intermediate representation of an embodiment of the present application;
[0064] Figure 6 This is an illustration of query sharding in an embodiment of the present application;
[0065] Figure 7 The data processing method of the embodiment of the present application is shown in FIG.
[0066] Figure 8 This is a diagram of a database according to an embodiment of the present application;
[0067] Fig. 9 The data processing method of the embodiment of the present application is shown in FIG.
[0068] Figures 10 to 13 The structure of the data processing device according to the embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0069] The following describes the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of the present application, rather than all embodiments. It is known to those skilled in the art that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0070] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0071] The method provided in the embodiment of the present application can be applied to a database system. Figure 1A Shows a typical logical architecture of a database system. Figure 1A The database system 100 includes a database 110 and a database management system (DBMS) 130 .
[0072] Among them, the database 110 is an organized data set stored in the data storage 120, that is, a collection of associated data organized, stored and used according to a specific data model. According to the different data models used to organize data, data can be divided into multiple types, such as relational data, graph data, time series data, etc. Relational data is data modeled using a relational model, usually represented as a table, and the rows in the table represent a set of related values of an object or entity. Graph data, referred to as "graph", is used to represent the relationship between objects or entities, such as social relationships. Time series data, referred to as time series data, is a data column recorded and indexed in chronological order, used to describe the state change information of an object in the time dimension.
[0073] The database management system 130 is the core of the database system and is a system software for organizing, storing and maintaining data. The client 200 can access the database 110 through the database management system 130, and the database administrator also performs database maintenance work through the database management system. The database management system 130 provides multiple functions for the client 200 to establish, modify and query the database, wherein the client 200 can be an application or a user device. The functions provided by the database management system 130 may include but are not limited to the following: (1) Data definition function. The database management system 130 provides a data definition language (DDL) to define the structure of the database 110. DDL is used to describe the database framework and can be saved in the data dictionary; (2) Data access function. The database management system 130 provides a data manipulation language (DML) to implement basic access operations on the database 110, such as retrieval, insertion, modification and deletion; (3) Database operation management function. The database management system 130 provides a data control function to effectively control and manage the operation of the database 110 to ensure that the data is correct and valid; (4) Database establishment and maintenance function, including loading of initial database data, database dumping, recovery, reorganization, system performance monitoring, analysis and other functions; (5) Database transmission. The database management system provides transmission of processed data to achieve communication between the client and the database management system, which is usually coordinated with the operating system.
[0074] The data storage 120 includes, but is not limited to, solid state drives (SSDs), disk arrays, cloud storage, or other types of non-transient computer-readable storage media. Figure 1A Fewer or more components than those shown in the figure, or including Figure 1A The components shown are different components, Figure 1A Only components more relevant to the implementation disclosed in the embodiment of the present invention are shown.
[0075] The database system provided in the embodiment of the present application may be a distributed database system (DDBS). In the process of transaction processing, in order to achieve concurrency control between transactions, the DDBS usually uses a global transaction manager (GTM) to manage transactions. Figure 1B and Figure 1C Introducing DDBS.
[0076] Figure 1B The schematic diagram of a distributed database system using a shared-storage architecture includes one or more coordinator nodes (CN), multiple data nodes (DN), one or more GTMs (such as Figure 1B The first GTM and the second GTM in the distributed database system are used to back up the data of the first GTM. The second GTM is used to back up the data of the first GTM and take over the work of the first GTM when the first GTM fails, so as to ensure the high reliability of the DDBS. The CN and the DN communicate through a network channel. In one embodiment, the network channel can be composed of network devices such as switches, routers and gateways. The CN, DN and GTM jointly implement the functions of the database management system and provide database retrieval, insertion, modification and deletion services to the client. In one embodiment, a database management system is deployed on each CN, DN and GTM. The shared data storage device stores data that can be shared by multiple DNs, and the DN can perform read and write operations on the data in the data storage device through the network channel. The shared data storage device can be a shared disk array. The CN, DN, the first GTM or the second GTM in the distributed database system can be a physical machine, such as a database server, or a virtual machine (VM) or a container running on abstract hardware resources. In one embodiment, the CN, DN, the first GTM or the second GTM is a virtual machine or a container, and the network channel is a virtual switching network, which includes a virtual switch. The database management system deployed in CN, DN, the first GTM or the second GTM is a DBMS instance, which can be a process or a thread, and these DBMSs cooperate to complete the functions of the database relational system. In another embodiment, CN, DN, the first GTM or the second GTM is a physical machine, and the network channel includes one or more switches, and the switch is a storage area network (SAN) switch, an Ethernet switch, a fiber switch or other physical switching equipment.
[0077] Figure 1C Schematic diagram of a distributed database system using a shared-nothing architecture. Each DN has its own exclusive hardware resources (such as data storage), operating system and database. The CN, DN, first GTM or second GTM communicate through a network channel. The network channel can be referred to in the above Figure 1BUnder this system, data will be distributed to each DN according to the database model and application characteristics. The query task will be divided into several parts by CN and executed in parallel on all DNs. They will coordinate calculations with each other and provide database services as a whole. All communication functions are implemented on a high-bandwidth network interconnection system. Figure 1B Like the distributed database system of the shared-storage architecture described above, the CN, DN, first GTM or second GTM here can be either a physical machine or a virtual machine.
[0078] In all embodiments of the present application, the data storage of the database system includes but is not limited to solid state drives (SSDs), disk arrays, or other types of non-transitory computer-readable media. Figure 1B-1C Although the database is not shown in the figure, it should be understood that the database is stored in the data storage. A person skilled in the art can understand that a database system may include Figures 1A-1C Fewer or more components than those shown in the figure, or including Figures 1A-1C The components shown in the figure are different components. Figures 1A-1C Only components more relevant to the implementation disclosed in the embodiments of the present application are shown. However, those skilled in the art can understand that a distributed database system can include any number of CNs and DNs. The database management system functions of each CN and DN can be implemented by a suitable combination of software, hardware and / or firmware running on each CN and DN.
[0079] Above Figure 1B and Figure 1C The described distributed database system includes multiple DNs and multiple CNs, wherein the function of each DN is substantially the same, and the function of each CN is also substantially the same.
[0080] Figure 1D An application data management system provided by an embodiment of the present application is shown. Specifically, as shown, a data node cluster may include multiple data node DNs. Application developers may deploy relevant data of developed applications in corresponding data node DNs. Applications may send service requests to application servers, and application servers may convert service requests into data operation requests; application servers may send data operation requests to distributed database systems. Then, one or more data node DNs in the distributed database system perform operations corresponding to the operation requests on data in the database memory. Specifically, the application server may send data operation requests to coordination nodes in the coordination node cluster, and the coordination nodes may forward the operation requests to relevant DNs, and the relevant DNs may perform data operations corresponding to the operation requests.
[0081] Reference Figure 2A , Figure 2A A system architecture diagram of an embodiment of the present application: wherein the present application can be applied to the architecture of a heterogeneous database system, such as Figure 2A As shown, the heterogeneous database includes a database system A and a database system B. The database system A and the database system B are heterogeneous databases. The so-called heterogeneous databases can be understood as databases with different processing tasks, structures or characteristics.
[0082] For example, Gauss (DWS) is a powerful distributed database suitable for large-scale data processing and analysis; SparkSQL is a database system built on Apache Spark, focusing on big data processing; Clickhouse is a columnar database suitable for high-speed data query and analysis; and Velox is an emerging scalable database designed to process large-scale graph data.
[0083] There are many differences and features between different database systems. They can be divided into relational databases (such as MySQL, PostgreSQL) and non-relational databases (such as MongoDB, Cassandra) based on the different data storage models. In addition, databases can also be divided into transactional databases, which are used to support complex data transaction processing, and analytical databases, which are used for high-performance data analysis and report generation, based on their main uses.
[0084] like Figure 2B In a common system with heterogeneous databases, if you want to implement the inter-database transfer of subquery plans, there will be a Double parsing optimization problem and unnecessary computational overhead. Figure 2B In the example, a query sent by a user to database system B (①) is found to be more suitable for execution in database system A only after query compilation and optimization. Then, the query is forwarded to database system A in the form of original SQL (②), and query compilation and optimization are performed on database system A. In database system A, CN forwards the query task in the form of query plan or SQL to a specific DN for execution (③). Compared with the ideal situation where the user directly sends the query to database system A (④), this architecture has the overhead of repeated query compilation and query optimization. In addition, this architecture cannot support more fine-grained optimal database selection such as subqueries or operations.
[0085] In order to solve the above problems, the present application embodiment provides a data processing method, referring to Figure 3 , Figure 3 A schematic diagram of a data processing method provided in an embodiment of the present application includes:
[0086] 301. The first database system obtains a first query plan for executing the first query according to a first query input by a user. The first query plan is described in a language that the first database system supports parsing.
[0087] In a possible implementation, query compilation and query optimization may be performed on a first query (e.g., an SQL query) input by a user to obtain a first query plan. For example, query compilation and query optimization may be performed on the first query input by a user through a CN node in a database system to obtain a first query plan.
[0088] 302. When the first database system determines, based on the first query plan, that the second database system is more suitable for executing the first query plan than the first database system, the first database system converts the first query plan into an intermediate representation IR of the first query plan, where the IR of the first query plan is described in a language that both the first database system and the second database system support parsing.
[0089] The query plans obtained by compilation and optimization of heterogeneous database systems (for example, the first database system and the second database system in the embodiment of the present application) are often not interoperable. That is, the query plan generated by the first database system cannot be accurately parsed by the second database system (or can be described as not supporting parsing). When sharing and collaborating query plans between heterogeneous database systems, for example, the first database system determines through analysis (for example, analysis from factors such as performance and cost) that it is better to execute the query plan (partially or completely) in the second database system (that is, it is more suitable for execution in the second database system), and the first database system needs to pass the corresponding information to the second database system. Since the second database system cannot parse the query plan already obtained by the first database system, the original user query is passed to the second database system in the prior art. After receiving the query, the second database system needs to perform repeated compilation and optimization to obtain the query plan, which will result in a waste of computing resources. In the embodiment of the present application, the first database system can convert the query plan into a language description that can be parsed by both the first database system and the second database system (that is, the intermediate representation of the first query plan in the embodiment of the present application). When it is determined that the query plan is more suitable for execution on other database systems (such as the second database system), the corresponding intermediate representation can be passed to the second database system. Since the intermediate representation is a language description that can be parsed by the second database system, the second database system does not need to recompile and optimize, but can directly convert the intermediate representation back to the query plan and execute it (or, if it can be determined that there are other database systems that are more suitable for execution, the intermediate representation can be passed to other database systems). In the above manner, the waste of computing power overhead caused by repeated compilation and optimization can be reduced.
[0090] For example, see Figure 8 , Figure 8 The two heterogeneous database systems (DB1 and DB2) may each include a module for generating a query plan, a module for converting into an intermediate representation, a parser, and an optimizer. The database systems may transfer the intermediate representation to collaborate in query execution.
[0091] The language used to describe the intermediate representation may be, but is not limited to, substrait or other common languages of heterogeneous database systems.
[0092] Taking substrait as an example, substrait's universal query plan representation allows users and applications to define and describe queries in a unified way without having to worry about the differences in the underlying databases. This not only simplifies the query and data operation process, but also promotes collaboration and interoperability between different database systems. It provides users with greater flexibility, allowing them to take advantage of the respective strengths of different database systems without having to pay too much attention to the underlying technical details. This helps to improve the consistency and collaboration of data in a multi-database environment, allowing users to better meet the needs of different application scenarios and improve the performance and efficiency of the overall system.
[0093] Optionally, the embodiment of the present application may use substrait to uniformly express query plans generated by multiple databases. Figure 4 , Figure 4 This is the conversion process from SQL to Substrait. For the query (SQL) input by the user, after completing the query compilation and query optimization of the database, a query plan is generated, which can then be converted to Substrait.
[0094] in, Figure 4 The substrait shown describes a query plan through four types of objects. Their names and functions are as follows:
[0095] Type: used to describe data types. Through common type definitions such as i8 and i32, the type and word length of data can be clearly defined to avoid incompatibility of data types between different databases.
[0096] Relation: used to describe the operation type and its connection method. An operation node in a database query plan generally corresponds to a Relation node of a Substrait, which records the type of the operation and the columns it acts on. In addition, it also records which Relations are the input of the operation. A tree-like query plan is represented as a Relation tree in Substrait.
[0097] Expression: defines the expression. Some operation expressions in the query plan are recorded through Expression.
[0098] Function: defines the functions used and their parameters and return value types.
[0099] In addition to the objects defined above, fields for describing database-specific information may be added to the intermediate representation to facilitate parsing of heterogeneous databases. In one possible implementation, the intermediate representation includes fields and associated identifiers, where the identifiers are used to indicate that the fields are data of a unique type in the first database. The second database includes a parser, which is used to parse the intermediate representation, and the parser is configured to recognize the identifiers and parse the unique type of data.
[0100] For example, the associated identifier can be "Node", see Figure 5 , Figure 5 This is a schematic diagram of an intermediate representation, in which the Node field is added on the basis of Substrait to support recording some information unique to some databases and improve the support capability of unified expression across databases. The database can design an extensible parser for the unified IR expression to parse and deparse some unique information, thereby supporting efficient execution between heterogeneous databases.
[0101] 303. The first database system sends the IR of the first query plan to the second database system.
[0102] In a diverse database ecosystem, it is critical to choose the database system that best suits a specific scenario. Specifically, when a database receives a query, the query may be more suitable for execution on other database systems. This decision (that is, which database system to choose) requires a comprehensive consideration of multiple factors, two of which are performance and cost.
[0103] First, performance is one of the key factors in database selection. Different database systems exhibit different performance characteristics on various computing tasks. Some database systems may excel in transaction processing, while others are more advantageous in large-scale data analysis. In addition, the difference in database performance may also increase significantly under different load conditions. Therefore, it is crucial to choose the most appropriate database system based on the performance requirements and load conditions of the application.
[0104] Secondly, cost is another factor that needs to be carefully considered. The computing resources (e.g., CPU, memory) and storage resources required by different database systems vary greatly, which directly affects the operating costs. Some database systems may require high-performance hardware and a large amount of memory, resulting in significant hardware costs. Others may run well on relatively low hardware configurations, thereby reducing overall costs. Choosing the most cost-effective database system can effectively control the project budget.
[0105] Considering performance and cost, you need to find a balance to ensure that the database selection can meet performance requirements while running within a reasonable cost range. This may require in-depth performance testing and cost analysis to determine the best database choice.
[0106] Existing technologies have provided some useful modeling tools for database selection. These modeling methods can evaluate database systems based on multiple dimensions, including cloud environment billing models, performance estimation models, and resource estimation models. These models can help decision makers better understand the performance and cost characteristics of different database systems, so as to more wisely choose a database system that is suitable for their application scenarios.
[0107] In one possible implementation, the first database may determine that the query plan (or part of it) is better executed on the second database (for example, when the execution engine of the first database is highly loaded and the engine of the second database is idle, the query shards may be forwarded to the engine of the second database for execution), and therefore an intermediate representation that is more suitable for execution on the second database needs to be passed to the second database.
[0108] In a possible implementation, when the execution performance of the first query plan executed on the second database system is higher than the execution performance of the first query plan executed on the first database system, it can be considered that the second database system is more suitable for executing the first query plan than the first database system.
[0109] In a possible implementation, when the execution cost of executing the first query plan on the second database system is lower than the execution cost of executing the first query plan on the first database system, it can be considered that the second database system is more suitable for executing the first query plan than the first database system.
[0110] For example, the first database system is in an overloaded state, while the second database system is in an idle state. For another example, the efficiency of the engine of the second database system in parsing and executing the first query plan is higher than that of the engine of the first database system.
[0111] For example, if the query submitted by the user to database system A is more suitable for execution on database system B, the query plan generated by database system A can be passed to database system B for execution through Substrait. In this way, repeated parsing optimization costs can be avoided: in the scenario of selecting the optimal database, the optimization cost of traditional Double SQL parsing is avoided.
[0112] In a possible implementation, the method further includes: the first database system obtains a third query plan for executing the second user query according to the received second user query, and the third query plan is described in a language supported by the first database system for parsing; the first database system divides the third query plan into at least two plan shards, and the at least two plan shards include a first plan shard and a second plan shard; when the first database system determines based on the first plan shard that the second database system is more suitable for executing the first plan shard than the first database system, the first database system converts the first plan shard into an IR of the first plan shard, and the IR of the first plan shard is described in a language supported by both the first database system and the second database system for parsing; the first database system sends the IR of the first plan shard to the second database system; the second database system converts the IR of the first plan shard into a fourth query plan, and the fourth query plan is described in a language supported by the second database system for parsing; the second database system executes the fourth query plan to obtain a second query result; when the first database system determines based on the second plan shard that the first database system is more suitable for executing the second plan shard than the second database system, the first database system executes the second plan shard to obtain a third query result.
[0113] Through the above method, when it is determined that part of the query plan (that is, the first plan shard) is more suitable for execution on the second database system, and part of the query plan (that is, the second plan shard) is more suitable for execution on itself (that is, the first database system), the plan suitable for execution on the second database system can be converted into a language supported by the second database system for parsing (that is, the IR of the first plan shard), and sent to the second database system, and the plan suitable for execution on itself (that is, the second plan shard) is executed.
[0114] In a possible implementation, the fusion database system also includes a third database system, and the third database system and the first database system have different language descriptions of the query plan. The method also includes: the first database system obtains a third query plan for executing the second user query based on the received second user query, and the third query plan is described in a language that the first database system supports parsing; the first database system divides the third query plan into at least two plan slices, and the at least two plan slices include a first plan slice and a second plan slice; when the first database system determines that the second database system is more suitable for executing the first plan slice than the first database system based on the first plan slice, the first database system converts the first plan slice into an IR of the first plan slice, and the IR of the first plan slice is described in a language that both the first database system and the second database system support parsing; the first database system converts the IR of the first plan slice Send to the second database system; the second database system converts the IR of the first plan shard into a fourth query plan, and the fourth query plan is described in a language that the second database system supports for parsing; the second database system executes the fourth query plan to obtain a second query result; when the first database system determines that the third database system is more suitable for executing the second plan shard than the first database system based on the second plan shard, the first database system converts the second plan shard into the IR of the second plan shard, and the IR of the second plan shard is described in a language that both the first database system and the third database system support for parsing; the first database system sends the IR of the second plan shard to the third database system; the third database system converts the IR of the second plan shard into a fifth query plan, and the fifth query plan is described in a language that the third database system supports for parsing; the third database system executes the fifth query plan to obtain a third query result.
[0115] Through the above method, when it is determined that part of the query plan (that is, the first plan shard) is more suitable for execution on the second database system, and part of the query plan (that is, the second plan shard) is more suitable for execution on the third database system, the plan suitable for execution on the second database system can be converted into a language supported by the second database system for parsing (that is, the IR of the first plan shard) and sent to the second database system, and the plan suitable for execution on the third database system can be converted into a language supported by the third database system for parsing (that is, the IR of the second plan shard) and sent to the third database system.
[0116] In one possible implementation, the third query plan includes multiple sub-query plans, and one plan slice among the multiple plan slices corresponds to at least one sub-query plan among the multiple sub-query plans; or, the query plan includes multiple operators, and one plan slice among the multiple plan slices corresponds to at least one operator among the multiple operators.
[0117] A query can be split into multiple shards. For example, if the coarse-grained division is based on the subquery plan, the query can be split into multiple subquery plans, and each plan shard can correspond to one or more subquery plans. For another example, the fine-grained division can be based on the operator, that is, the query can be split into multiple operators, and each plan shard can include one or more operators.
[0118] The first database system can reasonably split the query and allocate the database systems based on the performance or operation cost of other database systems, so as to improve the overall performance.
[0119] In a possible implementation, CN can perform a unified compilation and optimization to obtain a query plan and directly send it to multiple DNs. In this way, only one compilation and optimization is required to forward the SQL to DN for execution. This not only omits the query compilation and query optimization overhead of DN, but also ensures that each DN executes the same query plan.
[0120] Taking Substrait as an example, the query plan sharding technology in the embodiments of the present application can support more fine-grained optimal database execution selection for subqueries or operations. By performing more fine-grained query plan sharding on the query plan of Substrait's unified IR, the most appropriate database execution engine can be selected for each sub-plan or even each operation.
[0121] For example, refer to Figure 6 , Figure 6 This is a schematic diagram of query sharding, which includes execution engines of three databases. The same query can be executed by the execution engines of the three databases.
[0122] In a possible implementation, the coordination node CN of the first database converts the query plan into an intermediate representation (IR), and the data node DN of the first database executes the query plan corresponding to the second sub-representation to obtain a second query result.
[0123] For example, refer to Figure 7 , CN needs to distribute the generated query plan to multiple DNs for execution. The client submits the SQL to the CN, which obtains the intermediate representation and then sends it to the corresponding DN for execution. After receiving the intermediate representation, the DN converts it into a query plan for execution. Compared with the traditional method of directly forwarding SQL to the DN for execution, it not only omits the query compilation and query optimization overhead of the DN, but also ensures that each DN executes the same query plan.
[0124] For example, refer to Fig. 9 , Fig. 9It is a specific conversion and execution process, in which the query plan can be converted into the intermediate representation described by the substrait through "to substrait", and the intermediate representation described by the substrait can be converted into the query plan through "from substrait".
[0125] 304. The second database system converts the IR of the first query plan into a second query plan, where the second query plan is described in a language that the second database system supports parsing.
[0126] 305. The second database system executes the second query plan to obtain the first query result.
[0127] In a possible implementation, the first database system may receive the first query result from the second database system and return it to the user; or the second database system returns the first query result to the user.
[0128] Reference Fig.10 The embodiment of the present application also provides a fusion database system 1000, which can be used to implement the data processing method provided by any possible implementation method of the above method embodiment of the present application. The fusion database system 1000 includes a first database system 1001 and a second database system 1002. The second database system 1002 and the first database system 1001 have different language descriptions of query plans:
[0129] The first database system 1001 is used to obtain a first query plan for executing the first query according to a first query input by a user, wherein the first query plan is described in a language that the first database system 1001 supports parsing;
[0130] The first database system 1001 is configured to convert the first query plan into an intermediate representation IR of the first query plan when the first database system 1001 determines that the second database system 1002 is more suitable for executing the first query plan than the first database system 1001 based on the first query plan, wherein the IR of the first query plan is described in a language that both the first database system 1001 and the second database system 1002 can parse;
[0131] The first database system 1001 is used to send the IR of the first query plan to the second database system 1002;
[0132] The second database system 1002 is used to convert the IR of the first query plan into a second query plan, where the second query plan is described in a language that the second database system 1002 supports parsing;
[0133] The second database system 1002 is used to execute the second query plan to obtain the first query result.
[0134] In a possible implementation, the first database system 1001 is further configured to receive the first query result from the second database system 1002 and return the result to the user; or,
[0135] The second database system 1002 is also used to return the first query result to the user.
[0136] In a possible implementation, the first database system 1001 is further configured to obtain, according to the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that the first database system 1001 supports parsing;
[0137] The first database system 1001 is further configured to divide the third query plan into at least two plan slices, where the at least two plan slices include a first plan slice and a second plan slice;
[0138] The first database system 1001 is further configured to convert the first planned sharding into an IR of the first planned sharding when the first database system 1001 determines that the second database system 1002 is more suitable for executing the first planned sharding than the first database system 1001 based on the first planned sharding, wherein the IR of the first planned sharding is described in a language that both the first database system 1001 and the second database system 1002 support parsing;
[0139] The first database system 1001 is further used to send the IR of the first plan shard to the second database system 1002;
[0140] The second database system 1002 is further used to convert the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system 1002 supports parsing;
[0141] The second database system 1002 is further used to execute the fourth query plan to obtain the second query result;
[0142] The first database system 1001 is further configured to execute the second planned sharding to obtain a third query result when the first database system 1001 determines that the first database system 1001 is more suitable for executing the second planned sharding than the second database system 1002 based on the second planned sharding.
[0143] In a possible implementation, the third query plan includes multiple sub-query plans, and one plan slice among the multiple plan slices corresponds to at least one sub-query plan among the multiple sub-query plans; or,
[0144] The query plan includes multiple operators, and one plan slice among the multiple plan slices corresponds to at least one operator among the multiple operators.
[0145] In a possible implementation, the fusion database system further includes a third database system, and the third database system and the first database system 1001 have different language descriptions of the query plan;
[0146] The first database system 1001 is further configured to obtain, according to the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that the first database system 1001 supports parsing;
[0147] The first database system 1001 is further configured to divide the third query plan into at least two plan slices, where the at least two plan slices include a first plan slice and a second plan slice;
[0148] The first database system 1001 is further configured to convert the first planned sharding into an IR of the first planned sharding when the first database system 1001 determines that the second database system 1002 is more suitable for executing the first planned sharding than the first database system 1001 based on the first planned sharding, wherein the IR of the first planned sharding is described in a language that both the first database system 1001 and the second database system 1002 support parsing;
[0149] The first database system 1001 is further used to send the IR of the first plan shard to the second database system 1002;
[0150] The second database system 1002 is further used to convert the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system 1002 supports parsing;
[0151] The second database system 1002 is further used to execute the fourth query plan to obtain the second query result;
[0152] The first database system 1001 is further configured to convert the second plan sharding into an IR of the second plan sharding when the first database system 1001 determines that the third database system is more suitable for executing the second plan sharding than the first database system 1001 based on the second plan sharding, wherein the IR of the second plan sharding is described in a language that both the first database system 1001 and the third database system can parse;
[0153] The first database system 1001 is further used to send the IR of the second plan shard to the third database system;
[0154] The third database system is further used to convert the IR of the second plan shard into a fifth query plan, where the fifth query plan is described in a language that the third database system supports parsing;
[0155] The third database system executes the fifth query plan to obtain a third query result.
[0156] In one possible implementation, the intermediate representation includes a field and an associated identifier, where the identifier is used to indicate that the field is a unique type of data in the first database system 1001. The second database system 1002 includes a parser, which is used to parse the intermediate representation, and the parser is configured to recognize the identifier and parse the unique type of data.
[0157] In one possible implementation, the language of the IR is substrait.
[0158] In a possible implementation, the idleness of the load of the second database system 1002 is more suitable for executing the first query plan than the idleness of the load of the first database system 1001; or,
[0159] The execution performance of the second database system 1002 is more suitable for executing the first query plan than the execution performance of the first database system 1001 .
[0160] Specifically, the specific implementation of the fusion database system 1000 performing various operations of the data processing method can refer to the description of the relevant content in the above method embodiment, which will not be repeated here.
[0161] In the embodiment of the present application, the first database system and the second database system can be implemented by software or hardware. The implementation of the first database system is introduced below by taking the first database system as an example, and the implementation of the second database system can be used as a reference.
[0162] The first database system is an example of a software functional unit. The first database system may include code running on a computing instance. The computing instance may be at least one of a physical host (computing device), a virtual machine, a container, and other computing devices. Furthermore, the computing device may be one or more. For example, a node may include code running on multiple hosts, or virtual machines, or containers. It should be noted that multiple hosts, or virtual machines, or containers for running the code may be distributed in the same AZ or in different AZs, and each AZ includes a data center or multiple data centers with similar geographical locations. Multiple hosts, or virtual machines, or containers for running the code may be distributed in the same region or in different regions. Generally, a region may include multiple AZs, and a VPC is set in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.
[0163] Similarly, multiple hosts or virtual machines or containers for running the code may be distributed in the same VPC or in multiple VPCs, wherein usually a region may include multiple AZs.
[0164] The first database system is an example of a hardware functional unit, and the first database system may include at least one computing device, such as a server, etc. Alternatively, the module may also be a device implemented by ASIC or PLD, etc. The PLD may be implemented by CPLD, FPGA, GAL or any combination thereof.
[0165] The multiple computing devices included in the first database system can be distributed in the same AZ or in different AZs. The multiple computing devices included in the module can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the module can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs. The present application also provides a computing device 1000. Fig.11 As shown, computing device 1000 includes: bus 1002, processor 1004, memory 1006 and communication interface 1008. Processor 1004, memory 1006 and communication interface 1008 communicate through bus 1002. Computing device 1000 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in computing device 1000.
[0166] The bus 1002 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.11 The bus 1002 may include a path for transmitting information between various components of the computing device 1000 (eg, the memory 1006, the processor 1004, and the communication interface 1008).
[0167] The processor 1004 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0168] The memory 1006 may include a volatile memory, such as a random access memory (RAM). The memory 1006 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid state drive (SSD). The memory 1006 stores executable program code, and the processor 1004 executes the executable program code to implement the method executed by the aforementioned database system. Specifically, the memory 1006 stores instructions for the method executed by the database system (e.g., the first database system and the second database system).
[0169] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 and other devices or a communication network.
[0170] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0171] like Fig.12 As shown, the computing device cluster includes at least one computing device 1000. Instructions for executing the method executed by the database system (eg, the first database system and the second database system) may be stored in the memory 1006 in one or more computing devices 1000 in the computing device cluster.
[0172] In some possible implementations, one or more computing devices 1000 in the computing device cluster may also be used to execute part of the instructions of the method executed by the database system (e.g., the first database system and the second database system). In other words, the combination of one or more computing devices 1000 may jointly execute the instructions of the method executed by the database system (e.g., the first database system and the second database system).
[0173] It should be noted that the memory 1006 in different computing devices 1000 in the computing device cluster may store different instructions for executing partial functions of the heterogeneous database system 100 .
[0174] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.13 A possible implementation is shown. Fig.13 As shown, two computing devices 120A and 120B are connected via a network. Specifically, the two computing devices are connected to the network via a communication interface in each computing device. In this type of possible implementation, the memory 116 in the computing device 120A may store instructions for executing the functions of the second database system. At the same time, the memory 116 in the computing device 120B may store instructions for executing the functions of the first database system. Alternatively, the memory 116 in the computing device 120A may store instructions for executing part of the functions of the first database system. At the same time, the memory 116 in the computing device 120B may store instructions for executing another part of the functions of the first database system.
[0175] It should be understood that Fig.13 The functions of the computing device 120A shown in FIG. 1 may also be performed by multiple computing devices. Similarly, the functions of the computing device 120B may also be performed by multiple computing devices.
[0176] The present application embodiment also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to that of Fig.12 and Fig.13 Connection mode of computing device cluster. The difference is that the memory 116 in one or more computing devices in the computing device cluster may store the same instructions for executing the template generation method.
[0177] In some possible implementations, the memory 116 of one or more computing devices in the computing device cluster may also store partial instructions for executing the template generation method. In other words, a combination of one or more computing devices may jointly execute instructions for executing the template generation method.
[0178] It should be noted that the memory 116 in different computing devices in the computing device cluster can store different instructions for executing part of the functions of the template generation method. That is, the instructions stored in the memory 116 in different computing devices can implement the functions of one or more modules in the second database system and the first database system.
[0179] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by the computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned method, training sample generation method, and model training method applied to the optimization problem solving system 100 for executing the database system execution.
[0180] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be a software or program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the method executed by the above-mentioned database system, the training sample generation method, and the model training method.
[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data processing method, characterized in that: Applied to a fusion database system, the fusion database system includes a first database system and a second database system, the second database system and the first database system have different language descriptions of query plans, the method includes: The first database system obtains a first query plan for executing the first query according to a first query input by a user, wherein the first query plan is described in a language that the first database system supports parsing; When the first database system determines, based on the first query plan, that the second database system is more suitable for executing the first query plan than the first database system, the first database system converts the first query plan into an intermediate representation IR of the first query plan, where the IR of the first query plan is described in a language that both the first database system and the second database system can parse; The first database system sends the IR of the first query plan to the second database system; The second database system converts the IR of the first query plan into a second query plan, where the second query plan is described in a language that the second database system supports parsing; The second database system executes the second query plan to obtain the first query result.
2. The method according to claim 1, characterized in that The method further comprises: The first database system receives the first query result from the second database system and returns it to the user; or, The second database system returns the first query result to the user.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: The first database system obtains, according to the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that the first database system supports parsing; The first database system divides the third query plan into at least two plan slices, where the at least two plan slices include a first plan slice and a second plan slice; When the first database system determines that the second database system is more suitable for executing the first plan sharding than the first database system based on the first plan sharding, the first database system converts the first plan sharding into an IR of the first plan sharding, where the IR of the first plan sharding is described in a language that both the first database system and the second database system support parsing; The first database system sends the IR of the first plan shard to the second database system; The second database system converts the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system supports parsing; The second database system executes the fourth query plan to obtain a second query result; When the first database system determines, based on the second plan sharding, that the first database system is more suitable for executing the second plan sharding than the second database system, the first database system executes the second plan sharding to obtain a third query result.
4. The method according to claim 3, characterized in that The third query plan includes a plurality of sub-query plans, and one plan slice of the plurality of plan slices corresponds to at least one sub-query plan of the plurality of sub-query plans; or, The query plan includes multiple operators, and one plan slice among the multiple plan slices corresponds to at least one operator among the multiple operators.
5. The method according to claim 1 or 2, characterized in that: The fusion database system further includes a third database system, and the third database system and the first database system have different language descriptions of query plans. The method further includes: The first database system obtains, according to the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that the first database system supports parsing; The first database system divides the third query plan into at least two plan slices, where the at least two plan slices include a first plan slice and a second plan slice; When the first database system determines that the second database system is more suitable for executing the first plan sharding than the first database system based on the first plan sharding, the first database system converts the first plan sharding into an IR of the first plan sharding, where the IR of the first plan sharding is described in a language that both the first database system and the second database system support parsing; The first database system sends the IR of the first plan shard to the second database system; The second database system converts the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system supports parsing; The second database system executes the fourth query plan to obtain a second query result; When the first database system determines that the third database system is more suitable for executing the second plan sharding than the first database system based on the second plan sharding, the first database system converts the second plan sharding into an IR of the second plan sharding, where the IR of the second plan sharding is described in a language that both the first database system and the third database system support parsing; The first database system sends the IR of the second plan shard to the third database system; The third database system converts the IR of the second plan shard into a fifth query plan, where the fifth query plan is described in a language that the third database system supports parsing; The third database system executes the fifth query plan to obtain a third query result.
6. The method according to any one of claims 1 to 5, characterized in that: The intermediate representation includes a field and an associated identifier, wherein the identifier is used to indicate that the field is a unique type of data in the first database system. The second database system includes a parser, which is used to parse the intermediate representation, and the parser is configured to recognize the identifier and parse the unique type of data.
7. The method according to any one of claims 1 to 6, characterized in that: The language of the IR is substrait.
8. The method according to any one of claims 1 to 7, characterized in that: The second database system is more suitable for executing the first query plan than the first database system, including: The execution performance of the first query plan executed on the second database system is higher than the execution performance of the first query plan executed on the first database system; or, An execution cost of executing the first query plan on the second database system is lower than an execution cost of executing the first query plan on the first database system.
9. A fusion database system, characterized in that: The fusion database system includes a first database system and a second database system, wherein the second database system and the first database system have different language descriptions of query plans: The first database system is used to obtain a first query plan for executing the first query according to a first query input by a user, wherein the first query plan is described in a language that the first database system supports parsing; the first database system being configured to convert the first query plan into an intermediate representation IR of the first query plan when the first database system determines that the second database system is more suitable for executing the first query plan than the first database system based on the first query plan, wherein the IR of the first query plan is described in a language that both the first database system and the second database system can parse; The first database system is used to send the IR of the first query plan to the second database system; The second database system is used to convert the IR of the first query plan into a second query plan, where the second query plan is described in a language that the second database system supports parsing; The second database system is used to execute the second query plan to obtain the first query result.
10. The system according to claim 9, characterized in that The first database system is further configured to receive the first query result from the second database system and return the result to the user; or, The second database system is further used to return the first query result to the user.
11. The system according to claim 9 or 10, characterized in that: The first database system is further configured to obtain, according to the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that the first database system supports parsing; The first database system is further used to divide the third query plan into at least two plan slices, wherein the at least two plan slices include a first plan slice and a second plan slice; The first database system is further configured to convert the first plan shard into an IR of the first plan shard if the first database system determines that the second database system is more suitable for executing the first plan shard than the first database system based on the first plan shard, wherein the IR of the first plan shard is described in a language that both the first database system and the second database system support parsing; The first database system is further used to send the IR of the first plan shard to the second database system; The second database system is further used to convert the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system supports parsing; The second database system is further used to execute the fourth query plan to obtain a second query result; The first database system is further configured to execute the second plan sharding to obtain a third query result if the first database system determines, based on the second plan sharding, that the first database system is more suitable for executing the second plan sharding than the second database system.
12. The system according to claim 11, characterized in that The third query plan includes a plurality of sub-query plans, and one plan slice of the plurality of plan slices corresponds to at least one sub-query plan of the plurality of sub-query plans; or, The query plan includes multiple operators, and one plan slice among the multiple plan slices corresponds to at least one operator among the multiple operators.
13. The system according to claim 9 or 10, characterized in that The fusion database system further includes a third database system, wherein the third database system and the first database system have different language descriptions of query plans; The first database system is further configured to obtain, according to the received second user query, a third query plan for executing the second user query, wherein the third query plan is described in a language that the first database system supports parsing; The first database system is further used to divide the third query plan into at least two plan slices, wherein the at least two plan slices include a first plan slice and a second plan slice; The first database system is further configured to convert the first plan shard into an IR of the first plan shard if the first database system determines that the second database system is more suitable for executing the first plan shard than the first database system based on the first plan shard, wherein the IR of the first plan shard is described in a language that both the first database system and the second database system support parsing; The first database system is further used to send the IR of the first plan shard to the second database system; The second database system is further used to convert the IR of the first plan shard into a fourth query plan, where the fourth query plan is described in a language that the second database system supports parsing; The second database system is further used to execute the fourth query plan to obtain a second query result; The first database system is further configured to, when the first database system determines based on the second plan sharding that the third database system is more suitable for executing the second plan sharding than the first database system, convert the second plan sharding into an IR of the second plan sharding, wherein the IR of the second plan sharding is described in a language that both the first database system and the third database system support parsing; The first database system is further used to send the IR of the second plan shard to the third database system; The third database system is further used to convert the IR of the second plan shard into a fifth query plan, where the fifth query plan is described in a language that the third database system supports parsing; The third database system executes the fifth query plan to obtain a third query result.
14. The system according to any one of claims 9 to 13, characterized in that: The intermediate representation includes a field and an associated identifier, wherein the identifier is used to indicate that the field is a unique type of data in the first database system. The second database system includes a parser, which is used to parse the intermediate representation, and the parser is configured to recognize the identifier and parse the unique type of data.
15. The system according to any one of claims 9 to 14, characterized in that: The language of the IR is substrait.
16. The system according to any one of claims 9 to 15, characterized in that: The execution performance of executing the first query plan on the second database system is higher than the execution performance of executing the first query plan on the first database system; or, An execution cost of executing the first query plan on the second database system is lower than an execution cost of executing the first query plan on the first database system.
17. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 8.
18. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 8.
19. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 8.
Citation Information
Cited By
Level Structure in Query Plan
US20240281439A1