Data processing method and database system

By pushing the predicate condition of the parent query into the subquery in a distributed database and using the constraints of the distributed keys to perform predicate filtering, the problem of subquery scanning all data is solved, and query performance and resource utilization efficiency are improved.

CN120256473APending Publication Date: 2025-07-04HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410008097.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In a distributed database, subqueries need to scan all base table data and perform predicate conditional filtering, resulting in high overhead of broadcast data, operators process large amounts of data, consume resources, and affect query performance.

Method used

Push the predicate condition of the parent query into the subquery, and use the constraints of the distribution key to reduce the amount of data processed by the subquery, and reduce the amount of data transmission and processing by performing predicate filtering at the data node level.

Benefits of technology

It improves query performance, reduces the scan volume and resource consumption of data nodes, and improves the efficiency of execution plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256473A_ABST
    Figure CN120256473A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method which comprises the steps that a first query and a second query are obtained, the second query is a sub-query of the first query, the first query comprises a first predicate condition, the second query comprises a second predicate condition, and if a parent query comprises one sub-query, the parent query comprises one sub-query; if the distribution conditions of predicates in the parent query and the child query are consistent, pushing the predicate condition of the parent query into the child query to obtain a third query; and sending the third query to a data node. In the distributed database, for one query, if the parent query contains one sub-query and the distribution conditions of the predicates in the parent query and the sub-query are consistent, the predicate condition of the parent query can be pushed down to the sub-query, so that the data volume processed by the sub-query is reduced, and the query performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of database technology, and in particular to a data processing method, a database system, and a device. Background Art

[0002] A distributed database is a logically unified database formed by connecting multiple physically dispersed database units through a computer network. Each connected database unit is called a site or a node.

[0003] In a distributed database system, the base table of a subquery and the base table of a parent query are distributed on multiple data nodes DN. In the prior art, for a subquery, when executing, all the base table data is scanned, passed to the upper-layer operator, and predicate condition filtering is performed. Finally, the results of the subquery and the parent query are returned to the CN, and the CN returns the aggregated result to the user.

[0004] However, in the prior art, for a subquery, all the data in the base table of the subquery needs to be scanned and passed to the upper-layer operator, and the predicate condition filtering operation is not performed until after passing through multiple layers of operators. When the data volume is very large, on the one hand, the overhead of broadcasting data is very large, and the operator needs to process a large amount of data, consuming a large amount of resources and affecting the query performance. Summary of the Invention

[0005] In a first aspect, this application provides a data processing method, the method includes: obtaining a first query and a second query, where the second query is a subquery of the first query, the first query includes a first predicate condition, the second query includes a second predicate condition, the first predicate condition indicates data in a first data range of a first table that satisfies a first constraint, the second predicate condition indicates data in a second data range of a second table whose relationship with the data in the first data range satisfies a second constraint, and both the first data range and the second data range correspond to distribution keys in the table; obtaining a third query according to the first query and the second query; the third query includes a third predicate condition, the third predicate condition indicates data in a second data range of the second table whose relationship with data in a third data range of the first table satisfies the second constraint, and the data in the third data range is data in the first data range that satisfies the first constraint; sending the third query to a data node where the second table is deployed.

[0006] In an embodiment of this application, in a distributed database, for a query, if the query contains a subquery and the distribution conditions of the predicates in the parent query and the subquery are the same, then the predicate condition of the parent query can be pushed down to the subquery to reduce the amount of data processed by the subquery, thereby improving the query performance.

[0007] The predicate condition of the first query as the parent query is a constraint on the data area (first data range) where the distribution key of the first table is located (that is, the first constraint in the embodiment of the present application). This constraint can indicate which data that needs to be queried in the parent query is in the data area where the distribution key of the first table is located.

[0008] The predicate condition of the second query as a subquery is also a constraint on the data area (second data range) where the distribution key of the second table is located (that is, the second constraint in the embodiment of the present application). The second constraint can indicate which data in the data area where the distribution key of the second table is located that need to be queried in the subquery, and the second constraint is related to the distribution key indicated by the predicate condition of the first query, that is, the second predicate condition indicates the data in the second data range of the second table whose relationship with the data in the first data range satisfies the second constraint.

[0009] There is a first constraint in the parent query for the data area where the distribution key of the first table is located, and the second constraint in the subquery is related to the data area where the distribution key of the first table is located. Since the predicate condition in the parent query must be met in the final query result, the first constraint can be directly placed in the subquery, that is, the first constraint is directly added to the content in the subquery related to the data area where the distribution key of the first table is located (that is, the third query in the embodiment of the present application), thereby greatly reducing the number of data nodes scanned.

[0010] In a possible implementation, the first constraint indicates data in the first data range of the first table that is equal to a preset constant. For example, data in column a of the first table that is equal to 1. In this case, "the third predicate condition indicates data that the relationship between the data in the second data range of the second table and the third data range of the first table satisfies the second constraint" can be understood as "the third predicate condition indicates data that the second data range of the second table satisfies the first constraint, that is, data that is equal to the preset constant".

[0011] In one possible implementation, the third predicate condition indicates data in the second data range of the second table that is equal to a preset constant, and the condition of "equal to the preset constant" is specified from a direct or indirect parent query of the second query, "an indirect parent query refers to the relationship between a deep subquery and a parent query when the parent query contains multiple layers of nested subqueries."

[0012] In a possible implementation, the third query does not include the first query.

[0013] That is, the third query is obtained by delegating the first constraint in the first query to the second query, rather than completely fusing the first query and the second query.

[0014] In a possible implementation, the first query is a sub-query of the fourth query, the fourth query includes a fourth predicate condition, and the fourth predicate condition indicates data in a fourth data range of a third table that satisfies a third constraint.

[0015] In a possible implementation, obtaining a third query according to the first query and the second query includes:

[0016] Based on the obtained indication information, obtaining a third query according to the first query and the second query; the indication information is used to indicate pushing the first predicate condition down into the second predicate condition included in the second query.

[0017] In a possible implementation, the method further includes:

[0018] Receiving data in the second table transmitted by the data node according to the third query.

[0019] In a possible implementation, the first data range is a column in the first table, and the second data range is a column in the second table.

[0020] In a second aspect, the present application provides a database system, including a coordination node and a data node:

[0021] The coordination node is used to execute the method according to any one of the first aspect;

[0022] The data node is used to scan the second table according to the third query to obtain data in the second table;

[0023] Transmitting the data in the second table to the coordination node.

[0024] In a third aspect, the present application provides a data processing device, and the device includes:

[0025] A transceiver module, configured to obtain a first query and a second query, where the second query is a sub-query of the first query, the first query includes a first predicate condition, the second query includes a second predicate condition, the first predicate condition indicates data in a first data range of a first table that satisfies a first constraint, the second predicate condition indicates a relationship between data in a second data range of a second table and the data in the first data range that satisfies a second constraint, and both the first data range and the second data range correspond to distribution keys in the table;

[0026] A processing module, configured to obtain a third query according to the first query and the second query; the third query includes a third predicate condition, and the third predicate condition indicates data in the second data range of the second table and data in the third data range of the first table, where the relationship between the data satisfies the second constraint, and the data in the third data range is data in the first data range that satisfies the first constraint;

[0027] The transceiver module is further configured to send the third query to a data node where the second table is deployed.

[0028] In a possible implementation, the first constraint indicates data in the first data range of the first table that is equal to a preset constant.

[0029] In a possible implementation, the third query does not include the first query.

[0030] In a possible implementation, the first query is a subquery of a fourth query, and the fourth query includes a fourth predicate condition, and the fourth predicate condition indicates data in the fourth data range of the third table that satisfies the third constraint.

[0031] In a possible implementation, the processing module is specifically configured to:

[0032] Based on the obtained indication information, obtain a third query according to the first query and the second query; the indication information is used to indicate pushing down the first predicate condition to the second predicate condition included in the second query.

[0033] In a possible implementation, the transceiver module is further configured to:

[0034] Receive data in the second table transmitted by the data node according to the third query.

[0035] In a possible implementation, the first data range is a column in the first table, and the second data range is a column in the second table.

[0036] In a fourth aspect of the present application, a data processing device is provided. The device may include at least one processor, a memory, and a communication interface. The processor is coupled to the memory and the communication interface. The memory is used to store instructions, the processor is used to execute the instructions, and the communication interface is used to communicate with other network elements under the control of the processor. When the instructions are executed by the processor, the processor executes the method in any possible implementation manner of the first aspect.

[0037] In a fifth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a program, and the program causes the processor to execute any one of the data processing methods in the first aspect and its various implementation manners.

[0038] In the sixth aspect of the present application, a computer program product is provided. The computer program product includes computer-executable instructions stored in a computer-readable storage medium. At least one processor of the device can read the computer-executable instructions from the computer-readable storage medium, and the execution of the computer-executable instructions by the at least one processor causes the device to implement the method provided in the first aspect or any possible implementation manner of the first aspect.

[0039] In the seventh aspect of the present application, a chip system is provided. The chip system includes a processor for supporting a data processing device to implement the functions involved in the first aspect or any possible implementation manner of the first aspect. In a possible design, the chip system may further include a memory for storing the necessary program instructions and data for the device for managing transactions. The chip system may be composed of chips or may include chips and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figures 1A to 1D Schematic diagram of the application framework of the present application;

[0041] Figure 2A and Figure 2B Schematic diagram of the application framework of the present application;

[0042] Figure 3 Schematic diagram of the process of the data processing method in the embodiment of the present application;

[0043] Figure 4 , Figure 6 , Figure 8 , Figure 10 Schematic diagram of the base in the embodiment of the present application;

[0044] Figure 5 , Figure 7 , Figure 9 , Figure 11 Schematic diagram of the execution plan in the embodiment of the present application;

[0045] Figure 12 Schematic diagram of the process of the data processing method in the embodiment of the present application;

[0046] Figures 13 to 15 Schematic diagram of the structure of the data processing device in the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The embodiments of the present application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Those of ordinary skill in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0048] In the description, claims and the above-mentioned drawings of this application, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0049] For ease of understanding, the relevant terms involved in the embodiments of this application will be introduced first below.

[0050] Distributed database: A distributed database is a logically unified database formed by connecting multiple physically dispersed database units with a computer network. Each connected database unit is called a site or node.

[0051] Key–value database: A key–value database, or key–value store, is a data storage paradigm designed to store, retrieve, and manage associative arrays, which is a data structure more commonly referred to as a "dictionary" or hash table today. A dictionary contains a collection of objects or records. In turn, each record has multiple different "fields" or columns. Again, each field contains data. These records are stored and retrieved using a "key" that uniquely identifies this record, and the key is also used to quickly find data in the database.

[0052] Base table: That is, a table. A table (TABLE) is an object used to store data in a database. It is a collection of structured data and is the foundation of the entire database system. A table is a database object that contains all the data in the database. A table is defined as a collection of columns.

[0053] Index: It is a sorted data structure in a database management system to assist in quickly querying and updating data in a database table.

[0054] NULL value: A null value. It is a special marker used in Structured Query Language and is an identifier for unknown or missing logarithmic attributes in a database, used to indicate uncertain values in the database. It was introduced by E.F. Codd, the creator of the relational database model. SQL null values are used to meet the requirements of supporting "missing information and inapplicable information" in a real relational database management system (RDBMS);

[0055] Primary key: (PRIMARY KEY) Often in a table, there is a combination of one or more columns whose values can uniquely identify each row in the table. Such a column or columns is called the primary key of the table, through which the entity integrity of the table can be enforced. When creating or altering a table, a primary key can be created by defining a PRIMARY KEY constraint. A table can have only one PRIMARY KEY constraint, and the columns in the PRIMARY KEY constraint cannot accept NULL values. Since the PRIMARY KEY constraint ensures unique data, it is often used to define an identity column.

[0056] Distribution key: In a distributed database, a column (or a combination of columns) used to determine the database partition where a specific data row is stored.

[0057] Subquery: A subquery is a SELECT query that returns a single value and is nested in a SELECT, INSERT, UPDATE, DELETE statement, or another subquery. The upper query block is the parent query. Subqueries can be nested multiple levels.

[0058] The method provided by the embodiments of this application can be applied to a database system. Figure 1A shows a typical logical architecture of a database system. According to Figure 1A , the database system 100 includes a database 110 and a database management system (DBMS) 130.

[0059] Among them, the database 110 is an organized collection of data stored in a data storage 120, that is, a collection of related data organized, stored, and used according to a specific data model. According to the different data models used to organize the data, the data can be divided into multiple types, such as relational data, graph data, time series data, etc. Relational data is data modeled using a relational model, usually represented as a table, and the rows in the table represent a set of related values of an object or entity. Graph data, simply referred to as "graph", is used to represent the relationships between objects or entities, such as social relationships. Time series data, abbreviated as time series data, is a data column recorded and indexed in chronological order, used to describe the state change information of an object in the time dimension.

[0060] The database management system 130 is the core of the database system and is system software used to organize, store, and maintain data. The client 200 can access the database 110 through the database management system 130, and the database administrator also performs database maintenance work through the database management system. The database management system 130 provides various functions for the client 200 to create, modify, and query the database. Among them, the client 200 can be an application program or a user device. The functions provided by the database management system 130 may include but are not limited to the following: (1) Data definition function: The database management system 130 provides a data definition language (DDL) to define the structure of the database 110. The DDL is used to describe the database framework and can be saved in the data dictionary; (2) Data access function: The database management system 130 provides a data manipulation language (DML) to implement basic access operations on the database 110, such as retrieval, insertion, modification, and deletion; (3) Database operation management function: The database management system 130 provides a data control function to effectively control and manage the operation of the database 110 to ensure the correctness and effectiveness of the data; (4) Database establishment and maintenance function: including functions such as loading of initial database data, database dump, recovery, reorganization, system performance monitoring, and analysis; (5) Database transmission: The database management system provides data transmission processing to achieve communication between the client and the database management system, usually coordinated with the operating system to complete.

[0061] The data storage 120 includes but is not limited to solid state drives (SSDs), disk arrays, cloud storage, or other types of non-transitory computer-readable storage media. Those skilled in the art can understand that a database system may include fewer or more components than those Figure 1A shown in Figure 1A or include components different from those Figure 1A shown, and

[0062] only the components more relevant to the implementation disclosed in the embodiments of the present invention are shown.

[0062] The database system provided by the embodiments of the present application can be a distributed database system (DDBS). During the transaction processing of the DDBS, in order to implement concurrent control between transactions, a global transaction manager (GTM) is usually used to manage the transactions. The following will introduce the DDBS in combination with Figure 1B and Figure 1C for introduction.

[0063] Figure 1B Schematic diagram of a distributed database system adopting a shared-storage architecture, including one or more coordinator nodes (CNs), multiple data nodes (DNs), and one or more GTMs (such as Figure 1B the first GTM and the second GTM in). The first GTM serves as the primary GTM, and the second GTM is used to back up the data of the first GTM and take over the work of the first GTM in case of its failure, which can ensure the high reliability of the DDBS. The CNs and DNs communicate through a network channel. In one embodiment, the network channel can be composed of network devices such as switches, routers, and gateways. The CNs, DNs, and GTMs jointly implement the functions of the database management system and provide services such as database retrieval, insertion, modification, and deletion for clients. In one embodiment, a database management system is deployed on each CN, DN, and GTM. The shared data storage stores data that can be shared by multiple DNs, and the DNs can perform read and write operations on the data in the data storage through the network channel. The shared data storage can be a shared disk array. The CNs, DNs, the first GTM, or the second GTM in the distributed database system can be physical machines, such as database servers, or virtual machines (VMs) or containers running on abstract hardware resources. In one embodiment, the CNs, DNs, the first GTM, or the second GTM are virtual machines or containers, and the network channel is a virtual switching network, which includes a virtual switch. The database management system deployed in the CNs, DNs, the first GTM, or the second GTM is a DBMS instance, and the DBMS instance can be a process or a thread, and these DBMSs cooperate to complete the functions of the database relational system. In another embodiment, the CNs, DNs, the first GTM, or the second GTM are physical machines, and the network channel includes one or more switches, and the switches are storage area network (SAN) switches, Ethernet switches, fiber optic switches, or other physical switching devices.

[0064] Figure 1C Schematic diagram of a distributed database system adopting a shared-nothing architecture. Each DN has its own exclusive hardware resources (such as data storage), operating system, and database. The CNs, DNs, the first GTM, or the second GTM communicate through a network channel, and the network channel can refer to the above Figure 1BUnderstand the corresponding introduction of the part. Under this system, data will be allocated to each DN according to the database model and application characteristics. The query task will be split into several parts by the CN and executed in parallel on all DNs, collaborating with each other for calculation, and providing database services as a whole. All communication functions are implemented on a high-bandwidth network interconnection system. Just like Figure 1B the distributed database system with the shared-storage architecture described, the CN, DN, the first GTM or the second GTM here can be either a physical machine or a virtual machine.

[0065] In all embodiments of the present application, the data storage of the database system includes but is not limited to solid state drives (SSDs), disk arrays, or other types of non-transitory computer-readable media. Figures 1B - 1C Although the database is not shown in [reference], it should be understood that the database is stored in the data storage. Those skilled in the art can understand that a database system may include fewer or more components than Figures 1A - 1C the components shown in [reference], or include components different from Figures 1A - 1C the components shown in [reference], Figures 1A - 1C only the components more relevant to the implementation disclosed in the embodiments of the present application are shown. However, those skilled in the art can understand that a distributed database system can include any number of CNs and DNs. The database management system functions of each CN and DN can be respectively implemented by an appropriate combination of software, hardware, and / or firmware running on each CN and DN.

[0066] The above Figure 1B and Figure 1C described distributed database system includes multiple DNs and multiple CNs, where the functions of each DN are basically the same, and the functions of each CN are also basically the same.

[0067] Figure 1D Shows an application data management system provided by an embodiment of the present application. Specifically, as shown, the data node cluster may include multiple data nodes DN. Application developers can deploy the relevant data of the developed application in the corresponding data node DN. The application can send a service request to the application server, and the application server can convert the service request into a data operation request; the application server can send the data operation request to the distributed database system. Then, one or more data nodes DN in the distributed database system execute the operation corresponding to the operation request on the data in the database storage. Specifically, the application server can send the data operation request to the coordinator in the coordinator cluster, and the coordinator then forwards the operation request to the relevant DN, and the relevant DN executes the data operation corresponding to the operation request.

[0068] Refer to Figure 2A , Figure 2A , which is a system architecture diagram of an embodiment of the present application: It includes one or more coordination nodes (CNs) and multiple data nodes (DNs). Among them, the CN receives the SQL statement with subqueries sent by the user, generates an execution plan for predicate pushdown according to the distribution information of the parent query and subqueries in the query statement, and distributes the plan to the data node DN. On the DN, for subqueries, when scanning the base table data, a predicate condition filtering operation will be synchronously performed, and only the data that meets the predicate conditions will be passed to the upper-layer operator. The same applies to the parent query. Finally, the DN returns the data to the Gather operator of the CN, and after receiving the data, the CN returns the result to the user.

[0069] Refer to Figure 2B , Figure 2B , which is a system architecture diagram of an embodiment of the present application: The present application can be program code included in a distributed database and deployed on a server. Taking the application scenario shown in Figure 2B as an example, the program code of the present application exists inside the database parser, optimizer, and executor of the coordination node CN and data node DN of the distributed database. The base table data is stored on different DNs according to the distribution key of the base table through a distribution algorithm. As shown in Figure 2B , the user sends an SQL statement with subqueries on the CN. The CN is responsible for parsing and checking the user input statement and generating an execution plan, connecting to the DN through the network, sending data and instructions to the DN, and finally summarizing the statistical information of the DN. The DN is responsible for the specific execution process of the execution plan. Except for different stored data, each DN has no difference in architecture and supports various different data distribution methods. The CN can be selected through an election algorithm or adopt an architecture different from that of the DN. According to the specific deployment of the distributed database, there can be multiple CNs. Each module of the database is equally deployed on the CN and each DN, and the role assumed by the node in the distributed database is set through a configuration file.

[0070] A distributed database is a logically unified database formed by connecting multiple physically dispersed database units through a computer network. Each connected database unit is called a site or node.

[0071] In a distributed database system, the base tables of subqueries and the base tables of parent queries are distributed on multiple data nodes DN. In the prior art, for subqueries, all base table data is scanned during execution, passed to the upper-layer operator, and predicate condition filtering is performed. Finally, the results of the subquery and the parent query are returned to the CN, and the CN returns the aggregated result to the user.

[0072] In the prior art, the BROADCAST operator is used. The function of this operator is to broadcast the data of the current data node to all data nodes, and then perform subsequent operations. For scenarios where the predicate condition and the base table distribution key are inconsistent, this operator is required. For example, table t1 contains two columns a and b, and the data is distributed according to column a. Table t2 contains two columns, and the data is distributed according to column b. Both tables contain two rows of data [1,3] and [2,4]. All the data of table t1 is on DN1, and all the data of table t2 is on DN2. At this time, for the following query statement: select t1.b,(select t2.b from t2 where t2.a=t1.a)from t1, this statement contains a subquery. In the subquery, the b column of table t2 is returned for the data where the value of column a in table t2 is the same as the value of column a in the parent query of table t1. For this query, when executed on the data node, if BROADCAST is not performed, then since the data of table t1 and t2 is distributed on different data nodes, no matching data can be found within the data node, so the returned result is empty, which is an incorrect result. After adding the BROADCAST operator, the data on each data node is broadcast to other data nodes, and then when performing predicate filtering, the matching data can be found and the correct result can be returned.

[0073] However, in the prior art, for subqueries, all the data of the base table in the subquery needs to be scanned and passed to the upper-layer operator, and the predicate condition filtering operation is only performed after passing through multiple layers of operators. When the data volume is very large, on the one hand, the overhead of broadcasting data is very large, and the operator needs to process a large amount of data, consuming a large amount of resources and affecting the query performance.

[0074] To solve the above problems, the embodiments of the present application provide a data processing method, referring to Figure 3 , Figure 3 is a flow diagram of a data processing method provided by the embodiments of the present application, including:

[0075] 301. Obtain a first query and a second query, where the second query is a subquery of the first query. The first query includes a first predicate condition, and the second query includes a second predicate condition. The first predicate condition indicates the data that satisfies the first constraint in the first data range of the first table, and the second predicate condition indicates the data whose relationship between the data in the second data range of the second table and the data in the first data range satisfies the second constraint. Moreover, both the first data range and the second data range correspond to the distribution key in the table.

[0076] 302. Obtain a third query according to the first query and the second query; the third query includes a third predicate condition, and the third predicate condition indicates that the relationship between the data in the second data range of the second table and the data in the third data range of the first table satisfies the second constraint, and the data in the third data range is the data in the first data range that satisfies the first constraint;

[0077] In a possible implementation, a user may send a query statement (such as an SQL statement) with a subquery to the coordination node CN. The query statement may include a first query and a second query, where the second query is a subquery of the first query.

[0078] For example, the query statement includes a single-layer subquery, that is, the parent query (query A) includes the subquery (query B), where the second query may be query B and the first query may be query A.

[0079] For example, the query statement includes a multi-layer subquery, that is, the subquery (query B) of the parent query (query A) further includes a subquery (query C), where the second query may be query B, the first query may be query A, or the second query may be query C, the first query may be query B, and the fourth query may be query A.

[0080] For example, the query statement includes two single-layer subqueries, that is, the parent query (query A) includes the subquery (query B) and the subquery (query C), where the second query may be query B, the first query may be query A, or the second query may be query C, the first query may be query A.

[0081] For example, the query statement includes two multi-layer subqueries, that is, the parent query (query A) includes the subquery (query B) and the subquery (query C), the subquery (query B) may further include a subquery, and the subquery (query C) may further include a subquery.

[0082] In the embodiments of the present application, if the parent query includes a subquery and the distribution conditions of the predicates in the parent query and the subquery are the same, then the predicate condition of the parent query may be pushed down to the subquery to reduce the amount of data processed by the subquery, thereby improving the performance of the query.

[0083] The predicate condition of the first query as the parent query is a constraint on the data area (the first data range) where the distribution key of the first table is located (that is, the first constraint in the embodiments of the present application), and this constraint may indicate which data in the data area where the distribution key of the first table is located needs to be queried in the parent query.

[0084] The predicate condition of the second query as a subquery is also a constraint on the data area (second data range) where the distribution key of the second table is located (that is, the second constraint in the embodiment of the present application). The second constraint can indicate which data in the data area where the distribution key of the second table is located that need to be queried in the subquery, and the second constraint is related to the distribution key indicated by the predicate condition of the first query, that is, the second predicate condition indicates the data in the second data range of the second table whose relationship with the data in the first data range satisfies the second constraint.

[0085] There is a first constraint in the parent query for the data area where the distribution key of the first table is located, and the second constraint in the subquery is related to the data area where the distribution key of the first table is located. Since the predicate condition in the parent query must be met in the final query result, the first constraint can be directly placed in the subquery, that is, the first constraint is directly added to the content in the subquery related to the data area where the distribution key of the first table is located (that is, the third query in the embodiment of the present application), thereby greatly reducing the number of data nodes scanned.

[0086] For example, the base table t1 contains two columns, a and b, where column a is the distribution key of table t1. The base table t2 contains two columns, a and b, where column a is the distribution key of table t2. The semantics of the parent query is to obtain the value of column b from the data in table t1 where the value of column a is 1, and the semantics of the subquery is to obtain the value of column b from the data in table t2 where the value of column a is the same as the value of column a in table t1 in the parent query. Since the constraint "column a value is 1" already exists for column a in table t1 in the parent query, the semantics of the subquery can be directly modified to "obtain the value of column b from the data in table t2 where the value of column a is 1".

[0087] Here are a few application scenarios:

[0088] 1. Single single-level subquery.

[0089] The base table schema definition used in the first embodiment can be as follows: Figure 4 As shown, the example contains two base tables t1 and t2. Base table t1 contains two columns, a and b, where column a is the distribution key of table t1. Base table t2 contains two columns, a and b, where column a is the distribution key of table t2. Both base tables t1 and t2 contain 1 million rows of data, distributed on two data nodes DN1 and DN2, with 500,000 rows of data on each data node. The query statement used contains a single-layer subquery. The semantics of the parent query statement is to obtain the value of column b in the data where the value of column a is 1 from table t1, and the semantics of the subquery is to obtain the value of column b in the data where the value of column a is the same as the value of column a in table t1 in the parent query.

[0090] When the execution plan generated by the database optimizer before using subquery parameter passing is considered, for a subquery, the Seq Scan operator is first used to perform a full table scan on t2, and then the scanned data is sent to all data nodes DN using the BROADCAST operator. After the DN receives the data, the Materialize operator saves the data in memory, and finally the Result operator performs a predicate filtering operation. Since all the data in table t2 needs to be passed to the upper-layer operator, this consumes a large amount of resources and results in a very low efficiency of the execution plan.

[0091] Since the predicate conditions of the subquery and the parent query are both column a, which is the same as the distribution key of the base table, the predicate conditions of the parent query can be pushed down to the subquery. The execution plan generated by the database optimizer after using subquery parameter passing is as Figure 5 shown. For the subquery, during the process of scanning table t2 using the Seq Scan operator, a predicate filtering operation is performed synchronously, and only the data that meets the predicate conditions is passed to the upper-layer operator. In this way, the amount of data processed by the upper-layer operator is one row, which occupies much less resources compared to the one million rows of data above, greatly improving the execution performance.

[0092] 2. Multi-layer subqueries.

[0093] The definition of the base table schema can be as Figure 6 shown. The example contains three base tables, t1, t2, and t3. Base table t1 contains two columns, a and b, where column a is the distribution key of table t1. Base table t2 contains two columns, a and b, where column a is the distribution key of table t2. Base table t3 contains two columns, a and b, where column a is the distribution key of table t3. Each of the three base tables contains 1 million rows of data, which are distributed on two data nodes, DN1 and DN2, with 500,000 rows of data on each data node. The query statement used contains a multi-layer subquery, that is, a subquery contains another subquery. The semantic of the parent query is to obtain the value of column b from the data in table t1 where the value of column a is 1. The semantic of the first-layer subquery is to obtain the data from table t2 where the value of column a is the same as the value of column a in table t1 in the parent query. The semantic of the second-layer subquery is to obtain the value of column b from the data in table t3 where the value of column a is the same as the value of column a in table t2 in the first-layer subquery.

[0094] When the execution plan generated by the database optimizer before using subquery parameter passing, for the second-layer subquery, first use the Seq Scan operator to perform a full table scan on t3, and then send the scanned data to all data nodes DN using the BROADCAST operator. After the DN receives the data, the Materialize operator saves the data in memory, and finally the Result operator performs the predicate filtering operation. The same applies to the t2 table in the first-layer subquery. Since all the data of tables t2 and t3 needs to be passed to the upper-layer operator, this will consume a large amount of resources and result in a very low efficiency of the execution plan.

[0095] Note that since the predicate conditions of the second-layer subquery, the first-layer subquery, and the parent query are all column a, which is the same as the distribution key of the base table, the predicate condition of the parent query can be pushed down to the subquery. The execution plan generated by the database optimizer after using subquery parameter passing is as Figure 7 shown. For the second-layer subquery, during the process of scanning table t3 using the Seq Scan operator, the predicate filtering operation will be performed synchronously, and only the data that meets the predicate condition will be passed to the upper-layer operator. The same applies to table t2 in the first-layer subquery. In this way, the amount of data processed by the upper-layer operator is one row. Compared with the one million rows of data above, the resources occupied are greatly reduced, and the execution performance is greatly improved.

[0096] 3. Two single-layer subqueries.

[0097] The definition of the base table schema used can be as Figure 8 shown. The example contains three base tables t1, t2, and t3. Base table t1 contains two columns, a and b, where column a is the distribution key of table t1. Base table t2 contains two columns, a and b, where column a is the distribution key of table t2. Base table t3 contains two columns, a and b, where column a is the distribution key of table t3. Each of the three base tables contains 1 million rows of data, which are distributed on two data nodes DN1 and DN2, with 500,000 rows of data on each data node. The query statement used contains two single-layer subqueries. The semantics of the parent query statement is to obtain the value of column b from the data in table t1 where the value of column a is 1. The semantics of the first subquery is to obtain the data from table t2 where the value of column a is the same as the value of column a in table t1 in the parent query. The semantics of the second subquery is to obtain the value of column b from the data in table t3 where the value of column a is the same as the value of column a in table t1 in the parent query.

[0098] When the execution plan generated by the database optimizer before using subquery parameter passing, for the first subquery, the Seq Scan operator is first used to perform a full table scan on t2, and then the scanned data is sent to all data nodes DN using the BROADCAST operator. After receiving the data, the Materialize operator in DN saves the data in memory, and finally the Result operator performs the predicate filtering operation. The same applies to the t3 table in the first subquery. Since all the data in tables t2 and t3 needs to be passed to the upper-layer operator, this consumes a large amount of resources and results in a very low efficiency of the execution plan.

[0099] Note that since the predicate conditions of the two subqueries and the predicate condition of the parent query are all column a, which is consistent with the distribution key of the base table, the predicate condition of the parent query can be pushed down to the two subqueries. The execution plan generated by the database optimizer after using subquery parameter passing is as Figure 9 shown. For the first subquery, during the process of scanning table t2 using the Seq Scan operator, the predicate filtering operation will be performed synchronously, and only the data that meets the predicate condition will be passed to the upper-layer operator. The same applies to table t3 in the second subquery. In this way, the amount of data processed by the upper-layer operator is one row. Compared with the one million rows of data above, the occupied resources are greatly reduced, and the performance of the execution is greatly improved.

[0100] 4. Two multi-level subqueries.

[0101] The defined base table schema used can be as Figure 10 shown. The example contains five base tables t1, t2, t3, t4, and t5. Each base table contains two columns, a and b, where column a is the distribution key of the table. All five base tables contain 1 million rows of data, which are distributed on two data nodes DN1 and DN2, with 500,000 rows of data on each data node. The query statement used contains two multi-level subqueries. In the first subquery, the semantic of the first-level subquery is to obtain the data from table t2 where the value of column a is the same as the value of column a in table t1 in the parent query. The semantic of the second-level subquery is to obtain the value of column b from the data where the value of column a is the same as the value of column a in table t2 in the first-level subquery from table t3. In the second subquery, the semantic of the first-level subquery is to obtain the data from table t4 where the value of column a is the same as the value of column a in table t1 in the parent query. The semantic of the second-level subquery is to obtain the value of column b from the data where the value of column a is the same as the value of column a in table t4 in the first-level subquery from table t5.

[0102] When the execution plan generated by the database optimizer before using subquery parameter passing, in the first subquery, for the second-level subquery, first use the Seq Scan operator to perform a full table scan on the t3 table, then use the BROADCAST operator to send the scanned data to all data nodes DN. After DN receives the data, the Materialize operator saves the data in memory, and finally the Result operator performs the predicate filtering operation. The same applies to the first-level subquery. In the second subquery, for the second-level subquery, first use the Seq Scan operator to perform a full table scan on the t5 table, then use the BROADCAST operator to send the scanned data to all data nodes DN. After DN receives the data, the Materialize operator saves the data in memory, and finally the Result operator performs the predicate filtering operation. The same applies to the first-level subquery.

[0103] Note that since the predicate conditions of the two multi-level subqueries and the predicate condition of the parent query are all column a, which is consistent with the distribution key of the base table, the predicate condition of the parent query can be pushed down to the two multi-level subqueries. The execution plan generated by the database optimizer after using subquery parameter passing is as Figure 11 shown. In the first subquery, for the first-level subquery, during the process of using the Seq Scan operator to scan the table t3, the predicate filtering operation will be performed synchronously, and only the data that meets the predicate condition will be passed to the upper-level operator. The same applies to the first-level subquery. In the second subquery, for the first-level subquery, during the process of using the Seq Scan operator to scan the table t5, the predicate filtering operation will be performed synchronously, and only the data that meets the predicate condition will be passed to the upper-level operator. The same applies to the first-level subquery. In this way, in each subquery, the amount of data processed by the upper-level operator is one row. Compared with the one million rows of data above, the occupied resources are greatly reduced, and the execution performance is greatly improved.

[0104] It should be understood that in a possible implementation, when confirming whether to push down the parent query, the CN checks whether there is parameter passing information in the statement; or whether the predicate conditions and distribution information of the parent query and subquery in the statement are consistent. The parameter passing information can be called indication information, and the indication information is used to indicate pushing down the first predicate condition to the second predicate condition included in the second query. If there is parameter passing promotion information or the predicate distribution is consistent, a parameter passing path is directly generated. Otherwise, a BROADCAST operator is added, and then a parameter passing path is generated; the CN converts the path into an execution plan. The executor executes according to the generated execution plan and returns the execution result.

[0105] 303. Send the third query to the data node where the second table is deployed.

[0106] It should be understood that, as introduced above, the first query may be a sub-query of other queries (such as the fourth query in the embodiments of the present application). The fourth query includes a fourth predicate condition, and the fourth predicate condition indicates data that satisfies the third constraint in the fourth data range of the third table. It should be understood that the third table may be the same table as the first table or the second table.

[0107] Refer to Figure 12 , Figure 12 is a schematic flowchart of a data processing method according to an embodiment of the present application, including:

[0108] The user sends an SQL statement with a sub-query to the coordination node CN;

[0109] CN checks whether the statement has parameter passing information; or whether the predicate conditions and distribution information of the parent query and the sub-query in the statement are consistent;

[0110] If there is parameter passing promotion information or the predicate distribution is consistent, a parameter passing path is directly generated; otherwise, a BROADCAST operator is added, and then a parameter passing path is generated;

[0111] CN converts the path into an execution plan.

[0112] The executor executes according to the generated execution plan and returns the execution result.

[0113] Refer to Figure 13 , Figure 13 is a schematic structural diagram of a data processing device provided by an embodiment of the present application. As Figure 13 shown, a data processing device 1300 provided by an embodiment of the present application includes:

[0114] A transceiver module 1301, configured to obtain a first query and a second query. The second query is a sub-query of the first query. The first query includes a first predicate condition, and the second query includes a second predicate condition. The first predicate condition indicates data that satisfies the first constraint in the first data range of the first table. The second predicate condition indicates that the relationship between the data in the second data range of the second table and the data in the first data range satisfies the second constraint, and both the first data range and the second data range correspond to distribution keys in the table;

[0115] The transceiver module 1301 is further configured to send the third query to a data node where the second table is deployed.

[0116] Among them, the specific description of the transceiver module 1301 can refer to the introduction in steps 301 and 303 in the above embodiments, and will not be elaborated here.

[0117] A processing module 1302 is configured to obtain a third query according to the first query and the second query; the third query includes a third predicate condition, and the third predicate condition indicates data in a second data range of the second table and data in a third data range of the first table, where a relationship between the data satisfies the second constraint, and the data in the third data range is data in the first data range that satisfies the first constraint;

[0118] For a specific description of the processing module 1302, reference may be made to the description of step 302 in the foregoing embodiment, which will not be elaborated here.

[0119] In a possible implementation, the first query is a subquery of a fourth query, and the fourth query includes a fourth predicate condition, and the fourth predicate condition indicates data in a fourth data range of a third table that satisfies a third constraint.

[0120] In a possible implementation, the processing module 1302 is specifically configured to:

[0121] Based on the obtained indication information, obtain a third query according to the first query and the second query; the indication information is used to indicate pushing the first predicate condition down to the second predicate condition included in the second query.

[0122] In a possible implementation, the transceiver module 1301 is further configured to:

[0123] Receive data in the second table transmitted by the data node according to the third query.

[0124] In a possible implementation, the first data range is a column in the first table, and the second data range is a column in the second table.

[0125] Next, a data processing device provided in an embodiment of the present application will be introduced. Please refer to Figure 14 , Figure 14 which is a schematic structural diagram of a data processing device provided in an embodiment of the present application. Specifically, the data processing device 1400 includes: a receiver 1401, a transmitter 1402, a processor 1403, and a memory 1404 (where the number of processors 1403 in the data processing device 1400 may be one or more, Figure 14 taking one processor as an example), where the processor 1403 may include an application processor 14031 and a communication processor 14032. In some embodiments of the present application, the receiver 1401, the transmitter 1402, the processor 1403, and the memory 1404 may be connected through a bus or other means.

[0126] The memory 1404 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1403. A part of the memory 1404 may also include a non-volatile random access memory (NVRAM). The memory 1404 stores processor and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, where the operation instructions may include various operation instructions for implementing various operations.

[0127] The processor 1403 controls the operation of the data processing device. In a specific application, the various components of the data processing device are coupled together through a bus system, where the bus system may include a power bus, a control bus, a status signal bus, etc. in addition to the data bus. However, for the sake of clarity, all kinds of buses are referred to as the bus system in the figure.

[0128] The method disclosed in the embodiments of the present application above can be applied to the processor 1403 or implemented by the processor 1403. The processor 1403 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 1403 or instructions in the form of software. The above-mentioned processor 1403 may be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and may further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 1403 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1404, and the processor 1403 reads the information in the memory 1404 and combines its hardware to complete the execution of the above method for the actions of CN or DN.

[0129] The receiver 1401 can be used to receive input digital or character information, and generate signal inputs related to the relevant settings and function controls of the data processing device. The transmitter 1402 can be used to output digital or character information through the first interface; the transmitter 1402 can also be used to send instructions to the disk array through the first interface to modify the data in the disk array.

[0130] An embodiment of the present application also provides a data processing device. Please refer to Figure 15 , Figure 15 FIG. is a schematic structural diagram of a data processing device provided by an embodiment of the present application. The data processing device 1500 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1515 (for example, one or more processors) and a memory 1532, and one or more storage media 1530 (for example, one or more mass storage devices) for storing application programs 1542 or data 1544. Among them, the memory 1532 and the storage media 1530 can be transient storage or persistent storage. The program stored in the storage media 1530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the data processing device. Further, the central processor 1515 may be configured to communicate with the storage media 1530 and execute a series of instruction operations in the storage media 1530 on the data processing device 1500.

[0131] The data processing device 1500 may further include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, and one or more input / output interfaces 1558; or, one or more operating systems 1541, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0132] In the embodiment of the present application, the central processor 1515 is used to execute the execution actions of CN or DN in the above embodiment.

[0133] An embodiment of the present application also provides a computer program product, which when running on a computer, causes the computer to execute the steps performed by the aforementioned data processing device.

[0134] An embodiment of the present application also provides a computer-readable storage medium, in which a program for signal processing is stored, and when it runs on a computer, it causes the computer to execute the steps performed by the aforementioned data processing device.

[0135] The data processing device, data processing apparatus or terminal device provided by the embodiments of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit, etc. The processing unit may execute the computer-executable instructions stored in the storage unit to cause the chip in the data processing device to execute the data processing method described in the above embodiments, or to cause the chip in the data processing device to execute the data processing method described in the above embodiments. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit may also be a storage unit outside the chip in the radio access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0136] Among them, the processor mentioned anywhere above may be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.

[0137] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which may specifically be implemented as one or more communication buses or signal lines.

[0138] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, software program implementation is a better embodiment in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that makes contributions to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, a data processing device, or a network device, etc.) to execute the methods described in various embodiments of this application.

[0139] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0140] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a data processing device, or a data center to another website, computer, data processing device, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a data processing device or a data center that integrates one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

Claims

1. A data processing method, characterized in that, The method includes: Obtaining a first query and a second query, where the second query is a sub-query of the first query. The first query includes a first predicate condition, and the second query includes a second predicate condition. The first predicate condition indicates data in a first data range of a first table that satisfies a first constraint, and the second predicate condition indicates data in a second data range of a second table whose relationship with the data in the first data range satisfies a second constraint. Both the first data range and the second data range correspond to distribution keys in the table; Obtaining a third query according to the first query and the second query; the third query includes a third predicate condition, and the third predicate condition indicates data in a second data range of the second table whose relationship with the data in a third data range of the first table satisfies the second constraint, and the data in the third data range is the data in the first data range that satisfies the first constraint; Sending the third query to a data node where the second table is deployed.

2. The method according to claim 1, wherein The first constraint indicates data in the first data range of the first table that is equal to a preset constant.

3. The method according to claim 1 or 2, characterized in that, The third query does not include the first query.

4. The method according to any one of claims 1 to 3, characterized in that, The first query is a sub-query of a fourth query, and the fourth query includes a fourth predicate condition, and the fourth predicate condition indicates data in a fourth data range of a third table that satisfies a third constraint.

5. The method according to any one of claims 1 to 4, characterized in that, The obtaining the third query according to the first query and the second query includes: Based on the obtained indication information, obtaining a third query according to the first query and the second query; the indication information is used to indicate pushing down the first predicate condition into the second predicate condition included in the second query.

6. The method according to any one of claims 1 to 5, characterized in that The method further includes: Receiving data in the second table transmitted by the data node according to the third query.

7. The method according to any one of claims 1 to 6, characterized in that, The first data range is a column in the first table, and the second data range is a column in the second table.

8. A database system, characterized in that, Including a coordination node and a data node: The coordination node is configured to execute the method according to any one of claims 1 to 7; The data node is configured to scan the second table according to the third query to obtain data in the second table; Transmitting the data in the second table to the coordination node.

9. A data processing device, characterized in that, The apparatus includes: A transceiver module, configured to obtain a first query and a second query, where the second query is a sub-query of the first query. The first query includes a first predicate condition, and the second query includes a second predicate condition. The first predicate condition indicates data in a first data range of a first table that satisfies a first constraint, and the second predicate condition indicates data in a second data range of a second table whose relationship with the data in the first data range satisfies a second constraint. Both the first data range and the second data range correspond to distribution keys in the table; A processing module, configured to obtain a third query according to the first query and the second query; the third query includes a third predicate condition, and the third predicate condition indicates that the relationship between the data in the second data range of the second table and the data in the third data range of the first table satisfies the second constraint, and the data in the third data range is the data in the first data range that satisfies the first constraint; The transceiver module is further configured to send the third query to a data node where the second table is deployed.

10. The device according to claim 9, characterized in that, The first constraint indicates the data in the first data range of the first table that is equal to a preset constant.

11. The device according to claim 9 or 10, characterized in that, The third query does not include the first query.

12. The device according to any one of claims 9 to 11, characterized in that The first query is a subquery of a fourth query, and the fourth query includes a fourth predicate condition, and the fourth predicate condition indicates the data in the fourth data range of the third table that satisfies the third constraint.

13. The device according to any one of claims 9 to 12, characterized in that, The processing module is specifically configured to: Based on the obtained indication information, obtain a third query according to the first query and the second query; the indication information is used to indicate pushing down the first predicate condition to the second predicate condition included in the second query.

14. The device according to any one of claims 9 to 13, characterized in that The transceiver module is further configured to: Receive the data in the second table transmitted by the data node according to the third query.

15. The device according to any one of claims 9 to 14, characterized in that The first data range is a column in the first table, and the second data range is a column in the second table.

16. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions, and when the instructions are executed by one or more computers, the one or more computers perform the operations of the method according to any one of claims 1 to 7.

17. A computer program product, characterized in that, Including computer-readable instructions, when the computer-readable instructions run on a computer device, the computer device executes the method according to any one of claims 1 to 7.

18. A system, including at least one processor and at least one memory; the processor and the memory are connected through a communication bus and complete communication with each other; The at least one memory is used to store code; The at least one processor is used to execute the code to execute the method according to any one of claims 1 to 7.

19. A chip, characterized in that, Including at least one processing unit and an interface circuit, the interface circuit is used to provide program instructions or data for the at least one processing unit, and the at least one processing unit is used to execute the program instructions to implement the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Data processing method and database system

    WO2025145652A1