Data redistribution method and database
By introducing a data redistribution method in the database, the coordinator nodes include redistribution plans when generating the optimal execution plan, and actively push data from the storage node to the computing node, solving the problem of low data query performance under the storage and computing separation architecture, and achieving the effect of reducing network transmission volume and improving query performance.
Patent Information
- Application Number
- CN202311574829.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
Under the database architecture of separation of storage and computing, the computing layer needs to use a large number of network interactions when obtaining data from the storage layer, resulting in excessive network transmission volume and reducing data query performance.
By introducing a data redistribution method in the database, the coordinator nodes include redistribution plans when generating the optimal execution plan, actively pushing the data from the storage node to the computing node, reducing the number of network transmission hops and resource consumption.
It reduces the network transmission volume during data query, improves data query performance, and reduces the consumption of network and computing resources.
Smart Images

Figure CN120030042A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information technology (IT) technology, and in particular to a data redistribution method and a database. Background Art
[0002] As databases move toward cloud computing, database architectures have gradually evolved into storage-computing separation architectures, that is, the storage layer is separated from the computing layer. In this storage-computing separation architecture, storage resources and computing resources are applied for and released on demand. However, the action of the computing layer obtaining data from the storage layer often requires a large amount of network interaction, which results in excessive network transmission and reduces data query performance. Summary of the invention
[0003] The present application provides a data redistribution method, a database, a computer storage medium and a computer product, which can reduce the network transmission volume during the data query process and improve the data query performance.
[0004] In a first aspect, the present application provides a data redistribution method, which is applied to a database, wherein the database includes: a coordinator node, multiple computing nodes, and at least one storage node. The method includes: the coordinator node determines that the optimal execution plan generated by the coordinator node includes a redistribution plan, and the redistribution plan is used to redistribute first data, wherein the first data is stored in a first storage node among at least one storage node; the coordinator node transmits the redistribution plan to the first storage node; and the first storage node transmits the first data to at least one computing node among the multiple computing nodes based on the redistribution plan.
[0005] In this way, when there is a redistribution plan in the optimal execution plan, the data transmission is changed from pulling from the destination to actively pushing from the source, which can reduce the number of network transmission hops, reduce the amount of data transmitted in the network, and reduce the consumption of network and computing resources, thereby improving data query performance.
[0006] In a possible implementation, the first storage node transmits the first data to at least one computing node among the multiple computing nodes based on the redistribution plan, specifically including: the first storage node screens out the first data based on the redistribution plan; the first storage node screens out at least one computing node from the multiple computing nodes based on the first data; the first storage node transmits the first data to the screened computing node. In this way, the storage node can screen out the computing node that needs to process the data and transmit the data to the computing node, so that the computing node can process the data.
[0007] In a possible implementation, before the first storage node transmits the first data to at least one of the multiple computing nodes, it further includes: the first storage node establishes a data transmission channel with each of the at least one of the multiple computing nodes. In this way, data can be transmitted between the storage node and the computing node through this data transmission channel.
[0008] In a possible implementation, the coordinator node determines that the optimal execution plan generated by the coordinator node includes a redistribution plan, specifically including: the coordinator node determines that the aggregation key included in the query statement is inconsistent with the distribution key used when storing data in the database; or, the coordinator node determines that the join key included in the query statement is inconsistent with the distribution key used when storing data in the database.
[0009] In a second aspect, the present application provides a database, which includes: a coordinator node, multiple computing nodes, and at least one storage node. Among them, the coordinator node is used to determine that the optimal execution plan generated by the coordinator node includes a redistribution plan, and the redistribution plan is used to redistribute the first data, where the first data is stored in the first storage node of the at least one storage node. The coordinator node is further used to transmit the redistribution plan to the first storage node. The first storage node is used to transmit the first data to at least one of the multiple computing nodes based on the redistribution plan.
[0010] In a possible implementation, when the first storage node transmits the first data to at least one of the multiple computing nodes based on the redistribution plan, it is specifically used to: screen out the first data based on the redistribution plan; screen out at least one of the multiple computing nodes based on the first data; and transmit the first data to the screened computing nodes.
[0011] In a possible implementation, before the first storage node transmits the first data to at least one of the multiple computing nodes, it is further used to: establish a data transmission channel with each of the at least one of the multiple computing nodes.
[0012] In a possible implementation, when the coordinator node determines that the optimal execution plan generated by the coordinator node includes a redistribution plan, it is specifically used to: determine that the aggregation key included in the query statement is inconsistent with the distribution key used when storing data in the database; or, determine that the join key included in the query statement is inconsistent with the distribution key used when storing data in the database.
[0013] In a third aspect, the present application provides a computer-readable storage medium, including computer program instructions, when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method described in the first aspect or any possible implementation of the first aspect. The computing device cluster may include one or more computing devices.
[0014] In a fourth aspect, the present application provides a computer program product including instructions, which, when executed by a computing device cluster, enables the computing device cluster to perform the method described in the first aspect or any possible implementation of the first aspect. The computing device cluster may include one or more computing devices.
[0015] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic diagram of the logical architecture of a database management system provided in an embodiment of the present application;
[0017] Figure 2 It is a schematic diagram of the physical architecture of a database provided in an embodiment of the present application;
[0018] Figure 3 This is a schematic diagram of a data redistribution process provided by an embodiment of the present application;
[0019] Figure 4 It is a flowchart of a data redistribution method provided in an embodiment of the present application;
[0020] Figure 5 It is a schematic diagram of the structure of a database provided in an embodiment of the present application. DETAILED DESCRIPTION
[0021] The term "and / or" in this article is a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The symbol " / " in this article indicates that the associated objects are in an or relationship, for example, A / B means A or B.
[0022] The terms "first" and "second" in the specification and claims herein are used to distinguish different objects rather than to describe a specific order of the objects. For example, a first response message and a second response message are used to distinguish different response messages rather than to describe a specific order of the response messages.
[0023] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0024] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two. For example, multiple processing units refer to two or more processing units, etc.; multiple elements refer to two or more elements, etc.
[0025] First, some technical terms involved in this application are introduced.
[0026] (1) Query
[0027] Query refers to the operation of retrieving data from a database. It can obtain data from a single table or from multiple tables through a join operation. Queries can be written using SQL statements. Through queries, users can easily obtain the required data for data analysis, report generation, decision support, and other tasks.
[0028] (2) Operator
[0029] Operators are also called operators, which are used to process data. Operators are divided into logical operators and physical operators. Logical operators describe the semantics of operations but do not involve specific implementations. Physical operators describe the specific execution methods. For example, the join operation is a logical operator, and the corresponding hash join, nested loop join, and sort-merge join are physical operators.
[0030] (3) Query Optimizer
[0031] The query optimizer is an important component in the database management system. It is used to parse, analyze, optimize, and generate execution plans for SQL statements. The goal of the query optimizer is to find the best execution plan to meet the user's query requirements with the least time and resource cost. The query optimizer is mainly responsible for converting logical operators into physical operators and generating an efficient execution plan. Among them, the physical operators are displayed in the execution plan.
[0032] (4) Execution plan
[0033] The execution plan describes the specific steps and execution order of the database management system to execute the query statement. The basic operation unit in the execution plan is called a (physical) operator, which represents a specific operation, such as table scan, hash join, etc. The operators in the execution plan form a tree structure in the order of execution. The root node of the tree is the outermost operator, and the leaf node is the innermost operator.
[0034] (5) Connection operation
[0035] A join operation is to match the data in two or more tables according to the join conditions to generate a new table. Through the join operation, you can query related information from multiple tables. The join operation is usually implemented using SQL statements, including keywords such as INNER JOIN, LEFT JOIN, and RIGHT JOIN.
[0036] (6) Connection conditions
[0037] The join condition can specify the conditions for joining two tables, usually based on the comparison of certain columns in the two tables, so as to join the related data rows in the two tables together. For example, the query statement SELECT * FROM t1 INNER JOIN t2 ON t1.c1=t2.c2, the join condition is t1.c1=t2.c2, when the value of the c1 column of the t1 table is equal to the value of the c2 column of the t2 table, the related data rows of t1 and t2 are joined together. The join conditions are usually divided into equijoin and non-equijoin. Equijoin only uses the equal sign, and non-equijoin usually uses comparison operators, such as greater than, greater than or equal to, less than, less than or equal to, and not equal to.
[0038] (7) Connection key
[0039] A join key is a column or a combination of columns used to connect two tables. It can associate rows with the same or related values in the two tables to generate a result set after the join. In SQL queries, the JOIN clause is usually used to specify the join key. For example, in the join condition t1.c1=t2.c2, the join key of the t1 table is c1, and the join key of the t2 table is c2; in the join condition t1.c1>t2.c2 AND t1.c3>t2.c4, the join key of the t1 table is (c1,c3), and the join key of the t2 table is (c2,c4).
[0040] (8) Aggregation Key
[0041] Aggregation key is the column or attribute used to perform aggregation operation. Aggregation operation is an operation used to calculate summary values of data such as sum, average, count, maximum, minimum, etc.
[0042] (9) Distribution Key
[0043] A distribution key is a concept used in distributed databases to determine how data is distributed to different nodes. A distribution key is a key attribute or column in a distributed database that affects how data is distributed and stored.
[0044] (10) Redistribution
[0045] Redistribution refers to redistributing data to different nodes to improve query performance or data management.
[0046] (11) PushDown
[0047] Push-down is an optimization method in data query, which means putting the actions originally executed in the upper-level components into the lower-level components for early execution to improve query efficiency.
[0048] Next, the technical solution provided by this application is introduced.
[0049] For example, Figure 1 FIG. 1 shows a schematic diagram of the logical architecture of a database management system provided by an embodiment of the present application. Figure 1 As shown, the database management system may include: a client 100 and a database 200. The database 200 may include: an SQL engine 210 and a storage engine 220. The client 100 refers to various forms of connecting to a database, such as activeX data objects (ADO) connection used in .Net, Java database connection (JDBC) connection used in Java, etc.
[0050] The SQL engine 210 is mainly responsible for generating an efficient execution plan for the SQL statement input by the client 110 under the current load scenario, and running the execution plan. The SQL engine 210 may include: a connector 211, a query cache 212, a parser 213, an optimizer 214 and an executor 215. The connector 211 is mainly responsible for communicating with the client 110, and is responsible for business logic processing such as connection authentication, connection number judgment, and connection pool processing. The main function of the query cache 212 is to improve the efficiency of the query. The cache is stored in the form of a hash table of key and value. The key is a specific SQL statement, and the value is a collection of results. When a SQL statement arrives, if the query cache function is turned on, the SQL engine 210 can first check whether there is a data match in the query cache 212. If it matches, the matching data is directly returned to the client 110 without parsing the corresponding SQL statement. However, if there are user-defined functions, stored functions, user variables or temporary tables in the SQL statement, it will not pass through the query cache 212. If there is no match in the query cache 212, the parser 213 will be used to parse the corresponding SQL statement. The parser 213 is mainly responsible for parsing the SQL statement according to the grammatical rules, etc., and generating an internally recognizable parse tree. The optimizer 214 is responsible for optimizing the parse tree generated by the parser 213 to find an optimal execution plan. The executor 215 is mainly responsible for calling the interface of the storage engine 220 to execute the query or other operations according to the optimal execution plan after the optimizer 214 finds the optimal execution plan, and finally returns the query result set to the client 110.
[0051] The storage engine 220 is mainly responsible for the storage, retrieval and management of data. It defines important characteristics of the database management system such as how to organize data, how to execute queries and transactions, and the security and reliability of data.
[0052] For example, Figure 2 FIG. 1 shows a schematic diagram of the physical architecture of a database provided in an embodiment of the present application. Figure 2As shown, the database 200 may include at least one coordinator node, several computing nodes and several storage nodes. Each coordinator node and computing node in the database 200 may be any device or apparatus with computing capabilities, such as a server or a virtual machine running on general hardware. The storage node may be any device or apparatus with storage capabilities, such as memory, hard disk, disk array, cloud storage pool (such as object storage, distributed file system, cloud block storage, etc.). A computing node is associated with a storage node, and there may be a binding relationship between the two; of course, there may be no binding relationship between the computing node and the storage node. Each computing node interacts through a network or other communication methods, and has high parallel processing and expansion capabilities. Each computing node processes its own data separately, and summarizes the processed results to the upper layer or circulates between other nodes. In addition, the coordinator node and each computing node can interact through a network or other communication methods. In this embodiment, the coordinator node is mainly used to communicate with the client 110, and to generate an optimal execution plan, and send the optimal execution plan to the computing node. The computing node is mainly used to execute the execution plan received from the coordinator node. When executing the plan, it can read data from the storage node bound to it and process the data, or forward the data to other nodes to achieve data redistribution. The storage node is mainly used to store data. In some embodiments, the coordinator node can be integrated into the computing node or separated from the computing node, which is not limited here. In some embodiments, the coordinator node can be configured with Figure 1 The connector 211, query cache 212, parser 213 and optimizer 214 shown in the figure; each computing node can be configured with an executor 215; each storage node can be configured with a storage engine 220.
[0053] Figure 2 The physical architecture of the database 200 shown in the figure can be understood as a database architecture with separated storage and computing. This database architecture with separated storage and computing is derived from the traditional database architecture with integrated storage and computing. In the traditional database architecture with integrated storage and computing, storage resources and computing resources are bound together. When evolving to the database architecture with separated storage and computing, there is also a certain binding relationship between computing nodes and storage nodes. This means that when the plan executed by a computing node includes a redistribution plan, the computing node needs to first read data from the storage node with which it has a binding relationship, and then transmit the read data to other computing nodes. For example, Figure 3As shown, computing node A and storage node A have a binding relationship, computing node B and storage node B have a binding relationship, storage node A stores the data of column P1 in table T1, and storage node B stores the data of column P2 in table T1; if computing node A needs to operate on the data of column P2 in table T1, and computing node B needs to operate on the data of column P1 in table T1, then both computing nodes A and B need to perform redistribution operations. At this time, computing node A needs to first read the data of column P1 in table T1 from storage node A, and then transfer the data to computing node B; similarly, computing node B needs to first read the data of column P2 in table T1 from storage node B, and then transfer the data to computing node A. It can be seen that when executing the redistribution plan, the number of network transmission hops is large, and the data transmission volume of the entire network is also large, which increases the consumption of network and computing resources and reduces data query performance.
[0054] In view of this, an embodiment of the present application provides a data redistribution method. When the redistribution plan is included in the optimal execution plan, the redistribution plan can be pushed down to the storage node, so that the storage node can directly send the data to the corresponding computing node, thereby changing the data transmission from pulling from the destination end to actively pushing from the source end, reducing the number of network transmission hops, reducing the amount of data transmission in the network, and reducing the consumption of network and computing resources, thereby improving data query performance.
[0055] The following is an introduction to the data redistribution method provided in an embodiment of the present application.
[0056] For example, Figure 4 A flow chart of a data redistribution method provided by an embodiment of the present application is shown. The method can be applied to a database, which may include: a coordinator node, multiple computing nodes, and at least one storage node. Exemplarily, the database may be, but is not limited to, a cloud database. Figure 4 As shown, the data redistribution method may include the following steps:
[0057] S401: The coordinator node generates an optimal execution plan based on the query statement received from the client.
[0058] In this embodiment, when the coordinator node in the database receives a query statement from the client, it can process the query statement to generate an optimal execution plan. Figure 2 The coordinator node described in .
[0059] S402: The coordinator node determines that the optimal execution plan includes a redistribution plan, where the redistribution plan is used to redistribute the first data stored in the first storage node.
[0060] In this embodiment, in a scenario where data needs to be redistributed, the coordinator node can determine that the optimal execution plan it generates includes a redistribution plan. For example, when the aggregation key included in the query statement is inconsistent with the distribution key used when storing data in the database, the coordinator node can determine that the optimal execution plan includes a redistribution plan, or when the join key included in the query statement is inconsistent with the distribution key used when storing data in the database, the coordinator node can determine that the optimal execution plan includes a redistribution plan. Exemplarily, the redistribution plan can be used to redistribute the first data stored in the first storage node. Exemplarily, the first storage node can be, but is not limited to, Figure 2 A storage node described in .
[0061] S403: The coordinator node transmits the redistribution plan to the first storage node.
[0062] In this embodiment, the coordinator node can transmit the redistribution plan to the first storage node to push the redistribution plan down to the first storage node. Exemplarily, the coordinator node can learn the data that needs to be redistributed through the redistribution plan; further, by querying the metadata in the database, it can be known on which storage node the data is stored. In some embodiments, the coordinator node can directly send the redistribution plan to the first storage node, or forward the redistribution plan to the first storage node through a computing node. The specific method can be determined according to the actual situation and is not limited here.
[0063] S404: The first storage node transmits the first data to at least one computing node among the multiple computing nodes based on the redistribution plan.
[0064] In this embodiment, after the first storage node obtains the redistribution plan, it can execute the redistribution plan, and transmit the first data to at least one computing node among the multiple computing nodes, so that the first data can be processed by at least one computing node among the multiple computing nodes. Exemplarily, after obtaining the redistribution plan, the first storage node executes the redistribution plan and traverses and scans the data in each partition thereon to obtain the first data. Then, the first storage node can perform calculations according to a partitioning algorithm (such as a hash partitioning algorithm, etc.) to filter out at least one computing node for processing the first data from multiple computing nodes. In this way, at least one computing node is filtered out from multiple computing nodes in the database. Finally, the first storage node can transmit the first data to the filtered computing nodes. In some embodiments, the first storage node can first establish a data transmission channel with each filtered computing node, and then transmit data to each filtered computing node. Among them, the first storage node can perform a handshake operation with the computing node to establish a data transmission channel between the two. After the computing node obtains the first data, or obtains data from other storage nodes, it can calculate the obtained data. Exemplarily, the computing node can be but is not limited to Figure 2 A compute node described in .
[0065] In this way, when there is a redistribution plan in the optimal execution plan, the data transmission is changed from pulling from the destination to actively pushing from the source, which can reduce the number of network transmission hops, reduce the amount of data transmitted in the network, and reduce the consumption of network and computing resources, thereby improving data query performance.
[0066] In some embodiments, the database may include multiple storage nodes, and one storage node is associated with one computing node, for example, having a binding relationship, etc. In other embodiments, the computing nodes and storage nodes included in the database may not have an association relationship.
[0067] In some embodiments, Figure 4In the example, the storage node in the database can be a node logically, for example, it is composed of some thread pools. At this time, there may be no binding relationship between the data and the storage node, that is, the storage node can read the data of all partitions. In this scenario, the storage node is still used to push data, which can also improve the efficiency of data query. Since the data is stored according to the distribution key rule, such as a partition stores a file or the same group of files, but the computing node needs to redistribute the data and then aggregate them together for calculation. At this time, it will be more efficient for the storage node to read the files of a certain partition and then distribute them to different computing nodes. Compared with the computing nodes pulling separately, the storage node actively pushes the solution can effectively reduce the scanning of data. For example, when there are 3 computing nodes that need to pull data from the storage node, the computing node actively pulls the data, which requires each computing node to scan the data from the storage node once, while the solution in this embodiment only requires the storage node to scan once, which greatly reduces the number of data scans and reduces the input / output (I / O) of the disk.
[0068] It is understandable that the order of execution of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In addition, the various embodiments described above can be combined according to actual conditions, and the combined solutions are still within the scope of protection of the present application.
[0069] Based on the method in the above embodiment, an embodiment of the present application provides a database.
[0070] For example, Figure 5 The schematic diagram of the structure of a database provided by the embodiment of the present application is shown. Figure 5As shown, the database 500 may include: a coordinator node 510, multiple computing nodes 520 and at least one storage node 530. The coordinator node 510 is used to determine that the optimal execution plan generated by the coordinator node 510 includes a redistribution plan, and the redistribution plan is used to redistribute the first data, wherein the first data is stored in the first storage node 530 of at least one storage node 530, and the first data is processed by the first computing node 520 of the multiple computing nodes 520. The coordinator node 510 is also used to transmit the redistribution plan to the first storage node 530. The first storage node 530 is used to transmit the first data to at least one computing node 520 of the multiple computing nodes based on the redistribution plan. Exemplarily, the coordinator node 510 may communicate with the storage node 530 indirectly through the computing node 520, or may communicate with the storage node 530 directly. In addition, there may be a binding relationship between the computing node 520 and the storage node 530, or there may not be a binding relationship, which may be determined according to actual conditions. Exemplarily, the storage node 530 may be a memory, a hard disk, a disk array, or a cloud storage pool (eg, object storage, a distributed file system, a cloud block storage, etc.).
[0071] In some embodiments, when the first storage node 530 transmits the first data to at least one computing node 520 based on the redistribution plan, it is specifically used to: filter out the first data based on the redistribution plan; filter out at least one computing node 520 from multiple computing nodes 520 based on the first data; and transmit the first data to at least one computing node 520.
[0072] In some embodiments, before transmitting the first data to at least one computing node 520 , the first storage node 530 is further configured to: establish a data transmission channel with each of the screened computing nodes 520 .
[0073] In some embodiments, when the coordinator node 510 determines that the optimal execution plan generated by the coordinator node 510 includes a redistribution plan, it is specifically used to: determine that the aggregation key included in the query statement is inconsistent with the distribution key used when storing data in the database; or determine that the connection key included in the query statement is inconsistent with the distribution key used when storing data in the database.
[0074] It should be understood that the above-mentioned database is used to execute the methods in the above-mentioned embodiments. The implementation principles and technical effects of the corresponding nodes in the database are similar to those described in the above-mentioned methods. The working process of the database can refer to the corresponding process in the above-mentioned method and will not be repeated here.
[0075] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a computing device cluster including at least one computing device, the computing device cluster executes the method described in the above embodiment. Exemplarily, the computer-readable storage medium can be any available medium that can be stored by the computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), a semiconductor medium (e.g., a solid-state hard disk), or a cloud storage pool (e.g., an object storage, a distributed file system, a cloud block storage, etc.), etc.
[0076] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product including instructions. When the computer program product is run on a computing device cluster including at least one computing device, the computing device cluster executes the method in the above embodiment.
[0077] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0078] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0079] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.
[0080] It should be understood that the various numerical numbers involved in the embodiments of the present application are only used for the convenience of description and are not used to limit the scope of the embodiments of the present application.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A data redistribution method, It is characterized in that Applied to a database, the database includes: a coordinator node, multiple computing nodes and at least one storage node, the method includes: The coordinator node determines that the optimal execution plan generated by the coordinator node includes a redistribution plan, where the redistribution plan is used to redistribute the first data, wherein the first data is stored in a first storage node among the at least one storage node; The coordinator node transmits the redistribution plan to the first storage node; The first storage node transmits the first data to at least one computing node among the plurality of computing nodes based on the redistribution plan.
2. The method according to claim 1, It is characterized in that The first storage node transmits the first data to the at least one computing node based on the redistribution plan, specifically including: The first storage node filters out the first data based on the redistribution plan; The first storage node selects the at least one computing node from the plurality of computing nodes based on the first data; The first storage node transmits the first data to the at least one computing node.
3. The method according to claim 1 or 2, It is characterized in that Before the first storage node transmits the first data to the first computing node, the method further includes: The first storage node establishes a data transmission channel with each computing node in the at least one computing node.
4. The method according to any one of claims 1 to 3, It is characterized in that The coordinator node determines that the optimal execution plan generated by the coordinator node includes a redistribution plan, specifically including: The coordinator node determines that the aggregation key included in the query statement is inconsistent with the distribution key used when storing data in the database; Alternatively, the coordinator node determines that the connection key included in the query statement is inconsistent with the distribution key used when storing data in the database.
5. A database, It is characterized in that include: A coordinator node, multiple computing nodes, and at least one storage node; The coordinator node is used to determine that the optimal execution plan generated by the coordinator node includes a redistribution plan, and the redistribution plan is used to redistribute the first data, wherein the first data is stored in a first storage node among the at least one storage node; The coordinator node is further used to transmit the redistribution plan to the first storage node; The first storage node is used to transmit the first data to at least one computing node among the multiple computing nodes based on the redistribution plan.
6. The database according to claim 5, It is characterized in that When the first storage node transmits the first data to at least one computing node among the plurality of computing nodes based on the redistribution plan, the first storage node is specifically configured to: filter out the first data based on the redistribution plan; Based on the first data, selecting the at least one computing node from the plurality of computing nodes; The first data is transmitted to the at least one computing node.
7. A database according to claim 5 or 6, It is characterized in that Before transmitting the first data to the at least one computing node, the first storage node is further configured to: A data transmission channel is established with each computing node in the at least one computing node.
8. A database according to any one of claims 5 to 7, It is characterized in that When the coordinator node determines that the optimal execution plan generated by the coordinator node includes a redistribution plan, the coordinator node is specifically used to: Determining that an aggregation key included in the query statement is inconsistent with a distribution key used when storing data in the database; Alternatively, it is determined that the join key included in the query statement is inconsistent with the distribution key used when storing data in the database.
9. A computer-readable storage medium storing a computer program, which, when executed on a computing device cluster comprising at least one computing device, enables the computing device cluster to execute the method according to any one of claims 1 to 4.
10. A computer program product, It is characterized in that When the computer program product is run on a computing device cluster including at least one computing device, the computing device cluster is enabled to execute the method according to any one of claims 1 to 4.
Citation Information
Cited By
Data query method and device for distributed database, equipment and medium
CN121560998A