Partition hash connection method, device and storage medium
By dynamically determining the construction table and probing table in hash connection, and using the target hash connection algorithm for connection, the repartition problem caused by insufficient memory is solved and the connection performance is improved.
Patent Information
- Application Number
- CN202111254599.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-10-27
AI Technical Summary
During the hash connection process, insufficient memory space causes all data of the build table to be loaded, and it needs to be repartitioned and additional operations are added, resulting in a degradation of connection performance.
By obtaining the construction table of each partition table, determining the elements, comparing, determining the construction table and probing table based on preset conditions, using the target hash connection algorithm for connection, avoiding repartition.
Reduces the possibility of repartitioning, reduces additional operations, and ensures that connection performance does not decline.
Smart Images

Figure CN113986919B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to data table connection technology, and in particular, to a partitioned hash connection method, device, and storage medium. Background Art
[0002] In the database field, join is a common operation, which is to combine two or more tables in the database. According to different implementation methods, join can be divided into nested loop join, hash join and sort merge join. Among them, hash join has become one of the most commonly used join methods due to its high execution efficiency and support for massive data join.
[0003] Hash join is mainly divided into two stages, namely the establishment stage and the detection stage. In the establishment stage, each record in one of the tables to be connected is hashed to obtain a hash table (also called the construction table); in the detection stage, the other tables to be connected (also called the detection table) are scanned, the hash value of each record in the table is calculated, and compared with the hash table. If the connection conditions are met, the corresponding two records are connected until all records in the other tables to be connected are scanned.
[0004] In the implementation of hash join, the memory space may be too small to load all the data in the build table. Therefore, the build table and the detection table are partitioned. However, this does not guarantee that all the data in any partition of the build table can be loaded into the memory. In this case, the number of partitions needs to be expanded and re-partitioned, which will increase additional operations and lead to a decrease in connection performance. Summary of the invention
[0005] The embodiments of the present application provide a partition hash connection method, device and storage medium to reduce the possibility of re-partitioning, reduce additional calculations, and ensure that the connection performance is not reduced.
[0006] In a first aspect, an embodiment of the present application provides a partitioned hash join method, comprising:
[0007] In the case of performing a partitioned hash join on two tables to be joined, for any pair of partitioned tables obtained by partitioning, obtaining table construction determining elements of each partitioned table;
[0008] Comparing the construction table determination elements of the partition tables, determining the construction table according to the partition table corresponding to the construction table determination element that meets the preset conditions, and determining another partition table as the detection table;
[0009] Acquire the primary-secondary relationship of the two tables to be connected and the attribution relationship between the construction table and the detection table and the two tables to be connected, and determine the target hash connection algorithm according to the primary-secondary relationship and the attribution relationship;
[0010] Using the target hash join algorithm to connect the construction table and the detection table to obtain a connection result;
[0011] When all the connections of the partition tables are completed, the connection result of performing the hash connection on the two tables to be connected is obtained according to the connection results corresponding to each pair of partition tables.
[0012] In a second aspect, an embodiment of the present application further provides a computer device, including a processor and a memory, wherein the memory is used to store instructions, and when the instructions are executed, the processor performs the following operations:
[0013] In the case of performing a partitioned hash join on two tables to be joined, for any pair of partitioned tables obtained by partitioning, obtaining table construction determining elements of each partitioned table;
[0014] Comparing the construction table determination elements of the partition tables, determining the construction table according to the partition table corresponding to the construction table determination element that meets the preset conditions, and determining another partition table as the detection table;
[0015] Acquire the primary-secondary relationship of the two tables to be connected and the attribution relationship between the construction table and the detection table and the two tables to be connected, and determine the target hash connection algorithm according to the primary-secondary relationship and the attribution relationship;
[0016] Using the target hash join algorithm to connect the construction table and the detection table to obtain a connection result;
[0017] When all the connections of the partition tables are completed, the connection result of performing the hash connection on the two tables to be connected is obtained according to the connection results corresponding to each pair of partition tables.
[0018] In a third aspect, an embodiment of the present application further provides a storage medium, the storage medium is used to store instructions, and the instructions are used to execute:
[0019] In the case of performing a partitioned hash join on two tables to be joined, for any pair of partitioned tables obtained by partitioning, obtaining table construction determining elements of each partitioned table;
[0020] Comparing the construction table determination elements of the partition tables, determining the construction table according to the partition table corresponding to the construction table determination element that meets the preset conditions, and determining another partition table as the detection table;
[0021] Acquire the primary-secondary relationship of the two tables to be connected and the attribution relationship between the construction table and the detection table and the two tables to be connected, and determine the target hash connection algorithm according to the primary-secondary relationship and the attribution relationship;
[0022] Using the target hash join algorithm to connect the construction table and the detection table to obtain a connection result;
[0023] When all the connections of the partition tables are completed, the connection result of performing the hash connection on the two tables to be connected is obtained according to the connection results corresponding to each pair of partition tables.
[0024] The technical solution of the embodiment of the present application is as follows: when performing a partitioned hash connection on two tables to be connected, for any pair of partition tables obtained by partitioning, the construction table determination elements of each partition table are obtained; the construction table determination elements of each partition table are compared, and the construction table is determined according to the partition table corresponding to the construction table determination element that meets the preset conditions, and the other partition table is determined as the detection table; the primary and secondary relationships of the two tables to be connected and the attribution relationships of the construction table and the detection table with the two tables to be connected are obtained, and a target hash connection algorithm is determined according to the primary and secondary relationships and the attribution relationships; the construction table and the detection table are connected using the target hash connection algorithm to obtain a connection result; each time the partition tables are connected, the construction table and the detection table are re-determined according to the construction table determination elements, and if one of the partition tables meets the preset conditions, it can be determined as the construction table, thereby avoiding the possibility of re-partitioning. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1a A flowchart of a partitioned hash join method provided in Example 1 of the present application;
[0026] Figure 1b This is a specific schematic diagram of a hash connection provided in Example 1 of the present application;
[0027] Figure 2a A schematic diagram of a process of performing a hash join on a partition table using a preset hash right join algorithm provided in Embodiment 2 of the present application;
[0028] Figure 2b A specific schematic diagram of performing hash join on partition tables using a preset hash right join algorithm;
[0029] Figure 3a A schematic diagram of a process of performing a hash join on a partition table using a preset hash right join algorithm provided in Embodiment 3 of the present application;
[0030] Figure 3b A specific schematic diagram of performing a hash join on a partition table using a preset hash left join algorithm;
[0031] Figure 3c A connection diagram for tilted data provided in Embodiment 3 of the present application;
[0032] Figure 4 A schematic diagram of the structure of a partitioned hash connection device provided in Embodiment 4 of the present application;
[0033] Figure 5 A schematic diagram of the structure of a computer device provided in Example 5 of the present application. DETAILED DESCRIPTION
[0034] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for ease of description, only the parts related to the present application, rather than all structures, are shown in the accompanying drawings.
[0035] It should be mentioned before discussing the exemplary embodiments in more detail that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the operations (or steps) as sequential processes, many of the operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0036] The term "table to be connected" used in this article refers to a data table in the database that needs to be connected.
[0037] The term "partition" used in this article refers to dividing a data table into a preset number of sub-tables, and each sub-table area is called a partition.
[0038] The term "hash join" used in this article refers to a connection method for connecting data tables. Hash join is mainly divided into two stages, namely the establishment stage and the detection stage. In the establishment stage, each record in one of the tables to be connected is hashed to obtain a hash table (also called a build table); in the detection stage, the other tables to be connected (also called the detection table) are scanned, the hash value of each record in the table is calculated, and compared with the hash table. If the connection conditions are met, the corresponding two records are connected until all records in the other tables to be connected are scanned.
[0039] The term "partition table" used in this article refers to the subtable mentioned in the above introduction of "partition".
[0040] The term "table building determination factor" used in this article is used to determine information based on which a table is built from a pair of partition tables, such as data volume and data distribution information.
[0041] The term "preset condition" used in this document refers to a pre-set condition used to determine which partition table is a build table.
[0042] The term "construction table" used in this article refers to the hash table composed of hash values after the hash operation of all data needs to be performed at one time during the hash connection process.
[0043] The term "probe table" used in this article refers to the table that needs to be probed during the hash join process.
[0044] The term "primary-secondary relationship" used in this article refers to the correspondence between the two tables to be connected and the primary table mark and the secondary table mark when performing a hash join on the two tables to be connected. That is, one of the two tables to be connected will be marked as the primary table, and the other table will be marked as the secondary table. Generally, the table with the smaller amount of data is marked as the primary table.
[0045] The term "attribution relationship" used in this article refers to the inclusion relationship between the detection table, the construction table and the table to be connected. Generally, the partition table belongs to the table to be connected, that is, the partition table is obtained after the table to be connected is partitioned. If the two tables to be connected are table a and table b, then the detection table may be a part of table a, and correspondingly, the construction table is a part of table b, then the detection table belongs to table a, and the construction table belongs to table b.
[0046] The term "target hash join algorithm" used in this article refers to a hash join algorithm determined from a plurality of pre-set hash join algorithms that needs to be used in the current scenario, and is used to perform hash joins on the detection table and the construction table.
[0047] The term "join result" used in this article refers to the result obtained after performing a hash join on the build table and the probe table.
[0048] The term "data volume" as used herein refers to the size of the data a table contains.
[0049] The term “data distribution information” used herein refers to the repetitiveness information of data in a table, which is used to determine whether there is a skewed key in the table.
[0050] The term "number of skewed keys" used in this article refers to the specific number of skewed keys in a table. For example, there are two skewed keys in Table A, where skewed key A involves 8 data items and skewed key B involves 15 data items. Then the number of skewed keys in Table 2 is 8 skewed keys A and 15 skewed keys B.
[0051] The terms "preset hash right join algorithm" and "preset hash left join algorithm" used in this article are different from the traditional left outer join and right outer join. Left outer join and right outer join are join types at the logical level, and only require the results to meet the definition. Hash left join and hash right join are algorithm types at the implementation level, which are used to implement all logical join types. The specific join process will be introduced later, so I will not repeat it here.
[0052] The term "matching flag" used in this article refers to a flag added to the data in the build table when using the preset hash right join algorithm.
[0053] The term "row number" as used herein refers to the numbering of a row in a table.
[0054] To facilitate understanding, the main inventive concepts of the embodiments of the present application are briefly described.
[0055] In the prior art, when performing a hash connection, the table with the smaller data volume in the two tables to be connected is first determined as the main table, and the table with the larger data volume is determined as the secondary table. Then, the main table is loaded into the memory for hashing to obtain a hash table (i.e., a constructed table). However, if the current memory space is smaller than the size of the main table, the main table cannot be fully loaded into the memory at this time, and the main table and the secondary table need to be partitioned into the same number of partition tables.
[0056] After obtaining the partition table, the prior art will load the partition table of the main table into the memory in sequence, perform hashing, obtain the hash table corresponding to the partition table, and then detect the partition table of the corresponding secondary table to achieve the connection of the partition table of the corresponding main table and secondary table. However, since the partition uses a hash function, the size of the partition table obtained by the partition is different, which may cause the size of the partition table to still be larger than the memory size. At this time, it is necessary to reselect the hash function for repartitioning, which will increase additional operations and cause the connection performance to deteriorate.
[0057] In response to the above situation, the inventor delayed the process of determining the build table (i.e., the partition table that needs to be loaded into the memory) and the detection table until the partition table is loaded into the memory. That is to say, in the scheme of the present application, the partition table of the main table is not necessarily loaded into the memory for building the build table. Instead, the partition table of the main table or the partition table of the sub-table that meets the preset conditions is loaded into the memory to construct the corresponding build table.
[0058] Based on the above considerations, the inventor creatively proposed that, in the case of performing partitioned hash join on two tables to be connected, for any pair of partition tables obtained by partitioning, the construction table determination elements of each partition table are obtained; the construction table determination elements of each partition table are compared, and the construction table is determined according to the partition table corresponding to the construction table determination element that meets the preset conditions, and the other partition table is determined as the detection table; the primary and secondary relationships of the two tables to be connected and the attribution relationships of the construction table and the detection table with the two tables to be connected are obtained, and the target hash join algorithm is determined according to the primary and secondary relationships and the attribution relationships; the construction table and the detection table are connected using the target hash join algorithm to obtain a connection result; each time the partition tables are connected, the construction table and the detection table are re-determined according to the construction table determination elements, and if one of the partition tables meets the preset conditions, it can be determined as the construction table, thereby avoiding the possibility of re-partitioning.
[0059] Embodiment 1
[0060] Figure 1a The present invention provides a flowchart of a partition hash join method provided in the first embodiment of the present invention. The present invention is applicable to the case where the size of the partition table is larger than the size of the memory space during the partition hash join. The method can be executed by the partition hash join device provided in the present invention. The device can be implemented in software and / or hardware and can generally be integrated in a computer device. Figure 1a As shown, the method of the embodiment of the present application specifically includes:
[0061] Step 101: When performing a partitioned hash join on two tables to be joined, for any pair of partitioned tables obtained by partitioning, obtain table construction determination elements of each partitioned table.
[0062] It should be noted that when performing a hash connection on two tables to be connected, the table to be connected with a smaller amount of data will be determined as the main table, and the table with a larger amount of data will be determined as the secondary table. The main table will then be loaded into memory, and the data in the main table will be hashed using a hash function to obtain the hash values corresponding to each data. All hash values constitute the construction table.
[0063] If the current memory space is insufficient to load all the data of the primary table, the primary table and the secondary table will be partitioned. At this time, the partitioned hash join of the two tables to be joined is performed as mentioned in this step.
[0064] Generally speaking, the number of partitions for the main table and the secondary table is the same. In a specific example, if the main table is partitioned into 32 partition tables, the secondary table will also be partitioned into 32 partition tables, and each partition is in a one-to-one correspondence. The details can be shown in Table 1 below:
[0065] Table 1
[0066] Partition Number Main table Sub-table Partition 1 Partition table 1_main table Partition table 1_sub-table Partition 2 Partition table 2_main table Partition table 2_sub-table Partition 3 Partition table 3_main table Partition table 3_sub-table ....... ...... ......
[0067] As shown in Table 1, partition tables with the same partition number are considered a pair of partition tables, that is, "Partition Table 1_Main Table" and "Partition Table 1_Sub Table" are a pair of partition tables. Any pair of partition tables referred to in this step refers to two partition tables corresponding to any partition number.
[0068] It should be noted that after the main table and the secondary table are partitioned, a pair of partition tables corresponding to each partition number are generally connected in the order of the partition numbers. For the sake of convenience, this embodiment only takes "Partition Table 1_Main Table" and "Partition Table 1_Secondary Table" as examples to illustrate the connection process of the partition tables in the partition hash connection.
[0069] In this step, when any pair of partition tables are connected, the construction table determination elements of each partition table are first obtained, wherein the construction table determination elements are used to determine which partition table of the two partition tables is used as the table for determining the construction table and loaded into the memory.
[0070] In a specific example, the obtained construction table determination element of "Partition Table 1_Main Table" may be construction table determination element 1, and the obtained construction table determination element of "Partition Table 1_Sub-Table" may be construction table determination element 2.
[0071] Step 102: compare the construction table determination elements of each partition table, determine the construction table according to the partition table corresponding to the construction table determination element that meets the preset conditions, and determine another partition table as the detection table.
[0072] In this step, the construction table determination elements of each partition table are compared, and then based on the comparison results, the partition table corresponding to the construction table determination element that meets the preset conditions is selected, and then the construction table is obtained based on the selected partition table, and the other partition table is determined as the detection table.
[0073] Specifically, the elements for constructing a table may include data volume and data distribution information, wherein the data volume refers to the size of the space occupied by all the data contained in the partition table, and the data distribution information refers to the duplicate data existing in the partition table and the number of each duplicate data.
[0074] In a specific example, the specific data in “Partition Table 1_Main Table” may be as shown in Table 2, and the specific data in “Partition Table 1_Secondary Table” may be as shown in Table 3.
[0075] Table 2
[0076] Line number C1 C2 0 1 10 1 1 10 2 1 10 3 1 10 4 1 20 5 1 30 6 1 10 7 3 20
[0077] Table 3
[0078] Line number C1 C2 0 1 15 1 3 40 2 5 0
[0079] Assuming that the space occupied by a data unit in the data table is fixed, the data volume of "Partition Table 1_Main Table" is 16 data units, and the data volume of "Partition Table 1_Sub-Table" is 6 data units; the data distribution information of "Partition Table 1_Main Table" is that there is 1 duplicate data "1" in column C1, and the number of duplicates is "7", and there are 2 duplicate data "10" and "20" in column C2, and the number of duplicates is "5" "10" and "2" "20", and the data distribution information of "Partition Table 1_Sub-Table" is that there is 0 duplicate data in column C1, and the number of duplicates is "0", and there is 0 duplicate data in column C2, and the number of duplicates is "0".
[0080] In this step, the same parameter scores need to be compared. First, the number of skew keys of each partition table can be determined based on the data distribution information. Then, the partition table with the least number of skew keys and the smallest data volume is determined as the construction table, and the other partition table is determined as the detection table.
[0081] The number of tilted keys of each partition table is determined based on the data distribution information, that is, the repeated data in the data distribution data can be determined as the tilted key, and the corresponding number of repeated data is determined as the number of data items involved in the tilted key. In a specific example, as shown in Table 2 and Table 3, "Partition Table 1_Main Table" has 3 tilted keys, and the number of data items involved in the tilted keys is 7, 5, and 2 respectively; "Partition Table 1_Secondary Table" has 0 tilted keys, and the number of data items involved is 0.
[0082] It should be noted that the aforementioned minimum number of tilted keys and minimum amount of data are the preset conditions in this step. In a specific example, "Partition Table 1_Sub-Table" meets the preset conditions, so it is necessary to load "Partition Table 1_Sub-Table" into the memory, perform hash operations, obtain the hash values of each data, and the hash values constitute the corresponding construction table. And "Partition Table 1_Main Table" is determined as the detection table.
[0083] Step 103: Obtain the primary and secondary relationships of the two tables to be connected and the attribution relationships between the construction table and the detection table and the two tables to be connected, and determine the target hash connection algorithm according to the primary and secondary relationships and the attribution relationships.
[0084] It should be noted that this embodiment first needs to obtain the primary and secondary relationships. Specifically, the table to be connected with the smallest amount of data in the two tables to be connected can be determined as the primary table, and the other table to be connected can be determined as the secondary table.
[0085] This embodiment pre-sets two hash join algorithms, which are process algorithms for specifically implementing hash joins. Since it is unknown which partition table of the primary table and the secondary table is used as the construction table and which is used as the detection table in this embodiment, but the primary and secondary tables of the hash join need to be determined, in order to ensure the consistency of the hash join, this embodiment specially sets a preset hash left join algorithm and a preset hash right join algorithm.
[0086] In this step, if the build table belongs to the main table and the detection table belongs to the secondary table, the preset hash right join algorithm is determined as the target hash join algorithm; if the build table belongs to the secondary table and the detection table belongs to the main table, the preset hash left join algorithm is determined as the target hash join algorithm.
[0087] Therefore, when the partition table of the main table is determined to be a build table, a preset hash right join algorithm is used, and when the partition table of the secondary table is determined to be a build table, a preset hash left join algorithm is used. It should be noted that determining the partition table as a build table mentioned in this application means that determining the partition table is the table on which the build table is based.
[0088] In a specific example, for partition 1 in Table 1, "Partition Table 1_Sub-Table" is loaded into the memory through Table 2 and Table 3, and a hash operation is performed to obtain the hash value of each data, and the hash value constitutes the corresponding construction table. And "Partition Table 1_Main Table" is determined as the detection table.
[0089] Therefore, "Partition Table 1_Sub-Table" corresponds to the build table, and the build table belongs to the sub-table. At this time, the preset hash left join algorithm is determined as the target hash join algorithm.
[0090] Step 104: Use the target hash join algorithm to join the construction table and the detection table to obtain a join result.
[0091] In this step, different hash connection algorithms have different specific connection processes. The specific connection process will be introduced in subsequent embodiments and will not be repeated here.
[0092] In addition, when connecting the build table and the detection table, it is necessary to traverse the detection table. Every time a piece of data is traversed, the hash value of the connection attribute of the data needs to be calculated, and then compared with the hash value in the build table to find the result that meets the connection conditions, and the result is combined with the traversed data to obtain the connection data corresponding to the traversed data.
[0093] After the traversal of the detection table is completed, all connection data constitute the connection result in this step.
[0094] Still taking Table 2 and Table 3 as an example, the connection condition can be that c1 of Table 2 = c1 of Table 3 and c2 of Table 2 < c2 of Table 3. Then the specific connection process can be referred to Figure 1b , Figure 1b This is a specific schematic diagram of a hash connection provided in Example 1 of the present application.
[0095] like Figure 1b As shown, the leftmost table (i.e., Table 2) is traversed row by row. The first row is "1, 10". The row "1, 15" where c1 is also 1 is found in the middle table. Then the sizes of the two groups of c2 are compared. "10" < "15", which meets the condition and the comparison result is "true output". If not, the two groups of data are combined to obtain "1, 10, 1, 15", which is a connected data. The comparison result is "false skip", and the next data can be traversed.
[0096] After traversing Table 2, 6 connection data are obtained, and the 6 connection data constitute the connection result of the hash connection of Table 2 and Table 3.
[0097] Step 105: When all the partition table connections are completed, a connection result of performing a hash connection on two tables to be connected is obtained according to the connection results corresponding to each pair of partition tables.
[0098] In this step, after all the connection results are spliced, the connection result of the hash connection of the two tables to be connected can be obtained.
[0099] The embodiment of the present application provides a partitioned hash join method. When performing a partitioned hash join on two tables to be joined, for any pair of partitioned tables obtained by partitioning, the construction table determination elements of each partition table are obtained; the construction table determination elements of each partition table are compared, and the construction table is determined according to the partition table corresponding to the construction table determination element that meets the preset conditions, and the other partition table is determined as the detection table; the primary and secondary relationships of the two tables to be joined and the attribution relationships of the construction table and the detection table with the two tables to be joined are obtained, and a target hash join algorithm is determined according to the primary and secondary relationships and the attribution relationships; the construction table and the detection table are connected using the target hash join algorithm to obtain a connection result; each time the partition tables are connected, the construction table and the detection table are re-determined according to the construction table determination elements, and if one of the partition tables meets the preset conditions, it can be determined as the construction table, thereby avoiding the possibility of re-partitioning.
[0100] Embodiment 2
[0101] Figure 2a A schematic diagram of a flow chart of performing a hash join on a partition table using a preset hash right join algorithm is provided for the second embodiment of the present application. The embodiment of the present application can be combined with each optional solution in one or more of the above embodiments.
[0102] like Figure 2a As shown, the method of the embodiment of the present application specifically includes:
[0103] Step 201: traverse the data in the detection table, and when traversing to any data, query from the construction table whether there is data matching the data.
[0104] In this step, the detection table is traversed, specifically, it can be traversed row by row. For details, please refer to Figure 2b , Figure 2b This is a specific schematic diagram of performing hash join on a partition table using a preset hash right join algorithm.
[0105] like Figure 2bAs shown, the detection table is the table on the far left of the diagram, that is, the table including the data of "1a, 4c, 2b, 5d", and the construction table is the data table in the middle. When traversing to "1" in the detection table, check from the construction table whether there is data consistent with "1".
[0106] Step 202: If it exists, connect the matching data in the build table with the corresponding data in the detection table to obtain a connection data, and add a matching flag to the matching data in the build table.
[0107] In this step, a matching flag may be added in advance to the data in the construction table, with the initial value of the matching flag set to "false", and then in this step the matching flag corresponding to the matched data is updated to "true".
[0108] Of course, it is also possible not to add in advance, but only add the matched data to distinguish the unmatched data.
[0109] like Figure 2b As shown, the matching flags are the "T" and "F" shown in the middle table.
[0110] Step 203: until all the data in the detection table are traversed, the data in the construction table that do not have a matching identification bit are connected to the preset fields respectively to obtain the connection data corresponding to each piece of data that do not have a matching identification bit.
[0111] This step is based on the method mentioned in step 202 that "it is also possible not to add in advance, only add the matched data to distinguish the unmatched data", that is, there is no data with a matching identification bit, that is, there is no data matching the detection table.
[0112] In addition, if the method mentioned in step 202 is adopted, that is, "pre-adding a matching flag to the data in the build table, setting the initial value of the matching flag to "false", and then updating the matching flag corresponding to the matched data to "true" in this step", this step can be to connect the data in the build table with a matching flag of "false" to the preset field.
[0113] It should be noted that the preset field in this step can be "null".
[0114] Step 204: determine all the obtained connection data as the connection result of performing a hash connection between the construction table and the detection table.
[0115] In this embodiment, the preset hash right join algorithm is used, and the result consistent with the main table being the construction table and the secondary table being the detection table can still be obtained.
[0116] In addition, during the detection phase, it is necessary to traverse each piece of data in the detection table and obtain matching data from the construction table. This process usually results in a large number of function calls. If there is skewed data in the detection table, repeated function calls may occur.
[0117] In the case of skewed data, this embodiment can connect the skewed data before step 201. Specifically, the skewed data in the detection table can be determined first, and the row numbers of all the skewed data can be collected; data matching the skewed data can be determined from the construction table, and the matching data can be subjected to columnar operations with the data corresponding to the collected row numbers to obtain the connection data corresponding to the data of each collected row number; and then a traversal process is performed on the data in the detection table except for the skewed data, i.e., step 201.
[0118] It should be noted that the processing of tilted data will be described in detail in subsequent embodiments and will not be repeated here.
[0119] Embodiment 3
[0120] Figure 3a A schematic diagram of a flow chart of performing a hash join on a partition table using a preset hash right join algorithm is provided for the third embodiment of the present application. The embodiment of the present application can be combined with each optional solution in one or more of the above embodiments.
[0121] like Figure 3a As shown, the method of the embodiment of the present application specifically includes:
[0122] Step 301: traverse the detection table, and when traversing to any data, query from the construction table whether there is data matching the data.
[0123] In this step, the detection table is traversed, specifically, it can be traversed row by row. For details, please refer to Figure 3b , Figure 3b This is a specific schematic diagram of performing hash join on a partition table using a preset hash left join algorithm.
[0124] like Figure 3b As shown, the detection table is the table on the far left of the diagram, that is, the table containing the data "1, 1, 2, 3", and the construction table is the data table in the middle. When traversing to "1" in the detection table, check from the construction table whether there is data consistent with "1".
[0125] Step 302: If it exists, connect the matching data in the construction table with the corresponding data in the detection table to obtain a connection data.
[0126] Step 303: After traversing all the data in the detection table, all the obtained connection data are determined as the connection result of performing a hash connection between the construction table and the detection table.
[0127] It should be noted that, in this embodiment, for any traversal data in the detection table, if there is no corresponding matching data in the construction table, the unmatched data in the detection table can be combined with the preset field to form connection data, where the preset field is also "null".
[0128] In the case of skewed data, this embodiment can connect the skewed data before step 301. Specifically, the skewed data in the detection table can be determined first, and the row numbers of all the skewed data can be collected; data matching the skewed data can be determined from the construction table, and the matching data can be subjected to columnar operations with the data corresponding to the collected row numbers to obtain the connection data corresponding to the data of each collected row number; a process of traversing the data in the detection table except the skewed data can be performed, i.e., step 301.
[0129] For details, please refer to Figure 3c , Figure 3c A connection diagram for tilted data provided in Example 3 of the present application.
[0130] like Figure 3c As shown, the tilted data in Table 2 are now determined. Generally, the largest number is taken, that is, "1", and the corresponding row numbers are "0, 1, 2, 3, 4, 5, 6". They are stored in blocks, and then the data consistent with "1" in Table 3 is found, that is, "1, 15", and then the block-stored data is vectorized with "1, 15" (i.e., column-based operation), and then it is determined whether the connection conditions are met, that is, true and false as shown in the figure, and the true connection data can be output.
[0131] It should be noted that, to determine the skewed data in the construction table, sampling may be performed from the construction table, and data in which the number of identical data in the sampled data is greater than a preset number threshold is determined as skewed data.
[0132] Embodiment 4
[0133] Figure 4 This is a schematic diagram of the structure of a partition hash connection device provided in the fourth embodiment of the present application. The device can be implemented in software and / or hardware, and can generally be integrated in a computer device. Figure 4 As shown, the device includes: an element acquisition module 401, a comparison module 402, an algorithm determination module 403, a connection module 404 and a connection result determination module 405.
[0134] Among them, the element acquisition module 401 is used to obtain the construction table determination elements of each partition table for any pair of partition tables obtained by partitioning when performing partitioned hash connection on two tables to be connected; the comparison module 402 is used to compare the construction table determination elements of each partition table, determine the construction table according to the partition table corresponding to the construction table determination element that meets the preset conditions, and determine the other partition table as the detection table; the algorithm determination module 403 is used to obtain the primary and secondary relationships of the two tables to be connected and the attribution relationship between the construction table and the detection table and the two tables to be connected, and determine the target hash connection algorithm according to the primary and secondary relationships and the attribution relationship; the connection module 404 is used to connect the construction table and the detection table using the target hash connection algorithm to obtain a connection result; the connection result determination module 405 is used to obtain the connection result of the hash connection of the two tables to be connected according to the connection results corresponding to each pair of partition tables when all the connections of the partition tables are completed.
[0135] The embodiment of the present application provides a partitioned hash connection device. When performing a partitioned hash connection on two tables to be connected, for any pair of partitioned tables obtained by partitioning, the construction table determination elements of each partition table are obtained; the construction table determination elements of each partition table are compared, and the construction table is determined according to the partition table corresponding to the construction table determination element that meets the preset conditions, and the other partition table is determined as the detection table; the primary and secondary relationships of the two tables to be connected and the attribution relationships of the construction table and the detection table with the two tables to be connected are obtained, and a target hash connection algorithm is determined according to the primary and secondary relationships and the attribution relationships; the construction table and the detection table are connected using the target hash connection algorithm to obtain a connection result; each time the partition tables are connected, the construction table and the detection table are re-determined according to the construction table determination elements, and if one of the partition tables meets the preset conditions, it can be determined as the construction table, thereby avoiding the possibility of re-partitioning.
[0136] Based on the above embodiments, the comparison module 402 may include: a skew key number determination unit, used to determine the number of skew keys of each partition table based on data distribution information; a comparison unit, used to determine the partition table with the least number of skew keys and the smallest data volume as a construction table, and determine another partition table as a detection table.
[0137] Based on the above embodiments, the algorithm determination module 403 may include: a primary-secondary relationship determination unit, configured to determine the table to be connected with the smallest amount of data in the two tables to be connected as the primary table and the other table to be connected as the secondary table;
[0138] A first algorithm determination unit is used to determine a preset hash right join algorithm as a target hash join algorithm if the build table belongs to the main table and the detection table belongs to the secondary table;
[0139] The second distribution determination unit is used to determine a preset hash left join algorithm as a target hash join algorithm if the construction table belongs to the secondary table and the detection table belongs to the primary table.
[0140] On the basis of the above embodiments, the connection module 404 may include: a first traversal unit, configured to traverse the data in the detection table, and when traversing to any data, query from the construction table whether there is data matching the data;
[0141] A matching identification bit adding unit is used to connect the matching data in the build table with the corresponding data in the detection table to obtain a connection data, and add a matching identification bit to the matching data in the build table;
[0142] A first connection unit is used to connect the data in the construction table that do not have a matching identification bit with the preset field until all the data in the detection table are traversed, so as to obtain connection data corresponding to each piece of data that does not have a matching identification bit;
[0143] The first connection result determining unit is used to determine all the obtained connection data as the connection result of performing a hash connection between the construction table and the detection table.
[0144] On the basis of the above embodiments, the connection module 404 may further include: a first tilted data determining unit, configured to determine the tilted data in the detection table and collect the row numbers of all the tilted data;
[0145] A first matching unit is used to determine data matching the tilted data from the construction table, and perform column operation on the matching data and the data corresponding to the collected row numbers to obtain connection data corresponding to the collected data of each row number;
[0146] The first execution unit is used to perform a traversal process on the data in the detection table except the skewed data.
[0147] On the basis of the above embodiments, the connection module 404 may include: a second traversal unit, configured to traverse the detection table, and when traversing to any data, query from the construction table whether there is data matching the data;
[0148] A second connection unit is used to connect the matching data in the construction table with the corresponding data in the detection table to obtain a connection data if it exists;
[0149] The second connection result determining unit is used to determine all the obtained connection data as the connection result of performing hash connection between the construction table and the detection table after traversing all the data in the detection table.
[0150] On the basis of the above embodiments, the connection module 404 may further include: a second tilted data determination unit, configured to determine the tilted data in the detection table and collect the row numbers of all the tilted data;
[0151] A second matching unit is used to determine data matching the tilted data from the construction table, and perform column operation on the matching data and the data corresponding to the collected row numbers to obtain connection data corresponding to the data of each row number collected;
[0152] The second execution unit is used to perform a traversal process on the data in the detection table except the skewed data.
[0153] Based on the above embodiments, the second skewed data determination unit may include: a skewed data determination subunit, configured to perform sampling from the construction table and determine data in which the number of identical data in the sampled data is greater than a preset number threshold as skewed data.
[0154] The partitioned hash join device can execute the partitioned hash join method provided in any embodiment of the present application, and has the corresponding functional modules and beneficial effects for executing the partitioned hash join method.
[0155] Embodiment 5
[0156] Figure 5 A schematic diagram of the structure of a computer device provided in Example 5 of the present application. Figure 5 An exemplary computer device suitable for implementing the embodiments of the present application includes a processor 510, a memory 520, an input device 530, and an output device 540; the number of the computer device 510 can be one or more. Figure 5 A processor 510 is taken as an example; the processor 510, the memory 520, the input device 530 and the output device 540 in the device / terminal / server can be connected via a bus or other means. Figure 5 The example of connecting through bus is taken in the following.
[0157] The memory 520, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the partitioned hash join method in the embodiment of the present invention (for example, the element acquisition module 401, the comparison module 402, the algorithm determination module 403, the connection module 404 and the connection result determination module 405 in the partitioned hash join device). The processor 510 executes various functional applications and data processing of the device / terminal / server by running the software programs, instructions and modules stored in the memory 520, that is, the method of the above embodiment is implemented:
[0158] The processor 510 executes various functional applications and data processing by running instructions stored in the memory 520, such as performing the following operations:
[0159] In the case of performing a partitioned hash join on two tables to be joined, for any pair of partitioned tables obtained by partitioning, obtaining table construction determining elements of each partitioned table;
[0160] Comparing the construction table determination elements of each partition table, determining the construction table according to the partition table corresponding to the construction table determination element that meets the preset conditions, and determining another partition table as the detection table;
[0161] Obtain the primary-secondary relationship of the two tables to be connected and the attribution relationship between the construction table and the detection table and the two tables to be connected, and determine the target hash connection algorithm according to the primary-secondary relationship and the attribution relationship;
[0162] Use the target hash join algorithm to connect the construction table and the detection table to obtain the connection result;
[0163] When all the connections to the partition tables are completed, the connection result of performing a hash connection on the two tables to be connected is obtained according to the connection results corresponding to each pair of partition tables.
[0164] Based on the above embodiments, the elements for determining the construction table include the data volume and data distribution information; the processor 510 is configured to determine the construction table and the detection table in the following manner:
[0165] Comparing the construction table determination elements of each partition table, determining the construction table according to the partition table corresponding to the construction table determination element that meets the preset conditions, and determining another partition table as the detection table, including:
[0166] Determine the number of skew keys in each partition table based on data distribution information;
[0167] A partition table having the least number of skewed keys and the smallest amount of data is determined as a build table, and another partition table is determined as a probe table.
[0168] Based on the above embodiments, the processor 510 is configured to determine the target hash join algorithm in the following manner:
[0169] The table with the smallest amount of data in the two tables to be connected is determined as the main table, and the other table to be connected is determined as the secondary table;
[0170] If the construction table belongs to the main table and the detection table belongs to the secondary table, the preset hash right join algorithm is determined as the target hash join algorithm;
[0171] If the construction table belongs to the secondary table and the detection table belongs to the primary table, the preset hash left join algorithm is determined as the target hash join algorithm.
[0172] Based on the above embodiments, the processor 510 is configured to obtain the connection result in the following manner:
[0173] Traverse the data in the detection table, and when traversing to any data, query from the construction table whether there is data matching the data;
[0174] If it exists, connect the matching data in the build table with the corresponding data in the detection table to obtain a connection data, and add a matching flag to the matching data in the build table;
[0175] After all the data in the detection table are traversed, the data in the construction table that do not have a matching identification bit are connected to the preset fields respectively to obtain the connection data corresponding to each piece of data that does not have a matching identification bit;
[0176] All the obtained connection data are determined as the connection result of performing a hash connection between the build table and the detection table.
[0177] Based on the above embodiments, the processor 510 is configured to connect the tilt data in the following manner:
[0178] Determine the tilted data in the detection table and collect the row numbers of all the tilted data;
[0179] Determine the data matching the tilted data from the build table, and perform column operations on the matching data and the data corresponding to the collected row numbers to obtain the connection data corresponding to the collected data of each row number;
[0180] The process of traversing the data in the detection table except the skewed data.
[0181] Based on the above embodiments, the processor 510 is configured to connect the tilted data in the following manner before traversing the construction table:
[0182] Determine the skewed data in the build table and collect the row numbers of all skewed data;
[0183] Determine the data matching the tilted data from the detection table, and perform column operation on the matching data and the data corresponding to the collected row numbers to obtain the connection data corresponding to the collected data of each row number;
[0184] The process of traversing the data in the build table except for the skewed data.
[0185] Based on the above embodiments, the processor 510 is configured to determine the tilted data in the construction table in the following manner:
[0186] Sampling is performed from the construction table, and data in which the number of identical data in the sampled data is greater than a preset number threshold is determined as skewed data.
[0187] The memory 520 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 520 may further include a memory remotely arranged relative to the processor 510, and these remote memories may be connected to the device / terminal / server via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0188] The input device 530 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device / terminal / server. The output device 540 may include a display device such as a display screen.
[0189] Embodiment 6
[0190] Embodiment 6 of the present application provides a computer-readable storage medium, which is used to store instructions, and the instructions are used to execute the partition hash connection method provided in any embodiment of the present application.
[0191] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, device, or device.
[0192] Computer-readable signal media may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0193] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0194] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0195] Note that the above are only preferred embodiments of the present application and the technical principles used. Those skilled in the art will understand that the present application is not limited to the specific embodiments herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application is described in more detail through the above embodiments, the present application is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the appended claims.
Claims
1. A partitioned hash join method, characterized in that: include: In the case of performing a partitioned hash join on two tables to be joined, for any pair of partitioned tables obtained by partitioning, obtaining table construction determining elements of each partitioned table; Comparing the construction table determination elements of the partition tables, determining the partition table corresponding to the construction table determination element that meets the preset conditions as the construction table, and determining the other partition table as the detection table; The preset conditions are that the number of tilted keys is the smallest and the amount of data is the smallest; Acquire the primary-secondary relationship of the two tables to be connected and the attribution relationship between the construction table and the detection table and the two tables to be connected, and determine the target hash connection algorithm according to the primary-secondary relationship and the attribution relationship; Using the target hash join algorithm to connect the construction table and the detection table to obtain a connection result; When all the connections to the partition tables are completed, a connection result of performing a hash connection on the two tables to be connected is obtained according to the connection results corresponding to each pair of partition tables; The obtaining of the primary-secondary relationship between the two tables to be connected and the attribution relationship between the construction table and the detection table and the two tables to be connected, and determining a target hash connection algorithm according to the primary-secondary relationship and the attribution relationship, includes: The table to be connected with the smallest amount of data in the two tables to be connected is determined as a main table, and the other table to be connected is determined as a secondary table; If the construction table belongs to the main table and the detection table belongs to the secondary table, a preset hash right join algorithm is determined as a target hash join algorithm; If the construction table belongs to the secondary table and the detection table belongs to the primary table, a preset hash left join algorithm is determined as a target hash join algorithm.
2. The method according to claim 1, characterized in that The elements for determining the construction table include data volume and data distribution information; The comparing the construction table determination elements of the partition tables, determining the partition table corresponding to the construction table determination element that meets the preset conditions as the construction table, and determining another partition table as the detection table, includes: Determine the number of skew keys of each partition table based on the data distribution information; A partition table having the least number of skewed keys and the smallest amount of data is determined as a build table, and another partition table is determined as a probe table.
3. The method according to claim 1, characterized in that If the target hash join algorithm is a preset hash right join algorithm, the use of the target hash join algorithm to join the build table and the detection table to obtain a join result includes: Traversing the data in the detection table, when traversing to any data, querying from the construction table whether there is data matching the data; If so, connect the matching data in the build table with the corresponding data in the detection table to obtain a connection data, and add a matching flag to the matching data in the build table; After all the data in the detection table are traversed, the data in the construction table that do not have a matching identification bit are respectively connected with the preset fields to obtain the connection data corresponding to each piece of data that does not have a matching identification bit; All the obtained connection data are determined as the connection result of performing a hash connection between the construction table and the detection table.
4. The method according to claim 3, characterized in that Before traversing the detection table, the target hash join algorithm is used to connect the construction table and the detection table to obtain a connection result, and further includes: Determine the tilted data in the detection table and collect the row numbers of all the tilted data; Determine data matching the tilted data from the construction table, and perform column operation on the matching data and the data corresponding to the collected row numbers to obtain connection data corresponding to the collected data of each row number; The process of traversing the data in the detection table except the skewed data.
5. The method according to claim 1, characterized in that: If the target hash join algorithm is a preset hash left join algorithm, the use of the target hash join algorithm to join the build table and the detection table to obtain a join result includes: Traversing the detection table, when traversing to any data, querying from the construction table whether there is data matching the data; If so, connect the matching data in the build table with the corresponding data in the detection table to obtain a connection data; After traversing all the data in the detection table, all the obtained connection data are determined as the connection result of performing a hash connection between the construction table and the detection table.
6. The method according to claim 5, characterized in that Before traversing the detection table, the target hash join algorithm is used to connect the construction table and the detection table to obtain a connection result, and further includes: Determine the tilted data in the detection table and collect the row numbers of all the tilted data; Determine data matching the tilted data from the construction table, and perform column operation on the matching data and the data corresponding to the collected row numbers to obtain connection data corresponding to the collected data of each row number; The process of traversing the data in the detection table except the skewed data.
7. The method according to claim 6, characterized in that The step of determining the tilted data in the construction table includes: Sampling is performed from the construction table, and data in which the number of identical data in the sampled data is greater than a preset number threshold is determined as skewed data.
8. A computer device comprising a processor and a memory, wherein the memory is used to store instructions, and when the instructions are executed, the processor performs the following operations: In the case of performing a partitioned hash join on two tables to be joined, for any pair of partitioned tables obtained by partitioning, obtaining table construction determining elements of each partitioned table; Comparing the construction table determination elements of the partition tables, determining the partition table corresponding to the construction table determination element that meets the preset conditions as the construction table, and determining the other partition table as the detection table; The preset conditions are that the number of tilted keys is the smallest and the amount of data is the smallest; Acquire the primary-secondary relationship of the two tables to be connected and the attribution relationship between the construction table and the detection table and the two tables to be connected, and determine the target hash connection algorithm according to the primary-secondary relationship and the attribution relationship; Using the target hash join algorithm to connect the construction table and the detection table to obtain a connection result; When all the connections to the partition tables are completed, a connection result of performing a hash connection on the two tables to be connected is obtained according to the connection results corresponding to each pair of partition tables; The processor is configured to determine a target hash join algorithm by: The table to be connected with the smallest amount of data in the two tables to be connected is determined as a main table, and the other table to be connected is determined as a secondary table; If the construction table belongs to the main table and the detection table belongs to the secondary table, a preset hash right join algorithm is determined as a target hash join algorithm; If the construction table belongs to the secondary table and the detection table belongs to the primary table, a preset hash left join algorithm is determined as a target hash join algorithm.
9. The computer device according to claim 8, characterized in that The elements for determining the construction table include data volume and data distribution information; The processor is configured to determine the build table and the probe table by: The comparing the construction table determination elements of the partition tables, determining the partition table corresponding to the construction table determination element that meets the preset conditions as the construction table, and determining another partition table as the detection table, includes: Determine the number of skew keys of each partition table based on the data distribution information; A partition table having the least number of skewed keys and the smallest amount of data is determined as a build table, and another partition table is determined as a probe table.
10. The computer device according to claim 8, characterized in that The processor is configured to obtain the connection result in the following manner: Traversing the data in the detection table, when traversing to any data, querying from the construction table whether there is data matching the data; If so, connect the matching data in the build table with the corresponding data in the detection table to obtain a connection data, and add a matching flag to the matching data in the build table; After all the data in the detection table are traversed, the data in the construction table that do not have a matching identification bit are respectively connected with the preset fields to obtain the connection data corresponding to each piece of data that does not have a matching identification bit; All the obtained connection data are determined as the connection result of performing a hash connection between the construction table and the detection table.
11. The computer device according to claim 10, characterized in that The processor is configured to concatenate the tilted data before traversing the probe table in the following manner: Determine the tilted data in the detection table and collect the row numbers of all the tilted data; Determine data matching the tilted data from the construction table, and perform column operation on the matching data and the data corresponding to the collected row numbers to obtain connection data corresponding to the collected data of each row number; The process of traversing the data in the detection table except the skewed data.
12. The computer device according to claim 8, characterized in that The processor is configured to obtain the connection result in the following manner: Traversing the construction table, when traversing to any data, querying from the detection table whether there is data matching the data; If so, connect the matching data in the detection table with the corresponding data in the construction table to obtain a connection data; After traversing all the data in the construction table, all the obtained connection data are determined as the connection result of performing a hash connection between the construction table and the detection table.
13. The computer device according to claim 12, characterized in that The processor is configured to connect the tilted data in the following manner before traversing the build table: Determine the skewed data in the build table and collect the row numbers of all skewed data; Determine data matching the tilted data from the detection table, and perform column operation on the matching data and the data corresponding to the collected row numbers to obtain connection data corresponding to the data of each row number collected; The process of traversing the data in the build table except for the skewed data.
14. The computer device according to claim 13, characterized in that The processor is configured to determine the skewed data in the build table by: Sampling is performed from the construction table, and data in which the number of identical data in the sampled data is greater than a preset number threshold is determined as skewed data.
15. A storage medium, wherein the storage medium is used to store instructions, wherein the instructions are used to execute the partitioned hash join method as described in any one of claims 1-7.
Citation Information
Patent Citations
Method and device for database to execute hash link
CN113297209A
Parallel Hash join acceleration method and system based on FPGA
CN113468181A