Hash connection method, computing node, storage medium and program product
By using dynamic Bloom filters to filter data tables in distributed databases, the problem of inefficient hash connections is solved, and efficient data table connections are achieved in large-scale distributed databases.
Patent Information
- Application Number
- CN202210825987.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-13
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-07-13
AI Technical Summary
In distributed databases, the network overhead of hash connections increases exponentially with the increase in cluster size, resulting in inefficient hash connections and inability to meet the needs.
By filtering tuples in data tables based on dynamic Bloom filters, the amount of data that needs to be transmitted and probed is reduced, thereby improving the efficiency of hash connections. The specific steps include hashing the first data table based on the first hash function, generating a Bloom filter, filtering the second data table based on the Bloom filter, and finally hash connection.
Through the use of dynamic Bloom filters, the data volume and transmission overhead of the second data table are reduced, the efficiency of hash connection is improved, and the data table connection tasks in large-scale distributed databases can be more efficiently handled.
Smart Images

Figure CN115062027B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of database technology, and in particular to a hash connection method, a computing node, a storage medium, and a program product. Background Art
[0002] Hash join is a common operation to establish the relationship between multiple tables in the database. Hash join builds a hash table for one table, scans the tuples of another table and compares them with the hash table to detect the connection between the two.
[0003] For distributed databases, especially large-scale distributed databases such as data warehouses, since data is redistributed to multiple nodes, when performing hash joins, tuples in the data tables stored in each node need to be transmitted over the network so that when the tuples and the corresponding tuples in the hash table meet the equal connection conditions, they can be connected and output. The network overhead of hash joins will increase exponentially with the increase of cluster size, resulting in low efficiency of hash joins and failure to meet demand. Summary of the invention
[0004] The present application provides a hash connection method, a computing node, a storage medium and a program product, which filters tuples in a data table based on a dynamic Bloom filter, thereby reducing the amount of data to be transmitted and detected and improving the efficiency of the hash connection.
[0005] In a first aspect, the present application provides a hash connection method, comprising:
[0006] Based on the first hash function, a hash operation is performed on the first data table to obtain a hash table; based on the hash table, a Bloom filter is generated; according to the Bloom filter, the second data table is filtered; the hash table and the filtered second data table are hash-connected to obtain and output a hash connection table.
[0007] Optionally, generating a Bloom filter based on the hash table includes:
[0008] Generate a sub-Bloom filter corresponding to the computing node according to the hash table corresponding to the computing node; obtain the sub-Bloom filters corresponding to each other computing node; and generate the Bloom filter according to the sub-Bloom filters corresponding to each of the computing nodes.
[0009] Optionally, generating a Bloom filter based on the hash table includes:
[0010] Generate a sub-Bloom filter corresponding to the computing node according to the hash table corresponding to the computing node; obtain setting information of the sub-Bloom filters corresponding to each other computing node; generate the Bloom filter according to the setting information and the sub-Bloom filter corresponding to the computing node; wherein the setting information is used to describe the bit of the corresponding sub-Bloom filter that is set.
[0011] Optionally, after generating the sub-Bloom filter corresponding to the computing node according to the hash table corresponding to the computing node, the method further includes:
[0012] Calculate the amount of data required to transmit the sub-Bloom filter in a first transmission mode and a second transmission mode respectively; determine the target transmission mode of the sub-Bloom filter from the first transmission mode and the second transmission mode according to the amount of data required for transmission; broadcast the sub-Bloom filter corresponding to the computing node to other computing nodes based on the target transmission mode; wherein, broadcasting the sub-Bloom filter corresponding to the computing node based on the first transmission mode includes: broadcasting the sub-Bloom filter corresponding to the computing node to other computing nodes; broadcasting the sub-Bloom filter corresponding to the computing node based on the second transmission mode includes: broadcasting the setting information of the sub-Bloom filter corresponding to the computing node to other computing nodes; wherein the setting information is used to describe the position where the corresponding sub-Bloom filter is set.
[0013] Optionally, generating a sub-Bloom filter corresponding to the computing node according to the hash table includes:
[0014] Obtain at least one second hash function; calculate the second hash value of each tuple in the first data table according to the at least one second hash function; construct a sub-Bloom filter corresponding to the computing node according to the first hash value and the second hash value; wherein the first hash value is the hash value in the hash table.
[0015] Optionally, the method further includes:
[0016] Determine a hash cost according to the first data table; determine the number of hash functions of the sub-Bloom filter corresponding to the computing node according to the hash cost and the connection cost, so as to obtain a corresponding number of second hash functions according to the number of hash functions.
[0017] Optionally, filtering the second data table according to the Bloom filter includes:
[0018] When the hash distribution of the second data table is different from that of the first data table, the Bloom filter is broadcast to the process corresponding to the second data table; through the process corresponding to the second data table, the tuples in the scanned second data table are filtered according to the Bloom filter to obtain a filtered second data table, and the filtered second data table is sent to the process corresponding to the first data table, so that the process corresponding to the first data table performs a hash connection on the hash table and the filtered second data table.
[0019] In a second aspect, the present application provides a hash connection device, comprising:
[0020] A hash operation module is used to perform a hash operation on the first data table based on a first hash function to obtain a hash table; a filter generation module is used to generate a Bloom filter based on the hash table; a filtering module is used to filter the second data table according to the Bloom filter; a hash connection module is used to perform a hash connection on the hash table and the filtered second data table to obtain and output a hash connection table.
[0021] In a third aspect, the present application provides a computing node, including:
[0022] A processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided in the first aspect of the present application.
[0023] In a fourth aspect, the present application provides a distributed database, comprising multiple computing nodes provided in the third aspect of the present application.
[0024] In a fifth aspect, the present application provides a computer-readable storage medium, in which computer-readable storage medium is stored computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method provided in the first aspect of the present application.
[0025] In a sixth aspect, the present application provides a computer program product, including a computer program, which implements the method provided in the first aspect of the present application when executed by a processor.
[0026] The hash connection method, computing node, storage medium and program product provided by the present application are aimed at application scenarios in which the connection relationship between a first data table and a second data table is to be determined through a hash connection. After performing a hash operation on the first data table based on a first hash function to obtain a corresponding hash table, a Bloom filter is dynamically constructed based on the hash table, and then the second data table is filtered based on the constructed Bloom filter. A hash connection is performed with the filtered second data table through the hash table to determine the relationship between the first data table and the second data table, that is, to obtain a hash connection table, so as to facilitate subsequent data processing based on the hash connection table, such as data statistics. By filtering the second data table based on a dynamic Bloom filter, the amount of data in the second data table is reduced, thereby reducing the overhead of transmission and detection of the second data table and improving the efficiency of the hash connection. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0028] Figure 1 A schematic diagram of an application scenario provided for an embodiment of the present application;
[0029] Figure 2 A flowchart of a hash connection method provided in an embodiment of the present application;
[0030] Figure 3 For this application Figure 2 A schematic diagram of the process of step S203 in the illustrated embodiment;
[0031] Figure 4 For this application Figure 2 A flowchart of an implementation method of step S202 in the illustrated embodiment;
[0032] Figure 5 For this application Figure 4 A schematic diagram of a full transmission method of a sub-Bloom filter in the illustrated embodiment;
[0033] Figure 6 For this application Figure 2 A flowchart of another implementation of step S202 in the illustrated embodiment;
[0034] Figure 7 For this application Figure 6 A schematic diagram of a sub-Bloom filter setting transmission method in the illustrated embodiment;
[0035] Figure 8 A flowchart of another hash connection method provided in an embodiment of the present application;
[0036] Fig. 9A schematic diagram of the structure of a computing node provided in an embodiment of the present application.
[0037] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0038] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0039] First, the terms involved in this application are explained:
[0040] Tuple: It is a basic concept in relational databases. Each row in a data table is a tuple.
[0041] Hash join: A table join using a hash algorithm is a database operator that is used to combine data in two or more tables based on the relationship between certain columns.
[0042] Bloom Filter: It is used to retrieve whether an element exists in a set. It mainly calculates the hash value of the element through K hash functions, and maps the calculated K hash values to K bits of the binary vector or bit array corresponding to the Bloom filter. If at least one of the values on the K bits is 0, it means that the element does not exist in the corresponding set; if the values on the K bits are all 1, it means that the element may exist in the corresponding set.
[0043] A distributed database is a logically identical database formed by connecting multiple physically dispersed database units using a computer network. Each connected database unit is called a node or computing node. A distributed database includes at least two nodes. The nodes can be physical nodes distributed in different places or logical nodes distributed in the same physical database.
[0044] Figure 1 A schematic diagram of an application scenario provided in an embodiment of the present application, such as Figure 1 As shown in Figure 2, a distributed database consists of multiple physical nodes, such as Figure 1Nodes 1 to n in the , one or more data tables are stored in nodes 1 to n, and different data tables store different data respectively. Figure 1 In the example, each node stores two data tables. Node 1 stores data table 1-1 and data table 2-1, node 2 stores data table 1-2 and data table 2-2, and so on. Node n stores data table 1-n and data table 2-n. Data tables 1-1 to 1-n are all part of data table 1, and data tables 2-1 to 2-n are all part of data table 2.
[0045] In some embodiments, some nodes may not store any data, or a node may store data table 3 or a portion of data table 3. The present application does not limit the data storage situation of each node.
[0046] Exemplarily, data table 1 may be an order table, including information such as member number, order time, and order details, and data table 2 may be a member table, including information such as member number, member name, and member time.
[0047] The main process of hash join includes the build process and the probe process. In the build process, for the data table with a smaller amount of data in the hash join (such as data table 1), recorded as the right table, the connection attribute of each tuple therein is calculated by hash operation, thereby building the corresponding hash table. In the probe process, for the other data table of the hash join, that is, the data table with a larger amount of data (such as data table 2), recorded as the left table, each row of the left table is scanned and the hash value of the connection attribute is calculated, and compared with the hash table constructed in the build process, the records that meet the connection conditions are found, thereby determining the connection relationship between the left table and the right table.
[0048] Taking the hash connection of data table 1 (right table) and data table 2 (left table) as an example, before detection, data table 2 needs to be transmitted, that is, data table 2-1 to data table 2-n are transmitted through the network, so that the hash tables corresponding to data table 2 and data table 1 are in the same process to perform the above-mentioned detection process.
[0049] For distributed databases, the left table and the right table may be distributed and stored in multiple physical nodes. When performing hash joins, the above hash join method is used, and the left table needs to be transmitted over the network, which will cause the network overhead to increase as the scale of the distributed database increases, resulting in low hash join efficiency.
[0050] When the left table and the right table are stored in the same node, the left table still needs to be transmitted to the hash join operator through local transmission for hash join. Since the left table has a large amount of data and the transmission overhead is large, the hash join efficiency is low.
[0051] In order to improve the efficiency of hash joins, that is, to reduce the amount of data transmitted during the hash join process, the present application provides a hash join method based on a dynamic Bloom filter. The main process of the method is to build a dynamic Bloom filter based on the right table of the hash join (that is, the subsequent first data table), which is usually a data table with less data during the hash join, and the hash table generated in the construction phase. Then, based on the dynamic Bloom filter, the tuples in the left table of the hash join (that is, the subsequent second data table) are filtered, thereby reducing the data volume of the left table. The filtered left table is sent to the physical node or process where the right table is located through the network, and then a hash join is performed based on the hash table corresponding to the right table and the filtered left table, thereby reducing the transmission overhead of the left table and reducing the amount of data detected during the hash join, thereby improving the efficiency of the hash join of data tables in the distributed database.
[0052] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0053] Figure 2 A flowchart of a hash connection method provided in an embodiment of the present application is provided, wherein the method is applied to a computing node of a database.
[0054] In one embodiment, the database may be a distributed database, or may be a centralized database.
[0055] like Figure 2 As shown, the hash connection method includes the following steps:
[0056] Step S201: performing a hash operation on a first data table based on a first hash function to obtain a hash table.
[0057] The first data table is the input of the hash connection establishment process, that is, one of the right tables, and is a data table that needs to create a corresponding hash table through the establishment process. The first data table can be stored in one or more nodes of the database. The corresponding first data table is hashed by one or more computing nodes to obtain the corresponding hash table.
[0058] A compute node can be a physical node or a logical node.
[0059] Specifically, for each computing node storing the first data table, the stored first data table is scanned via the computing node, and then based on the first hash function, a hash operation is performed on the tuples in the scanned first data table to obtain a corresponding hash table.
[0060] The hash table may be stored in the memory of the corresponding computing node. The first data table may be stored in the memory and / or disk of the corresponding computing node.
[0061] The first Hash function may be a Hash function constructed in any manner, and this application does not limit this.
[0062] Step S202: Generate a Bloom filter based on the hash table.
[0063] Specifically, after generating the corresponding hash table, the computing node may also construct a corresponding Bloom filter based on the hash table to filter the second data table.
[0064] When there are multiple computing nodes, for the sake of easy distinction, the Bloom filter corresponding to each computing node is recorded as a sub-Bloom filter. The computing node can also broadcast the constructed sub-Bloom filter to other computing nodes to integrate the various sub-Bloom filters to obtain a complete Bloom filter, so as to filter each second data table based on the complete Bloom filter.
[0065] Specifically, the Bloom filter may be set according to each first hash value in the hash table corresponding to the first sub-computing node, so as to obtain the Bloom filter or sub-Bloom filter corresponding to the first sub-computing node.
[0066] Initially, the Bloom filter (or sub-Bloom filter) is composed of a bit array of all 0s. Based on the hash values (recorded as the first hash value) in the hash table corresponding to the first data table, the bit array corresponding to the Bloom filter can be set, that is, the bit corresponding to the first hash value is set, that is, the value on it is set to 1, and the sub-Bloom filter corresponding to the computing node can be obtained by traversing the first hash values in the hash table corresponding to the computing node.
[0067] The Bloom filter is constructed by reusing the hash value in the hash table output by the hash connection establishment process, thereby reducing the cost of constructing the Bloom filter and improving the efficiency of the Bloom filter construction.
[0068] Since the hash table corresponding to the first sub-computation node is calculated only based on the first hash function, the number k of hash functions corresponding to the Bloom filter is only one, resulting in a high misjudgment rate. In order to reduce the misjudgment rate, at least one second hash function can be obtained, and the second hash value of the tuple in the first data table is calculated based on the second hash function, so as to construct the Bloom filter based on the first hash value and the second hash value. The more the number of second hash functions, the more second hash values need to be calculated, thereby increasing the overhead of constructing the Bloom filter. When constructing the Bloom filter, the number of second hash functions should be determined by comprehensively considering the overhead and misjudgment rate.
[0069] Step S203: Filter the second data table according to the Bloom filter.
[0070] The second data table is the input of the hash connection detection process, that is, one of the left tables. The second data table can be stored in one or more nodes of the database.
[0071] In one embodiment, the first data table and the second data table may be stored in the same physical node.
[0072] In one embodiment, the first data table and the second data table may be stored in different physical nodes.
[0073] In one embodiment, each computing node may correspond to a set of first data tables and second data tables.
[0074] Through each computing node, the second data table of the computing node is filtered according to the dynamically constructed Bloom filter to delete tuples in the second data table whose values of the Bloom filter positions of the corresponding hash value mappings are not all 1, thereby reducing the amount of data in the second data table and obtaining a filtered second data table.
[0075] In one embodiment, the hash distribution of the first data table and the second data table is the same, that is, the first data table and the second data table are stored in the same slice, that is, the first data table and the second data table are stored in the same process of the computing node, then the constructed Bloom filter can be directly passed by the process to the scanning operator of the second data table, and the tuples in the scanned second data table are filtered according to the Bloom filter through the scanning operator.
[0076] For example, Figure 3 For this application Figure 2 The flowchart of step S203 in the embodiment shown is as follows: Figure 3 As shown, Figure 3In the example, the Bloom filter is an 8-bit array, and the number of hash functions of the Bloom filter is one (that is, only the first hash function is the hash function of the Bloom filter), wherein the values of the 2nd, 4th, and 7th bits of the bit array corresponding to the Bloom filter constructed based on the hash table are 1, and the values of the remaining bits are 0. If a tuple in the second data table, such as x, has a bit position of the bit array corresponding to the hash value calculated by the first hash function of 3, since the value of the 3rd bit of the bit array corresponding to the Bloom filter is 0, it means that the tuple does not exist in the first data table, then delete the tuple x. If a tuple in the second data table, such as y, has a bit position of the bit array corresponding to the hash value calculated by the first hash function of 7, since the value of the 7th bit of the bit array corresponding to the Bloom filter is 1, it means that the tuple may exist in the first data table, then retain the tuple y. By analogy, the tuples whose bit positions of the bit array of the Bloom filter mapped in the second data table do not all have the value of 1 are deleted.
[0077] Through the above filtering operation, tuples in the second data table that are not related to the first data table are deleted. Without affecting the hash connection, the amount of data in the second data table is greatly reduced, and the efficiency of the second data table transmission is improved. At the same time, by reducing the amount of data in the second data table, the detection process during the hash connection is accelerated.
[0078] Step S204: performing a hash connection on the hash table and the filtered second data table to obtain and output a hash connection table.
[0079] A hash connection detection process is performed on the filtered second data table and the hash table obtained based on the hash operation of the first data table, so as to obtain a hash connection table describing the connection relationship between the first data table and the second data table.
[0080] Specifically, traverse each row or each record in the filtered second data table, and search the hash table for a record that meets the connection condition with the record; if it exists, store the record and the record in the first data table in the hash table that meets the connection condition with the record in the hash connection table. After traversing the filtered second data table, a hash connection table corresponding to the first data table and the second data table is obtained.
[0081] In one embodiment, if there are multiple computing nodes in the database, that is, there are multiple groups of first data tables and second data tables, the hash connection tables generated by the various computing nodes can be merged to obtain the final output hash connection table.
[0082] Furthermore, the hash connection table may be sent to the user terminal so that the user can know the connection relationship between each group of the first data table and the second data table, so as to facilitate subsequent data analysis processes such as data statistics and policy customization based on the connection relationship.
[0083] For example, the first data table may be an order table of store A within a set time period (such as one day, one week, etc.), the second data table may be a customer table, and the hash connection table may be used to represent a table of customers who place orders at store A within the set time period, that is, through hash connection, the customers and their attributes corresponding to each order within the set time period of store A are determined. Furthermore, based on the hash connection table, the customer tags corresponding to each product of store A can be determined, or the attributes of customers who purchase each product of store A can be counted.
[0084] The hash connection method provided by the present application is aimed at an application scenario in which a connection relationship between a first data table and a second data table is to be determined through a hash connection. After a hash operation is performed on the first data table based on a first hash function to obtain a corresponding hash table, a Bloom filter is dynamically constructed based on the hash table, and then the second data table is filtered based on the constructed Bloom filter. A hash connection is performed with the filtered second data table through the hash table to determine the relationship between the first data table and the second data table, that is, to obtain a hash connection table, so as to facilitate subsequent data processing based on the hash connection table, such as data statistics. By filtering the second data table based on a dynamic Bloom filter, the amount of data in the second data table is reduced, thereby reducing the overhead of transmission and detection of the second data table, and improving the efficiency of the hash connection.
[0085] Figure 4 For this application Figure 2 The flowchart of an implementation method of step S202 in the embodiment shown in the figure is directed to a distributed database including multiple computing nodes, each computing node stores a set of first data tables and second data tables, and the transmission method of the sub-Bloom filter provided in the embodiment is a full transmission method, such as Figure 4 As shown, the above step S202 may specifically include the following steps:
[0086] Step S401: Generate a sub-Bloom filter corresponding to the computing node according to the hash table corresponding to the computing node.
[0087] The hash table corresponding to the computing node is a hash table obtained by the computing node performing a hash operation on the scanned first data table of the computing node based on the first hash function.
[0088] For each computing node, a sub-Bloom filter corresponding to the computing node is generated through the computing node based on the corresponding hash table. The computing node broadcasts the generated sub-Bloom filter to other computing nodes.
[0089] Step S402, obtaining the sub-Bloom filter corresponding to each other computing node.
[0090] Each computing node broadcasts the generated sub-Bloom filter to other computing nodes, so that each computing node can obtain a complete Bloom filter.
[0091] Step S403: Generate the Bloom filter according to the sub-Bloom filters corresponding to the computing nodes.
[0092] Specifically, for each computing node, an OR operation or addition is performed on the bit arrays corresponding to each sub-Bloom filter to obtain a global and complete Bloom filter, so as to filter the second data table of each computing node based on the Bloom filter.
[0093] For example, Figure 5 For this application Figure 4 The schematic diagram of the full transmission mode of the sub-Bloom filter in the embodiment shown is as follows: Figure 5 As shown, Figure 5 Take two computing nodes, Seg1 and Seg2, as an example. Figure 5 In the example, the bit array of the sub-Bloom filter is 11 bits. Each computing node, namely Seg1 and Seg2, constructs the corresponding sub-Bloom filter, namely Bloom1 (01010010010) and Bloom2 (10001010010), and each computing node broadcasts the sub-Bloom filter constructed or generated by the node to other computing nodes, and then each computing node adds the sub-Bloom filters, namely Bloom1+Bloom2, to obtain a global and complete Bloom filter (11011010010).
[0094] Through the above-mentioned full transmission method, the sub-Bloom filter generated by each computing node is broadcast to other computing nodes, so that each computing node obtains a global complete Bloom filter, and filters the second data table of this node based on the complete Bloom filter, thereby improving the accuracy of filtering and avoiding accidental deletion of records that are connected to records in the first data table.
[0095] Figure 6 For this application Figure 2 The flowchart of another implementation of step S202 in the embodiment shown in the figure is directed to a distributed database including multiple computing nodes, each computing node stores a set of first data tables and second data tables, and Figure 4 The transmission mode of the sub-Bloom filter shown in FIG. 1 is different. The transmission mode of the sub-Bloom filter provided in this embodiment is a setting transmission mode, such as Figure 6 As shown, the above step S202 may specifically include the following steps:
[0096] Step S601: Generate a sub-Bloom filter corresponding to the computing node according to the hash table corresponding to the computing node.
[0097] Step S602: Acquire the setting information of the sub-Bloom filters corresponding to each other computing node.
[0098] The setting information is used to describe the bits of the corresponding sub-Bloom filter that are set, that is, to describe the bits in the sub-Bloom filter whose values are 1. For example, the setting information (2, 4, 8, 11) indicates that the 2nd, 4th, 8th and 11th bits of the bit array of the corresponding sub-Bloom filter are set to 1, and the remaining bits are 0.
[0099] Each computing node broadcasts the setting information of the generated sub-Bloom filter to other computing nodes, so that each computing node can obtain a complete Bloom filter.
[0100] In order to increase the rate of determining the setting information, the setting information of each sub-Bloom filter may be determined based on the AVX instruction set accelerated computing technology.
[0101] Step S603: Generate the Bloom filter according to the setting information and the sub-Bloom filter corresponding to the computing node.
[0102] Specifically, for each computing node, the sub-Bloom filter generated or constructed by this node is set based on the setting information sent by other computing nodes, that is, the value of the bit corresponding to each setting information is set to 1, so that a global and complete Bloom filter can be obtained to filter the second data table of each computing node based on the Bloom filter.
[0103] Furthermore, the setting information of other computing nodes can be merged to delete repeated bits in the setting information to obtain merged setting information, and based on the merged setting information, the sub-Bloom filter corresponding to the computing node can be set or updated to obtain a global and complete Bloom filter.
[0104] For example, taking the sub-Bloom filter generated by the current computing node as "01101001", the setting information sent by the other two computing nodes are (2, 3) and (2, 4, 7) respectively, then the merged setting information is (2, 3, 4, 7). After updating the sub-Bloom filter based on the merged setting information, the resulting Bloom filter is "01111011", that is, the values of the 4th and 7th bits of the sub-Bloom filter are set to 1.
[0105] In one embodiment, "0" or "1" may be used as the first bit of the bit array.
[0106] For example, Figure 7 For this application Figure 6 The schematic diagram of the sub-Bloom filter setting transmission method in the embodiment shown is as follows: Figure 7 As shown, Figure 7Take three computing nodes, Seg1 to Seg3, for example. Figure 7 The bit array of the sub-Bloom filter is 8 bits. Each computing node broadcasts the setting information of the sub-Bloom filter constructed or generated by this node to other computing nodes. Then, each computing node obtains a global and complete Bloom filter based on the setting information sent by other computing nodes and the sub-Bloom filter generated by this node. Figure 7 The setting information of Seg1, Seg2 and Seg3 are (1, 4), (4, 6) and (2, 4) respectively. Figure 7 Only the process of Seg1 generating a complete Bloom filter is shown, specifically: Seg1 sets the sub-Bloom filter generated by Seg1 based on the setting information (4, 6) sent by Seg2 and the setting information (2, 4) sent by Seg2 to obtain a complete Bloom filter, that is, "01101010", where the first bit of the bit array is bit 0. Similarly, other computing nodes, namely Seg2 and Seg3, obtain the completed Bloom filter in a similar manner.
[0107] Through the above-mentioned setting transmission method, only the setting information of the sub-Bloom filter generated by each computing node is broadcast to other computing nodes, so that each computing node obtains a globally complete Bloom filter. In some cases, the setting information is determined faster. Since the amount of setting information data is smaller than that of the corresponding sub-Bloom filter, the amount of data required to be transmitted when generating the Bloom filter is reduced by only sending the setting information, thereby improving the efficiency of Bloom filter generation.
[0108] Figure 8 A flowchart of another hash connection method provided in an embodiment of the present application. This embodiment is aimed at a scenario where multiple groups of first data tables and second data tables need to be hashed together. Each group of first data tables and second data tables is stored in a corresponding computing node. This embodiment is in Figure 2 Based on the illustrated embodiment, step S202 and step S203 are further refined, and the method is applied to each computing node storing a set of first data tables and second data tables in a distributed database.
[0109] like Figure 8 As shown, the hash connection method may include the following steps:
[0110] Step S801: Based on a first hash function, a hash operation is performed on the scanned first data table to obtain a hash table.
[0111] Step S802: Generate a sub-Bloom filter corresponding to the computing node according to the hash table corresponding to the computing node.
[0112] In order to improve the efficiency of sub-Bloom filter generation, the sub-Bloom filter can be generated based on only one hash function, namely, the first hash function, that is, the sub-Bloom filter corresponding to the computing node is generated based only on each first hash value in the hash table corresponding to the computing node.
[0113] In order to reduce the misjudgment rate of the sub-Bloom filter and filter more data in the second data table, the sub-Bloom filter can be constructed based on multiple hash functions, one of which is the first hash function.
[0114] Optionally, generating a sub-Bloom filter corresponding to the computing node according to a hash table corresponding to the computing node includes:
[0115] Obtain at least one second hash function; calculate a second hash value of each tuple in the first data table according to the at least one second hash function; and construct a sub-Bloom filter corresponding to the computing node according to the first hash value and the second hash value.
[0116] The first hash value is a hash value in the hash table output in the establishment phase. The second hash function can be any hash function different from the first hash function.
[0117] In one embodiment, in order to improve efficiency, the at least one second Hash function may be generated based on the first Hash function.
[0118] Specifically, at least one second Hash function may be obtained by adjusting the output domain, mapping rule, etc. of the first Hash function.
[0119] A new second Hash function may be generated based on the generated second Hash function and at least one of the first Hash function.
[0120] Specifically, different offset values may be passed to the first Hash function to construct multiple second Hash functions in a linear manner.
[0121] Based on the second hash function, the second hash value of each tuple in the first data table is calculated, and then based on the second hash value and the previously calculated first hash value (the hash value in the hash table), the sub-Bloom filter corresponding to the computing node is constructed. The specific method is similar to step S202, and only the hash value based on it is replaced from "first hash value" to "first hash value and second hash value".
[0122] Through the above method, N+1 hash operations need to be performed on each tuple in the first data table, where N is the number of the second hash function, so as to obtain a first hash value and N second hash values corresponding to each tuple in the first data table, and then based on the first hash value and N second hash values corresponding to each tuple in the first data table, the corresponding sub-Bloom filter is set to obtain the sub-Bloom filter corresponding to the computing node.
[0123] The more the number of second hash functions, the greater the cost of constructing the sub-Bloom filter; the fewer the number of second hash functions, the higher the misjudgment rate of the sub-Bloom filter. In order to balance the misjudgment rate and the cost, the number of second hash functions needs to be reasonably set. The present application also provides a method for determining the number of second hash functions, which is specifically:
[0124] Determine the hash cost according to the first data table; determine the connection cost according to the second data table; determine the number of hash functions of the sub-Bloom filter corresponding to the computing node according to the hash cost and the connection cost, so as to obtain the corresponding number of second hash functions according to the number of hash functions.
[0125] The hash cost is used to describe the cost required for performing a hash operation on each row or tuple in the first data table, and the join cost is used to describe the cost of performing a hash join on the second data table if no filtering is performed.
[0126] Specifically, the hash cost may be determined based on the number of rows and structure of the first data table, and the connection cost may be determined based on the number of rows and structure of the second data table.
[0127] In one embodiment, the number of second hash functions is inversely correlated with the hash cost and positively correlated with the connection cost.
[0128] Specifically, for each computing node corresponding to a group of first data tables and second data tables, before generating the sub-Bloom filter of the computing node, it is necessary to first determine the number of hash functions of the sub-Bloom filter, specifically: based on the structure and number of rows of the first data table corresponding to the computing node, calculate the hash cost, based on the structure and number of rows of the second data table corresponding to the computing node, calculate the connection cost; and then determine the number of hash functions of the sub-Bloom filter corresponding to the computing node based on the connection cost, the misjudgment rate and the hash cost.
[0129] In one embodiment, the number N of the second hash functions of the sub-Bloom filter can be calculated based on the following expression:
[0130] C(i)=(p(i)-p(i+1))×C join -i×C hash
[0131] Starting from i=1 with a step size of 1, calculate the value of each C(i), C(N) is the minimum value of C(i), that is, the number N of the second hash function is selected when C(i) takes the minimum value, that is, when i=N, C(N) is the minimum value of each C(·) calculated.
[0132] Wherein, p(i) is the misjudgment rate of the sub-Bloom filter when the number of the second hash function is i, and i is a natural number; C join is the above connection cost, C hash is the hash cost mentioned above.
[0133] In one embodiment, an upper limit value of N may be set, such as 2, 3, 5, 7, 9 or other values.
[0134] Step S803 , respectively calculating the amount of data required to transmit the sub-Bloom filter in the first transmission mode and the second transmission mode.
[0135] Among them, the first transmission mode is the above-mentioned full transmission mode, and the second transmission mode is the above-mentioned setting transmission mode.
[0136] The amount of data required by the first transmission mode and the second transmission mode may be determined based on the data distribution of the sub-Bloom filter.
[0137] Step S804: determining a target transmission mode of the sub-Bloom filter from the first transmission mode and the second transmission mode according to the amount of data required for transmission.
[0138] Specifically, a transmission mode requiring the least amount of data between the first transmission mode and the second transmission mode may be determined as the target transmission mode to reduce the overhead of sub-Bloom filter broadcasting.
[0139] Step S805: if the target transmission mode is the first transmission mode, broadcast the sub-Bloom filter corresponding to the computing node to other computing nodes.
[0140] Step S806, obtaining the sub-Bloom filter corresponding to each other computing node.
[0141] Each computing node related to the hash connection, that is, the computing nodes corresponding to a group of first data tables, only needs to generate the sub-Bloom filter corresponding to the computing node based on the above method, and then broadcast it to other computing nodes corresponding to the first data table, so that each computing node obtains the sub-Bloom filter corresponding to each computing node.
[0142] Step S807: Generate the Bloom filter according to the sub-Bloom filters corresponding to each of the computing nodes. Jump to step S811.
[0143] Step S808: If the target transmission mode is the second transmission mode, broadcast the setting information of the sub-Bloom filter corresponding to the computing node to other computing nodes.
[0144] Step S809: Acquire the setting information of the sub-Bloom filters corresponding to each other computing node.
[0145] Step S810: Generate the Bloom filter according to the setting information and the sub-Bloom filter corresponding to the computing node.
[0146] After obtaining a complete Bloom filter based on the branch corresponding to the first transmission mode, i.e., step S805 to step S807, or based on the branch corresponding to the second transmission mode, i.e., step S808 to step S810, the second data table is filtered based on the Bloom filter, i.e., step S811 is executed.
[0147] Step S811: Filter the second data table according to the Bloom filter.
[0148] When the hash distribution of the first data table and the second data table corresponding to the same computing node is different, since the operator corresponding to the hash operation of the first data table and the operator corresponding to the hash connection are in the same post-slicing process, it is necessary to send and receive the second data table and the Bloom filter through the network. In order to reduce the amount of data when the second data table is transmitted, the second data table can be filtered through the Bloom filter before the second data table is sent and received.
[0149] Specifically, the Bloom filter may be first sent to the process where the second data table is located to filter the second data table, and then the filtered second data table may be sent to the process corresponding to the first data table to perform a hash connection with the hash table to obtain a hash connection table.
[0150] Optionally, filtering the second data table according to the Bloom filter includes:
[0151] When the hash distribution of the second data table is different from that of the first data table, the Bloom filter is broadcast to the process corresponding to the second data table; through the process corresponding to the second data table, the tuples in the scanned second data table are filtered according to the Bloom filter to obtain a filtered second data table, and the filtered second data table is sent to the process corresponding to the first data table, so that the process corresponding to the first data table performs a hash connection on the hash table and the filtered second data table.
[0152] In view of the different hash distributions between the first data table and the second data table, the second data table is filtered by a Bloom filter before being sent or received, thereby reducing the amount of data sent or received over the network and improving the efficiency of sending or receiving the second data table.
[0153] Step S812: Perform hash connection on the hash table and the filtered second data table to obtain and output a hash connection table corresponding to the computing node.
[0154] Specifically, the hash connection tables generated by the various computing nodes may be sent to the same computing node, so that the computing node integrates the hash connection tables to obtain and output the hash connection result.
[0155] In one embodiment, the hash connection result may also be fed back to a target terminal, which may be a user terminal, an online analysis terminal, or the like.
[0156] In this embodiment, for the application scenario of performing hash connection on multiple groups of first data tables and second data tables stored in multiple computing nodes in distributed data, for each computing node, after performing hash operation on the first data table corresponding to the computing node based on the first hash function to obtain the corresponding hash table, a sub-Bloom filter corresponding to the computing node is dynamically constructed based on the hash table, and based on the distribution situation, a transmission method with less data volume is selected to broadcast the sub-Bloom filter to other required computing nodes, so that each computing node obtains a complete Bloom filter to improve the accuracy of filtering; then, the second data table of each computing node is filtered based on the complete Bloom filter, and a hash connection is performed with the filtered second data table through the hash table to determine the connection between the first data table and the second data table, that is, to obtain a hash connection table, and by filtering the second data table based on the dynamic Bloom filter, the amount of data in the second data table is reduced, thereby reducing the overhead of network transmission and detection of the second data table, and improving the efficiency of hash connection.
[0157] An embodiment of the present application provides a hash connection device, which includes: a hash operation module, a filter generation module, a filtering module and a hash connection module.
[0158] Among them, the hash operation module is used to perform a hash operation on the scanned first data table based on the first hash function to obtain a hash table; the filter generation module is used to generate a Bloom filter based on the hash table; the filtering module is used to filter the second data table according to the Bloom filter; the hash connection module is used to perform a hash connection on the hash table and the filtered second data table to obtain and output a hash connection table.
[0159] Optionally, the database is a distributed database, including multiple computing nodes, each of which stores a first data table, and the filter generation module includes:
[0160] A sub-filter generation unit is used to generate a sub-Bloom filter corresponding to the computing node according to the hash table corresponding to the computing node; an acquisition unit is used to acquire the sub-Bloom filters corresponding to each other computing node; and a first filter generation unit is used to generate the Bloom filter according to the sub-Bloom filters corresponding to each of the computing nodes.
[0161] Optionally, the database is a distributed database, including multiple computing nodes, each of which stores a first data table, and the filter generation module includes:
[0162] A sub-filter generation unit, used to generate a sub-Bloom filter corresponding to the computing node according to the hash table corresponding to the computing node; a setting information acquisition unit, used to obtain the setting information of the sub-Bloom filters corresponding to each other computing node; a second filter generation unit, used to generate the Bloom filter according to the setting information and the sub-Bloom filter corresponding to the computing node; wherein the setting information is used to describe the bits of the corresponding sub-Bloom filter that are set.
[0163] Optionally, the device further comprises:
[0164] A filter broadcast module, for respectively calculating the amount of data required to transmit the sub-Bloom filter in a first transmission mode and a second transmission mode after generating a sub-Bloom filter corresponding to the computing node according to a hash table corresponding to the computing node; determining a target transmission mode of the sub-Bloom filter from the first transmission mode and the second transmission mode according to the amount of data required for transmission; broadcasting the sub-Bloom filter corresponding to the computing node to other computing nodes based on the target transmission mode; wherein, broadcasting the sub-Bloom filter corresponding to the computing node based on the first transmission mode includes: broadcasting the sub-Bloom filter corresponding to the computing node to other computing nodes; broadcasting the sub-Bloom filter corresponding to the computing node based on the second transmission mode includes: broadcasting the setting information of the sub-Bloom filter corresponding to the computing node to other computing nodes; wherein the setting information is used to describe the position where the corresponding sub-Bloom filter is set.
[0165] Optionally, the sub-filter generating unit is specifically used to:
[0166] Obtain at least one second hash function; calculate the second hash value of each tuple in the first data table according to the at least one second hash function; construct a sub-Bloom filter corresponding to the computing node according to the first hash value and the second hash value; wherein the first hash value is the hash value in the hash table.
[0167] Optionally, the device further comprises:
[0168] A hash function number determination module is used to determine the hash cost according to the first data table; determine the number of hash functions of the sub-Bloom filter corresponding to the computing node according to the hash cost and the connection cost, so as to obtain a corresponding number of second hash functions according to the number of hash functions.
[0169] Optional, filtering module, specifically used for:
[0170] When the hash distribution of the second data table is different from that of the first data table, the Bloom filter is broadcast to the process corresponding to the second data table; through the process corresponding to the second data table, the tuples in the scanned second data table are filtered according to the Bloom filter to obtain a filtered second data table, and the filtered second data table is sent to the process corresponding to the first data table, so that the process corresponding to the first data table performs a hash connection on the hash table and the filtered second data table.
[0171] The training device of the hash connection model provided in the embodiment of the present application can be used to perform the above Figures 2 to 8 The technical solutions provided in any corresponding embodiment have similar implementation principles and technical effects, which will not be described in detail in this embodiment.
[0172] Fig. 9 A schematic diagram of a computing node structure provided in an embodiment of the present application is shown in FIG. Fig. 9 As shown, the computing nodes provided in this embodiment include:
[0173] At least one processor 910; and a memory 920 communicatively connected to the at least one processor; wherein the memory 920 stores computer-executable instructions; and the at least one processor 910 executes the computer-executable instructions stored in the memory so that the electronic device performs a method as provided in any of the foregoing embodiments.
[0174] Optionally, the memory 920 may be independent or integrated with the processor 910 .
[0175] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the aforementioned embodiments and will not be described in detail here.
[0176] The present application also provides a distributed database, including a plurality of computing nodes. The plurality of computing nodes include Fig. 9 The illustrated embodiment provides a computing node.
[0177] An embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the method provided in any of the aforementioned embodiments can be implemented.
[0178] An embodiment of the present application also provides a computer program product, including a computer program, which implements the method provided in any of the above embodiments when executed by a processor.
[0179] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0180] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform some steps of the method described in each embodiment of the present application.
[0181] It should be understood that the above-mentioned processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The memory may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disk, etc.
[0182] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general or special purpose computer.
[0183] An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a main control device.
[0184] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0185] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0186] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods provided in each embodiment of the present application.
[0187] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0188] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A hash connection method, It is characterized in that include: Based on the first hash function, a hash operation is performed on the first data table to obtain a hash table; Based on the hash table, generate a Bloom filter; Filtering the second data table according to the Bloom filter; Performing hash connection on the hash table and the filtered second data table to obtain and output a hash connection table; Based on the hash table, a Bloom filter is generated, including: Generate a sub-Bloom filter corresponding to the computing node according to the hash table corresponding to the computing node; Calculating the amount of data required to transmit the sub-Bloom filter in a first transmission mode and a second transmission mode respectively; determining a target transmission mode of the sub-Bloom filter from the first transmission mode and the second transmission mode according to the amount of data required for transmission; Based on the target transmission mode, the sub-Bloom filter corresponding to the computing node is broadcast to other computing nodes, and the Bloom filter is generated according to the sub-Bloom filters corresponding to each of the computing nodes.
2. The method according to claim 1, It is characterized in that Based on the target transmission mode, broadcasting the sub-Bloom filter corresponding to the computing node to other computing nodes, and generating the Bloom filter according to the sub-Bloom filters corresponding to each of the computing nodes, including: If the target transmission mode is the first transmission mode, the sub-Bloom filter corresponding to the computing node is broadcast to other computing nodes; the sub-Bloom filters corresponding to each other computing node are obtained; and the Bloom filter is generated according to the sub-Bloom filters corresponding to each of the computing nodes.
3. The method according to claim 1, It is characterized in that Based on the target transmission mode, broadcasting the sub-Bloom filter corresponding to the computing node to other computing nodes, and generating the Bloom filter according to the sub-Bloom filters corresponding to each of the computing nodes, including: If the target transmission mode is the second transmission mode, broadcast the setting information of the sub-Bloom filter corresponding to the computing node to other computing nodes; obtain the setting information of the sub-Bloom filter corresponding to each other computing node; generate the Bloom filter according to the setting information and the sub-Bloom filter corresponding to the computing node; The setting information is used to describe the set bits of the corresponding sub-Bloom filter.
4. The method according to claim 1, It is characterized in that Generating a sub-Bloom filter corresponding to the computing node according to the hash table corresponding to the computing node includes: Obtaining at least one second hash function; Calculate a second hash value for each tuple in the first data table according to at least one second hash function; Constructing a sub-Bloom filter corresponding to the computing node according to the first hash value and the second hash value; The first hash value is a hash value in the hash table.
5. The method according to claim 4, It is characterized in that The method further comprises: Determine a hash cost according to the first data table; Determining a connection cost according to the second data table; The number of hash functions of the sub-Bloom filter corresponding to the computing node is determined according to the hash cost and the connection cost, so as to obtain a corresponding number of second hash functions according to the number of hash functions.
6. The method according to any one of claims 1 to 3, It is characterized in that Filtering the second data table according to the Bloom filter includes: When the hash distribution of the second data table is different from that of the first data table, broadcasting the Bloom filter to the process corresponding to the second data table; Through the process corresponding to the second data table, the tuples in the scanned second data table are filtered according to the Bloom filter to obtain a filtered second data table, and the filtered second data table is sent to the process corresponding to the first data table, so that the process corresponding to the first data table performs a hash connection on the hash table and the filtered second data table.
7. A computing node, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.
8. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.
9. A computer program product, It is characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.
Citation Information
Patent Citations
Filter transmission method, device and system for database table connection
CN113360507A