Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

22 results about "Hash join" patented technology

The hash join is an example of a join algorithm and is used in the implementation of a relational database management system. Hash joins are typically more efficient than nested loops joins, except when the probe side of the join is very small. However, hash joins can only be used to compute equijoins.

Probe-side annotated join decision making

A join decision manager (JDM) generates a data processing pipeline. The data processing pipeline includes at least one join operation associated with build-side row data and probe-side row data. The JDM determines the maximum cardinality associated with the probe-side row data. The JDM determines size of the build-side row data at a decision node of the at least one join operation. The JDM configures execution of the at least one join operation as one of a broadcast join or a hash-hash join based on the size of the build-side row data and the maximum cardinality.
Owner:SNOWFLAKE INC

Storage-side filtering for performing hash joins

Storage-side filtering may be implemented for performing hash join operations with respect to data stored in distributed data storage. When a query to a database stored in a distributed data store is received, a hash join operation may be identified for performing the query. As part of performing the hash join operation, a query engine may cause storage nodes in the distributed data storing data for the database to filter data before sending the data to the query engine according to a join predicate for the hash join operation. A result of the query may then be provided to a user, using the filtered data provided from the storage nodes.
Owner:AMAZON TECH INC

A method for accessing hash tables in OpenGauss hash joins

The present invention relates to an optimized access method and access system for a hash table in an OpenGauss hash connection. The method comprises the steps of constructing a hash table, completing the filling of a hash bucket array and a hash collision array; traversing the hash bucket array to detect whether a hash collision has occurred in the data stored at each location; and filling prompt information with the free space of the location information stored in the hash bucket array and the hash collision array based on the detection result. The method combines the advantages of the chain address method and the open addressing method, optimizes the implementation of the hash table from multiple perspectives, and distinguishes whether each conflicting data uses the chain address method or the open addressing method by filling in prompt information, thereby greatly improving the access performance of the hash table. In addition, the prompt information provided by the method can be used to determine whether a hash collision has occurred in the data without accessing the HashNext array, thereby improving the memory access and hash connection performance of the hash table.
Owner:广州海量数据库技术有限公司

Hash join processing method, system and device, and storage medium and program product

PCT designated stageWO2026021114A1Resource allocationSpecial data processing applicationsDatasheetHash join
Provided in the embodiments of the present disclosure are a hash join processing method, system and device, and a storage medium and a program product. The method comprises: scanning a plurality of rows of data in a first data table within a target scanning range corresponding to a target worker thread, and on the basis of a first hash join mode, writing the plurality of rows of data into a first hash table in a memory; if it is determined that an available memory capacity corresponding to a hash join task is lower than a set threshold, sending the scanned plurality of rows of data to a data partitioning management thread, and switching to use a second hash join mode; and receiving first target row data distributed by the data partitioning management thread, and on the basis of the second hash join mode, writing the first target row data into a second hash table corresponding to the target worker thread in the memory, so as to perform persistence processing of the data in the second hash table. The present disclosure can flexibly switch between different hash join modes on the basis of the memory occupancy of a data table for hash table creation.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD +1

Mechanisms for reducing probe-side spill in hash joins

In an embodiment, a computer system performs a hash join. During a build phase, the computer system constructs a hash table in memory based on rows of a first table. The constructing may result in build batches of rows, including one build batch stored in memory and multiple build batches stored in a storage. The computer system determines whether any of the multiple build batches is skewed according to a data skew condition. In response to determining that there is at least one build batch that is skewed, the computer system loads one or more of the multiple build batches into memory such that there are at least two build batches stored in memory. During a probe phase, the computer system identifies, based on the at least two build batches stored in memory, rows of a second table to join with the rows of the first table.
Owner:SALESFORCE INC

Method and apparatus for detecting scheduling deadlock in data query, and device

Discloses are a method for detecting a scheduling deadlock in data query. A query plan tree generated by a query statement includes a common temporary table production operator node, a common temporary table consumption operator node, and a hash join operator node. At least one right table chain in the tree is determined. Operator nodes in the right table chain are traversed from bottom to top starting from the leaf node in the right table chain, and the other operator nodes in the tree are traversed, to analyze an execution order relationship between the operator nodes traversed in the tree. When it is found based on the execution order relationship that there are two identical common temporary table consumption operator nodes in the tree that have scheduling dependency, and their upstream operators are the same common temporary table production operator, it is determined that there is a scheduling deadlock.
Owner:BEIJING VOLCANO ENGINE TECH CO LTD

A method, device and equipment for detecting deadlocks in data query scheduling

The present application discloses a method, apparatus, and device for detecting data query scheduling deadlocks. The query plan tree generated from the query statement includes a common temporary table production operator node, a common temporary table consumption operator node, and a hash join operator node. Determine at least one right table chain in the tree. The right table chain includes leaf nodes. When the parent node of the operator node in the right table chain is a hash join operator node, the operator node is the right child node of its parent node. The right table chain includes the parent node, and the right table chain also includes the root node of the non-hash join operator node. Traverse the operator nodes in the right table chain from the leaf nodes of the right table chain bottom-up, and traverse other operator nodes in the tree, and analyze the execution order relationship between the operator nodes traversed in the tree. When two identical common temporary table consumption operator nodes with scheduling dependencies are found in the tree based on the execution order relationship, and their upstream operator is the same common temporary table production operator, it is determined that there is a scheduling deadlock.
Owner:BEIJING VOLCANO ENGINE TECH CO LTD

Build-side skew handling for hash-partitioning hash joins in distributed database query execution

ActiveUS12675478B2Hash joinData set
Provided herein are systems, methods, and computer-storage media for managing data skew in hash join operations. A skew manager partitions build-side row data into multiple sets corresponding to hash-join-build (HJB) instances based on hash values. The skew manager detects skew in a build-side row set associated with a first HJB instance by analyzing the number of rows. Upon detecting skew, the skew manager redirects data rows to at least a second HJB instance. The method involves configuring skew caches, generating histograms, and detecting frequent hash values to identify skew. It also includes communicating skew notifications, broadcasting probe-side row data, and adjusting partitioning of probe-side data. The disclosed techniques further include buffering build-side row sets in streams and performing join operations based on these streams, enhancing efficiency in distributed computing environments.
Owner:SNOWFLAKE INC

Efficient Opcode-Driven Pipelined Execution Of Multi-Level Hash Joins

An efficient join processing technique is provided that improves multi-level hash join performance by decomposing join operations into a set of opcodes that describe any complex hash join, pipelining execution of these opcodes across join levels, and sharing join operand metadata across opcodes. Complex multi-level joins are easier to describe and execute when decomposed into opcodes. The join technique decomposes multi-level join operations into a minimal set of opcodes such that the join work at each node of the multi-level join can be fully described as an execution of a sequence of opcodes. Operand metadata is shared across the opcodes of all join levels that reference the operand, thereby obviating the need to copy or transmit rows between the join nodes.
Owner:ORACLE INT CORP

A bloom filter configuration method for hash join

The application provides a Bloom filter configuration method for a hash connection, comprising: sampling initial data, analyzing the initial data, dynamically analyzing a variety of change trends, fitting a change trend model, constructing a Bloom filter, evaluating a Bloom filter size, and adjusting a Bloom filter configuration. The application combines probability statistics and mathematical modeling to dynamically analyze part of data sets and estimate code value inclination, uses a change slope and a coefficient of variation to judge a stable trend of data, uses Hausdorff Distance to select an optimal model to estimate a Bloom filter, finally proposes different configuration Bloom filter strategies, efficiently configures a Bloom filter size, fully utilizes memory resources, and improves overall query efficiency of the hash connection.
Owner:SOUTH CHINA UNIV OF TECH

Method and apparatus for dynamic filtering of distributed database that performs hash join operations on data distributed and stored on multiple servers

ActiveKR103022328B1Execution planHash join
A distributed database dynamic filtering method according to various embodiments of the present invention may include the steps of: identifying a hash join operation node by traversing a query execution plan tree; identifying a filter target node by searching the right subtree of the hash join operation node; generating a filter identifier corresponding to the filter target node; recording the filter identifier in the left input table of the hash join operation node; distributing the execution plan tree to a plurality of servers; generating or updating a hash value-based dynamic filter based on records in the left input table; applying the generated dynamic filter to the filter target node by progressively propagating it downward along the right subtree of the hash join operation; configuring a global filter by mutually transmitting and merging filters generated at each server when the filter target node includes a communication node that performs inter-server communication; and removing join target records by applying the dynamic filter prior to performing a physical data scan on the right input table of the hash join operation node.
Owner:SEASPHERE CO LTD

A hash join parallel query method and system based on the openGauss database

This invention relates to the field of database query technology, providing a method and system for parallel hash join queries based on the OpenGauss database. The method includes: determining whether a shared hash table can store all inner table data; creating initialization conditions for the construction of the shared hash table and the joining of inner and outer tuples; when the shared hash table can store all inner table data, multi-threaded parallel scanning of the inner tables is performed, with a barrier structure controlling the construction of the shared hash table; after the shared hash table construction is completed, parallel scanning of the outer table is performed, and tuples from the inner and outer tables are joined; when the shared hash table cannot store all inner table data, the shared hash table is constructed in batches; after the current batch of shared hash table construction is completed, parallel scanning of the outer table is performed; after joining the current batch of tuples from the inner and outer tables, the next batch of shared hash tables is constructed, and the next batch of tuple joins from the inner and outer tables are executed. This method and system can improve CPU utilization and increase the efficiency of hash join queries.
Owner:广州海量数据库技术有限公司

Hash connection method, device, equipment and medium

The present application discloses a hash join method, apparatus, device, and medium, relating to the field of hardware acceleration of database query operations, including: obtaining first data of a first data tuple to be joined and second data of a second data tuple to be joined, respectively calculating the first data and the second data using a cuckoo algorithm to obtain a hash result of the first data and a hash result of the second data; determining a first target hash table corresponding to the hash result of the first data and a second target hash table corresponding to the hash result of the second data; dividing the first data and the second data into a plurality of groups of first sub-data and second sub-data, respectively storing the first sub-data and the second sub-data in the first target hash table and the second target hash table; reading the first sub-data and the second sub-data from the first target hash table and the second target hash table, and performing a merge join on the first sub-data and the second sub-data that meet an equal value condition to obtain the joined data. The present application also discloses a hash join method, apparatus, device, and medium, relating to the field of hardware acceleration of database query operations, including: obtaining first data of a first data tuple to be joined and second data of a second data tuple to be joined, respectively using a cuckoo algorithm to calculate the first data and the second data to obtain a hash result of the first data and the second data; determining a first target hash table corresponding to the hash result of the first data and a second target hash table corresponding to the hash result of the second data; dividing the first data and the second data into a plurality of groups of first sub-data and second sub-data, respectively storing the first sub-data and the second sub-data in the first target hash table and the second target hash table; reading the first sub-data and the second sub-data from the first target hash table and the second target hash table, and performing a merge join on the first sub-data and the second sub-data that meet an equal value condition to obtain the joined data. The present application also discloses a hash join method, apparatus, and medium, relating to the field of hardware acceleration of database query operations, including obtaining first data of a first data tuple to be joined and second data to obtain a hash result of the first data and the second data to
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Method for constructing large-scale data set based on multi-source data aggregation

The invention discloses a method for constructing a large-scale data set based on multi-source data aggregation. By providing an efficient multi-source data aggregation strategy, a set of complete data integration and cleaning method is constructed, and seamless joint and standardized processing of heterogeneous data are realized. And an automatic data cleaning tool and a rule engine are adopted, so that noise data can be accurately removed, missing values can be filled, duplicate removal can be realized, a data format can be standardized, and data consistency and integrity can be ensured. Besides, in combination with an efficient data aggregation method, such as hash connection, partition statistics and attribute fusion, cross-data-source intelligent integration is realized, and the accuracy and calculation efficiency of data processing are improved. The method is particularly suitable for large-scale data processing requirements, the quality and availability of a data set can be remarkably improved, reliable data support is provided for multiple application scenes such as big data analysis, machine learning model training and business intelligence, and data driving decision and intelligent analysis are assisted.
Owner:BEIJING BAIJU YIXING TECH CO LTD

HASH-join broadcast decision making in database systems

A system includes at least one hardware processor and memory storing instructions. The processor generates a query plan for a received query. The query plan includes multiple hash-join-build and hash-join-probe operations. A primary decision node is configured in the query plan. The primary decision node receives build-side data information from the hash-join-build operations. For each hash-join-build operation, a memory amount for performing a broadcast is determined. A subset of hash-join-build operations is selected for broadcast join distribution by comparing the memory amount to a broadcast memory threshold. The system selects a broadcast join distribution for the subset and a hash-hash join distribution for the remaining hash-join-build operations. The query plan is executed using the broadcast join distribution for the selected subset and the hash-hash join distribution for the remaining operations. This approach optimizes memory usage and join distribution during query execution.
Owner:SNOWFLAKE INC

Data query method and device, computer equipment and computer readable storage medium

The invention is suitable for the technical field of databases, and relates to a data query method and device, equipment and a medium. The method comprises the following steps: converting a received structured query statement into a physical execution plan, and if the physical execution plan comprises a Hash join operator, allocating a corresponding memory quota to the Hash join operator; if the first memory demand of the Hash table constructed by the Hash join operator is greater than the memory quota, executing the Hash join operator by adopting a memory elimination strategy; the memory elimination strategy comprises the steps that an execution quota is divided from memory quotas, the execution quota comprises a first sub-quota used for storing a hash table structure corresponding to a hash table and connection key data and a second sub-quota used for storing a data record of a small table related to the hash table, and the first sub-quota is larger than the second sub-quota; and respectively performing longest unused elimination management on the first sub-quota and the second sub-quota in the execution process of the Hash join operator. According to the invention, the execution efficiency of the Hash connection can be improved under the condition that the memory is limited.
Owner:SHENZHEN INST OF COMPUTING SCI

Hash join hardware acceleration method, device and equipment for large data table and medium

The application discloses a big data table hash connection hardware acceleration method and device, equipment and medium, and relates to the database heterogeneous acceleration field. The method is applied to an acceleration board card and comprises the following steps: receiving a connection operation type instruction; applying a first memory space of a first to-be-connected data table and a second memory space of a second to-be-connected data table respectively by using a connection table size instruction, so as to store each first shard of the first to-be-connected data table into the first memory space and store each second shard of the second to-be-connected data table into the second memory space; and performing hash connection calculation on each first shard in the first memory space and each second shard in the second memory space according to the connection operation type instruction and a preset parallel processing row number, so as to obtain corresponding connection table row data. Through the above scheme, the hash connection hardware acceleration effect can be effectively improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

A method for distributed database to perform hash join

ActiveCN115687357BImprove execution performanceExecution time balancingDatabase distribution/replicationSpecial data processing applicationsTable (database)Datasheet
The application discloses a method for performing hash connection of a distributed database, and relates to the technical field of distributed databases, which comprises the following steps: setting a relative inclination rate, calculating the product of the relative inclination rate and table data volume to obtain an inclination threshold; obtaining a heat map of table statistical information, screening elements in the table which exceed the inclination threshold to obtain an inclination value; expanding the table statistical information by using the inclination value; executing an SQL statement to obtain two input data tables, generating a hash connection physical plan by using the new table statistical information, performing hash distribution, average distribution or mirror distribution on tuples in the input data tables according to the plan, and after each node receives the tuples of the input data tables, establishing a hash table by using input data table data with small data volume, performing detection by using input data table data with large data volume, and finally performing a set union on the hash connection results of each node, so that the set union is the final hash connection result. The application can realize balanced distribution of tasks and shorten the total SQL execution time.
Owner:上海沄熹科技有限公司

Data page index-based hash join query method, apparatus and system, and medium

A data page index-based Hash Join query method, apparatus and system, and a medium. The method comprises: determining a build table and a probe table for a join query, as well as a join key field (S101); building a corresponding Hash table on the basis of data of the join key field in the build table, and determining whether a pre-created data page index of the join key field is present in the probe table, the data page index being an index structure using at least one data page as a unit (S102); if a pre-created data page index of the join key field is present in the probe table, associating the data page index of the join key field with the Hash table, and performing data page filtering for the probe table on the basis of an association result to generate a list of data pages to be accessed (S103); and reading corresponding data in the probe table on the basis of the list of data pages to be accessed, and performing a Hash Join with the Hash table to generate a corresponding data query result (S104).
Owner:JINZHUAN INFORMATION TECHNOLOGY CO LTD

A connection query method and system for Elasticsearch

The present invention discloses a connection query method and system for Elasticsearch. To overcome the problem that Elasticsearch cannot express and implement the requirements of SQL-style connection queries, the present invention includes the following steps: The user sends a connection query statement that conforms to the syntax to the query statement parsing layer through a RESTful client; the query statement parsing layer parses the received connection query statement, extracts information such as the target index and additional parameters, and sends this information to the Elasticsearch data processing framework; the Elasticsearch data processing framework configures parameters according to the received information, and obtains the data of the target index from Elasticsearch; uses the hash join algorithm to process the original data, and writes the processing result into Elasticsearch; after the writing is completed, the user directly queries the connection result from Elasticsearch through the RESTful client, and performs data analysis operations supported by Elasticsearch on the connection result. Solve the problem of efficient retrieval and analysis of multi-source data in scenarios such as security analysis.
Owner:SICHUAN UNIV

Method, device and storage medium for processing data table

There are provided a method, device, and storage medium for processing a data table. The method includes: performing an equivalent join on a left table and a right table in two data tables that are to be joined by a hash join, and acquiring first associated data of an association between the left table and the right table; filtering the first associated data according to a predetermined non-equivalent filtering condition, and determining second associated data satisfying the predetermined non-equivalent filtering condition in the first associated data; identifying, in a predetermined data structure according to a target join type of the two data tables; and processing the left table and / or the right table according to the target join type of the two data tables and the predetermined data structure to generate a target join table.
Owner:BEIJING VOLCANO ENGINE TECH CO LTD

A hash connection processing method, storage medium and device

The present invention relates to database technology, and in particular to a hash join processing method, storage medium, and device. The method comprises: calculating the cost of constructing a hash inner table using a first table, wherein the first table is any one of two tables to be joined in a pre-hash join; determining whether rehashing is required when constructing the hash inner table using the first table; if so, calculating an impact factor of rehashing; and correcting the cost of constructing the hash inner table using the first table based on the impact factor. The method of the present invention takes the impact of rehashing into account when constructing the cost of the hash inner table, thereby optimizing the inner and outer table selection algorithm for the hash join, avoiding the significant time consumption caused by rehashing, effectively shortening the overall operation time of the hash join, and thereby effectively improving the overall performance of the database.
Owner:CETC JINCANG (BEIJING) TECH CO LTD