Hash connection implementation method for relational data and Tag data

By realizing hash connection between relational data and Tag data in the timing engine, the problem of low multi-mode query performance in the prior art is solved, and the computing power and cross-mode query performance of the database are improved.

CN120123347APending Publication Date: 2025-06-10上海沄熹科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510275300.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In IoT-related business scenarios, it is difficult for the prior art to efficiently access relationship data and timing data in one query statement at the same time, resulting in low multimode query performance.

Method used

By introducing relational data management into the timing engine, we support the push of relational data to the timing engine, and realize hash connection between relational data and Tag data of the timing table in the timing engine, and use dynamic hash indexes for matching and filtering.

Benefits of technology

It improves the query analysis performance of the service database in multi-mode scenarios, enhances the computing power of the timing engine, reduces the amount of Tag data, and improves the cross-mode query performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123347A_ABST
    Figure CN120123347A_ABST
Patent Text Reader

Abstract

The invention discloses a Hash connection implementation method for relational data and Tag data, and relates to the technical field of data management and analysis. Comprising the following steps: step 1, based on an open service database, pushing relational data down to a time sequence engine in a data block format by utilizing a relational engine, and receiving the relational data by utilizing a relational batch processing queue through the time sequence engine; 2, materializing the relational data by using a Hash join operator in a time sequence engine, and respectively constructing dynamic Hash indexes of the relational data and the Tag data; 3, traversing the Tag data by using the Hash join operator according to the dynamic Hash indexes constructed by the relational data to search matched data, or a Hash connection operator is utilized to traverse the relational data according to a dynamic Hash index constructed by the Tag data to search the matched data, the successfully matched relational data and Tag data are combined to form a Hash connection result set, the Hash connection result set is packaged into a label line batch format, and the label line batch format is obtained. And 4, outputting the Hash connection result set to the table reader for subsequent data processing, and returning a processing result to the relation engine through the table reader.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a method for implementing hash join of relational data and Tag data, which relates to the technical field of data management and analysis. Background Art

[0002] In the business scenarios related to the Internet of Things, there are usually both time-series data and relational data, and the query operations need to access relational data and time-series data simultaneously in a query statement. In an open-source database, there is a relational engine for processing relational data and a time-series engine for processing time-series data. Due to the large differences between the relational model and the time-series model, the cost of data exchange between the relational model and the time-series model is relatively high. In addition, due to the incomplete computing power of the time-series engine, it is impossible to perform data processing in the time-series engine. Instead, a large amount of data has to be pulled into the relational engine, the time-series data is converted into relational data, and the relational engine processes the data. However, the method of pulling a large amount of data from the time-series engine to the relational engine consumes a lot of time, resulting in low multi-modal query performance.

[0003] To improve the query analysis performance of the open-source database in multi-modal scenarios, we need to enhance the computing power of the open-source database, support pushing relational data down to the time-series engine, and implement hash join in the time-series engine to support the hash join of relational data and Tag data in the time-series table, filter the Tag data as early as possible, reduce the computational amount of upstream operators, and thus improve the multi-modal query analysis performance of the open-source database. Summary of the Invention

[0004] Aiming at the problems of the existing technology, the present invention provides a method for implementing hash join of relational data and Tag data. To improve the query analysis performance of the open-source database in multi-modal scenarios, enhance the computing power of the open-source database, push relational data down to the time-series engine, and implement hash join of relational data and Tag data in the time-series table in the time-series engine, filter the Tag data as early as possible, reduce the computational amount of upstream operators, and thus improve the multi-modal query analysis performance of the open-source database.

[0005] The specific solution proposed by the present invention is as follows:

[0006] The present invention provides a method for implementing hash join of relational data and Tag data, including:

[0007] Step 1: Based on the open-source database, use the relational engine to push relational data to the time-series engine in the form of data blocks, and the time-series engine receives the relational data through the relational batch queue.

[0008] Step 2: Use the hash join operator in the time-series engine to materialize the relational data, and respectively construct dynamic hash indexes for the relational data and the Tag data.

[0009] Step 3: Use the hash join operator to traverse the Tag data through the dynamic hash index constructed according to the relational data to find matching data, or use the hash join operator to traverse the relational data through the dynamic hash index constructed according to the Tag data to find matching data.

[0010] Merge the successfully matched relational data and Tag data to form a hash join result set, and encapsulate the hash join result set into a tag row batch format.

[0011] Step 4: Output the hash join result set to a table reader for subsequent data processing, and return the processing result to the relational engine through the table reader.

[0012] Furthermore, in step 1 of the method for implementing hash join of relational data and Tag data, using the relational engine to push the relational data to the time series engine in the form of data blocks includes:

[0013] Use the relational engine to assemble the relational data into a batch mode in the form of data blocks.

[0014] Use the relational engine to push the relational data to the time series engine in batch mode through the API provided by the time series engine.

[0015] Each time the time series engine receives a batch of relational data, it is put into the relational batch queue.

[0016] When the relational engine finishes pushing the relational data, notify the time series engine through the API provided by the time series engine to change the no_more_data_chunk flag in the relational batch queue to false, so that the hash join operator can know that the relational data push is completed when reading the relational data.

[0017] Furthermore, in step 2 of the method for implementing hash join of relational data and Tag data, constructing the dynamic hash index of relational data and Tag data includes:

[0018] Use the hash join operator to select the access mode, and the access modes include hash relation scan mode hashRelScan, hash tag scan mode hashTagScan, and primary key hash tag scan mode primaryHashTagScan.

[0019] When the hash join operator selects the hash relation scan mode, use the relational data to construct a dynamic hash index.

[0020] When the hash join operator selects the hash tag scan mode, use the Tag data to construct a dynamic hash index.

[0021] When the hash join operator selects the primary key hash tag scan mode, the primary tag of the primary key is used as the hash index.

[0022] Furthermore, the dynamic hash index of the method for implementing hash join of relational data and Tag data includes three data structures, namely:

[0023] When the data on the hash table construction side is Tag data, the data format is the tag row batch format; when it is relational data, the data format is the data block format.

[0024] Linked List, used to link records with the same key value together.

[0025] Index, constructed based on the key value of each record in the data on the hash table construction side.

[0026] Furthermore, the index of the method for implementing hash join of relational data and Tag data includes two parts, namely:

[0027] Bucket instance, used to record the position of the first key value of each bucket in the hash table. The bucket number corresponding to each key value is:

[0028] bucket number = hash value % bucket count

[0029] where hash value is the value calculated by the key value through the hash function, and bucket count is the number of buckets.

[0030] Hash table, used to store all key values, including key value, hash value, nextindex row, and next row indice. Among them, next index row points to the next different key value with the same bucketnumber in the hash table, and next row indice points to the first record with the same key value in the data on the hash table construction side.

[0031] Further, when constructing the dynamic hash index in step 2 of the method for implementing the hash join of relational data and Tag data, the addOrUpdate method is used to insert or update the key value. When the key value does not exist, a record is inserted at the end of the corresponding bucket in the hash table, and the hash value and the key value are written. At the same time, the next rowindice is set to the value, and finally null is returned. When the key value already exists, the value corresponding to the key value, that is, the next rowindice, is directly updated, and finally the next row indice before the update is returned.

[0032] Further, when searching for matching data in step 3 of the method for implementing the hash join of relational data and Tag data, the hash join operator is used to read the probe-side data, and the matching data is searched according to the dynamic hash index:

[0033] The first record matching the hash table construction side data is obtained through the get method provided by the dynamic hash index, including the batch number and the offset;

[0034] According to the batch number and the offset, all matching data records in the hash table construction side data are found through the linked list Linked List.

[0035] The present invention also provides a device for implementing the hash join of relational data and Tag data, including a data push-down module, an index management module, a data matching module, and a subsequent processing module.

[0036] Based on the Kaiwu database, the data push-down module uses the relational engine to push down the relational data in the form of data blocks to the time series engine, and the time series engine receives the relational data through the relational batch queue.

[0037] The index management module uses the hash join operator in the time series engine to materialize the relational data, and constructs dynamic hash indexes for the relational data and the Tag data respectively.

[0038] The data matching module uses the hash join operator to traverse the Tag data according to the dynamic hash index constructed by the relational data to find matching data, or uses the hash join operator to traverse the relational data according to the dynamic hash index constructed by the Tag data to find matching data.

[0039] The successfully matched relational data and Tag data are merged to form a hash join result set, and the hash join result set is encapsulated into the label row batch format.

[0040] The subsequent processing module outputs the hash join result set to the table reader for subsequent data processing, and returns the processing result to the relational engine through the table reader.

[0041] The beneficial effects of the present invention are as follows:

[0042] The present invention introduces relational data management in the time series engine, supports receiving and managing relational data, makes it possible to implement cross-modal computing in the time series engine, provides an optimization selection space for query optimization, and thus makes it possible to select a better query plan.

[0043] The present invention implements the HashTagScan operator in the time series engine, supports hash join processing of relational data and Tag data. On the one hand, it can enhance the computing power of the time series engine, and on the other hand, it enables the time series engine to reduce the amount of Tag data through hash join, thereby improving cross-modal query performance.

[0044] The present invention proposes an efficient dynamic hash index method, which supports both establishing a hash index using Tag data and establishing a hash index using Rel data, enabling the optimization to determine which method to use based on the Tag and relational data volume to obtain the best query performance.

[0045] The present invention defines the output of HashTagScan as the output format TagRowBatch of the TagScan operator, enabling this operator to be used as the input of the tableReader without any modification to the tableReader, achieving more code reuse. Description of the Drawings

[0046] Figure 1 is the overall architecture schematic diagram of the hash join operator of the present invention.

[0047] Figure 2 is the schematic diagram of the data structure of the hash index Hash Index.

[0048] Figure 3 is the schematic diagram of inserting a non-existent key into the data structure of the hash index Hash Index.

[0049] Figure 4 is the schematic diagram of inserting an existing key into the data structure of the hash index Hash Index.

[0050] Figure 5 is the schematic diagram of key value lookup of the hash index Hash Index.

[0051] Figure 6It is a schematic diagram of the data structure that combines Tag data and relationship data and outputs them. Detailed implementation

[0052] Build Side data of the hash table construction side

[0053] Bucket instance

[0054] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the specific embodiments cited shall not be construed as limiting the present invention.

[0055] Embodiment 1

[0056] The present invention provides a method for implementing a hash join of relationship data and Tag data. In combination with Figure 1 , the main engine of the Kaiwu database is the relationship engine, which supports the processing of relationship data. The time series engine is an Agent engine of the Kaiwu database. These two engines can achieve cross-modal data query and analysis through data exchange. The process may include:

[0057] Step 1: Based on the Kaiwu database, use the relationship engine to push the relationship data to the time series engine in the format of data chunks DataChunk, and the time series engine uses the relationship batch queue RelBatchQueue to receive the relationship data.

[0058] Among them, the relationship engine assembles the relationship data into a batch mode in the format of data chunks, that is, assembles them into one batch after another.

[0059] The relationship engine uses the API provided by the time series engine to push one batch after another to the time series engine according to the batch mode.

[0060] Every time the time series engine receives a batch of relationship data, it is put into the relationship batch queue RelBatchQueue.

[0061] When the relationship engine finishes pushing the relationship data, it notifies the time series engine through the API provided by the time series engine to change the no_more_data_chunk flag in the relationship batch queue to false, so that the hash join operator can know that the relationship data push is completed when reading the relationship data. The flag in RelBatchQueue has an initial value of true.

[0062] Step 2: Use the hash join operator HashTagScan in the time series engine to materialize the relationship data and construct dynamic hash indexes for the relationship data and Tag data respectively.

[0063] Among them, HashTagScan supports building Hash Index using relational data and also supports building Hash Index using Tag data. If building Hash Index using relational data, it will traverse the Tag data and access the Hash Index to find matching data; if building Hash Index using Tag data, it will traverse the relational data and access the Hash Index to find matching data. The data used to build the dynamic Hash Index is called the hash table building side data, that is, Build Side data, and the data used to traverse and find matching data is called the probe side data, that is, Probe Side data.

[0064] The Tag table of Kaiwu Database will automatically build a hash index for the primary tag and persist it. When the join column in the query contains all primary tags, the optimizer of Kaiwu Database will choose to use the Hash Index of the primary tag to find matches, thus omitting the step of building the dynamic Hash Index. If the Hash Index of the primary tag cannot be used, it will choose the data with the smaller amount between the relational data and the Tag data to build the dynamic Hash Index.

[0065] HashTagScan provides three access modes:

[0066] When using the hashRelScan mode, it will build a dynamic Hash Index using relational data and traverse the Tag data to find matches.

[0067] When using the hashTagScan mode, it will build a dynamic Hash Index using Tag data and traverse the relational data to find matches.

[0068] When using the primaryHashTagScan mode, it will omit the construction of the dynamic Hash Index and directly use the already materialized Hash Index of the primary tag to traverse the relational data to find matches.

[0069] Step 3: Use the hash join operator HashTagScan to traverse the Tag data to find matching data according to the dynamic hash index built from the relational data, or use the hash join operator to traverse the relational data to find matching data according to the dynamic hash index built from the Tag data.

[0070] Merge the successfully matched relational data and Tag data to form a hash join result set, and encapsulate the hash join result set into the format of a batch of tag rows, TagRowBatch.

[0071] Step 4: Output the hash-joined result set to the table reader tableReader for subsequent data processing, and return the processing result to the relational engine through the table reader. Figure 1 The processing flow shown takes the relational data as the small table and the Tag data as the large table as an example. In the actual query process, it is possible that the relational data is the large table and the Tag data is the small table.

[0072] Embodiment 2

[0073] Based on Embodiment 1, combined with Figure 2 , in Step 2 when constructing the dynamic Hash Index, taking the hashRelScan mode as an example, the process of constructing the dynamic Hash Index is introduced in detail. The related data structures for dynamically creating the Hash Index include the BuildSide data DataChunk / TagRowBatch, the LinkedList, and the Hash Index.

[0074] Among them, for the data block DataChunk passed down from the relational engine, the number of records in each DataChunk may be different. The linked list LinkedList connects the same key values together. Each DataChunk corresponds to a LinkedList with the same number of records. That is to say, if there are n records in the DataChunk, then its corresponding LinkedList also has n records. The Hash Index consists of two parts: the bucket instance and the hash table.

[0075] Each dynamic Hash Index can contain multiple bucket instances. Each bucket instance contains the same number of buckets. Each bucket stores a pointer pointing to the first key stored in this bucket. The key is stored in the hash table introduced below. The key value is calculated into a hash value through the hash function, and then the bucket number of this key is calculated according to the number of buckets as:

[0076] bucket number = hash value % bucket count.

[0077] The hash table is a large array in the dynamic Hash Index, which is used to store the value and hash value of each key, connect the key values belonging to the same bucket through pointers, and point to the first record in the relational data. All records with the same key value can be found through the LinkedList.

[0078] As Figure 2 shown, the key value in the hash table stores the key value, the hash value stores the hash value corresponding to the key value, which is the value calculated by the key value through the hash function, the next index row stores the key of the next same bucket, and the next row indice stores the data record of the BuildSide with the same first key value.

[0079] The next row indice is the value in the Hash Index, defined as <batch number, offset>, which are the number of the batch and the record number of the current record within the batch respectively. Both Tag data and relational data are managed in batches. The batch format of Tag data is TagRowBatch, while the batch format of relational data is DataChunk.

[0080] During the process of constructing the Hash Index using the HashTagScan operator, for each record in the Build Side data, the addOrUpdate method of the Hash Index is called to insert the key value and value of the new record into the HashIndex. The addOrUpdate method requires two parameters. One is the key value, which is the key value corresponding to the current record in the Build Side, and the other is the value, which is the value of the Hash Index, that is, <batch number, offset>, which is the number of the batch where the current record in the Build Side is located and the offset of this record in that batch.

[0081] The addOrUpdate method will calculate the hash value of the key value, insert it into the corresponding bucket, and write its hash value and key value into the corresponding hash table. Among them, as Figure 3As shown, if the corresponding key value does not exist in the HashIndex, a record will be inserted at the end of the hash table, the hash value and the key value will be written, and the nextrow indice will be set to the value. Finally, addOrUpdate returns null. For example, Figure 4 As shown, if the corresponding key value already exists in the Hash Index, the value corresponding to the key value, that is, the next row indice, will be directly updated. Finally, addOrUpdate returns the next row indice before the update.

[0082] After the addOrUpdate in the HashTagScan operator returns in the Hash Index, the value is updated to the corresponding LinkedList of that record.

[0083] Embodiment 3

[0084] Based on Embodiment 2, when searching for matching data in Step 3, the hash join operator is used to read the probe side data, and the matching data is found according to the dynamic hash index:

[0085] In the stage of searching for matching data, the matching operation:

[0086] The HashTagScan operator obtains the first record of the Build Side with the matching key value in the Hash Index by calling the get method, including the batch number and the offset.

[0087] The get method of the Hash Index calculates the hash value according to the key value, finds the corresponding bucket, traverses all the key values in the hash table through the nextindex row, finds the key value that is exactly equal, and returns its corresponding nextrow indice.

[0088] If no matching key value is found, null will be returned, indicating that no matching Build Side record is found.

[0089] For example Figure 5As shown, during the matching process, each record on the Probe Side is traversed to generate its key value, and the get method of the Hash Index is called to return the first record corresponding to the key value in the Build Side. If the first record is not null, all matching Build Side records are traversed through the LinkedList. For each matching Build Side record, merging and output are performed. If the first record is null, it is skipped and the next record is processed.

[0090] In the primaryHashTagScan mode, the existing code provided by TagScan is directly used to scan the Tag data through the Hash Index of the primarytag.

[0091] When performing the merging and output of matching data, for each matching record on the Build Side and the Probe Side, merging and output are performed according to their respective formats and the output column requirements of the HashTagScan, as Figure 6 shown. Data merging is to merge the TagRowBatch and the DataChunk and output them in the format of the TagRowBatch. Since the tableReader only accepts Tag data in the TagRowBatch format, the result of the merged output by the HashTagScan must be in the TagRowBatch format.

[0092] For the case where the data source of the output column is the TagRowBatch, only a value needs to be added to the corresponding output column of the output TagRowBatch and its pointer to the value pointed to by the corresponding column of the original TagRowBatch is set, without data copying and conversion.

[0093] For the case where the data source of the output column is the DataChunk, the DataChunk uses a large memory to store all records. First, the storage address of the corresponding column data needs to be found in the DataChunk, and then a value is added to the corresponding output column of the output TagRowBatch and it is set to the storage address of the corresponding column data in the DataChunk.

[0094] Since the TagRowBatch uses addresses and does not copy data, all TagRowBatches and DataChunks, whether on the Build Side or the Probe Side, are retained until the HashTagScan operator finishes processing and closes.

[0095] Example 4

[0096] The present invention also provides an apparatus for implementing hash join of relational data and Tag data, including a data push-down module, an index management module, a data matching module, and a subsequent processing module.

[0097] Based on the Kaiwu database, the data push-down module uses the relational engine to push down relational data in the form of data blocks to the time-series engine, and the time-series engine receives the relational data through the relational batch queue.

[0098] The index management module materializes the relational data by using the hash join operator in the time-series engine, and constructs dynamic hash indexes for relational data and Tag data respectively.

[0099] The data matching module uses the hash join operator to traverse the Tag data according to the dynamic hash index constructed by the relational data to find matching data, or uses the hash join operator to traverse the relational data according to the dynamic hash index constructed by the Tag data to find matching data.

[0100] The successfully matched relational data and Tag data are merged to form a hash join result set, and the hash join result set is encapsulated into the batch format of tag rows.

[0101] The subsequent processing module outputs the hash join result set to the table reader for subsequent data processing, and returns the processing result to the relational engine through the table reader.

[0102] Regarding the information interaction and execution process among the above-mentioned modules in the apparatus, since they are based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention, and will not be elaborated here.

[0103] Similarly, the apparatus of the present invention introduces relational data management in the time-series engine, supports receiving and managing relational data, makes it possible to implement cross-modal calculation in the time-series engine, provides an optimization selection space for query optimization, and thus may select a better query plan.

[0104] The present invention implements the HashTagScan operator in the time-series engine, which supports the hash join processing of relational data and Tag data. On the one hand, it can enhance the computing power of the time-series engine, and on the other hand, it enables the time-series engine to reduce the amount of Tag data through hash join, thereby improving the cross-modal query performance.

[0105] The present invention proposes an efficient dynamic hash index method, which supports both establishing a hash index using Tag data and establishing a hash index using Rel data, enabling the optimization to decide which method to use according to the Tag and relational data volume to obtain the best query performance.

[0106] The present invention defines the output of HashTagScan as the output format TagRowBatch of the TagScan operator, enabling this operator to be used as the input of the tableReader without any modification to the tableReader, thus achieving more code reuse.

[0107] It should be noted that not all steps and modules in the above-mentioned processes and device structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as required. The system structure described in the above embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities separately, or some components in multiple independent devices may be jointly implemented.

[0108] The above-described embodiments are only preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention. The protection scope of the present invention is subject to the claims.

Claims

1. A method for implementing hash connection between relational data and tag data, characterized in that include: Step 1: Based on the open service database, the relational engine is used to push the relational data in the data block format to the time series engine, and the relational batch processing queue is used by the time series engine to receive the relational data; Step 2: Use the hash join operator in the time series engine to materialize the relational data and build dynamic hash indexes for the relational data and tag data respectively. Step 3: Use the dynamic hash index constructed by the hash join operator based on the relational data to traverse the Tag data to find matching data, or use the dynamic hash index constructed by the hash join operator based on the Tag data to traverse the relational data to find matching data. Merge the successfully matched relationship data and Tag data to form a hash join result set, and encapsulate the hash join result set into a tag row batch format. Step 4: Output the hash join result set to the table reader for subsequent data processing, and return the processing result to the relational engine through the table reader.

2. The method for implementing hash connection of relational data and tag data according to claim 1, characterized in that In step 1, the relational engine is used to push the relational data in the form of data blocks to the time series engine, including: Use the relational engine to assemble relational data into batch mode in data block format. The relational engine uses the API provided by the timing engine to push relational data to the timing engine in batch mode. Each time a batch of relational data is received by the timing engine, it is placed in the relational batch processing queue. When the relational engine completes pushing down relational data, it notifies the timing engine through the API provided by the timing engine to change the no_more_data_chunk flag in the relational batch processing queue to false, so that the hash join operator can be informed of the completion of pushing down relational data when reading relational data.

3. The method for implementing hash connection of relational data and tag data according to claim 1, characterized in that In step 2, dynamic hash indexes of relational data and tag data are constructed, including: Use the hash join operator to select the access mode, which includes the hash relation scan mode hashRelScan, the hash tag scan mode hashTagScan, and the primary key hash tag scan mode primaryHashTagScan. When the hash join operator selects the hash relation scan mode, the relational data is used to build a dynamic hash index. When the hash join operator selects the hash tag scan mode, the tag data is used to build a dynamic hash index. When the hash join operator selects the primary key hash tag scan mode, the primary key primary tag is used as the hash index.

4. The method for implementing hash connection of relational data and tag data according to claim 1, characterized in that The dynamic hash index contains three data structures: The data on the hash table construction side is in the tag row batch format when it is tag data, and in the data block format when it is relational data; Linked List, used to connect records with the same key value; The index is constructed based on the key value of each record in the data on the hash table construction side.

5. The method for implementing hash connection of relational data and tag data according to claim 4, characterized in that The index consists of two parts: Bucket instance is used to record the position of the first key value of each bucket in the hash table. The bucket number corresponding to each key value is: bucket number=hash value% bucket count The hash value is the value calculated by the hash function of the key value, and the bucket count is the number of buckets; Hash table is used to store all key values, including key value, hash value, next index row and next row indice, where next index row points to the next different key value with the same bucket number in the hash table, and next row indice points to the first record with the same key value in the data on the hash table construction side.

6. The method for implementing hash connection of relational data and tag data according to claim 5, characterized in that When building a dynamic hash index in step 2, use the addOrUpdate method to insert or update the key value. When the key value does not exist, a record will be inserted at the end of the corresponding bucket in the hash table, and the hash value and key value will be written. At the same time, the next row index will be set to the value value, and null will be returned. When the key value already exists, the value value corresponding to the key value, that is, the next row index, will be directly updated, and the next row index before the update will be returned.

7. The method for implementing hash connection of relational data and tag data according to claim 4, characterized in that When searching for matching data in step 3, the detection side data is read using the hash join operator, and matching data is searched based on the dynamic hash index: Use the get method provided by the dynamic hash index to obtain the first record of data matching the hash table construction side, including the batch number and offset. According to the batch number and offset, all matching data records in the hash table construction side data are found through the linked list.

8. A device for implementing hash connection between relational data and tag data, characterized in that It includes data push module, index management module, data matching module and subsequent processing module. The data push module is based on the open service database and uses the relational engine to push the relational data in the form of data blocks to the time series engine. The time series engine uses the relational batch queue to receive the relational data. The index management module uses the hash join operator in the time series engine to materialize the relational data and construct dynamic hash indexes for the relational data and tag data respectively. The data matching module uses the dynamic hash index constructed by the hash join operator based on the relational data to traverse the Tag data to find matching data, or uses the dynamic hash index constructed by the hash join operator based on the Tag data to traverse the relational data to find matching data. Merge the successfully matched relationship data and Tag data to form a hash join result set, and encapsulate the hash join result set into a tag row batch format. The subsequent processing module outputs the hash connection result set to the table reader for subsequent data processing, and returns the processing result to the relational engine through the table reader.