Data management method and device based on time sequence database Hash index and medium
By using the Tag Table and Hash Index architecture in the timing database, the Primary Tag field is used to create a Hash timing index, which solves the problem of low performance of the timing database in high concurrent large number of writes scenarios, and realizes fast query and write, improving the concurrency performance of the system.
Patent Information
- Application Number
- CN202510139254.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-16
AI Technical Summary
How to improve the performance of timing databases in high concurrent large-scale write scenarios, the existing Hash query method may lead to O(n) time complexity in the worst case.
Using an architecture including Tag Table and Hash Index, in the KaiwuDB database timing engine, the Primary Tag field is used to create a Hash timing index, and quickly locate whether the Primary Tag has a record through the Hash algorithm to decide whether to write the Primary Tag record.
Improve the performance of query and write, ensure the performance of time-series database in high concurrent large-scale write scenarios, reduce the impact of hash conflicts, and improve the concurrency performance of hash indexes.
Smart Images

Figure CN120011367A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of time series databases, and in particular to a method, device and medium for managing hash index data based on a time series database. Background Art
[0002] Time series databases are mainly used to process data with time tags (data that changes in time order, i.e., time serialization), which are usually collected and generated by various real-time monitoring, inspection, and analysis equipment. Time series databases have significant advantages in processing time series data and are widely used in many fields such as the Internet of Things, finance, and energy management. However, it also has certain challenges in terms of learning costs, hardware resource requirements, user freedom, and cost. Therefore, when selecting and using a time series database, it is necessary to weigh and consider the specific application scenarios and requirements.
[0003] The time complexity of the hash-based query method is O(1) in an ideal situation, but in practice it is affected by hash conflicts and conflict resolution methods. If the hash function is designed properly and the size of the hash table is moderate, the hash table can provide query operations with close to constant time complexity. However, in the worst case, due to the existence of hash conflicts, the time complexity of the search may rise to O(n).
[0004] Therefore, how to improve the performance of time series databases in high-concurrency and large-scale write scenarios is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The technical task of the present invention is to provide a method, device and medium for managing hash index data of a time series database to solve the problem of how to improve the performance of high-concurrency and large-scale write scenarios of a time series database.
[0006] The technical task of the present invention is achieved in the following way: a method for managing data based on a hash index of a time series database. The method adopts an architecture including a tag table and a hash index. In the KaiwuDB database time series engine, the primary tag field is used to create a hash time series index. In a large number of write scenarios, the hash algorithm is used to quickly locate whether a record of the primary tag already exists according to the value of the primary tag field, and then decide whether to write the primary tag record. Specifically:
[0007] For an existing Primary Tag, it means that the device already exists, and the existing device ID is used to write metrics;
[0008] If the Primary Tag does not exist, it means that the device does not exist. Use the new device ID to write metrics.
[0009] As a preferred embodiment, the Tag Table structure is used to store all Tag data, using a memory mapping architecture to directly access records of corresponding row numbers based on row IDs.
[0010] Preferably, the Hash Index structure uses the Primary Tag as the Key, and the Value corresponding to each Key points to the row ID of the Tag Table. The corresponding Key and the row ID of the TagTable corresponding to the Key can be quickly queried based on the Primary Tag, and then the record of the Tag Table can be accessed through the row ID.
[0011] Preferably, the Hash Index includes a memory structure and a file structure;
[0012] The memory structure is Bucket, which stores the address of the record pointing to the Index file. Multiple Buckets are used as a Segment logical object.
[0013] The file structure is used to store each Hash Index record, defining the stored Key, Hash Code, and each Hash Index record pointer.
[0014] Preferably, the structure of the Segment logical object is defined as follows:
[0015] class HashSegment{
[0016] private:
[0017] HashIndexRowlD*m mem bucket_;
[0018] TagHashBucketRWLock*m_bucket_rwlock_;
[0019] size tm bucket_count_;
[0020] public:
[0021] explicit HashSegment(size_t bucket_count=8);
[0022] };
[0023] Among them, HashIndexRowlD*m mem bucket_ represents the pointer to the hash bucket memory;
[0024] TagHashBucketRWLock*m_bucket_rwlock_ represents the read-write lock that controls access to the hash bucket;
[0025] size tm bucket_count_ indicates the number of hash buckets.
[0026] Preferably, the file structure is defined as follows:
[0027] struct HashIndexData{
[0028] HashCodehash_val;
[0029] TagTableRowlD tbl_row;
[0030] HashIndexRowlDnext_row;
[0031] };
[0032] Among them, HashCodehash_val represents the hash value;
[0033] TagTableRowlD tbl_row represents a pointer or identifier pointing to a table row;
[0034] HashIndexRowlDnext_row represents a pointer or identifier for the next hash index row.
[0035] As a preferred option, Hash Index defines a read and write interface, and improves the usability of Hash Index by encapsulating the write and query operations of Hash Index;
[0036] The specific steps of writing Hash Index are as follows:
[0037] (1) Calculate the corresponding Hash Code based on the Key;
[0038] (2) Determine whether the data in the hash table needs to be redistributed and reorganized (rehash):
[0039] ①If yes, proceed to step (3);
[0040] ②If not, jump to step (4);
[0041] (3) reallocate and reorganize data in the hash table;
[0042] (4) Calculate the corresponding Segment and Bucket based on the Hash Code;
[0043] (5) Write the Index Data record and update the Bucket value to the address of the latest Index Data.
[0044] Preferably, the Hash Index performs the query as follows:
[0045] (1) Calculate the corresponding Hash Code based on the Key, and obtain the corresponding Segment and Bucket based on the Hash Code;
[0046] (2) Get the address of the Index Data pointed to by the Bucket;
[0047] (3) Determine whether the address of Index Data is equal to the Key value:
[0048] ①If it is equal to the Key value, it will be returned directly;
[0049] ②If it is not equal to the Key value, continue to the next address of Index Data until the linked list access of Index Data ends.
[0050] An electronic device comprising: a memory and at least one processor;
[0051] Wherein, the memory stores a computer program;
[0052] The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the above-mentioned method for managing data based on hash index of time series database.
[0053] A computer-readable storage medium stores a computer program, and the computer program can be executed by a processor to implement the above-mentioned method for managing data based on hash index of a time series database.
[0054] The method, device and medium for managing data based on hash index of time series database of the present invention have the following advantages:
[0055] (1) The present invention can quickly query whether each Primary Tag record already exists through the Hash time series index, thereby improving the query and write performance, and further improving the performance of the high-concurrency and large-scale write scenario of the time series;
[0056] (ii) The Hash Index of the present invention uses the Hash algorithm to quickly locate the record corresponding to the Primary Tag under the condition of O(1) time complexity, and groups the Buckets to increase the lock concurrency granularity, thereby improving the concurrency performance of the Hash Index. At the same time, a unified read and write interface is abstracted, which is convenient for direct calling without having to worry about the internal calculation and implementation of the Hash Index.
[0057] (III) The content structure of the Hash Index of the present invention regards multiple Buckets as a Segment logical object, and a Segment has an independent lock to improve the lock granularity and performance;
[0058] (iv) The file structure of the Hash Index of the present invention improves the conflict handling efficiency of the Hash algorithm by adding and defining a Hash Index record pointer, thereby improving the query efficiency;
[0059] (V) The Hash Index of the present invention defines a read and write interface, and improves the usability of the Hash Index by encapsulating the write and query operations of the Hash Index. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The present invention is further described below in conjunction with the accompanying drawings.
[0061] Attached Figure 1 It is a structural diagram of the Hash index data management method based on the time series database;
[0062] Attached Figure 2 Flow chart for writing to Hash Index;
[0063] Attached Figure 3 A flowchart of executing a query for a Hash Index. DETAILED DESCRIPTION
[0064] The method, device and medium for managing data based on hash index of a time series database of the present invention are described in detail below with reference to the accompanying drawings and specific embodiments of the specification.
[0065] Embodiment 1:
[0066] In this embodiment, primary Tag: Tag is a label field in time series data, and Primary Tag represents a Tag field that can identify a unique device. Multiple Tag fields can be defined as a Primary Tag, and the Primary Tag is used to identify a unique device in the KaiwuDB database.
[0067] In this embodiment, Hash query: KaiwuDB time series engine uses column storage and append to record Tag data when storing Tag data. When it is necessary to check whether a certain device already has a Tag record, the Hash query algorithm can be used to query the result within O(1) time complexity, reducing the scan query Tag record time and CPU and IO consumption, thereby improving query and write efficiency.
[0068] In this embodiment, Bucket: is the value of Hash Code corresponding to different Keys divided by the Hash algorithm. The buckets corresponding to different Keys may be the same.
[0069] In this embodiment, Segment: multiple buckets are used as a Segment logical object, each Segment object has independent lock management, the concurrency granularity of the Hash is increased, and the concurrency performance is improved.
[0070] This embodiment provides a method for managing data based on a hash index of a time series database. The method adopts an architecture including a tag table and a hash index. In the KaiwuDB database time series engine, the primary tag field is used to create a hash time series index. In a large number of write scenarios, a hash algorithm is used to quickly locate whether a record of the primary tag already exists according to the value of the primary tag field, and then decide whether to write the primary tag record. Specifically:
[0071] For an existing Primary Tag, it means that the device already exists, and the existing device ID is used to write metrics;
[0072] If the Primary Tag does not exist, it means that the device does not exist. Use the new device ID to write metrics.
[0073] As attached Figure 1 As shown, the Tag Table structure in this embodiment is used to store all Tag data, adopts a memory mapping architecture, and directly accesses the record of the corresponding row number according to the row ID.
[0074] The Hash Index structure in this embodiment uses the Primary Tag as the Key, and the Value corresponding to each Key points to the row ID of the Tag Table. The corresponding Key and the row ID of the TagTable corresponding to the Key are quickly queried according to the Primary Tag, and then the record of the Tag Table is accessed through the row ID.
[0075] The Hash Index in this embodiment includes a memory structure and a file structure;
[0076] The memory structure is Bucket, which stores the address of the record pointing to the Index file. Multiple Buckets are used as a Segment logical object.
[0077] The file structure is used to store each Hash Index record, defining the stored Key, Hash Code, and each Hash Index record pointer.
[0078] The structure of the Segment logical object in this embodiment is defined as follows:
[0079] class HashSegment{
[0080] private:
[0081] HashIndexRowlD*m mem bucket_;
[0082] TagHashBucketRWLock*m_bucket_rwlock_;
[0083] size tm bucket_count_;
[0084] public:
[0085] explicit HashSegment(size_t bucket_count=8);
[0086] };
[0087] Among them, HashIndexRowlD*m mem bucket_ represents the pointer to the hash bucket memory;
[0088] TagHashBucketRWLock*m_bucket_rwlock_ represents the read-write lock that controls access to the hash bucket;
[0089] size tm bucket_count_ indicates the number of hash buckets.
[0090] The file structure in this embodiment is defined as follows:
[0091] struct HashIndexData{
[0092] HashCodehash_val;
[0093] TagTableRowlD tbl_row;
[0094] HashIndexRowlDnext_row;
[0095] };
[0096] Among them, HashCodehash_val represents the hash value;
[0097] TagTableRowlD tbl_row represents a pointer or identifier pointing to a table row;
[0098] HashIndexRowlDnext_row represents a pointer or identifier for the next hash index row.
[0099] The Hash Index in this embodiment defines a read and write interface, and improves the usability of the Hash Index by encapsulating the write and query operations of the Hash Index. The key codes are as follows:
[0100] class MMapHashindex{
[0101] protected:
[0102] TagHashindexMutex*m_rehash_mutex_;
[0103] TagHashindexRWLock*m_file_rwlock_;
[0104] HashIndexData*mem hash_;
[0105] HashFunc hash func_;
[0106] std::vector<HashSegment*> buckets_;
[0107] public:
[0108] explicit MMapHashindex(size_t bkt_instances=1, size_tper bkt_count=8);
[0109] *@brief store a key in the hash table.
[0110] @paramkeykey to be found.
[0111] @paramlength of the keylen
[0112] @return 0success.
[0113] int put(const char*key,int len,TagTableRowlD tag table_rowid);
[0114] *@brief find the value stored in hash table for a given key.
[0115] key to be found.@paramkey
[0116] @paramlen
[0117] length of the key.
[0118] @return the stored value in the hash table if key is found; 0otherwise.
[0119] uint32_t get(const char*key,int len).
[0120] As attached Figure 2 As shown, the Hash Index write in this embodiment is specifically as follows:
[0121] (1) Calculate the corresponding Hash Code based on the Key;
[0122] (2) Determine whether the data in the hash table needs to be redistributed and reorganized (rehash):
[0123] ①If yes, proceed to step (3);
[0124] ②If not, jump to step (4);
[0125] (3) reallocate and reorganize data in the hash table;
[0126] (4) Calculate the corresponding Segment and Bucket based on the Hash Code;
[0127] (5) Write the Index Data record and update the Bucket value to the address of the latest Index Data.
[0128] As attached Figure 3 As shown, the Hash Index query in this embodiment is specifically as follows:
[0129] (1) Calculate the corresponding Hash Code based on the Key, and obtain the corresponding Segment and Bucket based on the Hash Code;
[0130] (2) Get the address of the Index Data pointed to by the Bucket;
[0131] (3) Determine whether the address of Index Data is equal to the Key value:
[0132] ①If it is equal to the Key value, it will be returned directly;
[0133] ②If it is not equal to the Key value, continue to the next address of Index Data until the linked list access of Index Data ends.
[0134] Embodiment 2:
[0135] This embodiment also provides an electronic device, including: a memory and a processor;
[0136] Wherein, the memory stores computer-executable instructions;
[0137] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the method for managing data based on hash index of a time series database in any embodiment of the present invention.
[0138] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or any conventional processor, etc.
[0139] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, at least one disk storage period, a flash memory device, or other volatile solid-state storage devices.
[0140] Embodiment 3:
[0141] This embodiment also provides a computer-readable storage medium, which stores a plurality of instructions, which are loaded by a processor, so that the processor executes the method for managing data based on a hash index of a time series database in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code for implementing the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.
[0142] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.
[0143] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer via a communication network.
[0144] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.
[0145] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.
[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for managing hash index data based on a time series database, characterized in that: This method uses an architecture including TagTable and Hash Index. In the KaiwuDB database time series engine, the Primary Tag field is used to create a Hash time series index. In a large number of write scenarios, the Hash algorithm is used to quickly locate whether the Primary Tag record already exists based on the value of the Primary Tag field, and then decide whether to write the Primary Tag record. Specifically: For an existing Primary Tag, it means that the device already exists, and the existing device ID is used to write metrics; If the Primary Tag does not exist, it means that the device does not exist. Use the new device ID to write metrics.
2. The method for managing data based on hash index of time series database according to claim 1, characterized in that: The TagTable structure is used to store all Tag data. It adopts a memory mapping architecture and directly accesses the records of the corresponding row number based on the row ID.
3. The method for managing data based on hash index of a time series database according to claim 1, characterized in that: The HashIndex structure uses the Primary Tag as the Key. The Value corresponding to each Key points to the row ID of the Tag Table. The corresponding Key and the row ID of the Tag Table corresponding to the Key can be quickly queried based on the Primary Tag, and then the record of the Tag Table can be accessed through the row ID.
4. The method for managing data based on hash index of a time series database according to claim 1 or 3, characterized in that: Hash Index includes memory structure and file structure; The memory structure is Bucket, which stores the address of the record pointing to the Index file. Multiple Buckets are used as a Segment logical object. The file structure is used to store each Hash Index record, defining the stored Key, Hash Code, and each HashIndex record pointer.
5. The method for managing data based on hash index of time series database according to claim 4 is characterized in that: The structure of the Segment logical object is defined as follows: class HashSegment{ private: HashIndexRowlD*m mem bucket_; TagHashBucketRWLock*m_bucket_rwlock_; size tm bucket_count_; public: explicit HashSegment(size_t bucket_count=8); }; Among them, HashIndexRowlD*m mem bucket_ represents the pointer to the hash bucket memory; TagHashBucketRWLock*m_bucket_rwlock_ represents the read-write lock that controls access to the hash bucket; size tm bucket_count_ indicates the number of hash buckets.
6. The method for managing data based on hash index of a time series database according to claim 4, characterized in that: The file structure is defined as follows: struct HashIndexData{ HashCodehash_val; TagTableRowlD tbl_row; HashIndexRowlDnext_row; }; Among them, HashCodehash_val represents the hash value; TagTableRowlD tbl_row represents a pointer or identifier pointing to a table row; HashIndexRowlDnext_row represents a pointer or identifier for the next hash index row.
7. The method for managing data based on hash index of a time series database according to claim 1, characterized in that: HashIndex defines the read and write interface, and improves the usability of Hash Index by encapsulating the write and query operations of Hash Index; The specific steps of writing Hash Index are as follows: (1) Calculate the corresponding Hash Code based on the Key; (2) Determine whether the data in the hash table needs to be redistributed and reorganized: ①If yes, proceed to step (3); ②If not, jump to step (4); (3) reallocate and reorganize data in the hash table; (4) Calculate the corresponding Segment and Bucket based on the Hash Code; (5) Write the Index Data record and update the Bucket value to the address of the latest Index Data.
8. The method for managing data based on hash index of time series database according to claim 7, characterized in that: HashIndex performs queries as follows: (1) Calculate the corresponding Hash Code based on the Key, and obtain the corresponding Segment and Bucket based on the Hash Code; (2) Get the address of the Index Data pointed to by the Bucket; (3) Determine whether the address of Index Data is equal to the Key value: ①If it is equal to the Key value, it will be returned directly; ②If it is not equal to the Key value, continue to the next address of Index Data until the linked list access of Index Data ends.
9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores a computer program; The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the method for managing data based on hash index of a time series database as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which can be executed by a processor to implement the method for managing data based on a hash index of a time series database as claimed in any one of claims 1 to 8.