Database dynamic hash index data access method

By introducing dynamic hash functions and bucket address mapping tables into the database, the hash index bucket is dynamically split, which solves the problem of performance degradation of traditional hash indexes when the data volume increases, and realizes automatic index scaling and efficient storage.

CN120030010APending Publication Date: 2025-05-23SHENZHEN TINYSOFT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311563620.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Traditional hash indexes are prone to hash conflicts when the data volume increases, resulting in performance degradation, and the index scale cannot be expanded. The index needs to be rebuilt to ensure efficiency, but this is difficult to achieve when the data volume changes dramatically.

Method used

The dynamic hash function h(k,I)=m%(b*nI), where I is the maximum number of splits and n is the expansion coefficient. By dynamically splitting the hash index bucket and updating the bucket address mapping table, ensuring that the index automatically expands when the data grows, avoiding the reconstruction of the index.

Benefits of technology

It realizes that when the data volume changes dramatically, there is no need to reconstruct the hash index, maintaining the time complexity O(1), improving the storage space utilization efficiency, and avoiding concurrent conflicts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030010A_ABST
    Figure CN120030010A_ABST
Patent Text Reader

Abstract

The invention relates to a data access method for a dynamic hash index of a database. The expansion method for the dynamic hash index of the database comprises the following steps: acquiring a key value k and a maximum splitting frequency I of a data insertion process; calculating a hash index bucket number bn based on the dynamic hash function h (k, I); judging whether the physical address storage space is sufficient based on the bucket address mapping table; if yes, storing the key value k to a physical address; if the hash index buckets are not enough, splitting all the hash index buckets, distributing physical addresses for the new hash index buckets, I = I + 1, and expanding the bucket address mapping table; moving an original key value kx in the database into the corresponding physical address based on a result of the dynamic hash function h (kx, I); and storing a result of the key value k based on the dynamic hash function h (k, I) into a corresponding physical address. By using the method provided by the invention, even if the data volume changes violently, the hash index does not need to be reconstructed, and the access time complexity basically keeps the effect of O (1).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of hash indexes, and in particular to a data access method of a database dynamic hash index. Background Art

[0002] Database hash index, also known as hash index, is widely used in various existing database products because of its high storage speed and time complexity of O(1). Currently, database products such as Oracle and MySQL support hash index. Hash index is a hash-based data structure used to efficiently find and access data in large-scale data. The data keyword (Key) is converted into a hash value through a hash function, and the hash value is used as an index to store and access data.

[0003] The usual database hash function is h(k)=m%b, where k is the key value and b is the number of hash index buckets. Depending on the data type of k, different hash algorithms are used, such as the BKDR algorithm, the CRC algorithm, etc. A function is used to map the data in the domain where k is located to the m corresponding to the integer domain, and then the mapping value m is divided by the number of hash index buckets b, and the remainder is taken as the hash index bucket number. After being mapped by the hash function, the key points to a hash index bucket. If the hash function is a one-to-one mapping, then the number of hash index buckets is fixed.

[0004] With respect to the above-mentioned related technologies, the inventors believe that there are the following problems: once an index is established, its scale cannot be expanded. If the amount of data grows beyond expectations, serious hash conflicts will occur, resulting in a sharp drop in performance. To ensure the efficiency of the hash index when the amount of data increases, the hash index must be rebuilt, and the index is invalid during the reconstruction process. Therefore, before creating a traditional hash index, the user needs to accurately estimate the storage scale. In many systems, this is difficult to do. If the user overestimates the storage scale, storage resources will be wasted. Summary of the invention

[0005] In order to make hash indexes applicable to changing data, even if the amount of data changes dramatically, the hash index does not need to be rebuilt, and the time complexity of access is basically maintained at O(1). The present application provides a data access method for a database dynamic hash index, which is applicable to various database products using hash indexes, especially to database products whose total data volume cannot be accurately estimated and whose data is constantly growing.

[0006] The present application provides a data access method for a database dynamic hash index, which adopts the following technical solution: A method for expanding a database dynamic hash index, wherein the database dynamic hash index includes a dynamic hash function h(k,I)=m%(b*n I) and a bucket address mapping table, wherein k is a key value, b is an initial number of buckets, I is a maximum number of splits (I>=0), an initial value of I is 0, and n is an expansion coefficient of the hash index bucket split (n>1); the bucket address mapping table is used to record the physical address corresponding to each hash index bucket; the expansion method specifically comprises the following steps: Obtain the key value k and the maximum number of split times I of the data insertion process; Calculate the hash index bucket number bn of the target hash index bucket based on the dynamic hash function h(k, I); Based on the bucket address mapping table, determine whether the physical address storage space corresponding to the target hash index bucket is sufficient; if sufficient, store the key value k to the physical address; If it is insufficient, all hash index buckets are split, and the physical address is assigned to each new hash index bucket, I=I+1, and the bucket address mapping table is expanded; the original key value kx in the database is moved to the corresponding physical address based on the result of the dynamic hash function h(kx,I); the key value k is stored in the corresponding physical address based on the result of the dynamic hash function h(k,I).

[0007] By adopting the above technical solution, a physical storage structure is designed for database products with growing data. A new hash function is introduced in this physical storage structure: h(k,I). In addition to the key value k, the parameters of this hash function also include the maximum number of splits of the hash index bucket I. When the database hash index is insufficient in storage space, the number of hash index buckets can be dynamically expanded, b*n I That is, the total number of hash index buckets after each split, which solves the problem of the hash value calculation method constantly changing when the hash index is constantly split. As the amount of data stored in the database grows, the database dynamically splits the hash index, allocates physical addresses to the new hash index buckets, and expands the bucket address mapping table to establish a mapping relationship. In the early stage of the database, the number of hash index buckets is small, and the storage space occupied is small. As the amount of data stored in the database continues to grow, the hash index buckets are dynamically split, gradually expanding the storage space occupied, and improving the utilization efficiency of the storage space.

[0008] Preferably: each of the hash index bucket numbers bn corresponds to a split counter, the split counter records the number of splits i of the corresponding hash index bucket, the initial value of i is 0, and the maximum number of splits I is the maximum value of the number of splits i recorded by each split counter; If it is determined that the physical address storage space of the target hash index bucket is insufficient, the following steps are performed: Splitting all hash index buckets, atomically setting the number of splits i=i+1 of the target hash index bucket, and initializing and setting a new hash index bucket split from the target hash index bucket, wherein the initialization setting refers to allocating the physical address, and atomically setting the number of splits i based on the number of splits i of the target hash index bucket; Determine whether the atomic setting of the split times i=i+1 is successful; if the atomic setting fails, the data insertion process is rolled back and re-executed; Expanding the bucket address mapping table; The original key value kx in the target hash index bucket is stored to the corresponding physical address based on the result of the dynamic hash function h(kx,i); Based on the result of the dynamic hash function h(k,i), the key value k is stored in the physical address.

[0009] By adopting the above technical solution, in order to avoid concurrent conflicts caused by multiple data insertion processes splitting the same hash index bucket at the same time, the number of splits i is introduced as a marker. If a process fails to atomically set the number of splits i=i+1, it means that another process is also inserting into the hash index bucket at this time. Due to the characteristics of atomicity, the addition operation is indivisible. All operations of a process are either executed uninterruptedly or not executed at all. Therefore, the data insertion process will wait in a loop, avoiding concurrent conflicts caused by multiple processes splitting the same hash index bucket at the same time.

[0010] As the amount of data stored in the database increases, the number of hash index buckets in the database hash index can be locally and dynamically expanded. If a hash index bucket is full, the hash index bucket is dynamically split and a physical address is assigned to the new hash index bucket. The number of splits i is updated at the same time to complete the initialization of the new hash index bucket. Other hash index buckets are not initialized for the time being and are in an uninitialized state. This can correctly determine the split level at which the data should be inserted, ensuring the normal operation of the database dynamic hash index without rebuilding the entire hash index to re-hash the original keys in the database. This allocation method further improves the utilization efficiency of storage space.

[0011] Preferably, the step of judging whether the physical address storage space corresponding to the target hash index bucket is sufficient based on the bucket address mapping table further includes the following steps: Determine whether the target hash index bucket is initialized; If not initialized, the target hash index bucket of the key value k when the maximum splitting number I=I-1 is recalculated based on the dynamic hash function h(k,I), and this judgment is repeated.

[0012] By adopting the above technical solution, if the target hash index bucket is not initialized, it means that under the current maximum number of splits I, the target hash index bucket has not been assigned a physical address, and the hash index bucket corresponding to the previous split level has not been split. Therefore, I=I-1 is substituted to recalculate the target hash index bucket of the key value k at the previous split level and judge again. When the hash function is reasonably designed, that is, after the key is hashed, the stored data can be evenly distributed in each bucket, and the average number of times to find the target hash index bucket will be less than twice.

[0013] In theory, the use scale of the database's dynamic hash index is unlimited, and the access performance is basically the same as that of the traditional hash index.

[0014] Preferably, the step of determining whether the atomization setting of the splitting times i=i+1 is successful or not, if the atomization setting is successful, further comprises the following steps: Determine whether the number of splits i of the new hash index bucket split from the target hash index bucket is greater than the maximum number of splits I; if so, atomically set the maximum number of splits I=I+1; if not, the maximum number of splits I remains unchanged.

[0015] By adopting the above technical solution, when the number of splits i is greater than the maximum number of splits I, the maximum number of splits I needs to be updated, and the maximum number of splits I is atomically set to I=I+1, which can ensure the correctness of the maximum number of splits I and avoid incorrect judgment of the relationship between the number of splits i and the maximum number of splits I when other hash index buckets need to be split.

[0016] Preferably: each hash index bucket includes a key quantity counter, the key quantity counter records the key quantity kn in the corresponding hash index bucket, and a single hash index bucket can accommodate a maximum number of keys N; based on the bucket address mapping table, judging whether the physical address storage space corresponding to the target hash index bucket is sufficient is specifically the following steps: Compare whether the number of keys kn of the target hash index bucket is less than N; if so, the result is sufficient; if not, the result is insufficient.

[0017] By adopting the above technical solution and introducing the key number kn, the count of kn can be changed before inserting or deleting the key value k in the hash index bucket, so as to correctly judge whether there are other processes inserting or deleting in the hash index bucket, thereby avoiding erroneous judgment of whether the hash index bucket needs to be split in a concurrent environment.

[0018] Preferably, the step of storing the key value k to the physical address specifically comprises the following steps: Atomically setting the number of keys kn of the target hash index bucket to kn+1; Determine whether the atomic setting is successful; if not, the data insertion process is rolled back and re-executed; If successful, the key value k is stored in the corresponding physical address, and it is determined whether the insertion is successful; if the insertion is not successful, kn=kn-1, and the insertion result outputs false.

[0019] By adopting the above technical solution, when the hash index bucket is not full, the number of keys kn is set atomically to kn+1. When the atomic setting fails, it means that another process is also inserting or deleting in the hash index bucket and is operating on kn, so the current insertion process will enter a loop waiting, avoiding concurrent conflicts of dynamic hash indexes. When the insertion fails, it means that before the process inserts the key value k, another process inserts the same key value k, avoiding the existence of the same key value k in a bucket, and at the same time setting the number of keys kn=kn-1, so that the number of keys kn is correctly indicated.

[0020] Preferably, the data stored in each hash index bucket is stored in a linked list structure at the physical address, and the target hash index bucket enters the cache.

[0021] By adopting the above technical solution, the data in the hash index bucket is stored using a lock-free linked list algorithm. The linked list consists of a series of nodes, each node contains data and a pointer to the next node. Insertion and deletion operations do not need to move the positions of other nodes, but simply adjust the pointer. This means that even in highly concurrent situations, other threads can continue to insert and delete in the linked list without being blocked, making insertion and deletion operations more efficient, unlike arrays that require locks to protect concurrent access and avoid concurrency conflicts.

[0022] Because each split is based on the bucket number of the target hash index bucket, plus the initial bucket number b with the expansion factor n-1 times, the bucket number of the new hash index bucket is obtained, that is, bn+b, bn+2b, bn+3b...bn+(n-1)b. Therefore, no matter how it is split, the size of each hash bucket remains fixed. The target hash index bucket entering the memory is stored in the cache to improve the access efficiency. The cache can be set to a certain size and managed by first-in-first-out or other strategies.

[0023] In a second aspect, the present application provides a method for querying a database dynamic hash index, which adopts the following technical solution: a method for querying a database dynamic hash index, based on the dynamic hash index in the method for expanding the database dynamic hash index, specifically comprising the following steps: Obtain the maximum number of split times I and the key value k of the data query process; Calculating the hash index bucket number bn of the target hash index bucket based on the dynamic hash function h(k,I); Determine whether the target hash index bucket is initialized; if not, recalculate the hash index bucket number bn of the new target hash index bucket when I=I-1 based on the dynamic hash function h(k,I), and repeat this step of determination again; If it has been initialized, determine whether the number of splits i of the target hash index bucket is greater than the maximum number of splits I; if it is greater, recalculate the hash index bucket number bn of the new target hash index bucket when I=i based on the dynamic hash function h(k,I), and return to the previous step to re-determine whether the new target hash index bucket is initialized; If not, determine whether the key value k exists in the target hash index bucket; if so, the query result outputs true; if not, the query result outputs false.

[0024] For the initial state, there are only b initialized buckets. Later, when a hash index bucket is split, it is necessary to initialize the newly split hash index bucket of the hash index bucket to accommodate the data in the split hash index bucket. Under the current maximum number of splits I, if the calculated hash index bucket number bn corresponds to other unsplit hash index buckets, it is in an uninitialized state. By adopting the above technical solution, if the hash index bucket corresponding to the calculated hash index bucket number bn has not been initialized, it means that the hash index bucket has not been split, and it is necessary to query the hash index bucket number bn corresponding to the key at the previous split level.

[0025] If the number of splits i of the target hash index bucket is greater than the maximum number of splits I, it means that the target hash index bucket has just been split due to other insertion processes, and the number of splits i of the target hash index bucket has been updated, while the maximum number of splits I has not been updated. Therefore, the key value k can no longer be found in the bucket, and the hash index bucket corresponding to the key value k when I=i after the split must be searched again. If the split level i of the target hash index bucket has changed at this time, but the key in the target hash index bucket may not have been actually moved to the new hash index bucket, that is, the number of splits i of the new hash index bucket has not been updated, the new bucket is in an uninitialized state at this time, and the loop search continues. Concurrency conflicts between query and insertion in the hash index are avoided. Data query on dynamic hash structures is realized.

[0026] In a third aspect, the present application provides a method for deleting a dynamic hash index of a database, which adopts the following technical solution: a method for deleting a dynamic hash index of a database, based on the dynamic hash index in the method for extending the dynamic hash index of the database, specifically comprising the following steps: Obtain the key value k of the data deletion process; Based on the query method of the database dynamic hash index, query the key value k; If the query result outputs true, kn=kn-1 is set atomically; Determine whether the atomic setting is successful; if not, the data insertion process is rolled back and re-executed; if successful, delete the key value k and determine whether the deletion is successful; If the deletion is successful, the deletion result output is true; if the deletion is unsuccessful, kn=kn+1, and the deletion result output is false.

[0027] The data deletion process does not involve the splitting of the hash index bucket, and the logical steps are similar to the query. By adopting the above technical solution, the key number kn is set atomically = kn-1. When the atomic setting fails, it means that another process is also inserting or deleting in the hash index bucket at this time, so it enters a loop waiting to avoid concurrent conflicts that cause kn errors. When the deletion is not successful, it means that another process has deleted the key value k before the process has deleted it. In this case, kn should be increased by one to avoid kn errors. Concurrency conflicts in dynamic hash indexes are avoided. Data deletion of dynamic hash structures is realized.

[0028] In a fourth aspect, the present application provides a method for updating a dynamic hash index of a database, which adopts the following technical solution: a method for updating a dynamic hash index of a database, based on the dynamic hash index in the method for extending the dynamic hash index of the database, specifically comprising the following steps: Obtain the key value k1 of the update object of the data update process and the key value k2 of the update result; Based on the database dynamic hash index deletion method, delete the key value k1; If the deletion result outputs false, the update result outputs false; If the deletion result outputs true, insert the key value k2 based on the insertion method of the database dynamic hash index; If the insert result outputs false, the update result outputs false; If the insert result outputs true, the update result outputs true.

[0029] In most cases, the update of the dynamic hash structure cannot be performed in the original location, because once the key changes, it is usually migrated to another hash index bucket, so the update process is decomposed into the steps of first deleting and then inserting. By adopting the above technical solution, the data update of the dynamic hash structure is realized.

[0030] In summary, the present application includes at least one of the following beneficial technical effects: 1. This solution is a general database dynamic hash index solution, which is particularly suitable for dynamically changing data. It overcomes the disadvantage that the previous database hash index is only applicable to static data. While ensuring that the access efficiency of the hash index remains basically unchanged, it expands the application field of the hash index; 2. This dynamic hash indexing scheme proposes a solution to concurrent conflicts during data access and a lock-free update mechanism; 3. The data stored in this dynamic hash indexing scheme is stored in a linked list structure at the physical address. The data in the hash index bucket is inserted using a lock-free linked list algorithm. Insertion and deletion operations do not require moving the positions of other nodes, but simply adjust the pointer. Even in highly concurrent situations, other threads can continue to perform insertion and deletion operations in the linked list without being blocked, making insertion and deletion operations more efficient and avoiding concurrent conflicts; 4. The target hash index bucket in this dynamic hash index scheme enters the cache for data access. No matter how the hash index buckets are split, the size of each bucket remains fixed, and the buckets entering the memory are used as cache, which can further improve the access efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a schematic diagram of a dynamic hash index structure used as an example in this application; Figure 2 It is a block diagram of an extension method of a database dynamic hash index; Figure 3 yes Figure 2 A diagram of additional steps performed before step S3; Figure 4 yes Figure 2 The specific steps of S3A are shown in the figure below: Figure 5 It is a query method block diagram of a database dynamic hash index; Figure 6 It is a block diagram of a deletion method of a database dynamic hash index; Figure 7 The present invention is a block diagram of a method for updating a database dynamic hash index. DETAILED DESCRIPTION

[0032] The embodiment of the present application discloses a data access method of a database dynamic hash index.

[0033] To more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. However, it should be understood by those of ordinary skill in the art that the present application can be implemented without these details. In some cases, in order to avoid unnecessary descriptions that make various aspects of the present application obscure, well-known methods, processes, systems, components and / or circuits that have been described at a higher level will not be described in detail. For those of ordinary skill in the art, it is obvious that various changes can be made to the embodiments disclosed in the present application, and without departing from the principles and scope of the present application, the general principles defined in the present application can be applied to other embodiments and application scenarios. Therefore, the present application is not limited to the embodiments shown, but conforms to the broadest scope consistent with the scope claimed for protection of the present application.

[0034] It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention. In addition, the technical features involved in each embodiment of the present invention described below can be combined with each other as long as there is no conflict between them.

[0035] In order to enable those skilled in the art to understand the technical solution provided by the present application in detail, firstly, a brief description is given of the specific terms and concepts involved in the embodiments of the present application.

[0036] The initial state of the dynamic hash index structure is a number of initial hash index buckets, and the hash index bucket numbers are defined in sequence. For example, there are five initial hash index buckets, and the hash index bucket numbers are defined as 1, 2, 3, 4, and 5 respectively. Each hash index bucket corresponds to a physical block in the storage device, and each physical block has a unique physical address, which is used to identify the position of the physical block on the storage device. Each hash index bucket is provided with a split counter i of the bucket, i is greater than or equal to 0 and the initial value is 0. The dynamic hash index structure is also provided with a maximum split count counter I, I is greater than or equal to 0 and the initial value is 0, which records the maximum value of each split counter. Each hash index bucket is also provided with a corresponding key quantity counter, which records the number of keys kn of data stored in the physical block of the physical address corresponding to the hash index bucket number bn.

[0037] The database dynamic hash index includes a dynamic hash function and a bucket address mapping table.

[0038] The dynamic hash function is specifically h(k, I) = m%(b*n I ). Wherein, k is the key value of the data access process, I is the maximum value of the number of splits i recorded by each split counter, n is the expansion factor of the hash index bucket each time the hash index is expanded, n is an integer greater than 1, m is the mapping value of the key value k in the integer domain, and b is the initial number of buckets.

[0039] Calculating the mapping value m requires a function to map the data in the domain of k to the integer domain of m so that it can be brought into the dynamic hash function to obtain the integer remainder. For example, the BKDR algorithm, CRC algorithm, etc. can be used for strings.

[0040] The data in each hash index bucket is identified by the mapping value m and is stored in the physical block of the corresponding physical address through a linked list structure. A linked list is a common data structure used to store and organize data. A linked list consists of a series of nodes, each of which contains data and a pointer to the next node. Because the insertion and deletion operations of the linked list do not need to move the positions of other nodes, they only need to simply adjust the pointer. This means that even in highly concurrent situations, other threads can continue to insert and delete in the linked list without being blocked, making the insertion and deletion operations more efficient, unlike arrays that require locks to protect concurrent access and avoid concurrent conflicts. However, compared with the array structure, the linked list structure is less efficient for random data access, but for the dynamic hash index of this database, each hash index bucket only stores a small part of the data, and even if the search is traversed from the beginning, it will not increase the search time much.

[0041] The dynamic hash index of this database requires the user to first define the initial number of buckets b and the maximum number of keys N that a single hash index bucket can accommodate based on the initial data volume of the database. The number of initial buckets b does not need to be very accurate, as long as it can meet the storage requirements of the initial data. The initial number of buckets b multiplied by the maximum number of keys N that a single hash index bucket can accommodate is the amount of data that the database can initially accommodate. The expansion coefficient n of the hash index bucket split is defined based on the data growth rate. The expansion coefficient n represents the number of new hash index buckets split from a single hash index bucket each time the hash index bucket splits. If the expansion coefficient n is set large, the storage space expanded each time the split is larger; if the expansion coefficient n is set small, the storage space expanded each time the split is smaller. A moderate expansion coefficient n can gradually and progressively expand the storage space occupied by the dynamic hash index to improve the utilization efficiency of the storage space.

[0042] like Figure 1 As shown, it is a schematic diagram of a dynamic hash index structure as an example. On the left is the bucket address mapping table, and on the right is a physical block corresponding to each hash index bucket in the storage device. Each physical block has a unique physical address, which is used to identify its location on the storage device. The hash index bucket number bn and its corresponding physical address are mapped to form a bucket address mapping table. The physical address of the hash index bucket with no physical address assigned in the same address mapping table is 0, which means that the corresponding hash index bucket has not been initialized. The hash index bucket number bn is greater than the initial bucket number b, and the number of splits of the split counter of the bucket i=0, which also means that the corresponding hash index bucket has not been initialized.

[0043] Attached Figure 1It shows that the initial number of buckets b=5, the expansion coefficient n=3, and a single hash index bucket can accommodate a dynamic hash index with the number of keys N=3. It shows the process of the dynamic hash index separating from the upper state to the lower state after the dynamic hash index executes the data insertion process of the mapping value m=32 of the key value k.

[0044] The embodiment of the present application discloses a method for expanding a database dynamic hash index.

[0045] like Figure 2 As shown, the method for extending the database dynamic hash index specifically includes the following steps: S1: Get the key value k and the maximum number of splits I of the data insertion process.

[0046] The data insertion process includes the key value k and the data content, and the key value k is used to uniquely identify and access the data content. The maximum number of splits I is the record value of the highest number of splits counter, which records the maximum value of the number of splits i of each split counter.

[0047] S2: Calculate the hash index bucket number bn of the target hash index bucket based on the dynamic hash function h(k,I).

[0048] The dynamic hash function is h(k,I)=m%(b*n I ), the calculation result is the hash value of the key value k, and the hash value is used as the hash index bucket number bn. For example: Figure 1 As shown, the initial number of buckets of the dynamic hash index structure is b=5, the expansion coefficient is n=3, and a single hash index bucket can accommodate N=3 keys. Figure 1 When the highest split count counter I in the middle and lower half is in the state of 1, a data insertion process is received, and the mapping value m of the key value k of the data insertion process is 32, and h(32,1)=32 / (5*3 1 )=32 / 15=6 remainder 2, that is, h(32,1)=2, and the hash index bucket number bn=2 of the target hash index bucket where the data is inserted.

[0049] like Figure 3 As shown, based on the bucket address mapping table, it is determined whether the physical address storage space corresponding to the target hash index bucket is sufficient, and the following steps are also included: S2.5: Determine whether the target hash index bucket is initialized; if not, recalculate the target hash index bucket of key value k when the maximum number of splits I=I-1 based on the dynamic hash function h(k,I), and repeat this determination.

[0050] Example: When the dynamic hash index is in the Figure 1When the highest split count counter I=1 in the middle and lower half is in the state, when the mapping value m=35 of the key value k of the insertion process, h(35,1)=5, it is determined whether bucket 5 is initialized. The physical address of bucket 5 is 0 or the split count i=0 of bucket 5 is less than the current maximum split count I=1, which confirms that bucket 5 is not initialized, indicating that the hash index bucket of the previous split level of bucket 5 has not been split. Recalculate h(35,0)=0 and determine whether bucket 0 is initialized again. Because bucket 0 has been initialized, the insertion process should be inserted into bucket 0.

[0051] If initialized, S3: Based on the bucket address mapping table, determine whether the physical address storage space corresponding to the target hash index bucket is sufficient.

[0052] The specific judgment method is to compare whether the number of keys kn of the target hash index bucket is less than N; if it is less than, the judgment result is sufficient; if not less than, the judgment result is insufficient.

[0053] For example, when the data with the mapping value m=35 of the key value k of the insertion process in the above example is stored in bucket 0, it is found that the number of keys kn=3 for data stored in bucket 0, and a single hash index bucket can accommodate the maximum number of keys N=3, so the judgment result is insufficient. If the key mapping value m=14 stored by the insertion process, the hash index bucket number bn=h(14,0)=4 of the target hash index bucket finally stored, and the number of keys kn=2 for data stored in bucket 4, so the judgment result is sufficient.

[0054] It can be understood that similar methods of determining whether the physical address storage space is sufficient by judging the preset rule value are within the protection scope of this solution, such as the ratio of the preset capacity exceeding the bucket size or the number of repeated links in the hash index bucket exceeds the specified value.

[0055] If the storage space is sufficient, S3A: store the key value k to the physical address.

[0056] like Figure 4 As shown, storing the key value k to the physical address specifically includes the following steps: S3A-1: Atomically set the number of keys of the target hash index bucket kn=kn+1; determine whether the atomic setting is successful.

[0057] When the hash index bucket is not full, the number of keys kn is set atomically to kn + 1. If the atomic setting fails, it means that another process is also inserting or deleting in the hash index bucket and is operating on kn, so the current insert process will enter a loop waiting, avoiding concurrent conflicts of dynamic hash indexes.

[0058] If unsuccessful, S3A-2a: the data insertion process is rolled back and re-executed.

[0059] If successful, S3A-2b: store the key value k to the corresponding physical address and determine whether the insertion is successful; If the insertion is not successful, S3A-3a: kn = kn-1, and the insertion result output is false; If the insertion is successful, S3A-3b: the insertion result outputs true.

[0060] For example, following the above example, the data with the key mapping value m=14 can be inserted into the next node of the data with the mapping value m=29, and the linked list has a pointer to the next node with the mapping value m=29. The query result outputs true. The data insertion process is completed.

[0061] Among them, when the insertion fails, it is because another process inserts the same key value k before the process inserts the key value k, which avoids the existence of the same key value k in a bucket, and at the same time sets the key number kn=kn-1, so that the key number kn is correctly indicated.

[0062] If the storage space is insufficient, S3B: split all hash index buckets, atomically set the number of splits of the target hash index bucket i=i+1, and initialize the new hash index bucket split by the target hash index bucket, wherein the initialization setting refers to allocating a physical address and atomically setting the number of splits i based on the number of splits i of the target hash index bucket.

[0063] Among them, the expression of splitting all hash index buckets is only to make it easier for people to understand the maximum split count I. The final splitting method depends on which hash index buckets are assigned physical addresses. Only the hash index buckets assigned physical addresses are considered to be truly split, and the hash index buckets not assigned physical addresses are only virtually split.

[0064] For example, if the initial number of buckets b=5, the five initial hash index buckets are split, and the bucket numbers of the new hash index buckets obtained from each initial hash index bucket are bn+b, bn+2b, bn+3b...bn+(n-1)b, that is, the bucket numbers of the new hash index buckets split from bucket 2 are bn=7 and bn=12, and the five initial hash index buckets each split into two new hash index buckets, for a total of 10 new hash index buckets, plus the five initial hash index buckets, there are a total of 15 hash index buckets. The number of splits of bucket 2 is atomically set to i=i+1, that is, i=0+1=1, which means that bucket 2 has been split once. Among the 10 new hash index buckets split, only the new hash index buckets split from the target hash index bucket are initialized, that is, the hash index buckets with bn=7 and bn=12 are assigned physical addresses, and the number of splits i of bucket 7 and bucket 12 is updated to be consistent with bucket 2.

[0065] Among them, key x and key y stored in different original hash index buckets will not fall into the same bucket no matter how the original hash index bucket is split. That is, when n remains unchanged, for any positive integers i1 and i2, h(x,i1) is not equal to h(y,i2).

[0066] S4: Determine whether the atomic setting of the split times i=i+1 is successful; if the atomic setting fails, S4A: the data insertion process is rolled back and re-executed.

[0067] Explanatory: If a process fails to atomically set the split count i=i+1, it means that another process is also inserting data into the hash index bucket at this time. Due to the atomicity, the process is inseparable. All operations of a process are either executed uninterruptedly or not executed at all. Therefore, the data insertion process will wait in a loop, which can avoid concurrent conflicts caused by multiple processes splitting the same hash index bucket at the same time.

[0068] If the atomic setting is successful, S4B: determine whether the number of splits i of the new hash index bucket split by the target hash index bucket is greater than the maximum number of splits I; if so, S5A: atomically set the maximum number of splits I=I+1; if not, S5B: the maximum number of splits I remains unchanged.

[0069] Among them, when the number of splits i is greater than the maximum number of splits I, it means that the maximum level of the dynamic hash index structure needs to be further split. At this time, the maximum number of splits I needs to be updated. The atomic setting of the maximum number of splits I = I + 1 can ensure the correctness of the maximum number of splits I and avoid incorrect judgment of the relationship between the number of splits i and the maximum number of splits I when other hash index buckets need to be split.

[0070] S6: Expand the bucket address mapping table.

[0071] The bucket number of the new hash index bucket and its corresponding physical address are recorded in the bucket address mapping table.

[0072] S7: Store the original key value kx in the target hash index bucket to the corresponding physical address based on the result of the dynamic hash function h(kx,i).

[0073] Exemplary: See attached Figure 1 , when the dynamic hash index is in the attached Figure 1When the highest split count counter I in the upper and middle parts is in the state of 0, the mapping value m received by the insertion process is 32. At this time, h(32,0)=2, but bucket 2 is full and needs to be split. After the split, bucket 2 is split into bucket 7 and bucket 12. At this time, I=1, and the original key value in bucket 2 needs to be recalculated based on I=1. The hash value, that is, h(2,1)=2, h(27,1)=12, h(37,1)=7, the key with mapping value m=2 remains in bucket 2, the key with mapping value m=27 is moved to bucket 12, and the key with mapping value m=37 is moved to bucket 7. The number of keys kn in bucket 2, bucket 7, and bucket 12 changes accordingly, and the dynamic hash index becomes attached. Figure 1 The state of the lower middle part.

[0074] S8: Based on the result of the dynamic hash function h(k,i), the key value k is stored in the physical address.

[0075] Same as Figure 4 Step S3A-1: Atomically set the number of keys kn of the target hash index bucket to kn+1; and determine whether the atomic setting is successful.

[0076] If unsuccessful, S3A-2a: the data insertion process is rolled back and re-executed.

[0077] If successful, S3A-2b: store the key value k to the corresponding physical address and determine whether the insertion is successful; if the insertion is not successful, kn=kn-1, and the insertion result outputs false; if the insertion is successful, the insertion result outputs true.

[0078] Calculate h(32,1)=2, store the corresponding insertion process into the next node of the data with mapping value m=2, and the linked list has a pointer to the next node of the data with mapping value m=2. The query result outputs true. The data insertion process is completed.

[0079] Explanation: The read target hash index bucket enters the cache. The cache here is in memory, which is the cache from the data block in the file system to the memory. There are many ways to implement such a cache, such as using the operating system's memory and the file system's map mapping mechanism, or the application can write a cache management framework for management. The cache can be scaled and managed through first-in-first-out or other strategies. The core is to retain physical blocks with high access frequency and eliminate physical blocks that are not frequently accessed. This further improves the access efficiency of the index.

[0080] The embodiment of the present application also discloses a method for querying a database dynamic hash index.

[0081] Reference Figure 5 , the query methods of database dynamic hash index include: S1: Get the maximum number of splits I and the key value k of the data query process.

[0082] S2: Calculate the hash index bucket number bn of the target hash index bucket based on the dynamic hash function h(k,I).

[0083] S3: Determine whether the target hash index bucket is initialized; if not, recalculate the hash index bucket number bn of the new target hash index bucket when I=I-1 based on the dynamic hash function h(k,I), and repeat this step again.

[0084] If it has been initialized, S4: determine whether the number of splits i of the target hash index bucket is greater than the maximum number of splits I; if it is greater, recalculate the hash index bucket number bn of the new target hash index bucket when I=i based on the dynamic hash function h(k,I), and return to the previous step to re-determine whether the new target hash index bucket is initialized.

[0085] If not, S5: determine whether the key value k exists in the target hash index bucket; if so, the query result outputs true; if not, the query result outputs false.

[0086] For example, in conjunction with Figure 1 , when the dynamic hash index is in the attached Figure 1 When the highest split count counter I in the middle and lower half is in the state of 1, the mapping value m of the key value k received in the data query process is 21, and the hash index bucket number bn=h(21,1)=6 of the target hash index bucket can be obtained. Bucket 6 is checked and its physical address is found to be 0, indicating that bucket 6 is not initialized. It can also be judged that bucket 6 is not initialized based on the fact that bucket 6 is not the initial hash index bucket and the split count i of bucket 6 is 0. Recalculate h(21,0)=1, check bucket 1, bucket 1 has been initialized, and i=I=1, find the key in bucket 1, and output the query result true. The data query process is completed.

[0087] If the mapping value m of the key value k received in the data query process is 17, h(17,1)=2 is calculated, bucket 2 is checked, bucket 2 has been initialized, and i=I=1, the key is not found in bucket 2, and the query result outputs false. The data query process is completed.

[0088] The functions performed in the above-mentioned method for querying a database dynamic hash index and the technical details of each function are the same or similar to the corresponding features in the previously described method for extending a database dynamic hash index, so they are not repeated here.

[0089] The embodiment of the present application also discloses a method for deleting a dynamic hash index of a database.

[0090] Reference Figure 6 , the deletion methods of database dynamic hash index include: S1: Get the key value k of the data deletion process.

[0091] S2: A query method based on a database dynamic hash index to query the key value k.

[0092] S3: If the query result outputs true, atomically set kn=kn-1.

[0093] S4: Determine whether the atomic setting is successful; if not, the data insertion process is rolled back and re-executed; if successful, S5: delete the key value k and determine whether the deletion is successful.

[0094] If the deletion is successful, the deletion result output is true; if the deletion is unsuccessful, kn=kn+1, and the deletion result output is false.

[0095] For example, in conjunction with Figure 1 , when the dynamic hash index is in the attached Figure 1 When the highest split counter I in the lower half is 1, the mapping value m=21 of the key value k of the received data deletion process is found in bucket 1, the corresponding data of the deletion process mapping value m=21 is deleted from the node where it is located, and the pointer of the mapping value m=31 data in the linked list points to the node where the mapping value m=16 data is located. The deletion result output is true. The data deletion process is completed.

[0096] The functions performed in the above-mentioned method for deleting a dynamic hash index of a database and the technical details of each function are the same or similar to the corresponding features in the previously described method for extending a dynamic hash index of a database, so they are not repeated here.

[0097] The embodiment of the present application also discloses a method for updating a dynamic hash index of a database.

[0098] Reference Figure 7 , the update method of the database dynamic hash index includes: S1: Get the key value k1 of the update object and the key value k2 of the update result of the data update process.

[0099] S2: Based on the deletion method of the database dynamic hash index, delete the key value k1.

[0100] If the deletion result outputs false, the update result outputs false.

[0101] If the deletion result outputs true, S3: inserts the key value k2 based on the insertion method of the database dynamic hash index.

[0102] If the insert result outputs false, the update result outputs false.

[0103] If the insert result outputs true, the update result outputs true.

[0104] The functions performed in the above-mentioned method for updating a dynamic hash index of a database and the technical details of each function are the same or similar to the corresponding features in the previously described method for extending a dynamic hash index of a database, so they are not repeated here.

[0105] The above are all preferred embodiments of the present application, and the protection scope of the present application is not limited thereto. Therefore, any equivalent changes made according to the structure, shape, and principle of the present application should be included in the protection scope of the present application.

Claims

1. A method for extending database dynamic hash index, Features: The database dynamic hash index includes a dynamic hash function h(k, I)=m%(b*n I ) and a bucket address mapping table, wherein k is a key value, b is an initial number of buckets, I is a maximum number of splits (I>=0), an initial value of I is 0, and n is an expansion coefficient of the hash index bucket split (n>1); the bucket address mapping table is used to record the physical address corresponding to each hash index bucket; the expansion method specifically comprises the following steps: Obtain the key value k and the maximum number of split times I of the data insertion process; Calculate the hash index bucket number bn of the target hash index bucket based on the dynamic hash function h(k,I); Based on the bucket address mapping table, determine whether the physical address storage space corresponding to the target hash index bucket is sufficient; if sufficient, store the key value k to the physical address; If it is insufficient, all hash index buckets are split, and the physical address is assigned to each new hash index bucket, I=I+1, and the bucket address mapping table is expanded; the original key value kx in the database is moved to the corresponding physical address based on the result of the dynamic hash function h(kx,I); the key value k is stored in the corresponding physical address based on the result of the dynamic hash function h(k,I).

2. The method for expanding a database dynamic hash index according to claim 1, Features: Each of the hash index bucket numbers bn corresponds to a split counter, and the split counter records the number of splits i of the corresponding hash index bucket, where the initial value of i is 0, and the maximum number of splits I is the maximum value of the number of splits i recorded by each split counter; If it is determined that the physical address storage space of the target hash index bucket is insufficient, the following steps are performed: Splitting all hash index buckets, atomically setting the number of splits i=i+1 of the target hash index bucket, and initializing and setting a new hash index bucket split from the target hash index bucket, wherein the initialization setting refers to allocating the physical address, and atomically setting the number of splits i based on the number of splits i of the target hash index bucket; Determine whether the atomic setting of the split times i=i+1 is successful; if the atomic setting fails, the data insertion process is rolled back and re-executed; Expanding the bucket address mapping table; The original key value kx in the target hash index bucket is stored to the corresponding physical address based on the result of the dynamic hash function h(kx,i); Based on the result of the dynamic hash function h(k,i), the key value k is stored in the physical address.

3. The method for expanding a database dynamic hash index according to claim 2, Features: The step of determining whether the physical address storage space corresponding to the target hash index bucket is sufficient based on the bucket address mapping table also includes the following steps: Determine whether the target hash index bucket is initialized; If not initialized, the target hash index bucket of the key value k when the maximum splitting number I=I-1 is recalculated based on the dynamic hash function h(k,I), and this judgment is repeated.

4. The method for expanding a database dynamic hash index according to claim 2, Features: The step of determining whether the atomization setting of the splitting times i=i+1 is successful, if the atomization setting is successful, further includes the following steps: Determine whether the number of splits i of the new hash index bucket split from the target hash index bucket is greater than the maximum number of splits I; if so, atomically set the maximum number of splits I=I+1; if not, the maximum number of splits I remains unchanged.

5. The method for expanding a database dynamic hash index according to claim 1, Features: Each hash index bucket includes a key quantity counter, which records the number of keys kn in the corresponding hash index bucket. A single hash index bucket can accommodate a maximum number of keys N. Based on the bucket address mapping table, judging whether the physical address storage space corresponding to the target hash index bucket is sufficient is specifically the following steps: Compare whether the number of keys kn of the target hash index bucket is less than N; if so, the result is sufficient; if not, the result is insufficient.

6. The method for expanding a database dynamic hash index according to claim 5, Features: The step of storing the key value k to the physical address specifically includes the following steps: Atomically setting the number of keys kn of the target hash index bucket to kn+1; Determine whether the atomic setting is successful; if not, the data insertion process is rolled back and re-executed; If successful, the key value k is stored in the corresponding physical address, and it is determined whether the insertion is successful; if the insertion is not successful, kn=kn-1, and the insertion result outputs false.

7. The method for expanding a database dynamic hash index according to claim 1, Features: The data stored in each hash index bucket is stored in a linked list structure at the physical address, and the target hash index bucket enters the cache.

8. A query method for database dynamic hash index, Features: The dynamic hash index in the expansion method based on the database dynamic hash index specifically includes the following steps: Obtain the maximum number of split times I and the key value k of the data query process; Calculating the hash index bucket number bn of the target hash index bucket based on the dynamic hash function h(k, I); Determine whether the target hash index bucket is initialized; if not, recalculate the hash index bucket number bn of the new target hash index bucket when I=I-1 based on the dynamic hash function h(k, I), and repeat this step of determination again; If it has been initialized, determine whether the number of splits i of the target hash index bucket is greater than the maximum number of splits I; if it is greater, recalculate the hash index bucket number bn of the new target hash index bucket when I=i based on the dynamic hash function h(k, I), and return to the previous step to re-determine whether the new target hash index bucket is initialized; If not, determine whether the key value k exists in the target hash index bucket; if so, the query result outputs true; if not, the query result outputs false.

9. A method for deleting a database dynamic hash index. Features: The dynamic hash index in the expansion method based on the database dynamic hash index specifically includes the following steps: Obtain the key value k of the data deletion process; Based on the query method of the database dynamic hash index, query the key value k; If the query result outputs true, kn=kn-1 is set atomically; Determine whether the atomic setting is successful; if not, the data insertion process is rolled back and re-executed; if successful, delete the key value k and determine whether the deletion is successful; If the deletion is successful, the deletion result output is true; if the deletion is unsuccessful, kn=kn+1, and the deletion result output is false.

10. A method for updating a database dynamic hash index. Features: The dynamic hash index in the expansion method based on the database dynamic hash index specifically includes the following steps: Obtain the key value k1 of the update object of the data update process and the key value k2 of the update result; Based on the database dynamic hash index deletion method, delete the key value k1; If the deletion result outputs false, the update result outputs false; If the deletion result outputs true, insert the key value k2 based on the insertion method of the database dynamic hash index; If the insert result outputs false, the update result outputs false; If the insert result outputs true, the update result outputs true.