A data storage method, device, and storage medium based on dynamic HASH
By adopting a dynamic HASH-based data storage method in the database, using multi-level HASH tables and address pointers to organize data, the problem of increasing the number of splits caused by data insertion is solved, and the concurrency and insertion efficiency of data are improved.
Patent Information
- Application Number
- CN202211045070.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-30
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-08-30
AI Technical Summary
In the case of huge data volume, the number of splits caused by the prior art data insertion increases, resulting in the problem of degradation of data concurrent insertion performance.
Using a dynamic HASH-based data storage method, by setting HASH operation rules at each level, creating multi-level HASH table PAGE pages, and organizing data storage through address pointers to avoid split operations and improve concurrency.
It realizes that when storing massive data, no split operation is performed, and only dynamically expands multi-level HASH tables, which improves the concurrency of data, avoids PAGE page reorganization and split operations, and improves the efficiency of data insertion.
Smart Images

Figure CN115470208B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of databases, and in particular, to a data storage method, device, and storage medium based on dynamic HASH. Background Art
[0002] With the development of the Internet, more and more large information systems are applied in various industries. These large applications use various types of databases and generate a huge amount of data. When the database classifies and stores this data, a fast insertion and fast query method is required to improve the storage performance of the database. Currently, data storage mainly adopts a series of database data storage-related technologies such as using B-trees to organize database data, using B-tree splitting, B-tree reorganization to insert data, and binary search to query data. These existing mainstream database storage technologies are all based on B-trees to organize data. When a B-tree needs to store a large amount of data, due to the characteristics of the B-tree, that is, one table corresponds to a unified B-tree root node. As the amount of inserted data increases, the B-tree root node needs to be split first, and then more leaf nodes and in-page nodes are derived, and the depth of the B-tree increases. The greater the cost for the database to maintain the organizational structure of the B-tree, the more obvious the performance bottleneck. The internal nodes of the B-tree need to continuously split or add internal nodes to maintain the stability of the B-tree. When splitting, the entire B-tree needs to be locked. At this time, only the thread holding the lock can operate the B-tree, and other threads can only be suspended and wait, reducing the concurrency of data storage and affecting the data storage performance.
[0003] In view of this, overcoming the defects of this prior art is an urgent problem to be solved in this technical field.
[0004] Inventive Data
[0005] The technical problem to be solved by the present invention is: in the case of a huge amount of data, how to solve the problem that the increase in the number of splits caused by data insertion in the prior art leads to a decrease in the concurrent insertion performance of data.
[0006] The present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention proposes a data storage method based on dynamic HASH, including:
[0008] Setting HASH operation rules for each level according to a preset method to ensure that data obtains different HASH values through the HASH operation rules of each level;
[0009] Create a first-level HASH table PAGE page in the database. Using the first-level HASH operation rule, obtain the HASH value of the data corresponding to the first-level HASH operation rule, and obtain the bucket of the HASH table PAGE page corresponding to the data according to the HASH value of the first-level HASH operation rule;
[0010] Determine the pre-stored PAGE page according to the address pointer type and the pointed object of the bucket corresponding to the data;
[0011] Determine the target storage PAGE page according to the free space size of the pre-stored PAGE page and whether it contains the target data identical to the currently inserted data, and perform the data storage operation.
[0012] Preferably, the determining the target storage PAGE page according to the free space size of the pre-stored PAGE page and whether it contains the target data identical to the currently inserted data specifically includes:
[0013] If the space of the pre-stored PAGE page is sufficient to save the data and the pre-stored PAGE page does not contain the target data, then store the currently inserted data into the pre-stored PAGE page.
[0014] Preferably, the determining the target storage PAGE page according to the free space size of the pre-stored PAGE page and whether it contains the target data identical to the currently inserted data further includes:
[0015] If the space of the pre-stored PAGE page is not sufficient to save the data, then reclassify the data in the pre-stored PAGE page into the data table;
[0016] Create a second-level HASH table PAGE page, and point the address pointer of the bucket corresponding to the pre-stored PAGE page to the second-level HASH table PAGE page;
[0017] Through the second-level HASH operation rule, obtain the HASH value of the data that has not been inserted in the data table, obtain the bucket of the HASH table PAGE page corresponding to the data according to the HASH value, and point the address pointer of the bucket to the target storage PAGE page;
[0018] Store the data in the target storage PAGE page.
[0019] Preferably, the determining the target storage PAGE page according to the free space size of the pre-stored PAGE page and whether it contains the target data identical to the currently inserted data further includes:
[0020] If the pre-stored PAGE page contains the target data, then create a target storage PAGE page under the current-level HASH table;
[0021] Store the current inserted data and the target data within the target storage PAGE page;
[0022] In the pre-stored PAGE page, modify the position storing the target data to an address pointer pointing to the target storage PAGE page.
[0023] Preferably, determining the target storage PAGE page according to the free space size of the pre-stored PAGE page and whether it contains target data identical to the current inserted data further includes:
[0024] If the space of the pre-stored PAGE page is sufficient to store the data; and the pre-stored PAGE page does not contain the target data;
[0025] Then determine whether an address pointer is stored in the pre-stored PAGE page;
[0026] If an address pointer is stored in the pre-stored PAGE page, obtain the data stored in the first storage PAGE page pointed to by the address pointer, and determine whether the data is identical to the current inserted data;
[0027] If they are identical, store the current inserted data into the first storage PAGE page;
[0028] If they are not identical, store the current inserted data into the pre-stored PAGE page.
[0029] Preferably, storing the current inserted data into the first storage PAGE page includes:
[0030] Determine whether the first storage PAGE page has free space;
[0031] If the first storage PAGE page has free space, store the current inserted data into the first storage PAGE page;
[0032] If the first storage PAGE page has no free space, create a second storage PAGE page, and establish an association between the first storage PAGE page and the second storage PAGE page to form a linked list for storing identical data;
[0033] Store the current inserted data into the second storage PAGE page.
[0034] Preferably, obtaining the bucket of the HASH table PAGE page corresponding to the data according to the HASH value of the first-level HASH operation rule specifically includes:
[0035] Obtain the first-level HASH algorithm through the HASH operation rules at all levels, and calculate the first-level HASH value of the data using the first-level HASH algorithm;
[0036] The first-level HASH value is used as the first-level KEY value of the data, and the number of buckets of the first-level HASH table PAGE page is modulo calculated using the first-level KEY value to obtain the buckets of the HASH table PAGE page corresponding to the data.
[0037] Preferably, determining the pre-stored PAGE page according to the address pointer type and the pointed object of the bucket corresponding to the data specifically includes:
[0038] If the address pointer is a null pointer, create a data PAGE page corresponding to the first-level HASH table PAGE page, and use the data PAGE page as a pre-stored PAGE page;
[0039] If the address pointer is a non-null pointer and the address pointer points to a data PAGE page of the HASH table at this level, the data PAGE page is used as a pre-stored PAGE page;
[0040] If the address pointer is a non-null pointer and the address pointer points to the next-level HASH table PAGE page, then obtain the next-level HASH table PAGE page pointed to by the address pointer;
[0041] In the next-level HASH table PAGE page, the HASH value of the data is calculated using the HASH operation rule corresponding to the next-level HASH table, the bucket of the next-level HASH table PAGE page is determined using the HASH value, and the data PAGE page pointed to by the bucket of the next-level HASH table PAGE page is used as the pre-stored PAGE page.
[0042] In a second aspect, the present invention further provides a data storage device based on dynamic HASH, which is used to implement the data storage method based on dynamic HASH in the first aspect, and the device includes:
[0043] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to execute the dynamic HASH-based data storage method described in the first aspect.
[0044] In a third aspect, the present invention further provides a non-volatile computer storage medium, wherein the computer storage medium stores computer executable instructions, and the computer executable instructions are executed by one or more processors to complete the dynamic HASH-based data storage method described in the first aspect.
[0045] The present invention adopts a dynamic HASH storage method to organize data. Although the basic storage unit (PAGE) remains unchanged, when storing a large amount of data, there will be no splitting operation, but only dynamically expand the multi-level HASH table, and only lock one bucket of the upper-level HASH table, without affecting the operations of other buckets, improving the concurrency of data. Secondly, the present invention adopts a PAGE linked list organization for storing the same data. When the PAGE space is insufficient, a new PAGE is created and put into the linked list, thus avoiding the recombination and splitting operations of the PAGE. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 is a flowchart of a data storage method based on dynamic HASH provided by an embodiment of the present invention;
[0048] Figure 2 is a flowchart of a data storage method based on dynamic HASH provided by an embodiment of the present invention;
[0049] Figure 3 is a schematic diagram of the framework process of a data storage method based on dynamic HASH provided by an embodiment of the present invention;
[0050] Figure 4 is a data structure table of a specific application scenario of a data storage method based on dynamic HASH provided by an embodiment of the present invention;
[0051] Figure 5 is a data structure table of a specific application scenario of a data storage method based on dynamic HASH provided by an embodiment of the present invention, which obtains the data with the data HASH value through the first-level HASH algorithm;
[0052] Figure 6 is a memory structure diagram of the PAGE H1 of the first-level HASH table of a specific application scenario of a data storage method based on dynamic HASH provided by an embodiment of the present invention;
[0053] Figure 7 is a schematic diagram of the storage structure for inserting data in a specific application scenario of a data storage method based on dynamic HASH provided by an embodiment of the present invention;
[0054] Figure 8It is a schematic diagram of the storage structure for inserting data in a specific application scenario of the data storage method based on dynamic HASH provided by an embodiment of the present invention;
[0055] Figure 9 It is a data structured table with a data HASH value obtained by the second-level HASH algorithm in a specific application scenario of the data storage method based on dynamic HASH provided by an embodiment of the present invention;
[0056] Figure 10 It is a schematic diagram of the storage structure for inserting data in a specific application scenario of the data storage method based on dynamic HASH provided by an embodiment of the present invention;
[0057] Figure 11 It is a schematic diagram of the storage structure for inserting data in a specific application scenario of the data storage method based on dynamic HASH provided by an embodiment of the present invention;
[0058] Figure 12 It is a schematic diagram of the storage structure for inserting data in a specific application scenario of the data storage method based on dynamic HASH provided by an embodiment of the present invention;
[0059] Figure 13 It is a schematic diagram of the storage structure for inserting data in a specific application scenario of the data storage method based on dynamic HASH provided by an embodiment of the present invention;
[0060] Figure 14 It is a schematic diagram of the storage structure for inserting data in a specific application scenario of the data storage method based on dynamic HASH provided by an embodiment of the present invention;
[0061] Figure 15 It is a schematic diagram of the structure of a data storage device based on dynamic HASH provided by an embodiment of the present invention. Detailed implementation manners
[0062] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0063] In the description of the present invention, the orientation or positional relationship indicated by terms such as "inside", "outside", "longitudinal", "transverse", "upper", "lower", "top", "bottom", etc. is the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention rather than requiring the present invention to be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.
[0064] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0065] Embodiment 1:
[0066] The present invention provides a data storage method based on dynamic HASH, which is characterized by including:
[0067] Set the HASH operation rules at each level according to a preset method to ensure that different HASH values can be obtained for data through the HASH operation rules at each level;
[0068] Create the first-level HASH table PAGE page in the database, use the first-level HASH operation rules to obtain the HASH value of the data corresponding to the first-level HASH operation rules, and obtain the bucket of the HASH table PAGE page corresponding to the data according to the HASH value of the first-level HASH operation rules; determine the pre-storage PAGE page according to the address pointer type and the pointed object of the bucket corresponding to the data; determine the target storage PAGE page according to the free space size of the pre-storage PAGE page and whether it contains the target data identical to the currently inserted data, and perform the data storage operation.
[0069] The determination of the target storage PAGE page according to the free space size of the pre-storage PAGE page and whether it contains the target data identical to the currently inserted data specifically includes: if the space of the pre-storage PAGE page is sufficient to save the data and the pre-storage PAGE page does not contain the target data, then store the currently inserted data into the pre-storage PAGE page.
[0070] The determination of the target storage PAGE page according to the free space size of the pre-storage PAGE page and whether it contains the target data identical to the currently inserted data further includes: if the space of the pre-storage PAGE page is not sufficient to save the data, then return the data in the pre-storage PAGE page to the data table; create the second-level HASH table PAGE page, and point the address pointer of the bucket corresponding to the pre-storage PAGE page to the second-level HASH table PAGE page; obtain the HASH value of the data not inserted in the data table through the second-level HASH operation rules, obtain the bucket of the HASH table PAGE page corresponding to the data according to the HASH value, and point the address pointer of the bucket to the target storage PAGE page; store the data in the target storage PAGE page.
[0071] Determining the target storage PAGE page according to the free space size of the pre-stored PAGE page and whether it contains target data identical to the currently inserted data further includes: if the target data is included in the pre-stored PAGE page, creating a target storage PAGE page under the current-level HASH table; storing the currently inserted data and the target data in the target storage PAGE page; in the pre-stored PAGE page, modifying the position storing the target data to an address pointer pointing to the target storage PAGE page.
[0072] Determining the target storage PAGE page according to the free space size of the pre-stored PAGE page and whether it contains target data identical to the currently inserted data further includes: if the space of the pre-stored PAGE page is sufficient to save data; and the target data is not included in the pre-stored PAGE page; then determining whether an address pointer is stored in the pre-stored PAGE page; if an address pointer is stored in the pre-stored PAGE page, obtaining the data stored in the first storage PAGE page pointed to by the address pointer, and determining whether the data is identical to the currently inserted data; if they are identical, storing the currently inserted data into the first storage PAGE page; if they are not identical, storing the currently inserted data into the pre-stored PAGE page.
[0073] Storing the currently inserted data into the first storage PAGE page includes: determining whether the first storage PAGE page has free space; if the first storage PAGE page has free space, storing the currently inserted data into the first storage PAGE page; if the first storage PAGE page has no free space, creating a second storage PAGE page, and establishing an association between the first storage PAGE page and the second storage PAGE page to form a linked list for storing identical data; storing the currently inserted data into the second storage PAGE page.
[0074] Obtaining the bucket of the HASH table PAGE page corresponding to the data according to the HASH value of the first-level HASH operation rule specifically includes: obtaining the first-level HASH algorithm through the HASH operation rules of each level, and calculating the first-level HASH value of the data by using the first-level HASH algorithm; using the first-level HASH value as the first-level KEY value of the data, and performing a remainder calculation on the number of buckets of the first-level HASH table PAGE page through the first-level KEY value to obtain the bucket of the HASH table PAGE page corresponding to the data.
[0075] The method of determining the pre-storage PAGE page according to the address pointer type and the pointed object of the bucket corresponding to the data specifically includes: if the address pointer is a null pointer, creating a data PAGE page corresponding to the first-level HASH table PAGE page, and using the data PAGE page as the pre-storage PAGE page; if the address pointer is a non-null pointer, and the address pointer points to the data PAGE page of the current-level HASH table, then using the data PAGE page as the pre-storage PAGE page; if the address pointer is a non-null pointer, and the address pointer points to the next-level HASH table PAGE page, obtaining the next-level HASH table PAGE page pointed to by the address pointer; in the next-level HASH table PAGE page, calculating the HASH value of the data according to the HASH operation rule corresponding to the next-level HASH table, determining the bucket of the next-level HASH table PAGE page by the HASH value, and using the data PAGE page pointed to by the bucket of the next-level HASH table PAGE page as the pre-storage PAGE page.
[0076] The present invention adopts a dynamic HASH storage method to organize data. Although the basic storage unit (PAGE page) does not change, there will be no split operation when storing massive data. Only the multi-level HASH table will be dynamically expanded, and only one bucket of the upper-level HASH table will be locked, which will not affect the operation of other buckets, thereby improving the concurrency of data. Secondly, the present invention adopts a PAGE page linked list organization for the storage of the same data. When the PAGE page space is insufficient, a new PAGE page is created, and the newly created PAGE page is put into the linked list, thereby avoiding the reorganization and splitting operations of the PAGE page.
[0077] Embodiment 2:
[0078] Embodiment 2 of the present invention mainly explains the contents of embodiment 1 in detail. Embodiment 2 provides a data storage method based on dynamic HASH, such as Figure 1 As shown, including:
[0079] Step 201: setting HASH operation rules at each level according to a preset method, so as to ensure that data obtains different HASH values through the HASH operation rules at each level.
[0080] Among them, the embodiments of the present invention implement data storage by storing (inserting) data while dynamically expanding the HASH table. The PAGE page is the basic unit for storing database data in memory. The size of the PAGE page is related to the configuration of the database. When the configuration of the database is determined, the size of the PAGE page is a fixed value. Usually, the PAGE page stores tuple information of data. The PAGE page in the embodiments of the present invention stores data record information (hereinafter, data is used instead of data record information in the description of the following solutions) or HASH table information (i.e., the in-memory structure of the HASH table).
[0081] Since the HASH table information recorded by each level of the HASH table is limited, the data that can be stored in the corresponding data PAGE page is limited. To store as much information as possible, the embodiments of the present invention set the HASH operation rules for each level according to a preset method, so that different data obtain different HASH values through operations at each level. When the target storage PAGE page corresponding to the upper-level HASH table is full (insufficient to continue storing the inserted data), a lower-level HASH table and the target storage PAGE page corresponding to the lower-level HASH table are created. Through different-level HASH operation rules, it is ensured that the currently inserted data and the remaining uninserted data can be inserted as evenly as possible into the target storage PAGE page corresponding to the lower-level HASH table PAGE page through the HASH operation rules of the lower level.
[0082] The preset method in the embodiments of the present invention is mainly determined according to the specific situation of the data to be stored (inserted) in the database. The preset method described in the embodiments of the present invention needs to ensure that the HASH values calculated by the data to be stored through the operation rules at each level are different, and it is ensured that the currently inserted data and the remaining uninserted data can be inserted as evenly as possible into the target storage PAGE page corresponding to the lower level through the HASH operation rules of the lower level. The embodiments of the present invention perform data storage by gradually creating HASH table PAGE pages and data PAGE pages, creating multi-level HASH tables and data PAGE pages while storing until all data is inserted.
[0083] Step 202: Create a first-level HASH table PAGE page in the database, use the first-level HASH operation rule to obtain the HASH value of the data corresponding to the first-level HASH operation rule, and obtain the bucket of the data corresponding to the HASH table PAGE page according to the HASH value of the first-level HASH operation rule; determine the pre-storage PAGE page according to the address pointer type and the pointed object of the bucket corresponding to the data.
[0084] Among them, in the embodiment of the present invention, a first-level HASH table PAGE page is first created, and the first-level HASH table PAGE page is used as the first node stored in the database of the embodiment of the present invention. Each level of the HASH table PAGE page occupies a complete PAGE page storage unit. According to the previously set HASH operation rules at each level, the first-level HASH operation rule (algorithm) is obtained, and the first-level HASH values corresponding to all the data to be inserted are calculated. According to the HASH value of the first-level HASH operation rule, the bucket of the HASH table PAGE page corresponding to the data is obtained. Specifically: the first-level HASH algorithm is obtained through the HASH operation rules at each level, and the first-level HASH value of the data is calculated by using the first-level HASH algorithm; the first-level HASH value is used as the first-level KEY value of the data (since the KEY value of the data is used for modulo calculation with the number of HASH buckets when calculating the HASH bucket corresponding to the data by the HASH operation formula, the KEY value is introduced here for explanation). The modulo calculation is performed on the number of buckets of the first-level HASH table PAGE page by using the first-level KEY value (HASH value), and the bucket of the HASH table PAGE page corresponding to the data is obtained.
[0085] For example: the number of buckets of each level of the HASH table PAGE page of database A is M, the HASH value of data a after the first-level HASH algorithm is T, the value obtained by taking the remainder of T / M is S, and the bucket label of the first-level HASH table PAGE page corresponding to data a is S.
[0086] For the HASH table PAGE of a specific database, assume that the memory occupied by the PAGE is N, the size of the address pointer in the corresponding HASH bucket is R (R is a fixed value), and the number of buckets in each level of the HASH table PAGE is W. According to the formula for calculating the number of buckets in the HASH table PAGE: W = N ÷ R, it can be seen that the number of buckets in each level of the HASH table PAGE is the same (the number of buckets in each HASH table PAGE is fixed and equal). Based on the type of address pointer stored in the first-level HASH table PAGE and the object it points to, determine the PAGE where the data is pre-stored. The address pointer in the embodiment of the present invention mainly plays a role similar to a bridge, connecting the relationship between the upper-level PAGE and the lower-level PAGE during the entire data insertion process. According to the storage content in the PAGE, there are mainly two types of PAGEs in the embodiment of the present invention: the first type is used to store the inserted data, called the data PAGE; the second type is used to store HASH table information (i.e., the HASH table memory structure), called the HASH table PAGE. Based on the two types of PAGEs of the present invention, the address pointer may have the following three situations: the data PAGE points to the data PAGE, the data PAGE points to the HASH table PAGE, and the HASH table PAGE points to the HASH table PAGE (the principle of the HASH table PAGE pointing to the data PAGE is the same as that of the data PAGE pointing to the HASH table PAGE, counted as one situation).
[0087] The address pointer type of the present invention is divided into two types: null pointer and non-null pointer. A null pointer means that when the data corresponds to the current PAGE (HASH table PAGE or data PAGE) and points to the lower-level PAGE, there is no target to point to, that is, the lower-level PAGE does not exist, and the lower-level PAGE needs to be created before pointing can be performed; a non-null pointer means that when the data corresponds to the current PAGE and points to the lower-level PAGE, there is a target to point to.
[0088] According to the address pointer type and the pointed object, the present invention determines the pre-stored PAGE page, which specifically includes: if the address pointer is a null pointer, a data PAGE page corresponding to the first-level HASH table PAGE page is created, and the data PAGE page is used as the pre-stored PAGE page; if the address pointer is a non-null pointer and the address pointer points to a data PAGE page, the data PAGE page is used as the pre-stored PAGE page; if the address pointer is a non-null pointer and the address pointer points to the next-level HASH table PAGE page, the next-level HASH table PAGE page pointed to by the address pointer is obtained; in the next-level HASH table PAGE page, the HASH value of the data is calculated according to the HASH operation rule corresponding to the next-level HASH table, the bucket of the next-level HASH table PAGE page is determined by the HASH value, and the data PAGE page pointed to by the bucket of the next-level HASH table PAGE page is used as the pre-stored PAGE page. It should be noted that the data of the present invention can only be stored in the data PAGE page. The HASH table PAGE page is mainly used to expand the performance of storing and expanding concurrent insertions, and record the position of the data PAGE page where the data is inserted, which is convenient for query. In addition, the data PAGE page corresponding to the HASH table PAGE page described in the embodiment of the present invention actually refers to the data PAGE page corresponding to the specific bucket of the HASH table PAGE page. To facilitate understanding that the pre-stored PAGE page of the present invention is the data PAGE page pointed to by the bucket of the HASH table PAGE page corresponding to the data, whether the data is stored (inserted) in the pre-stored PAGE page also needs to be further determined according to specific circumstances (which will be described later and will not be elaborated here). Based on this, this data PAGE page is called the pre-stored PAGE page.
[0089] Step 203: Determine the target storage PAGE page according to the free space size of the pre-stored PAGE page and whether it contains the target data identical to the currently inserted data, and perform the data storage operation.
[0090] Among them, after obtaining the free space size of the pre-stored PAGE page, by comparing the space size occupied by the currently inserted data, it can be determined whether the space in the pre-stored PAGE page is sufficient to store the data that needs to be inserted currently. Since the inserted data can only be stored in the data PAGE page, by judging the free space size of the pre-stored PAGE page and whether the currently inserted data is the same as the data in the pre-stored PAGE page, a method of determining whether the pre-stored PAGE page is the target storage PAGE page or a new data PAGE page needs to be created as the target storage PAGE page, the target storage PAGE page of the currently inserted data is obtained, and the current data is inserted into the target storage PAGE page. It should be noted that the target storage PAGE page corresponding to step 302 in the embodiment of the present invention represents the data PAGE page where the currently inserted data is finally stored.
[0091] Based on the prior art, the present invention introduces a HASH storage method and adopts a non-splitting storage method, which to a certain extent solves the problem of the decline in data concurrent insertion performance caused by the increase in the number of splits during insertion in the prior art, and improves the efficiency of data insertion.
[0092] To elaborate on the complete solution of the present invention, the following will explain the specific content details of the present invention in detail. During the process of data storage, in order to successfully store the inserted data in the corresponding data PAGE page, the present invention sets up a pre-stored PAGE page and marks the storage direction of the data step by step through the address pointer in the HASH table PAGE page. After obtaining the data PAGE page of the currently inserted data, the specific PAGE page where the data is stored can be traced through the pointing record of the address pointer, thereby achieving the purpose of quickly querying data. The pre-stored PAGE page has been explained above and will not be elaborated here. After obtaining the pre-stored PAGE page, the target storage PAGE page is determined according to the free space size of the pre-stored PAGE page and whether the data PAGE page contains the same data as the data currently being inserted, specifically including:
[0093] If the space of the pre-stored PAGE page is sufficient to store the data and the pre-stored PAGE page does not contain the same data as the current data, the current data is stored in the pre-stored PAGE page. At this time, the pre-stored PAGE page and the target storage PAGE page where the data is finally stored are the same PAGE page. The pre-stored PAGE page and the data PAGE page where the data is finally stored in the present invention represent two different states of the PAGE page, which can be the same data PAGE page or different data PAGE pages; the current data represents the data currently being inserted.
[0094] If the space of the pre-stored PAGE page is not sufficient to store the data, it enters the step of creating the second-level HASH table PAGE page and the data PAGE page pointed to by the address pointer of the second-level HASH table PAGE page, stores the data in the data PAGE page pointed to by the second-level HASH table PAGE page, and inserts the data in the pre-stored PAGE page into the data PAGE page pointed to by the second-level HASH table PAGE page according to the second-level HASH operation rule, and deletes the original pre-stored PAGE page.
[0095] If the pre-stored PAGE page contains data identical to the current data, a new data PAGE page is created, and the current data is stored in the new data PAGE page. Specifically, if the pre-stored PAGE page contains data identical to the current data, a new data PAGE page is created, the stored data in the pre-stored PAGE page is traversed, and all data identical to the currently inserted data (including the current data) is stored in the newly created data PAGE page, and the pointing of the address pointer is adjusted. In addition, after the newly created data PAGE page is full, the present invention further creates a new data PAGE page, stores the identical data in the newly created data PAGE page, and adjusts the pointing of the address pointer. It should be noted that the data storage in the embodiments of the present invention is actually achieved by establishing a dynamic HASH table, obtaining the data PAGE page where the data is stored by using the HASH table and the address pointer in the data PAGE page, and associating through the pointing of the address pointer, so that the data can be quickly and efficiently queried after storage. After obtaining the data PAGE page of the data and storing the data in the corresponding data PAGE page, it is necessary to adjust the pointing direction of the address pointer in the data storage process. The entire data storage process and the adjustment process of the pointing of the address pointer will be explained by specific embodiments later, and no further explanation is made on the adjustment of the pointing of the address pointer here.
[0096] The embodiments of the present invention mainly increase the concurrency (concurrent insertion performance) of data through multi-level HASH dynamic expansion. Further, the steps of creating the second-level HASH table PAGE page and the data PAGE page pointed to by the address pointer of the second-level HASH table PAGE page are as Figure 2 shown and specifically include:
[0097] Step 301: Create a second-level HASH table PAGE page, and point the address pointer of the bucket corresponding to the pre-stored PAGE page to the second-level HASH table PAGE page. Through the second-level HASH operation rule, obtain the HASH value of the data not inserted in the data table, and obtain the bucket of the HASH table PAGE page corresponding to the data according to the HASH value.
[0098] Among them, when the space of the pre-stored PAGE page is not enough to save the currently inserted data, the steps of creating the second-level HASH table PAGE page and the data PAGE page pointed to by the address pointer of the second-level HASH table PAGE page are entered. At this time, it can be considered that the pre-stored PAGE page is full (not enough to store any of the data to be inserted). When the pre-stored PAGE page is full, a new data PAGE page is usually created, and the data is stored in the newly created PAGE page. When there is a large amount of data, although the data can be stored, it is very difficult to accurately and quickly query the data that needs to be used during the use process. The present invention uses a method of multi-level HASH dynamic expansion. When the space of the pre-stored PAGE page is not enough to save the currently inserted data, a second-level HSAH table PAGE page is created, and the data PAGE page pointed to by the address pointer of the bucket of the second-level HASH table PAGE page corresponding to the currently inserted data, and the second-level HASH operation rule is obtained. The data in the pre-stored PAGE page and the remaining data to be inserted are processed through the second-level HASH algorithm. The HASH value of the data is obtained through the second-level HASH algorithm, and the bucket of the second-level HASH table PAGE page corresponding to the currently inserted data and the address pointer where the bucket is stored are obtained by taking the remainder, so as to facilitate subsequent storage of the data in the data PAGE page pointed to by the address pointer of the second-level HASH table PAGE page. The present invention is carried out in a way of expanding while storing. After the data PAGE page corresponding to the bucket (specific bucket) of the first-level HASH table PAGE page is full, a second-level HASH table PAGE page and the data PAGE page pointed to by the address pointer of the bucket of the second-level HASH table PAGE page corresponding to the currently inserted data are newly created through the second-level HASH expansion method, and the currently inserted data is stored, and so on, expanding while storing until all the data is stored.
[0099] Step 302: Obtain the address pointer of the bucket of the HASH table PAGE page corresponding to the data according to the HASH value, point it to the target storage PAGE page, and store the data in the target storage PAGE page.
[0100] Among them, after the data PAGE page pointed to by the address pointer of the second-level HASH table PAGE page is created, the currently inserted data is stored in the created data PAGE page, and the address pointer during the data storage process is adjusted to facilitate subsequent quick query of the specific stored data.
[0101] The embodiments of the present invention adopt a method of storing while expanding. Through multi-level HASH dynamic expansion, data is stored level by level. Next, the process of multi-level HASH dynamic expansion in the embodiments of the present invention will be described. The present invention also includes adopting a method of storing level by level. Through the PAGE page of the second-level HASH table, the PAGE page of the Nth-level HASH table is obtained, where N≥2 (N is an integer). Specifically, it includes: if the space of the data PAGE page pointed to by the bucket of the PAGE page of the second-level HASH table is not enough to store the currently inserted data, then enter the step of creating the PAGE page of the third-level HASH table and the data PAGE page pointed to by the address pointer of the third-level HASH table; if the space of the data PAGE page pointed to by the bucket of the PAGE page of the third-level HASH table is not enough to store the currently inserted data, then enter the step of creating the PAGE page of the fourth-level HASH table and the data PAGE page pointed to by the address pointer of the fourth-level HASH table; through a recursive method, if the space of the data PAGE page pointed to by the bucket of the PAGE page of the (N−1)th-level HASH table is not enough to store the currently inserted data, then enter the step of creating the PAGE page of the Nth-level HASH table and the data PAGE page pointed to by the address pointer of the Nth-level HASH table. Among them, the size of N mainly depends on the storage occupied by the total amount of inserted data. It should be noted that when performing multi-level HASH dynamic expansion, when calculating the bucket of the HASH table PAGE page corresponding to the inserted data, the corresponding algorithm should use the HASH operation rule of the corresponding level set according to the preset method. Through the HASH operation rule of the corresponding level of the inserted data, the bucket of the HASH table PAGE page corresponding to the inserted data is calculated. For example: when the corresponding level of the currently inserted data is the PAGE page of the fifth-level HASH table, the fifth-level HASH operation rule is used to calculate the HASH value of the currently inserted data and the bucket of the fifth-level HASH table PAGE page corresponding to the currently inserted data, and the currently inserted data is stored in the data PAGE page corresponding to the bucket of the fifth-level HASH table PAGE page corresponding to the currently inserted data (it is necessary to ensure that the data PAGE page is sufficient to store the currently inserted data. If it is not enough to insert, then perform the next-level HASH dynamic expansion and store the data in the data PAGE page corresponding to the next-level HASH dynamic expansion).
[0102] The present invention adopts a dynamic HASH storage method to organize data. Although the basic storage unit (PAGE page) remains unchanged, when storing a large amount of data, there will be no splitting operation. Instead, it will only dynamically expand the multi-level HASH table, and only lock one bucket of the upper-level HASH table, without affecting the operations of other buckets, thus improving the concurrency of data. Secondly, the present invention organizes the storage of the same data using a PAGE page linked list. When the PAGE page space is insufficient, a new PAGE page is created and added to the linked list, avoiding PAGE page reorganization and splitting operations.
[0103] Embodiment 3:
[0104] To further elaborate on the complete solution of the present invention, Embodiment 3 explains the solution of the present invention in detail in combination with a specific application scenario. As Figure 3 shown, it is a flowchart representing the specific application scenario of the present invention.
[0105] First, obtain the data to be inserted and process the data. As Figure 4 shown, create a database table T (area code, city name) based on dynamic HASH to facilitate the subsequent rapid insertion and storage of data. To simplify the storage process of the embodiment of the present invention, when the data is modulo-calculated to obtain the bucket subscript of the corresponding HASH table, the HASH value of the data is used to represent the KEY value of the data for calculation.
[0106] Assume that the basic storage unit PAGE page in the memory of the database has a fixed length, and only 4 city names and their area codes can be stored in one PAGE page. The calculation formula of the HASH value corresponding to each piece of data has different algorithms in each level of the HASH table; according to the HASH value corresponding to each piece of data, find the corresponding bucket subscript in the HASH table. The calculation formula is to take the remainder after dividing the HASH value by the number of buckets. Now, the structured data needs to be saved to table T. Among them, the bucket subscript represents the specific bucket number within the PAGE page of the HASH table corresponding to the inserted data.
[0107] Create a HASH table PAGE page P1 in the memory, and create a first-level HASH table PAGE page H1 in P1. Calculate the number of buckets included in the HASH table PAGE page H1 according to the PAGE page size, and divide the PAGE page into multiple buckets to form the first-level HASH table PAGE page H1. The algorithm for calculating the HASH value of the first-level HASH table PAGE page H1 in Embodiment 3 of the present invention is to add the area code value and the number of strokes of the city name to obtain a structured data table with HASH values, as Figure 5 shown. Assume that the first-level HASH table PAGE page H1 contains 5 buckets. The structures of P1 and H1 in the memory are as Figure 6As shown in the figure. Among them, the number of buckets in each PAGE page of the HASH table in the embodiment of the present invention is 5.
[0108] Save the first row of data "010|Beijing". According to the calculated HASH value 23, find the bucket with index 3 in H1. It is found that the address pointer in the bucket is a null pointer. At this time, a new data PAGE page P2 needs to be created, and the address pointer points to P2. After calculating the free space of P2, it is found that this row of data can be saved in P2. The structure in memory after saving is as Figure 7 shown. Continue to save the remaining structured data. After saving "0371|Zhengzhou", the structure in memory is as Figure 8 shown.
[0109] Continue to save "0311|Shijiazhuang". It is found that the HASH value is 333, and the corresponding bucket index in H1 is 3. According to the address pointer in the bucket, directly locate the data PAGE page P2. After calculating the free space of P2, it is found that the space of this page has been used up (or is not enough to store the currently inserted data), and H1 needs to be dynamically expanded.
[0110] Create a new PAGE page P5. Create a next-level (second-level) HASH table PAGE page H2 in this page. The algorithm for calculating the HASH value of H2 is the area code value plus twice the number of strokes of the city name (for the nth-level HASH table PAGE page Hn, the algorithm for calculating the HASH value of Hn is the area code value plus n times the number of strokes of the city name). Recalculate the HASH values of the existing data in P2 and the currently to-be-inserted (remaining) data. The data is as Figure 9 shown, that is, form a new table T with the existing data in P2 and the remaining data in the original table T, and calculate the HASH value of the to-be-inserted data according to the HASH value calculation rule of the second-level HASH table.
[0111] Among them, the number of buckets in H2 is 5. The structure in memory after adding the new HASH table is as Figure 10 shown. Disconnect the pointer in the bucket with index 3 in H1 from P2 and point it to P5. And it is necessary to re-locate the position of the bucket for the data "0311|Shijiazhuang" in H2. The calculated HASH value in H2 is 353, and the corresponding bucket index is 3. Since the pointer stored in the bucket is null, a new data PAGE page P6 needs to be created to save the data. After judging the free space in P6, it is found that this row of data can be saved. The structure in memory after saving is as Figure 11 shown.
[0112] Relocate the bucket subscript for the 4 tuple information ("010|Beijing, 023|Chongqing, 020|Guangzhou, 0755|Shenzhen") in data PAGE P2 under the first-level HASH table in H2, and create a new data PAGE or directly locate the data PAGE according to the type of pointer in the bucket to complete data storage. After the storage is completed, the structure in memory is as Figure 12 shown. Continue to insert "010|Beijing" into table T. During the insertion process, it is found that the same data already exists in P7, so the original data in P7 needs to be processed. Create a new data PAGE P9, save the newly inserted data in P9, then save the same data queried from P7 in P9 as well, and modify the position in P7 that originally stored the same data to store a pointer pointing to P9. After the processing is completed, the memory structure is as Figure 13 shown.
[0113] Continue to insert "010|Beijing" into table T. When the space in P9 is insufficient to insert this data, it is necessary to continue to add a new data PAGE P10. P10 and P9 form a linked list to store the same data. The structure in memory is as Figure 14 shown. After all the data is stored, the remaining data is saved in sequence according to the above process. When the transaction is committed, these PAGEs in memory are flushed to the disk for persistent storage. It should be noted that, as Figure 3 shown, the embodiment of the present invention adopts a hierarchical storage method. The expansion method from the first-level HASH table PAGE to the Nth-level HASH table PAGE is the same as the expansion method from the first-level HASH table PAGE to the second-level HASH table PAGE in Embodiment 3 of the present invention. The difference is that the HASH operation rules at each level are different, and the HASH bucket subscripts corresponding to the data are also different, which will not be elaborated here.
[0114] Embodiment 4:
[0115] As Figure 15 shown, it is a schematic architecture diagram of the dynamic HASH-based data storage device according to the embodiment of the present invention. The dynamic HASH-based data storage device of this embodiment includes one or more processors 21 and a memory 22. Among them, Figure 15 One processor 21 is taken as an example here.
[0116] The processor 21 and the memory 22 can be connected through a bus or other means. Figure 15 Taking the connection through the bus as an example here.
[0117] The memory 22, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs and non-volatile computer-executable programs, such as the data storage method based on dynamic HASH in Embodiment 1. The processor 21 executes the data storage method based on dynamic HASH by running the non-volatile software programs and instructions stored in the memory 22.
[0118] The memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 22 optionally includes a memory remotely disposed relative to the processor 21, and these remote memories can be connected to the processor 21 through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0119] The program instructions / modules are stored in the memory 22 and, when executed by the one or more processors 21, execute the data storage method based on dynamic HASH in the above Embodiment 1. For example, execute each of the Figures 1 - 14 steps shown above.
[0120] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data storage method based on dynamic HASH, characterized in that, Including: Set the HASH operation rules for each level according to a preset method to ensure that different HASH values can be obtained for data through the HASH operation rules of each level; Create the first-level HASH table PAGE page in the database, use the first-level HASH operation rules to obtain the HASH value of the data corresponding to the first-level HASH operation rules, and obtain the bucket of the data corresponding HASH table PAGE page according to the HASH value of the first-level HASH operation rules; including: obtain the first-level HASH algorithm through the HASH operation rules of each level, and calculate the first-level HASH value of the data using the first-level HASH algorithm; use the first-level HASH value as the first-level KEY value of the data, perform a modulo operation on the number of buckets of the first-level HASH table PAGE page through the first-level KEY value to obtain the bucket of the data corresponding HASH table PAGE page; Determine the pre-stored PAGE page according to the address pointer type and the pointed object of the bucket corresponding to the data; including: if the address pointer is a null pointer, create a data PAGE page corresponding to the first-level HASH table PAGE page, and use the data PAGE page as the pre-stored PAGE page; if the address pointer is a non-null pointer and the address pointer points to the data PAGE page of the current-level HASH table, use the data PAGE page as the pre-stored PAGE page; if the address pointer is a non-null pointer and the address pointer points to the next-level HASH table PAGE page, obtain the next-level HASH table PAGE page pointed to by the address pointer; in the next-level HASH table PAGE page, calculate the HASH value of the data according to the HASH operation rules corresponding to the next-level HASH table, determine the bucket of the next-level HASH table PAGE page with this HASH value, and use the data PAGE page pointed to by the bucket of the next-level HASH table PAGE page as the pre-stored PAGE page; Determine the target storage PAGE page according to the free space size of the pre-stored PAGE page and whether it contains the target data identical to the currently inserted data, and perform the data storage operation.
2. The data storage method based on dynamic HASH according to claim 1, characterized in that, The determination of the target storage PAGE page according to the free space size of the pre-stored PAGE page and whether it contains the target data identical to the currently inserted data specifically includes: If the space of the pre-stored PAGE page is sufficient to store the data and the pre-stored PAGE page does not contain the target data, store the currently inserted data into the pre-stored PAGE page.
3. The data storage method based on dynamic HASH according to claim 1, characterized in that, The determination of the target storage PAGE page according to the free space size of the pre-stored PAGE page and whether it contains the target data identical to the currently inserted data further includes: If the space of the pre-stored PAGE page is not sufficient to store the data, restore the data in the pre-stored PAGE page to the data table; Create the second-level HASH table PAGE page, and point the address pointer of the bucket corresponding to the pre-stored PAGE page to the second-level HASH table PAGE page; Obtain the HASH value of the data not inserted in the data table through the second-level HASH operation rule, obtain the bucket of the HASH table PAGE page corresponding to the data according to the HASH value, and point the address pointer of the bucket to the target storage PAGE page; Store the data in the target storage PAGE page.
4. The data storage method based on dynamic HASH according to claim 1, characterized in that, The determining the target storage PAGE page according to the free space size of the pre-stored PAGE page and whether it contains target data identical to the currently inserted data further includes: If the target data is included in the pre-stored PAGE page, create a target storage PAGE page under the current-level HASH table; Store the currently inserted data and the target data in the target storage PAGE page; In the pre-stored PAGE page, modify the position storing the target data to the address pointer pointing to the target storage PAGE page.
5. The data storage method based on dynamic HASH according to claim 1, characterized in that, The determining the target storage PAGE page according to the free space size of the pre-stored PAGE page and whether it contains target data identical to the currently inserted data further includes: If the space of the pre-stored PAGE page is sufficient to store the data; and the pre-stored PAGE page does not contain the target data; Then determine whether an address pointer is stored in the pre-stored PAGE page; If an address pointer is stored in the pre-stored PAGE page, obtain the data stored in the first storage PAGE page pointed to by the address pointer, and determine whether the data is identical to the currently inserted data; If they are identical, store the currently inserted data in the first storage PAGE page; If they are not identical, store the currently inserted data in the pre-stored PAGE page.
6. The data storage method based on dynamic HASH according to claim 5, characterized in that, The storing the currently inserted data in the first storage PAGE page includes: Determine whether the first storage PAGE page has free space; If the first storage PAGE page has free space, store the currently inserted data in the first storage PAGE page; If the first storage PAGE page has no free space, create a second storage PAGE page, and establish an association between the first storage PAGE page and the second storage PAGE page to form a linked list for storing identical data; Store the currently inserted data in the second storage PAGE page.
7. A data storage device based on dynamic HASH, characterized in that, Includes: At least one processor; At least one memory; Wherein, the at least one processor and the at least one memory are communicatively connected to each other, the at least one memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the dynamic HASH-based data storage method provided in any one of claims 1-6.
8. A non-volatile computer storage medium, characterized in that, The computer storage medium stores computer-executable instructions, and the computer-executable instructions are executed by one or more processors to complete the dynamic HASH-based data storage method provided in any one of claims 1-6.
Citation Information
Patent Citations
Dynamic capacity expansion method and system of database table
CN107633097A
Entry data storage and query method, and device thereof
CN108255912A