Data storage method and device

By preconfiguring the partition queue in Hbase and generating index subtables when the partition reaches the threshold, the IO resource consumption problem caused by the increase in Region size is solved, and the write operation performance is improved.

CN115480710BActive Publication Date: 2025-06-10AGRICULTURAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211282285.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-19
Publication Date
2025-06-10
Estimated Expiration
2042-10-19

AI Technical Summary

Technical Problem

In the Hbase distributed storage system, as the number of data table rows increases, the size of the region increases, resulting in the consumption of a large amount of IO resources when the region is divided into two new regions, and the write operation performance is reduced.

Method used

By preconfiguring the partition queue corresponding to the data table to which the data to be stored belongs, and generating an indexed partition table when the first partition size reaches the threshold, adding it to the index list of the data table, and deleting the first partition from the partition queue, avoiding data segmentation and interaction and saving IO resources.

Benefits of technology

Improves the data writing performance of Hbase, reduces the consumption of IO resources, and avoids the performance degradation caused by Region segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115480710B_ABST
    Figure CN115480710B_ABST
Patent Text Reader

Abstract

The present application discloses a data storage method and apparatus. By pre-configuring a partition queue of a data table to which data to be stored belongs in a distributed storage system Hbase, the target storage location of the data to be stored in Hbase is determined. Since the arrangement order of the partitions in the partition queue can represent the order of storing the data table, the data to be stored can be stored in a first partition at the head position of the partition queue. When the size of the first partition reaches a threshold, an index sub-table of the generated first partition is added to an index list for subsequent data search, and the first partition is deleted from the partition queue, so that subsequent data can be stored in a new first partition at the head position, without the need to split the stored data on the first partition and without data interaction between partitions, saving IO resources and improving the data writing performance of Hbase.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data storage, and more specifically, to a data storage method and device. Background Art

[0002] In the distributed storage system Hbase, a partition Region is the smallest unit of distributed storage, and different Regions can be distributed on different server nodes. During the process of using Hbase to store a data table, initially, the data table is stored in one Region. As the amount of data to be stored increases, specifically, as the number of rows in the data table continuously increases, the Region continuously grows. When the Region grows to a threshold value, it will be equally divided into two new Regions to meet the data storage requirements.

[0003] However, the operation of splitting a large Region into two small Regions consumes a large amount of IO resources, resulting in poor write operation performance of Hbase. Summary of the Invention

[0004] In view of the above problems, this application is proposed to provide a data storage method and device to improve the write operation performance of Hbase.

[0005] The specific solutions are as follows:

[0006] In a first aspect, a data storage method is provided, which is applied to the distributed storage system Hbase. The method includes:

[0007] According to a data storage request sent by a client, determine a partition queue corresponding to the data table to which the data to be stored belongs. The partition queue includes a plurality of partitions arranged in a certain order.

[0008] Store the data to be stored in the first partition at the head of the partition queue, and monitor the size of the first partition.

[0009] When the size of the first partition reaches a preset threshold, generate an index sub-table of the first partition, add the index sub-table to the index list of the data table, where the index sub-table is used to represent the corresponding relationship between the row keys of each row of data stored in the first partition and the first partition, and delete the first partition from the partition queue.

[0010] Optionally, the method further includes enabling the Bloom filter of Hbase.

[0011] The generating of the index sub-table of the first partition includes:

[0012] Merge all the stored files in the first partition and the Bloom filter in the memory to obtain the first Bloom filter of the first partition, where the first Bloom filter is used to determine whether the row key of the data to be searched belongs to the first partition;

[0013] Generate a binary tuple from the first partition and the first Bloom filter to obtain the index sub-table of the first partition.

[0014] Optionally, the method further includes:

[0015] In the case of receiving a data update request sent by the client, based on the data update request, determine the target data table to which the target data to be updated belongs and the target row key of the target data;

[0016] According to the target index list of the target data table, determine whether there is a target partition corresponding to the target row key;

[0017] If there is, use the target data to update the data corresponding to the target row key in the target partition;

[0018] If not, write the target data into the first partition corresponding to the target data table, where the first partition corresponding to the target data table is the partition at the head position in the partition queue corresponding to the target data table.

[0019] Optionally, the step of determining whether there is a target partition corresponding to the target row key according to the target index list of the target data table includes:

[0020] Use each Bloom filter in the target index list to determine whether the target row key belongs to the corresponding partition to obtain the partition where the determination hits;

[0021] In the partition where the determination hits, perform a primary search on the target row key. If the primary search hits, determine the partition where the primary search hits as the target partition;

[0022] If there is no partition where the determination hits or the primary search does not hit, perform a secondary search on the target row key in the first partition corresponding to the target data table, determine the partition where the secondary search hits as the target partition. If the secondary search does not hit, it indicates that there is no target partition corresponding to the target row key.

[0023] Optionally, the method further includes:

[0024] In the case of adding a new partition, insert the newly added partition into the partition queues of several data tables.

[0025] Optionally, the newly added partitions are multiple newly added partitions;

[0026] Inserting the newly added partitions into the partition queues of several data tables respectively includes:

[0027] Randomly inserting the multiple newly added partitions into any positions of the partition queues of several data tables respectively.

[0028] Optionally, the method further includes:

[0029] Monitoring the data volume of the server nodes to which each partition belongs, and deleting each partition on the server nodes with the data volume exceeding the preset value from each partition queue.

[0030] In a second aspect, a data storage device is provided, which is applied to a distributed storage system Hbase. The device includes:

[0031] A partition queue acquisition unit, configured to determine a partition queue corresponding to a data table to which data to be stored belongs according to a data storage request sent by a client, where the partition queue includes several partitions arranged in a certain order;

[0032] A data storage unit, configured to store the data to be stored into a first partition at the head of the queue in the partition queue;

[0033] A partition monitoring unit, configured to monitor the size of the first partition, and generate an index sub-table of the first partition and add the index sub-table to an index list of the data table when the size of the first partition reaches a preset threshold, where the index sub-table is used to represent the corresponding relationship between the row keys of each row of data stored in the first partition and the first partition, and delete the first partition from the partition queue.

[0034] Optionally, the device further includes:

[0035] A partition queue management unit, configured to insert the newly added partitions into the partition queues of several data tables respectively when new partitions are added.

[0036] Optionally, the device further includes:

[0037] A partition monitoring unit, configured to monitor the data volume of the server nodes to which each partition belongs, and delete each partition on the server nodes with the data volume exceeding the preset value from each partition queue.

[0038] With the above technical solution, in the distributed storage system Hbase of the present application, a partition queue corresponding to the data table to which the data to be stored belongs is preconfigured, and the arrangement order of the partitions in the partition queue can represent the storage sequence of the data table by these partitions, thereby determining the target storage location of the data to be stored in Hbase. Thus, the data to be stored can be stored in the first partition at the head position of the partition queue. That is to say, only one partition stores the data to be stored at the same time. When the size of the first partition reaches the threshold, an index sub-table of the first partition is generated, and by adding the index sub-table to the index list, the index list of the data table is gradually generated. Through the index list, when performing data search subsequently, the partition to which the data to be searched belongs can be determined according to the row key of the data to be searched. When the size of the first partition reaches the threshold, the first partition is also deleted from the partition queue, that is, no new data to be stored is written into the first partition, so that in the subsequent data storage process, new data to be stored can be stored in the new first partition at the head position of the partition queue, without splitting the data in the first partition and without data interaction between the current partition and the next partition, saving IO resources and improving the data writing performance of Hbase. Description of the Drawings

[0039] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0040] Figure 1 It is a schematic flow chart of a data storage method provided by an embodiment of the present application;

[0041] Figure 2 It is a schematic structural diagram of a data storage device provided by an embodiment of the present application. Detailed Embodiments

[0042] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0043] The present application provides a data storage method and apparatus, which can be applied to a distributed storage system Hbase composed of several server nodes to implement the storage of data tables and improve the data writing performance of Hbase. It should be noted that a partition Region is the smallest unit of distributed storage in Hbase. Different partitions can be distributed on different server nodes, and a partition cannot be split into multiple server nodes. In addition, the data in a partition all belong to a data table.

[0044] Figure 1 is a schematic flowchart of a data storage method shown according to an embodiment of the present application. In combination with Figure 1 as shown, the data storage method provided by the present application may include the following steps:

[0045] Step S101, determine the partition queue of the data table to which the data to be stored belongs according to the data storage request.

[0046] It should be noted that for each data table to be stored in Hbase, before storing the data table, a partition queue can be pre-configured for the data table. Exemplarily, the configuration of the partition queue can be implemented by a functional component for providing a consistency service for Hbase. The functional component can be ZooKeeper, ETCD, braft, etc.; the partition queue can include several partitions arranged in a certain order. The partitions in the partition queue are partitions that allow the data of the data table to be written, and the arrangement order of the partitions in the partition queue can represent the storage order of the data table. Specifically, the partition at the head of the partition queue is the partition currently being or about to store the data table, and the other partitions are partitions that may store the data table subsequently. In addition, a data table only has a corresponding partition queue, but during the storage process of the data table, it is allowed to modify or configure the partition queue, including but not limited to adjusting the arrangement order of the partitions in the partition queue, adding new partitions to the partition queue, or deleting partitions in the partition queue.

[0047] On the above basis, the step S101, determining the partition queue of the data table to which the data to be stored belongs according to the data storage request, may refer to configuring the partition queue of the data table in the case of initially storing the data table, and obtaining the partition queue of the data table in the current situation in the case of continuing to store the data table.

[0048] Step S102, store the data to be stored into the first partition in the partition queue.

[0049] Among them, the first partition is the partition at the head position in the partition queue. It should be noted that the first partition can store data through the following steps:

[0050] Write the data to be stored into the memory MemStore of the first partition, and monitor the size of the memory;

[0051] When the size of the memory reaches the threshold, flush the data in the memory to the storage file of the first partition.

[0052] Exemplarily, a data flush operation can be performed when the size of the memory reaches 64MB. It should be noted that one flush corresponds to one storage file, and after each flush, the data in the memory will be cleared to facilitate subsequent data writing. Using the memory size of the partition as the basis for performing the flush operation balances the sizes of the storage files of the partition and facilitates the generation of storage files with balanced sizes.

[0053] Step S103: Determine whether the size of the first partition exceeds a preset threshold. If so, execute step S104.

[0054] It should be noted that the current version of Hbase can monitor the size of the first partition. Specifically, the size of the first partition can be obtained by monitoring the sizes of all the storage files and the memory of the first partition.

[0055] Step S104: Add the generated index sub-table of the first partition to the index list, and delete the first partition from the partition queue.

[0056] Specifically, the above step S104 may include:

[0057] S01: Generate the index sub-table of the first partition.

[0058] Among them, the index sub-table can be used to represent the corresponding relationship between the row keys of each row of data stored in the first partition and the first partition.

[0059] S02: Add the index sub-table to the index list of the data table.

[0060] Among them, the index list of the data table can be pre-generated before storing the data table. The initial index list is empty. By continuously adding the index sub-tables of each partition during the storage process of the data table, an index list that can represent the corresponding relationship between the row keys of each row of data in the data table and the partitions storing each row of data is generated for subsequent data lookup.

[0061] It should be noted that for each data table, each row of data therein has its own corresponding row key. That is to say, a row key can uniquely identify a row of data in the data table. Therefore, the row key can be used as the basis for looking up data.

[0062] S03. Delete the first partition from the partition queue.

[0063] By deleting the first partition from the partition queue, the replacement of the first partition for data storage to be performed is achieved. Taking the size of the partition as the basis for replacing the first partition balances the size of the partitions, enabling each partition to store the data table evenly.

[0064] Optionally, it is also possible to monitor the partition queue of the data table through the client, so that the client obtains the partition at the head position of the partition queue, determines the first partition, and sends a data storage request to the first partition for the first partition to store the data to be stored corresponding to the data storage request.

[0065] The above data storage method determines the target storage location of the data to be stored in Hbase according to the partition queue corresponding to the data table to which the data to be stored belongs, which is pre-configured in the distributed storage system Hbase. Since the arrangement order of each partition in the partition queue represents the storage sequence of the data table, the data to be stored can be stored in the first partition at the head position of the partition queue. That is to say, only one partition performs the storage task of the data to be stored at the same time. When the size of the first partition reaches the threshold, an index sub-table of the first partition is generated, and by adding the index sub-table to the index list, the index list of the data table is gradually generated. Through the index list, when performing data lookup later, it is possible to determine the partition to which the data to be looked up belongs according to the row key of the data to be looked up. When the size of the first partition reaches the threshold, a partition deletion operation is also performed, that is, the first partition is deleted from the partition queue, and no new data to be stored is written into the first partition, so that in the subsequent data storage process, the new data to be stored can be stored in the new first partition at the head position of the partition queue, without splitting the data in the first partition and without data interaction between the current partition and the next partition, saving IO resources and improving the data writing performance of Hbase.

[0066] In some embodiments provided by the present application, the data storage method may further include enabling the Bloom filter of Hbase.

[0067] It should be noted that each storage file and memory of each partition of the Hbase have their own Bloom filters, and the Bloom filters can determine whether a certain row key belongs to the storage file / memory of the partition.

[0068] On the basis of the above, the above step S01, generating the index sub-table of the first partition, may include:

[0069] Merge the Bloom filters of all storage files and memories in the first partition to obtain the first Bloom filter of the first partition, and the first Bloom filter is used to determine whether the row key of the data to be searched belongs to the first partition;

[0070] Generate a binary tuple from the first partition and the first Bloom filter to obtain the index sub-table of the first partition.

[0071] The process of data routing for the data stored according to the data storage method provided by the present application will be described below. If you want to determine the partition to which the target data to be searched belongs, where the target data table to which the target data belongs is stored in Hbase according to the Figure 1 storage method shown, the following steps can be executed:

[0072] S11. Determine the target data table to which the target data belongs and the target row key of the target data.

[0073] S12. According to the target index list of the target data table, determine whether there is a target partition corresponding to the target row key.

[0074] Specifically, the above step S12 may include: making a first judgment to determine whether there is a partition corresponding to the target row key in the target index list of the target data table; in the case where the result of the first judgment is negative, making a second judgment to determine whether the first partition of the target data table corresponds to the target row key; the partition corresponding to the target row key determined by the first judgment and the second judgment is the target partition, and if the results of both the first judgment and the second judgment are negative, it means that Hbase does not store the target data corresponding to the target row key.

[0075] In addition, data routing can occur during the process of the client attempting to read data, or during the process of the client attempting to modify or update data. In some embodiments provided by the present application, the data storage method may further include:

[0076] Step S21. When receiving a data update request sent by the client, based on the data update request, determine the target data table to which the target data to be updated belongs and the target row key of the target data.

[0077] Step S22: According to the target index list of the target data table, determine whether there is a target partition corresponding to the target row key. If there is, execute Step S23; if not, execute Step S24.

[0078] It should be noted that the non-existence mentioned above may indicate that the data to be updated has not actually been stored in the Hbase. That is to say, Step S24 is actually the storage process of the data to be stored; the related description of Step S22 can refer to Step S12 above.

[0079] Step S23: Use the target data to update the data corresponding to the target row key in the target partition.

[0080] Step S24: Write the target data into the first partition corresponding to the target data. The first partition corresponding to the target data is the partition at the head position in the partition queue corresponding to the target data table.

[0081] Optionally, when performing data routing, the client can also determine the partition corresponding to the data to be searched according to the index list of the data table, and then send a data search request to the determined partition; if the client wants to modify or update data, it can send a data update request to the determined partition when the client determines the partition to which the data to be modified belongs, so that the determined partition can perform data update.

[0082] The following describes the data routing process when the Bloom filter of Hbase is enabled. Optionally, the determining whether there is a target partition corresponding to the target row key according to the target index list of the target data table may include:

[0083] Step S31: Use each Bloom filter in the target index list to determine whether the target row key belongs to the corresponding partition. If the determination hits, execute Step S32 for each partition where the determination hits; if the determination does not hit, execute Step S34.

[0084] It should be noted that using the Bloom filter can improve the performance of determining whether a row key belongs to a partition. However, the Bloom filter has a certain false positive rate. Therefore, the determination hit in Step S31 cannot determine the target partition corresponding to the target row key. It is also necessary to search for the target row key in each partition where the determination hits.

[0085] Step S32: Conduct a primary search for the target row key in the partition where the determination hits. If there is a partition where the primary search hits, execute Step S33; if there is no partition where the primary search hits, execute Step S34.

[0086] Step S33: Determine the partition hit in the initial search as the target partition.

[0087] Step S34: In the first partition corresponding to the target data table, search for the target row key again. If the re-search hits, execute Step S35; if the re-search misses, it indicates that there is no target partition corresponding to the target row key.

[0088] Wherein, the first partition corresponding to the target data table is the partition at the head position in the partition queue corresponding to the target data table.

[0089] Step S35: Determine the partition hit in the re-search as the target partition.

[0090] The above data storage method enables the Bloom filter of Hbase, so that in subsequent data routing, it can be determined whether a certain data belongs to a partition through the Bloom filter, improving the data routing performance.

[0091] In some embodiments provided by the present application, the data storage method may further include:

[0092] In the case of adding a new partition, insert the newly added partition into the partition queues of several data tables respectively.

[0093] The data storage solution provided by the present application, when adding a server node or a new partition in Hbase, does not require data splitting or data migration for the data stored in the existing partitions. Instead, by adding the newly added partition to the partition queues of each data table, for example, the newly added partition can be added to the head position of the partition queue of the data table, so that subsequent data can be stored in the newly added partition.

[0094] In a possible implementation, the newly added partition is multiple newly added partitions.

[0095] Based on the above, the inserting the newly added partition into the partition queues of several data tables respectively may include:

[0096] Randomly insert the multiple newly added partitions into any position in the partition queues of several data tables respectively.

[0097] Optionally, the multiple newly added partitions can be randomly arranged, and the arranged multiple newly added partitions are added to the head position of the partition queue. In addition, when configuring the partition queue for different data tables, a randomly generated arrangement order can also be used, so that the data of different data tables can be written into the partitions belonging to different server nodes to balance the data volume stored in each server node.

[0098] In some embodiments provided by the present application, the data storage method may further include:

[0099] Monitoring the data volume of the server nodes to which each partition belongs, and deleting each partition on the server node with the data volume exceeding the preset value from each partition queue.

[0100] To further balance the data volume stored in each server node, specifically, to avoid a too large data volume stored in individual server nodes, each partition on the server node with the data volume exceeding the preset value can be deleted from each partition queue, and no more data is stored in the server node with a larger data volume. That is to say, the partition queues of each data table stored in Hbase can be managed and controlled. Specifically, the arrangement order of the partitions in the partition queue can be adjusted, so as to realize the control of the target storage location of the data to be stored. Partitions can be added to the partition queue, or partitions in the partition queue can be deleted to balance the data volume stored in the server nodes. In addition, new server nodes are also allowed to be accessed, that is, by adding the partitions of the new nodes to the partition queue, the expansion of the server cluster is realized.

[0101] The data storage device provided by the embodiments of the present application will be described below. The data storage device described below can be correspondingly referred to the data storage method described above. The device can be applied to the distributed storage system Hbase.

[0102] See Figure 2 , Figure 2 which is a schematic structural diagram of a data storage device disclosed in the embodiments of the present application.

[0103] As Figure 2 shown, the device may include:

[0104] A partition queue acquisition unit 11, configured to determine a partition queue corresponding to the data table to which the data to be stored belongs according to a data storage request sent by a client, where the partition queue includes a plurality of partitions arranged in a certain order;

[0105] A data storage unit 12, configured to store the data to be stored in a first partition at the head of the queue in the partition queue;

[0106] A partition monitoring unit 13, configured to monitor the size of the first partition, and when the size of the first partition reaches a preset threshold, generate an index sub-table of the first partition, add the index sub-table to an index list of the data table, where the index sub-table is used to represent the corresponding relationship between the row keys of each row of data stored in the first partition and the first partition, and delete the first partition from the partition queue.

[0107] In some embodiments provided by the present application, the device may further include a Bloom filter opening unit for opening the Bloom filter of the Hbase.

[0108] Based on the above, the process of the partition monitoring unit 13 generating the index sub-table of the first partition may include:

[0109] Merge all the stored files and the in-memory Bloom filter in the first partition to obtain the first Bloom filter of the first partition, where the first Bloom filter is used to determine whether the row key of the data to be searched belongs to the first partition;

[0110] Generate a binary tuple from the first partition and the first Bloom filter to obtain the index sub-table of the first partition.

[0111] In some embodiments provided by the present application, the device may further include a data query unit and a data update unit;

[0112] The data query unit is configured to, when receiving a data update request sent by a client, based on the data update request, determine the target data table to which the target data to be updated belongs and the target row key of the target data, and determine whether there is a target partition corresponding to the target row key according to the target index list of the target data table;

[0113] The data update unit is configured to, when the data query unit determines that there is a target partition corresponding to the target row key, use the target data to update the data corresponding to the target row key in the target partition.

[0114] Based on the above, the data storage unit 12 may further be configured to, when the data query unit determines that there is no target partition corresponding to the target row key, write the target data into the first partition corresponding to the target data table, where the first partition corresponding to the target data table is the partition at the head position in the partition queue corresponding to the target data table.

[0115] In some embodiments provided by the present application, the process of the data query unit determining whether there is a target partition corresponding to the target row key according to the target index list of the target data table may include:

[0116] Use each Bloom filter in the target index list to determine whether the target row key belongs to the corresponding partition to obtain the partition where the determination hits;

[0117] In the partition where the determination hits, perform a primary search on the target row key. If the primary search hits, determine the partition where the primary search hits as the target partition;

[0118] If there is no partition where the judgment hits or the initial search misses, then in the first partition corresponding to the target data table, the target row key is searched again to determine the partition where the re-search hits as the target partition. If the re-search misses, it indicates that there is no target partition corresponding to the target row key.

[0119] In some embodiments provided by the present application, the device may further include a partition queue management unit, which is used to insert the newly added partition into the partition queues of several data tables when a new partition is added.

[0120] In some embodiments provided by the present application, the device may further include a partition monitoring unit, which is used to monitor the data volume of the server nodes to which each partition belongs, and delete each partition on the server nodes whose data volume exceeds the preset value from each partition queue.

[0121] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0122] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0123] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data storage method, characterized in that, applied to the distributed storage system Hbase, the method includes: According to the data storage request sent by the client, determine the partition queue corresponding to the data table to which the data to be stored belongs, and the partition queue includes several partitions arranged in a certain order; Store the data to be stored in the first partition at the head of the partition queue, and monitor the size of the first partition; When the size of the first partition reaches a preset threshold, generate an index sub-table of the first partition, add the index sub-table to the index list of the data table, where the index sub-table is used to represent the corresponding relationship between the row keys of each row of data stored in the first partition and the first partition, and delete the first partition from the partition queue; Wherein, the method further includes enabling the Bloom filter of the Hbase; The generating the index sub-table of the first partition includes: Merge all the stored files and the in-memory Bloom filter in the first partition to obtain the first Bloom filter of the first partition, and the first Bloom filter is used to determine whether the row key of the data to be searched belongs to the first partition; Generate a binary tuple from the first partition and the first Bloom filter to obtain the index sub-table of the first partition; Wherein, the method further includes: When receiving the data update request sent by the client, based on the data update request, determine the target data table to which the target data to be updated belongs and the target row key of the target data; According to the target index list of the target data table, determine whether there is a target partition corresponding to the target row key; If it exists, use the target data to update the data corresponding to the target row key in the target partition; If it does not exist, write the target data into the first partition corresponding to the target data table, and the first partition corresponding to the target data table is the partition at the head of the partition queue corresponding to the target data table.

2. The method according to claim 1, characterized in that, The determining whether there is a target partition corresponding to the target row key according to the target index list of the target data table includes: Use each Bloom filter in the target index list to determine whether the target row key belongs to the corresponding partition, and obtain the partition where the determination hits; In the partition where the determination hits, perform a primary search on the target row key. If the primary search hits, determine the partition where the primary search hits as the target partition; If there is no partition where the determination hits or the primary search does not hit, perform a secondary search on the target row key in the first partition corresponding to the target data table, determine the partition where the secondary search hits as the target partition. If the secondary search does not hit, it means that there is no target partition corresponding to the target row key.

3. The method according to any one of claims 1-2, characterized in that, The method further includes: When a new partition is added, insert the newly added partition into the partition queues of each of the several data tables.

4. The method according to claim 3, wherein, the newly added partitions are multiple newly added partitions; the inserting the newly added partitions into the partition queues of several data tables respectively includes: randomly inserting the multiple newly added partitions into any positions of the partition queues of several data tables respectively.

5. The method according to any one of claims 1-2, wherein, the method further includes: monitoring the data volume of the server nodes to which each partition belongs, and deleting each partition on the server nodes with the data volume exceeding the preset value from each partition queue.

6. A data storage device, wherein, applied to the distributed storage system Hbase, the device includes: a partition queue acquisition unit, configured to determine a partition queue corresponding to a data table to which data to be stored belongs according to a data storage request sent by a client, where the partition queue includes several partitions arranged in a certain order; a data storage unit, configured to store the data to be stored into a first partition at the head position in the partition queue; a partition monitoring unit, configured to monitor the size of the first partition, and generate an index sub-table of the first partition and add the index sub-table to an index list of the data table when the size of the first partition reaches a preset threshold, where the index sub-table is used to represent the corresponding relationship between the row keys of each row of data stored in the first partition and the first partition, and delete the first partition from the partition queue; wherein, the device further includes a Bloom filter enabling unit, configured to enable the Bloom filter of the Hbase; the process of the partition monitoring unit generating the index sub-table of the first partition includes: merging all storage files and the in-memory Bloom filter in the first partition to obtain a first Bloom filter of the first partition, where the first Bloom filter is used to determine whether the row key of the data to be searched belongs to the first partition; generating a binary tuple from the first partition and the first Bloom filter to obtain the index sub-table of the first partition; the device further includes a data query unit and a data update unit; the data query unit is configured to, when receiving a data update request sent by the client, determine a target data table to which target data to be updated belongs and a target row key of the target data based on the data update request, and determine whether there is a target partition corresponding to the target row key according to a target index list of the target data table; the data update unit is configured to, when the data query unit determines that there is a target partition corresponding to the target row key, update the data corresponding to the target row key in the target partition with the target data; the data storage unit is further configured to, when the data query unit determines that there is no target partition corresponding to the target row key, write the target data into a first partition corresponding to the target data table, where the first partition corresponding to the target data table is the partition at the head position in the partition queue corresponding to the target data table.

7. The device according to claim 6, wherein, the device further comprises: a partition queue management unit, configured to insert the newly added partition into the partition queues of several data tables when a new partition is added.

8. The device according to claim 6, wherein, the device further comprises: a partition monitoring unit, configured to monitor the data volume of the server nodes to which each partition belongs, and delete each partition on the server nodes whose data volume exceeds a preset value from each partition queue.

Citation Information

Patent Citations

  • DOT in-fragment secondary index method and DOT in-fragment secondary index system

    CN104133867A

  • A data deletion method and a distributed storage system

    CN109558065A