HBase off-line data import optimization method based on key value pre-partitioning
By using MapReduce tasks to perform key-value pre-partitioning and global sorting during the offline data import process in HBase, and generating HFiles aligned with Region boundaries, the problem of resource waste and inefficiency caused by HFile splitting is solved, achieving efficient data import and fast queryability.
Patent Information
- Application Number
- CN202511311330.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-15
AI Technical Summary
In the existing technology, during the offline data import process of HBase, the mismatch between HFile and Region boundaries leads to resource waste and low efficiency. Especially when processing ultra-large data sets, the HFile splitting process consumes a lot of resources and affects the timeliness of data import.
By querying the metadata of the target HBase table, the key-value range of the Region is obtained. The data key-value of each generated HFile falls completely within the target Region range. MapReduce tasks are used for pre-partitioning and global sorting to generate HFiles aligned with the Region boundaries. These HFiles are then loaded into the corresponding Regions using a batch loading tool, bypassing the HBase write-ahead log path for file movement.
It effectively avoids HFile splitting operations, reduces disk I/O and CPU resource consumption, improves the efficiency of large-scale data import, shortens the latency from data import to queryability, ensures the stability and resource utilization of cluster services, and meets real-time analysis needs.
Smart Images

Figure CN120804106A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of big data storage and distributed computing, and particularly relates to a HBase offline data import optimization method based on key value pre-partitioning. BACKGROUND
[0002] With the evolution of data scale to PB or even EB level, distributed NoSQL database HBase has become a key component for real-time read and write of massive data. For data filling of HBase, generating HFile based on MapReduce framework and performing bulk loading (BulkLoad) is the mainstream offline import solution in the industry. The technical evolution of this solution aims to balance the concurrent overhead and data processing efficiency of the import task. The early solution starts an independent Reduce task for each HBase data partition (Region), which can ensure accurate correspondence between the generated file and the partition, but when the number of Regions is large, the time and resource overhead caused by the start and scheduling of Reduce tasks is huge, which limits the overall import performance. To solve this problem, the subsequent optimization introduces a merging mechanism, which allows multiple Regions to process data in a single Reduce task through a ratio parameter (ratio), significantly reducing the total number of Reduce tasks and reducing the task start overhead.
[0003] However, the above-mentioned merging optimization exposes a new performance bottleneck when processing super large scale data sets (such as tens of TB level). Since the data key value (RowKey) range processed by a single Reduce task may cross the boundaries of multiple Regions, the single HFile generated by it must be split (Split) by the HBase server when loaded into HBase, which is split and distributed to the correct Region. This splitting process involves complex data block migration, metadata update and storage space reallocation, which constitutes a high-load I / O intensive operation. This process not only consumes a large amount of cluster disk I / O and CPU resources, but also causes the Region to be temporarily unavailable during splitting, which seriously affects the timeliness of data import and may cause cluster load jitter, becoming a key obstacle to the efficiency of HBase offline data import.
[0004] CN117130561A discloses a data storage method, device and business system across Hbase cluster, which comprises, in the case of receiving data upload message, obtaining the position information, first time information and first identification information of the data to be stored, encoding the position information, first time information and first identification information to obtain the index number, reversing the index number to obtain the row key value of the data to be stored, storing the data to be stored into the first target partition in the target server according to the row key value, and the first target partition is the storage partition corresponding to the byte of the corresponding first identification information in the row key value. Although this scheme mentions key value partitioning, it does not consider the alignment of Region boundary HFile and the splitting process of HFile.
[0005] CN120277045A discloses a data processing method, device and electronic equipment based on a distributed file model, which comprises, obtaining source data, determining the current job type, configuring the target parameters of the data warehouse and the parameters of the memory computing engine, filtering the source data to obtain data fields, establishing the mapping relationship between the HBase data table and the data fields, and then processing the data file to form the original distributed data set, obtaining the storage unit interval in the HBase data table, and re-partitioning the original distributed data set according to the storage unit interval to obtain the partitioned distributed data set, mapping the key value pairs in the partitioned distributed data set to the data storage basic unit of HBase, inputting the source data into the data storage basic unit, marking the data storage basic unit to obtain the marked storage basic unit, and executing the write or delete operation on the source data. This scheme mainly solves the problem of high response delay caused by low system throughput when the data volume is large, and does not achieve the effect of offline import and shortening query delay.
[0006] Therefore, there is an urgent need for an import optimization method that can fundamentally avoid HFile splitting while taking into account task concurrency efficiency. SUMMARY
[0007] This section aims to summarize some aspects of the embodiments of the present application and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of the specification to avoid obscuring the purpose of this section, abstract and title, and such simplifications or omissions cannot be used to limit the scope of the present application.
[0008] In view of the above-mentioned existing problems, the present application is proposed. Therefore, the present application provides an HBase offline data import optimization method based on key value pre-partitioning, which is used to solve the problem of resource waste and low efficiency caused by HFile splitting when loading due to the mismatch between HFile and Region boundary in the prior art.
[0009] To solve the above technical problems, the application provides the following technical scheme: a HBase offline data import optimization method based on key value pre-partitioning, comprising: querying metadata of a target HBase table to obtain key value ranges of all regions of the table; based on the obtained key value ranges of the regions, processing source data and generating a plurality of HFiles, so that all data key values in each generated HFile are completely within a corresponding target region key value range; loading the generated plurality of HFiles into their respective corresponding regions by calling a batch loading tool of the HBase.
[0010] As a preferred scheme of the HBase offline data import optimization method based on key value pre-partitioning, wherein: the metadata of the target HBase table is obtained by accessing a system table storing region distribution information.
[0011] As a preferred scheme of the HBase offline data import optimization method based on key value pre-partitioning, wherein: the process of processing source data is completed by executing a MapReduce task.
[0012] As a preferred scheme of the HBase offline data import optimization method based on key value pre-partitioning, wherein: the MapReduce task includes a Map phase, through which the source data is read and globally sorted according to HBase row keys determined from each source data record.
[0013] As a preferred scheme of the HBase offline data import optimization method based on key value pre-partitioning, wherein: the MapReduce task includes a Reduce phase, in which a single Reduce task receives a sorted data stream and creates and generates an independent HFile for each target region covered by the data stream within the task in parallel.
[0014] As a preferred scheme of the HBase offline data import optimization method based on key value pre-partitioning, wherein: the Reduce phase comprises: parallel generation of HFiles is realized by maintaining a mapping relationship between a region identifier and an HFile writer; for each received data, first determine the region identifier corresponding to its row key, and then use the HFile writer associated with the region identifier in the mapping relationship for writing. If the corresponding HFile writer does not exist in the mapping relationship, the writer is created first and then the writing is performed.
[0015] As a preferred scheme of the HBase offline data import optimization method based on key-value pre-partitioning, the number of Reduce tasks started in the MapReduce task is determined according to the total number of Regions of the target HBase table and a preset proportion parameter.
[0016] As a preferred scheme of the HBase offline data import optimization method based on key-value pre-partitioning, the MapReduce task further performs a partition step between the Map phase and the Reduce phase, and the partition step sends data with continuous row keys to the same Reduce task.
[0017] As a preferred scheme of the HBase offline data import optimization method based on key-value pre-partitioning, the generated HFiles are loaded into the respective corresponding Regions, and the loading process is performed by bypassing the pre-write log path of HBase.
[0018] As a preferred scheme of the HBase offline data import optimization method based on key-value pre-partitioning, the method further includes: constructing a data structure supporting interval query according to the obtained key-value range of the Region, and then quickly locating the Region to which the row key belongs.
[0019] Compared with the prior art, the application has the following beneficial effects: 1. By pre-generating HFiles aligned with Region boundaries in the Reduce phase of MapReduce, the HFile splitting operation necessary when data is loaded to the HBase server is fundamentally eliminated, the disk I / O and CPU resource consumption caused by data migration and metadata update is reduced, and the overall efficiency of large-scale data offline import is improved. 2. The instantaneous peak of cluster load caused by HFile splitting is avoided, the data loading process is smoothly transitioned to a lightweight HDFS file moving operation, the impact on the HBase cluster is reduced, the stability of the cluster service is ensured, and the overall resource utilization of the system is improved. 3. By bypassing the temporary unavailable time of the Region caused by HFile splitting, and combining the bypassing of the pre-write log path by the batch loading tool, near real-time visibility after data loading is achieved, the delay from data import to query is greatly shortened, the timeliness of data import is effectively enhanced, and the needs of real-time analysis and other business scenarios can be better met. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort. Among them: Figure 1 This is a general flow chart of an HBase offline data import optimization method based on key value pre-partitioning according to an embodiment of the present invention; Figure 2 This is an offline import flow diagram for existing HBase versions before 1.2.6, according to an embodiment of the present invention, based on a key value pre-partitioning HBase offline data import optimization method. Figure 3 This is an offline import flow diagram of the existing HBase version after 1.2.6 for the HBase offline data import optimization method based on key value pre-partitioning according to an embodiment of the present invention; Figure 4 A data import flow diagram of an HBase offline data import optimization method based on key value pre-partitioning according to an embodiment of the present invention; Figure 5 This is a data import flow chart of an HBase offline data import optimization method based on key value pre-partitioning according to an embodiment of the present invention; Figure 6 This figure shows the scenario test results of the HBase offline data import optimization method based on key value pre-partitioning according to an embodiment of the present invention. DETAILED DESCRIPTION
[0021] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, but not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0022] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0023] Second, the "one embodiment" or "an embodiment" referred to herein means a specific feature, structure, characteristic, or combination of features and characteristics described herein that can be included in at least one implementation of the present application. The various appearances of "in one embodiment" or "in an embodiment" in the specification do not all refer to the same embodiment, although they can.
[0024] The present application is described in detail in conjunction with the schematic drawings, and in the detailed description of the embodiments of the present application, the sectional view of the device structure is partially enlarged without the general proportion for the convenience of illustration, and the schematic drawings are only examples, which should not limit the scope of protection of the present application herein. In addition, the three-dimensional spatial dimensions of length, width and depth should be included in actual production.
[0025] Meanwhile, in the description of the present application, it should be noted that the terms "upper, lower, inner and outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first, second or third" are only for the purpose of description, and cannot be understood as indicating or implying relative importance.
[0026] Unless otherwise expressly specified and limited, the terms "mounting, connecting, connection" in the present application should be understood broadly, for example: it can be fixed connection, detachable connection or integral connection; it can also be mechanical connection, electrical connection or direct connection, it can also be indirectly connected through intermediate medium, or it can be the internal communication of two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0027] Embodiment 1 Reference Figure 1 For the first embodiment of the present application, the embodiment provides a HBase offline data import optimization method based on key value pre-partitioning, comprising: S1, querying the metadata of the target HBase table to obtain the key value range of all Regions of the table; Further, the HBase client application program is initialized and connected to the target HBase cluster through the standard HBase configuration information (such as the ZooKeeper cluster address); the connection is the entrance for all interactions with the HBase cluster, and through the connection, a Table object instance representing the target data table can be obtained; from the Table object instance, a key interface instance-RegionLocator can be obtained; It needs to be explained that RegionLocator is a standard tool provided by HBase for locating Region information, which encapsulates all underlying details of accessing HBase metadata; Specifically, the core function of the RegionLocator interface is to query the system table storing Region distribution information, i.e., the hbase:meta table; and the query process is transparent to the user, i.e., when calling methods such as getAllRegionLocations() of RegionLocator, the RegionLocator interface will automatically initiate a scan operation to the hbase:meta table; Specifically, each row of the hbase:meta table records the metadata information of a Region of a user table; Furthermore, in the present application scheme, only by scanning the hbase:meta table, all rows related to the target table can be filtered out, and for each row, the following key information is extracted: Region Name: represents the unique identifier of the Region; Start Key: represents the lower bound of the RowKey of the data stored by the Region; it is a byte array, and for the first Region of the table, its start key is empty; End Key: represents the upper bound of the RowKey of the data stored by the Region; it is also a byte array, which defines the open interval boundary of the Region range; for the last Region of the table, its end key is empty; Furthermore, by traversing all entries about the target table in the hbase:meta table, a complete list containing all Region information of the table can be obtained; each element in the list defines the key value range of a Region, i.e. ; It needs to be noted that if a list containing all Region boundaries is directly used for query, it is inefficient; and in the traditional Reduce phase, since the data stream is continuous, the Reduce task needs to frequently judge whether the current processed RowKey crosses the Region boundary; if the entire list is traversed every time (i.e., the time complexity of linear scan is , N is the number of Regions), when the number of Regions reaches thousands or even tens of thousands, it will cause huge performance loss; therefore, the present application scheme constructs the information list of the key value range of the obtained Region into a data structure for interval query in the client memory; Specifically, the preferred construction scheme is obtained by using java.util.TreeMap or a functionally similar SkipList, and the specific construction process is as follows: A TreeMap<byte[], RegionInfo> instance is created, in which the key (Key) is the start key of each Region, and the value (Value) is an object containing the complete information (such as Start Key, End Key, etc.) of the Region. TreeMap will automatically sort all start keys according to the byte sequence. After sorting is completed, a data structure for interval queries is obtained; The TreeMap becomes a high-efficiency Region "navigation map". When the Reduce task needs to locate the Region to which any row key currentRowKey belongs, it only needs to call the floorEntry(currentRowKey) method of TreeMap. This method can quickly find the maximum start key less than or equal to the currentRowKey in logarithmic time complexity (O(log n)), and the start key corresponds to the Region where the currentRowKey is located; For example, assuming that we have RegionA [a, c) and RegionB [c, f), when the currentRowKey is d, floorEntry('d') will return the entry with the start key c, thereby immediately determining that the row key belongs to RegionB; Further, the constructed Region boundary data structure (such as TreeMap) that can be efficiently queried is serialized and distributed to each node executing the task through the MapReduce DistributedCache mechanism or job configuration, so that the Reduce task can be directly used in the local memory; S2, based on the key value range of the obtained Region, processing the source data and generating multiple HFiles, so that all data keys and values in each generated HFile fall completely within the key value range of a corresponding target Region; Further, the MapReduce task mainly includes three stages: the Map stage, the partitioning and sorting stage, and the Reduce stage; Further, for the Map stage, the stage is mainly responsible for reading the source data and converting it into the key value format required by HBase, that is, converting all source data records into the KeyValue format recognizable by HBase, and taking the row key as the basis to prepare for the next stage of partitioning and sorting; Specifically, a Map task reads the source data shard assigned to it, and for each source data record, the Mapper performs the following operations: According to the format of the source data (such as CSV, JSON or delimiter text), all fields required to constitute a record of HBase are parsed; then, from the parsed fields, the HBase row key of the record is constructed according to the business logic; the row key, column family, column qualifier, timestamp and value of the cell are encapsulated into one or more KeyValue objects of HBase; finally, the constructed KeyValue objects are output by the Mapper as its output; It should be noted that the HBase row key is the only basis for sorting and locating data in HBase; In addition, in order to take advantage of the sorting mechanism of MapReduce, we encourage the row key to be used as the key of the output and the KeyValue object to be used as the value of the output; It should be noted that between the Map phase and the Reduce phase, since MapReduce performs partitioning and sorting processes, the present application utilizes the global sorting idea similar to Hadoop TeraSort to perform global sorting according to the HBase row key determined from each source data record; Further, in order to ensure that the data with consecutive row keys can be sent to the same Reduce task, the present application also uses a partitioner based on the range of row keys, which determines the number of Reduce tasks according to the total number of Regions of the target HBase table and a user-configurable preset ratio parameter (ratio), and then divides the value range of the entire row key into several continuous intervals equal to the number of Reduce tasks. When a Map output key-value pair needs to be partitioned, the partitioner will determine which interval its row key belongs to and direct it to the Reduce task responsible for the interval, ensuring that the input data stream received by each Reduce task is not only strictly sorted by row key, but also covers a globally ordered and continuous segment of the row key range; Specifically, the formula for calculating the number of Reduce tasks is: Reduce task number = Region total number * ratio; It should be noted that since the above operation is performed, a continuous and ordered data stream is received by a single Reduce task, and the row key range of the data stream may fall within a Region or span the boundaries of multiple Regions. The goal of the Reduce task is to accurately divide and write the data stream into HFiles corresponding to each Region. The process of the Reduce task is as follows: At the beginning of the Reduce task, load and deserialize the Region boundary query data structure (e.g. TreeMap) built in step S1 from the distributed cache, and initialize a data structure (e.g. a HashMap<RegionID, HFile.Writer>) for maintaining the mapping between Regions and their corresponding HFile writers (HFile.Writer); The Reduce task processes each KeyValue object in its input data stream in order, and for each KeyValue object, performs the following logic: Extract the row key of the current KeyValue, and use the in-memory Region boundary query data structure (TreeMap) to quickly determine the target Region ID that the row key belongs to, in logarithmic time complexity, through query methods such as floorEntry(); Use the obtained Region ID as the key to look up the corresponding HFile.Writer instance in the HashMap; If the writer exists: it means that the current data and the previous data belong to the same Region, and the writer is directly used. If the writer does not exist, it means that the input data stream has crossed a Region boundary. At this time, the Reduce task will first check and close the writer of the previous Region (if it exists), and then create a new and independent HFile.Writer instance for the current new target Region, and specify a unique output path for the HFile on the HDFS that is related to the target Region (and its column family). The newly created writer is stored in the HashMap; Use the obtained or newly created writer to append the current KeyValue object to the corresponding HFile; When the Reduce task has processed all input data streams, traverse all remaining HFile writers in the HashMap and call their close() methods to ensure that all generated HFiles are correctly closed and persisted to the HDFS; It should be noted that through the execution process of each stage of the above MapReduce task, even if a single Reduce task processes data that spans multiple Regions, multiple independent HFiles can be intelligently generated, and the data within each HFile strictly belongs to the key-value range of a corresponding target Region, thereby achieving the pre-alignment of HFiles and Region boundaries; S3, loading the generated multiple HFiles into their respective corresponding Regions by calling the bulk loading tool of HBase; It should be noted that the step is intended to import the multiple HFiles aligned with the above Region boundary into the target HBase table, so that the data is visible to users, and the loading process is completed by calling the bulk loading (BulkLoad) tool provided by HBase officially, using the LoadIncrementalHFiles class, which can be called through command line or programming; Specifically, the specific implementation steps of the loading process are as follows: When the MapReduce task of S2 step is successfully executed, all generated HFiles are stored in a specified output directory of HDFS, at this time, the loading process is started, and two core parameters need to be provided to the LoadIncrementalHFiles tool; Specifically, the core parameters include the HFile source path: the path of the output directory of the MapReduce task on HDFS; target HBase table name: the name of the HBase table to which the data is imported; When the bulk loading tool is called, it does not directly perform data writing, but first connects to the HBase cluster and communicates with the HBase Master node. The loading tool will traverse all HFiles (i.e. HFile file list) under the source path, and submit the HFile file list to the Master; The Master node determines which Region on which RegionServer should "claim" each HFile according to the column family information of each HFile and the Region distribution information in the hbase: meta table managed by itself, and then the Master node distributes the loading instructions to each related RegionServer; The RegionServer receiving the instruction will perform the core operation of loading, which is not a traditional data writing, but a HDFS file moving operation. Specifically, the RegionServer will move the HFile assigned by it from the temporary output directory of HDFS to the corresponding column family storage directory under the HBase table directory of the Region. Because this process belongs to a file system level metadata operation, its speed will be much faster than transmitting data through the network and writing it piece by piece. And since it has been ensured in S2 step that the data of each HFile belongs to a single Region completely, the moving operation does not need any data level check or splitting, and can be completed directly; Once the file is moved successfully, the RegionServer updates the Region's metadata and adds the corresponding new HFile to its managed data file list. At this point, the data in the HFile is immediately visible to the client's read request. Furthermore, the entire loading process bypasses HBase's write-ahead log. In standard HBase write operations (such as put), data is first written to the in-memory MemStore and a log record is appended to the write-ahead log to prevent data loss. However, batch loading allows complete, persistent HFiles to be directly incorporated into Region management. Since the data already exists on HDFS, there's no need for redundancy protection through write-ahead logs. It should be noted that by bypassing the write-ahead log, a large amount of disk I / O overhead is avoided, the throughput of data import is improved, and the write pressure on the RegionServer is reduced. This not only makes the loading process faster, but also makes the load of the entire HBase cluster more stable during data import, minimizing the interference with online services.
[0028] Example 2 Reference Figures 2 to 6 , which is the second embodiment of the present invention, provides an HBase offline data import optimization method based on key value pre-partitioning, including: comparing the present invention solution with the existing HBase version before and after 1.2.6 offline solution; In versions prior to HBase 1.2.6, HBase data import relies on MapReduce. Each region starts a Map phase, where data preprocessing (sorting, reorganization, etc.) is completed. The processed data is then handed over to Reduce for processing. However, if a table has many regions, many Reduces will be started to process the data during data import. The Reduce startup phase takes a lot of time (mainly due to class loading and resource allocation), which greatly increases the total data import time. Figure 2 As shown; In the version after HBase-1.2.6, the performance problem caused by too much Reduce is optimized in the HBase source layer, a ratio parameter is added in the Importtsv tool class, so that the result data of Map is pre-merged before Reduce, and the bucket is introduced, so as to reduce the number of Reduce; for example, if a table has 10 regions, and the ratio is set to 0.5, then the entire MapReduce task starts 10 Maps and 5 Reduces, generates 5 HFile files, and imports them into the corresponding 10 regions, so that the number of Reduces is reduced by half, and the corresponding startup loss is reduced by half, as shown in Figure 3 Although the design can optimize the frequent startup of Reduce caused by too many partitions of the HBase table, since the number of HFile files does not match the number of regions, when processing a large amount of data (tens of TB, hundreds of TB), the HFile files need to be split into the same number of regions; therefore, the HFile splitting process also takes a long time (since the splitting of the region is essentially writing, it also occupies the disk IO resource, and the bottleneck of HBase is mainly on the disk IO); The scheme of the present application processes the source data through the Map stage, divides the tasks according to the preset number of regions, and globally sorts the RowKey; by establishing the mapping relationship of RowKey to Region; using the RegionLocator interface of HBase to obtain the Region distribution information in the hbase: meta table in real time, extracting the key value range of each Region, and using the skip list to realize Complexity interval query; according to the ratio setting ratio, Map can transmit data to the corresponding Reduce according to the mapping relationship; the Reduce divides the data into multiple groups according to the obtained Region boundary, and generates an independent HFile for each group (such as data spanning 2 regions, then output 2 HFiles), as shown in Figure 4 And Figure 5 ; In addition, for the versions before HBase-1.2.6, HBase-1.2.6 and the improved version of the scheme of the present application, the ratio of the above-mentioned versions is set to 0.5, and the time interval of the test results in 100GB / 200GB / 500GB three scenarios is as shown in Figure 6 As shown in Figure 6It can be clearly seen that when the data amount is small (100 GB), the method of the present application is close to the performance of HBase-1.2.6 version, and is superior to the versions before HBase-1.2.6, which reflects the advantage of reducing the Reduce startup overhead; as the data amount increases to 200 GB and 500 GB, the performance bottleneck of HBase-1.2.6 version is highlighted, and its time consumption increases sharply, especially in the 500 GB scenario, because of the huge I / O overhead caused by large-scale HFile splitting, its time consumption even exceeds the versions before HBase-1.2.6; in contrast, the method of the present application shows the optimal performance and good linear scalability under all data amount levels, because the HFile splitting is fundamentally avoided, the larger the data amount is, the more significant the performance advantage of the present application over HBase-1.2.6 version is; The test results fully prove that the method of the present application effectively solves the performance bottleneck problem caused by HFile splitting under large data amount in the prior art, and significantly improves the data import efficiency.
[0029] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program codes. The solutions in the embodiments of the present application can be implemented in various computer languages, such as object-oriented programming languages Java and interpreted scripting language JavaScript, etc.
[0030] The present application is described with reference to flowcharts and / or block diagrams according to the methods, devices (systems), and computer program products of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The device that implements the functions specified in one block or multiple blocks.
[0031] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the flow Figure 1 The flow or flows and / or blocks Figure 1 The flow or flows and / or blocks
[0032] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that execute on the computer or other programmable apparatus provide steps for implementing the flow Figure 1 The flow or flows and / or blocks Figure 1 The flow or flows and / or blocks
[0033] Although preferred embodiments of the application have been described herein, substitutions and alterations are possible in view of the disclosure of this application without departing from the spirit and scope of the present application. Therefore, it is the intent of the appended claims to cover all such variations and modifications as come within the scope of the present application. It is
[0034] It is apparent that a person skilled in the art can make various changes and modifications to the application without departing from the spirit and scope thereof. Therefore, if these modifications and changes fall within the scope of the claims and their equivalents, it is intended to include them in the application.
Claims
1. A method for optimizing HBase offline data import based on key value pre-partitioning, characterized in that: include: Query the metadata of the target HBase table and obtain the key value ranges of all regions in the table; Based on the key value range of the region, the source data is processed and multiple HFiles are generated, so that all data key values in each generated HFile fall completely within the key value range of a corresponding target region; By calling the HBase batch loading tool, the generated multiple HFiles are loaded into their respective corresponding Regions.
2. The HBase offline data import optimization method based on key value pre-partitioning according to claim 1 is characterized in that: The metadata of the query target HBase table is obtained by accessing the system table storing Region distribution information.
3. The HBase offline data import optimization method based on key value pre-partitioning according to claim 1 is characterized in that: The process of processing source data is completed by executing a MapReduce task.
4. The HBase offline data import optimization method based on key value pre-partitioning according to claim 3 is characterized in that: The MapReduce task includes a Map phase, in which source data is read and globally sorted according to the HBase row key determined from each source data record.
5. The HBase offline data import optimization method based on key value pre-partitioning according to claim 3 is characterized in that: The MapReduce task includes a Reduce phase. In the Reduce phase, a single Reduce task receives a sorted data stream and creates and generates an independent HFile in parallel for each target Region covered by its data stream within the task.
6. The HBase offline data import optimization method based on key value pre-partitioning according to claim 5 is characterized in that: The Reduce phase includes: The parallel generation of HFile is achieved by maintaining a mapping relationship between Region identifier and HFile writer; For each piece of data received, first determine the Region ID corresponding to its row key, and then use the HFile writer associated with the Region ID in the mapping relationship to write; If the corresponding HFile writer does not exist in the mapping relationship, create the writer first and then write.
7. The HBase offline data import optimization method based on key value pre-partitioning according to claim 5 is characterized in that: The number of Reduce tasks started in the MapReduce task is determined by the total number of Regions in the target HBase table and a preset ratio parameter.
8. The HBase offline data import optimization method based on key value pre-partitioning according to claim 3 is characterized in that: The MapReduce task further performs a partitioning step between the Map phase and the Reduce phase, and the partitioning step sends data with consecutive row keys to the same Reduce task.
9. The HBase offline data import optimization method based on key value pre-partitioning according to claim 1 is characterized in that: The generated multiple HFiles are loaded into their respective corresponding Regions, wherein the loading process is performed by bypassing the pre-write log path of HBase.
10. The HBase offline data import optimization method based on key value pre-partitioning according to claim 1, characterized in that: The method further includes: constructing a data structure supporting interval query based on the key value range of the obtained Region, thereby quickly locating the Region to which the row key belongs.
Citation Information
Patent Citations
Cross-Hbase cluster data storage method and device and service system
CN117130561A
Data processing method and device based on distributed file model and electronic equipment
CN120277045A
HBase database-based data batch loading method and device
CN105808577A
Reading-and-writing separation HBase warehousing method
CN105893521A
Data distribution method and device based on MapReduce as well as computer readable storage medium
CN108595268A