HBase data cleaning method and device

By automatically searching and cleaning timeout data tables during file merging of the HBase cluster, the problem of excessive pressure on the master node is solved, the reliability and efficiency of data cleaning are improved, and the overall performance of the HBase cluster is improved.

CN113515509BActive Publication Date: 2025-05-06INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110538879.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-18
Publication Date
2025-05-06
Estimated Expiration
2041-05-18

AI Technical Summary

Technical Problem

Due to the frequent use of truncate for cleaning operations, the existing HBase data cleaning methods lead to excessive pressure on the master node master, affecting the overall performance of the HBase cluster.

Method used

During the file merging of the HBase cluster, the target data table with the current stored data duration exceeding the timeout threshold is automatically searched, and the target batch ID of the target data table is obtained, and the data with the target batch ID is cleaned from the target data table.

Benefits of technology

It effectively reduces the operating pressure of the master nodes in the HBase cluster, improves the reliability and efficiency of the HBase data cleaning process, and improves the overall performance and operation stability of the HBase cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113515509B_ABST
    Figure CN113515509B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides an HBase data cleaning method and device, which can be used in the field of big data technology. The method includes: if the HBase cluster is currently in the file merging period, then searching the target data table whose current storage time exceeds the timeout threshold in the HBase cluster, and obtaining the current target batch identifier of the target data table; cleaning the data with the target batch identifier from the target data table, wherein the target batch identifier is pre-added to the data operated by the user in the target data table. The present application can effectively reduce the operating pressure of the master node in the HBase cluster, and can improve the reliability and efficiency of the HBase data cleaning process, and improve the overall performance and operating stability of the HBase cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, in particular to the field of big data technology, and specifically to an HBase data cleaning method and device. Background Art

[0002] The architecture of the distributed column storage database HBase mainly consists of two parts: the master node Master and the slave node RegionServer. When it comes to operations such as adding, deleting, modifying, and querying tables, the master node Master is required to manage and transmit external requests.

[0003] Currently, the main method for cleaning data in HBase is to use the truncate operation to clear the HBase table. However, since the master node Master has only two instances, one active and one standby, it is impossible to expand the capacity. Therefore, when the truncate operation is used to frequently clear the HBase table, it is easy to cause a lot of pressure on the master node Master, thereby affecting the overall performance of HBase. Summary of the invention

[0004] In response to the problems in the prior art, the present application provides an HBase data cleaning method and device, which can effectively reduce the operating pressure of the master node in the HBase cluster, and can improve the reliability and efficiency of the HBase data cleaning process, and improve the overall performance and operating stability of the HBase cluster.

[0005] In order to solve the above technical problems, this application provides the following technical solutions:

[0006] In a first aspect, the present application provides an HBase data cleaning method, comprising:

[0007] If the HBase cluster is currently in the process of merging files, then the target data table whose current storage time exceeds the timeout threshold is searched in the HBase cluster, and the current target batch identifier of the target data table is obtained;

[0008] The data with the target batch identifier is cleaned from the target data table, wherein the target batch identifier is pre-added to the data in the target data table operated by the user.

[0009] Furthermore, if the HBase cluster is currently in the file merging period, before searching the HBase cluster for a target data table whose current storage data duration exceeds the timeout threshold and obtaining the current target batch identifier of the target data table, the method further includes:

[0010] Set the current expired file automatic cleanup parameter state of the HBase cluster to an executable state;

[0011] According to the pre-acquired correspondence between the data table identifier, the column family name and the timeout threshold, and the pre-acquired table creation configuration information, a configuration table with the timeout threshold added thereto is created in the HBase cluster.

[0012] Furthermore, it also includes:

[0013] Obtain a write request for a data table in the HBase cluster, wherein the write request includes a data table identifier, a data location identifier, write data, and a batch identifier;

[0014] According to the data position identifier and the write data, data writing processing is performed on the write position in the data table corresponding to the data table identifier, and a corresponding relationship between the batch identifier in the write request and the write data is added.

[0015] Furthermore, it also includes:

[0016] Obtain a read request for a data table in the HBase cluster, wherein the read request includes a data table identifier, a data location identifier, and a batch identifier;

[0017] The data at the reading position in the data table corresponding to the data table identifier is retrieved according to the data position identifier for user reading, and a corresponding relationship between the batch identifier in the read request and the reading position is added.

[0018] Furthermore, it also includes:

[0019] Receive and store an HBase write configuration table, wherein the HBase write configuration table is used to store the corresponding relationship between the data table identifier, data location identifier, write data, and batch identifier input by the user;

[0020] Correspondingly, obtaining a write request for a data table in the HBase cluster includes:

[0021] The newly added tasks in the HBase write configuration table are scanned periodically to obtain write requests for the data tables in the HBase cluster.

[0022] Furthermore, it also includes:

[0023] Receive and store an HBase read configuration table, wherein the HBase read configuration table is used to store the corresponding relationship between the data table identifier, the data location identifier, and the batch identifier input by the user;

[0024] Correspondingly, obtaining a read request for a data table in the HBase cluster includes:

[0025] The newly added tasks in the HBase read configuration table are scanned periodically to obtain read requests for the data tables in the HBase cluster.

[0026] Furthermore, if the HBase cluster is currently in the file merging period, before searching the HBase cluster for a target data table whose current storage data duration exceeds the timeout threshold and obtaining the current target batch identifier of the target data table, the method further includes:

[0027] Based on the preset file merging time and frequency, the HBase cluster is subjected to file merging processing.

[0028] In a second aspect, the present application provides an HBase data cleaning device, comprising:

[0029] The timeout verification module is used to search for a target data table whose current storage time exceeds the timeout threshold in the HBase cluster if the HBase cluster is currently in the process of file merging, and obtain the current target batch identifier of the target data table;

[0030] The batch cleaning module is used to clean the data with the target batch identifier from the target data table, wherein the target batch identifier is pre-added to the data in the target data table operated by the user.

[0031] In a third aspect, the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the HBase data cleaning method when executing the program.

[0032] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the HBase data cleaning method when executed by a processor.

[0033] It can be seen from the above technical solution that the present application provides an HBase data cleaning method and device, the method comprising: if the HBase cluster is currently in a file merge period, then searching the target data table whose current storage time exceeds the timeout threshold in the HBase cluster, and obtaining the current target batch identifier of the target data table; cleaning the data with the target batch identifier from the target data table, wherein the target batch identifier is pre-added to the data operated by the user in the target data table, and by automatically searching for the target data table whose current storage time exceeds the timeout threshold during the file merge of the HBase cluster for table clearing, the operating pressure of the master node in the HBase cluster can be effectively reduced, and the reliability and effectiveness of the HBase data cleaning process can be improved, and the overall performance and operating stability of the HBase cluster can be improved; by obtaining the current target batch identifier of the target data table, and cleaning the data with the target batch identifier from the target data table, the data to be cleaned can be processed in batches, without the need to search and clear the entire table, thereby further reducing the operating pressure of the master node in the HBase cluster, and effectively improving the efficiency of HBase data cleaning and the use efficiency of HBase. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0035] Figure 1 It is a schematic diagram of the interaction between the HBase data cleaning device in the embodiment of the present application and the client device and the HBase cluster respectively.

[0036] Figure 2 This is a first flow chart of the HBase data cleaning method in the embodiment of the present application.

[0037] Figure 3 It is a flowchart of step 010 and step 020 in the HBase data cleaning method in the embodiment of the present application.

[0038] Figure 4 It is a flowchart of step 310 and step 320 in the HBase data cleaning method in the embodiment of the present application.

[0039] Figure 5 It is a flowchart of step 410 and step 420 in the HBase data cleaning method in the embodiment of the present application.

[0040] Figure 6 It is a flowchart of step 030 and step 311 in the HBase data cleaning method in the embodiment of the present application.

[0041] Figure 7 It is a flowchart of step 040 and step 411 in the HBase data cleaning method in the embodiment of the present application.

[0042] Figure 8 This is a second flow chart of the HBase data cleaning method in the embodiment of the present application.

[0043] Fig. 9 It is a structural diagram of the HBase data cleaning device in an embodiment of the present application.

[0044] Fig.10 It is a flowchart of the time-cycle-based HBase data cleaning method in the application example of this application.

[0045] Fig.11 It is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0047] It should be noted that the HBase data cleaning method and device disclosed in this application can be used in the field of big data technology, and can also be used in any field other than the field of big data technology. The application field of the HBase data cleaning method and device disclosed in this application is not limited.

[0048] In view of the existing HBase data cleaning method, since it is necessary to frequently use the cleaning instruction truncate to frequently clean the table, which will cause excessive pressure on the master node Master, and thus affect the overall performance of the HBase cluster, the embodiments of the present application respectively provide an HBase data cleaning method, an HBase data cleaning device, and an electronic device computer-readable storage medium. If the HBase cluster is currently in a file merge period, a target data table whose current storage time exceeds the timeout threshold is searched in the HBase cluster, and the current target batch identifier of the target data table is obtained; data with the target batch identifier is cleaned from the target data table, wherein the target batch identifier is pre-added to the target data table and operated by the user. In the data being processed, by automatically searching for the target data table whose current storage time exceeds the timeout threshold for table clearing during the file merging of the HBase cluster, the operating pressure of the master node in the HBase cluster can be effectively reduced, and the reliability and effectiveness of the HBase data cleaning process can be improved, and the overall performance and operating stability of the HBase cluster can be improved; by obtaining the current target batch identifier of the target data table and cleaning the data with the target batch identifier from the target data table, the data to be cleaned can be processed in batches without searching and clearing the entire table, which can further reduce the operating pressure of the master node in the HBase cluster and effectively improve the efficiency of HBase data cleaning and the efficiency of HBase use.

[0049] In one or more embodiments of the present application, HBase (Hadoop Database) refers to a high-reliability, high-performance, column-oriented, scalable distributed storage system, which consists of a master node Master and slave nodes Region Server, where Master can also be specifically written as HMaster, and Region Server can also be written as HRegionServer or RegionServer, etc.

[0050] Based on the above content, the present application also provides an HBase data cleaning device for implementing the HBase data cleaning method provided in one or more embodiments of the present application, see Figure 1, the HBase data cleaning device can communicate and connect with the client device held by the user and the master node of the HBase cluster by itself or through a third-party server, etc. The HBase data cleaning device can be a server that receives the user's operation instructions or requests for the HBase cluster from the client device, and can also obtain relevant configuration files pre-set by the user from the client device, a third-party database or locally, such as at least one of the timestamp-based critical value TTL (Time to Live) configuration table, batch configuration table, cleaning frequency configuration table, HBase data storage parameter table, HBase write configuration table and HBase read configuration table mentioned in one or more embodiments of the present application. After cleaning the data with the target batch identifier from the target data table of the HBase cluster, the HBase data cleaning device can also send the corresponding data cleaning results to the client device for display, so that the user can promptly know the data cleaning results in the HBase cluster.

[0051] It is understandable that the client device may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.

[0052] The client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and other implementation scenarios may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster consisting of multiple servers, or a server structure of a distributed device.

[0053] The server and the client device may communicate with each other using any suitable network protocol, including network protocols that have not yet been developed on the date of filing this application. The network protocols may include, for example, TCP / IP, UDP / IP, HTTP, HTTPS, etc. Of course, the network protocols may also include, for example, RPC (Remote Procedure Call Protocol) and REST (Representational State Transfer) protocols used on top of the above protocols.

[0054] The details are described in detail through the following embodiments and application examples.

[0055] In order to solve the problem that the existing HBase data cleaning method requires frequent use of the cleaning instruction truncate to frequently clean the table, which will cause excessive pressure on the master node Master, thereby affecting the overall performance of the HBase cluster, etc., this application provides an embodiment of an HBase data cleaning method, see Figure 2 The HBase data cleaning method performed by the HBase data cleaning device specifically includes the following contents:

[0056] Step 100: If the HBase cluster is currently in the file merging period, then search the HBase cluster for a target data table whose current data storage duration exceeds a timeout threshold, and obtain the current target batch identifier of the target data table.

[0057] In step 100, the file merge may refer to a major compacting process of the HBase database; if the HBase cluster is currently in the file merge period, searching the HBase cluster for a target data table whose current storage time exceeds a timeout threshold means: utilizing an automatic cleanup mechanism in the major compacting process of the HBase database to clean up overdue data during the periodic major compacting process of the HBase database.

[0058] It is understandable that the target data table refers to the data table in the HBase cluster whose current storage time exceeds the timeout threshold. Specifically, the critical value TTL (Time to Live, used to limit the timeout of data) based on the timestamp can be cited, that is, the critical value TTL based on the timestamp is used as a specific example of the timeout threshold mentioned in one or more embodiments of the present application.

[0059] Step 200: Cleaning the data with the target batch identifier from the target data table, wherein the target batch identifier is pre-added to the data in the target data table operated by the user.

[0060] In step 200, the internal management mechanism in the HBase cluster will clean up expired data during the periodic major compaction of files; tables that need to be frequently cleaned up and truncated need to be cleaned up according to the set TTL life cycle; at the same time, the user automatically adds a batch identifier when writing and reading the table, and only operates on the data of the corresponding batch each time, and no longer performs a full table cleanup and truncation operation. The data that needs to be cleaned up is deleted according to the TTL setting when the HBase database is merged by major compaction.

[0061] From the above description, it can be seen that the HBase data cleaning method provided in the embodiment of the present application can effectively reduce the operating pressure of the master node in the HBase cluster, and can improve the reliability and effectiveness of the HBase data cleaning process, and improve the overall performance and operating stability of the HBase cluster by automatically searching for the target data table whose current storage time exceeds the timeout threshold during the file merging of the HBase cluster for table clearing. By obtaining the current target batch identifier of the target data table and cleaning the data with the target batch identifier from the target data table, the data to be cleaned can be processed in batches without searching and clearing the entire table, thereby further reducing the operating pressure of the master node in the HBase cluster and effectively improving the efficiency of HBase data cleaning and the utilization efficiency of HBase.

[0062] In order to improve the reliability and effectiveness of the timed and quantitative cleaning of overdue data in HBase during file merging, an embodiment of the HBase data cleaning method provided in this application is provided. Figure 3 , the HBase data cleaning method further includes the following contents before step 100:

[0063] Step 010: Set the current expired file automatic cleanup parameter state of the HBase cluster to an executable state.

[0064] Specifically, you can pre-set the file merging and automatic cleanup parameters of the cluster, log in to the HBase cluster to read the parameter value of the expired file automatic cleanup parameter HBase.store.delete.expired.storefile. When the parameter is set to true, no operation is required; when the parameter is set to false, change the parameter value to true to ensure that the HBase database automatically deletes expired data files when performing an internal major compact merge.

[0065] Step 020: Based on the pre-acquired correspondence between the data table identifier, column family name and timeout threshold, and the pre-acquired table creation configuration information, a configuration table with the timeout threshold added thereto is created in the HBase cluster.

[0066] In step 020, the table construction configuration information refers to the relevant information used to configure table construction in HBase, such as table name, column name (field name), compression format, storage format, etc., which is configured on demand according to user needs.

[0067] Specifically, you can pre-set the TTL configuration table: used to configure the column family and the corresponding TTL time when creating the table; the column family and TTL time are set according to user needs. The storage format of the table is as follows:

[0068] "Table name, column family name for which TTL is set, and TTL time (in seconds)".

[0069] And set the batch configuration table: used to configure the batch information added when reading and writing data in the HBase table. The batch information is set according to user needs. The storage format of the table is as follows:

[0070] "Table name, batch field name, total number of batches".

[0071] Then set the cleanup frequency configuration table: configure the time and frequency for HBase to perform major compact merges and clean up expired data files.

[0072] And the HBase data storage parameter table: used to configure the relevant information for building tables in HBase, including:

[0073] "Table name, column name (field name), compression format, storage format, etc. are configured on demand according to user needs."

[0074] Then after step 010, connect to the HBase cluster to create a table. By reading the TTL configuration table in the management configuration module, obtain the table name, column family name, and TTL time of the column family set by the user; by reading the HBase data storage parameter table, obtain the compression format, storage format and other table creation parameters set by the user, introduce the HBase create command in the Java code, log in to the cluster to complete the table creation and TTL setting operations.

[0075] From the above description, it can be seen that the HBase data cleaning method provided in the embodiment of the present application can effectively improve the reliability and effectiveness of the timed and quantitative cleaning of overdue data in HBase during file merging by setting the current automatic cleaning parameter state of the HBase cluster's expired files to an executable state and performing table creation processing, thereby further reducing the operating pressure of the master node in the HBase cluster and improving the overall performance and operating stability of the HBase cluster.

[0076] In order to improve the efficiency and effectiveness of cleaning the data with the target batch identifier in the subsequent data cleaning stage, an embodiment of the HBase data cleaning method provided in this application is provided. Figure 4 , before step 100 and after step 020 in the HBase data cleaning method, or after step 100 in the HBase data cleaning method, may also specifically include the following content:

[0077] Step 310: Obtain a write request for a data table in the HBase cluster, wherein the write request includes a data table identifier, a data location identifier, write data, and a batch identifier.

[0078] Step 320: According to the data location identifier and the write data, data writing processing is performed on the write location in the data table corresponding to the data table identifier, and a corresponding relationship between the batch identifier in the write request and the write data is added.

[0079] From the above description, it can be seen that the HBase data cleaning method provided in the embodiment of the present application can effectively improve the efficiency and effectiveness of cleaning the data with the target batch identifier in the subsequent data cleaning stage by adding the correspondence between the batch identifier in the write request and the write data in the process of writing data to the write position in the data table corresponding to the data table identifier, thereby further reducing the operating pressure of the master node in the HBase cluster and improving the overall performance and operating stability of the HBase cluster.

[0080] In order to improve the efficiency and effectiveness of cleaning the data with the target batch identifier in the subsequent data cleaning stage, an embodiment of the HBase data cleaning method provided in this application is provided. Figure 5 , before step 100 and after step 020 in the HBase data cleaning method, or after step 100 in the HBase data cleaning method, may also specifically include the following content:

[0081] Step 410: Obtain a read request for a data table in the HBase cluster, wherein the read request includes a data table identifier, a data location identifier, and a batch identifier.

[0082] Step 420: retrieve the data at the reading position in the data table corresponding to the data table identifier according to the data position identifier for the user to read, and add the corresponding relationship between the batch identifier in the read request and the reading position.

[0083] From the above description, it can be seen that the HBase data cleaning method provided in the embodiment of the present application can effectively improve the efficiency and effectiveness of cleaning the data with the target batch identifier in the subsequent data cleaning stage by adding the correspondence between the batch identifier in the read request and the read position in the data table corresponding to the data table identifier during the data reading processing, thereby further reducing the operating pressure of the master node in the HBase cluster and improving the overall performance and operating stability of the HBase cluster.

[0084] In order to provide a specific method for obtaining written data, an embodiment of the HBase data cleaning method provided in this application is provided. Figure 6 , the HBase data cleaning method further includes the following contents before step 100:

[0085] Step 030: Receive and store an HBase write configuration table, wherein the HBase write configuration table is used to store the correspondence between the data table identifier, data location identifier, write data, and batch identifier input by the user.

[0086] Specifically, you can pre-build and store the HBase write configuration table, which is used to store the business data written by users. Users configure the business data information to be written into the table, including the table name, column family name, column values, batch number, etc.

[0087] Correspondingly, step 310 in the HBase data cleaning method specifically includes the following contents:

[0088] Step 311: regularly scan the newly added tasks in the HBase write configuration table to obtain write requests for the data tables in the HBase cluster.

[0089] In step 311, the newly added tasks in the HBase write configuration table can be scanned regularly and the corresponding data write operation can be completed. Among them, writing data: receiving the scheduled task scanning module instruction, reading the HBase write configuration table, obtaining the business data and batch information that the user needs to write, introducing the HBase put command in the Java code, and logging in to the cluster to complete the business data and batch data write operation.

[0090] From the above description, it can be seen that the HBase data cleaning method provided in the embodiment of the present application, by applying the HBase write configuration table, can improve the convenience and efficiency of obtaining write requests for the data table in the HBase cluster, and thus can further improve the efficiency and effectiveness of cleaning the data with the target batch identifier in the subsequent data cleaning stage.

[0091] In order to provide a specific method for obtaining read data, an embodiment of the HBase data cleaning method provided in this application is provided, see Figure 7 , the HBase data cleaning method further includes the following contents before step 100:

[0092] Step 040: receiving and storing an HBase read configuration table, wherein the HBase read configuration table is used to store the corresponding relationship between the data table identifier, the data location identifier, and the batch identifier input by the user;

[0093] Specifically, you can pre-build and store the HBase read configuration table, which is used to store the business data that users need to read. Users configure the business data information to be read into the table, including the table name, column family name, column names, batch number, etc.

[0094] Correspondingly, step 410 in the HBase data cleaning method specifically includes the following contents:

[0095] Step 411: regularly scan the newly added tasks in the HBase read configuration table to obtain read requests for the data tables in the HBase cluster.

[0096] In step 411, the newly added tasks in the HBase read configuration table can be scanned regularly and the corresponding data read operation can be completed. Among them, reading data: receiving the scheduled task scanning module instruction, reading the HBase read configuration table, obtaining the table and batch information that the user needs to read, introducing the HBase scan command in the Java code, and logging in to the cluster to complete the specific batch of business data reading operations.

[0097] From the above description, it can be seen that the HBase data cleaning method provided in the embodiment of the present application, by applying the HBase read configuration table, can improve the convenience and efficiency of obtaining read requests for the data table in the HBase cluster, and thus can further improve the efficiency and effectiveness of cleaning the data with the target batch identifier in the subsequent data cleaning stage.

[0098] In order to improve the reliability of file merging processing for HBase cluster, an embodiment of the HBase data cleaning method provided in this application is provided. Figure 8 Step 100 of the HBase data cleaning method further specifically includes the following contents:

[0099] Step 050: Based on the preset file merging time and frequency, perform file merging processing on the HBase cluster.

[0100] In step 050, the cleaning frequency configuration table can be read, and according to the set file merging and cleaning time and frequency, the data cleaning module is regularly triggered to complete the cluster file merging and cleaning operations. That is, receiving the scheduled task scanning module instruction, triggering the major compact mechanism inside the HBase cluster, and completing the major compact merging and overdue data cleaning operations of the cluster.

[0101] From the above description, it can be seen that the HBase data cleaning method provided in the embodiment of the present application can improve the reliability of file merging processing for the HBase cluster by detecting whether the file merging cycle is reached based on the preset file merging time and frequency, thereby effectively improving the reliability of data cleaning for the HBase cluster and improving the overall performance and operation stability of the HBase cluster.

[0102] From the software level, in order to solve the problem that the existing HBase data cleaning method requires frequent use of the cleaning instruction truncate to frequently clean the table, which will cause excessive pressure on the master node Master, thereby affecting the overall performance of the HBase cluster, etc., the present application provides an embodiment of an HBase data cleaning device for executing all or part of the content of the HBase data cleaning method, see Fig. 9 , the HBase data cleaning device specifically includes the following contents:

[0103] The timeout verification module 10 is used to search for a target data table whose current storage time exceeds a timeout threshold in the HBase cluster if the HBase cluster is currently in a file merge period, and obtain the current target batch identifier of the target data table.

[0104] In the timeout verification module 10, the file merge may refer to the major compact process of the HBase database; if the HBase cluster is currently in the file merge period, searching the HBase cluster for a target data table whose current storage time exceeds the timeout threshold means: utilizing the automatic cleanup mechanism in the major compact process of the HBase database file to clean up the overdue data during the regular major compact process of the file merge.

[0105] The batch cleaning module 20 is used to clean the data with the target batch identifier from the target data table, wherein the target batch identifier is pre-added to the data in the target data table operated by the user.

[0106] In the batch cleaning module 20, the internal management mechanism in the HBase cluster will clean up the expired data during the periodic major compaction of files; the tables that need to be frequently cleaned up and truncated need to be cleaned up according to the set TTL life cycle; at the same time, the user automatically adds a batch identifier when writing and reading the table, and only operates the data of the corresponding batch each time, and no longer performs a full table cleanup and truncate operation. The data that needs to be cleaned up will be deleted according to the TTL setting when the HBase database is merged in major compaction.

[0107] The embodiment of the HBase data cleaning device provided in the present application can be specifically used to execute the processing flow of the embodiment of the HBase data cleaning method in the above embodiment. Its functions are not repeated here, and reference can be made to the detailed description of the above method embodiment.

[0108] From the above description, it can be seen that the HBase data cleaning device provided in the embodiment of the present application can effectively reduce the operating pressure of the master node in the HBase cluster, and can improve the reliability and effectiveness of the HBase data cleaning process, and improve the overall performance and operating stability of the HBase cluster by automatically searching for the target data table whose current storage time exceeds the timeout threshold during the file merging of the HBase cluster for table clearing processing; by obtaining the current target batch identifier of the target data table and cleaning the data with the target batch identifier from the target data table, the data to be cleaned can be processed in batches without searching and clearing the entire table, thereby further reducing the operating pressure of the master node in the HBase cluster, and effectively improving the efficiency of HBase data cleaning and the utilization efficiency of HBase.

[0109] In order to further illustrate the present solution, the present application also provides a specific application example of an HBase data cleaning method. Regarding the field of distributed column storage HBase database, when data cleaning operations need to be frequently performed on HBase tables, a time-based HBase data cleaning method is provided. Fig.10 , which solves the problem of excessive pressure on the master node Master caused by frequent table clearing in the existing distributed column storage database HBase. By avoiding frequent truncate operations, the function of reducing the pressure on the HBase master node Master is achieved.

[0110] The implementation principle of this solution mainly utilizes the automatic cleanup mechanism in the major compaction process of HBase database files: a timestamp-based critical value TTL (Time toLive, used to limit the timeout period of data) can be set for column families in HBase database, and its internal management mechanism will clean up expired data during the regular major compaction process of files; tables that need to be frequently truncated need to be cleaned up according to the set TTL life cycle; at the same time, batch identifiers are automatically added when users write and read the table, and only the corresponding batch of data is operated each time, and the truncate operation of the entire table is no longer performed. The data that needs to be cleaned up will be deleted according to the TTL setting when the HBase database major compaction is merged.

[0111] The above technical implementation is divided into the following modules:

[0112] 1. Cluster status setting module: This module is used to set the file merging and automatic cleanup parameters of the cluster. Log in to the HBase cluster to read the parameter value of HBase.store.delete.expired.storefile (delete expired files in the HBase cluster). When the parameter is set to true, no operation is required; when the parameter is set to false, change the parameter value to true to ensure that the HBase database automatically deletes expired data files when performing internal major compact merges.

[0113] 2. Management configuration module: This module is used to manage various configurations and states when creating tables, reading and writing data, and cleaning data in the HBase cluster, including the following functions:

[0114] (1) TTL configuration table: used to configure the column family and the corresponding TTL time when creating a table. The column family and TTL time are set according to user requirements. The storage format of this table is as follows:

[0115] Table name, column family name for setting TTL, TTL time (in seconds);

[0116] Such as: ABC_DEF, column_family_info, 86400.

[0117] (2) Batch configuration table: used to configure the batch information added when reading and writing data in the HBase table. The batch information is set according to user needs. The storage format of this table is as follows:

[0118] Table name, batch field name, total batch number;

[0119] For example: ABC_DEF, batch_info, 12.

[0120] (3) Cleanup frequency configuration table: configure the time and frequency for HBase to perform major compaction and clean up expired data files;

[0121] (4) HBase data storage parameter table: used to configure the relevant information for building a table in HBase, including table name, column name (field name), compression format, storage format, etc., which can be configured on demand according to user needs.

[0122] (5) HBase write configuration table: used to store business data written by users. Users configure the business data information to be written into the table, including table name, column family name, column values, batch number, etc.

[0123] (6) HBase read configuration table: used to store the business data that users need to read. Users configure the business data information to be read into the table, including the table name, column family name, column names, batch number, etc.

[0124] 3. Table creation module: This module is used to connect to the HBase cluster to create a table. By reading the TTL configuration table in the management configuration module, the table name, column family name, and TTL time of the column family set by the user are obtained; by reading the HBase data storage parameter table, the compression format, storage format and other table creation parameters set by the user are obtained, the HBase create command is introduced in the Java code, and the cluster is logged in to complete the table creation and TTL setting operations.

[0125] 4. Scheduled task scanning module: This module has the following two scheduled task functions:

[0126] (1) Regularly scan the new tasks in the HBase write configuration table and the HBase read configuration table, and trigger the read and write data module to complete the corresponding write and read data operations;

[0127] (2) Read the cleaning frequency configuration table, and based on the set file merging and cleaning time and frequency, periodically trigger the data cleaning module to complete the cluster file merging and cleaning operations.

[0128] 5. Data reading and writing module: This module is used to connect to the HBase cluster for data writing and reading operations.

[0129] (1) Writing data: Receive the scheduled task scanning module instruction, read the HBase write configuration table, obtain the business data and batch information that the user needs to write, introduce the HBase put command in the Java code, log in to the cluster to complete the business data and batch data writing operation;

[0130] (2) Reading data: Receive the scheduled task scanning module instruction, read the HBase reading configuration table, obtain the table and batch information that the user needs to read, introduce the HBase scan command in the Java code, and log in to the cluster to complete the business data reading operation of a specific batch.

[0131] 6. Data cleaning module: This module receives instructions from the scheduled task scanning module, triggers the major compact mechanism inside the HBase cluster, and completes the major compact merge and overdue data cleaning operations of the cluster.

[0132] Based on the above technical solution, the time-cycle-based HBase data cleaning method provided in the application example of this application reduces the pressure on the HBase master node Master by avoiding frequent cleaning truncate operations, effectively reduces the impact of cleaning truncate operations on HBase performance, and improves the utilization efficiency of the HBase column database.

[0133] From the hardware level, in order to solve the problem that the existing HBase data cleaning method needs to frequently use the cleaning instruction truncate to frequently clean the table, which will cause excessive pressure on the master node Master, thereby affecting the overall performance of the HBase cluster, etc., the present application provides an embodiment of an electronic device for implementing all or part of the contents of the HBase data cleaning method, and the electronic device specifically includes the following contents:

[0134] Fig.11 FIG. 9 is a schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. Fig.11 As shown, the electronic device 9600 may include a central processor 9100 and a memory 9140; the memory 9140 is coupled to the central processor 9100. It is worth noting that Fig.11 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.

[0135] In one embodiment, the HBase data cleaning function may be integrated into the central processing unit. The central processing unit may be configured to perform the following control:

[0136] Step 100: If the HBase cluster is currently in the file merging period, then search the HBase cluster for a target data table whose current data storage duration exceeds a timeout threshold, and obtain the current target batch identifier of the target data table.

[0137] In step 100, the file merge may refer to a major compacting process of the HBase database; if the HBase cluster is currently in the file merge period, searching the HBase cluster for a target data table whose current storage time exceeds a timeout threshold means: utilizing an automatic cleanup mechanism in the major compacting process of the HBase database to clean up overdue data during the periodic major compacting process of the HBase database.

[0138] It is understandable that the target data table refers to the data table in the HBase cluster whose current storage time exceeds the timeout threshold. Specifically, the critical value TTL (Time to Live, used to limit the timeout of data) based on the timestamp can be cited, that is, the critical value TTL based on the timestamp is used as a specific example of the timeout threshold mentioned in one or more embodiments of the present application.

[0139] Step 200: Cleaning the data with the target batch identifier from the target data table, wherein the target batch identifier is pre-added to the data in the target data table operated by the user.

[0140] In step 200, the internal management mechanism in the HBase cluster will clean up expired data during the periodic major compaction of files; tables that need to be frequently cleaned up and truncated need to be cleaned up according to the set TTL life cycle; at the same time, the user automatically adds a batch identifier when writing and reading the table, and only operates on the data of the corresponding batch each time, and no longer performs a full table cleanup and truncation operation. The data that needs to be cleaned up is deleted according to the TTL setting when the HBase database is merged by major compaction.

[0141] From the above description, it can be seen that the electronic device provided in the embodiment of the present application can effectively reduce the operating pressure of the master node in the HBase cluster, and can improve the reliability and effectiveness of the HBase data cleaning process, and improve the overall performance and operating stability of the HBase cluster by automatically searching for the target data table whose current storage time exceeds the timeout threshold for table clearing during the file merging of the HBase cluster; by obtaining the current target batch identifier of the target data table and clearing the data with the target batch identifier from the target data table, the data to be cleaned can be processed in batches without searching and clearing the entire table, thereby further reducing the operating pressure of the master node in the HBase cluster, and effectively improving the efficiency of HBase data cleaning and the utilization efficiency of HBase.

[0142] In another embodiment, the HBase data cleaning device may be configured separately from the central processing unit 9100. For example, the HBase data cleaning device may be configured as a chip connected to the central processing unit 9100, and the HBase data cleaning function may be implemented under the control of the central processing unit.

[0143] like Fig.11 As shown, the electronic device 9600 may also include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily have to include Fig.11 In addition, the electronic device 9600 may also include Fig.11 For components not shown, reference may be made to the prior art.

[0144] like Fig.11 As shown, the central processing unit 9100 is sometimes also referred to as a controller or an operation control, and may include a microprocessor or other processor device and / or logic device. The central processing unit 9100 receives input and controls the operation of various components of the electronic device 9600.

[0145] The memory 9140 may be, for example, one or more of a cache, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory or other suitable devices. The above-mentioned information related to the failure may be stored, and a program for executing the relevant information may also be stored. The CPU 9100 may execute the program stored in the memory 9140 to implement information storage or processing, etc.

[0146] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to provide power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display may be, for example, an LCD display, but is not limited thereto.

[0147] The memory 9140 may be a solid-state memory, such as a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It may also be a memory that saves information even when the power is off, can be selectively erased, and is provided with more data, examples of which are sometimes referred to as EPROMs, etc. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142, which is used to store application programs and function programs or processes for executing the operation of the electronic device 9600 through the central processor 9100.

[0148] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).

[0149] The communication module 9110 is a transmitter / receiver 9110 that sends and receives signals via an antenna 9111. The communication module (transmitter / receiver) 9110 is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as the case of a conventional mobile communication terminal.

[0150] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module and / or a wireless LAN module, etc. The communication module (transmitter / receiver) 9110 is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, thereby realizing a common telecommunication function. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor 9130 is also coupled to the central processor 9100, so that recording can be performed on the local machine through the microphone 9132, and the sound stored on the local machine can be played through the speaker 9131.

[0151] The embodiments of the present application also provide a computer-readable storage medium capable of implementing all the steps in the HBase data cleaning method in the above embodiments. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, all the steps in the HBase data cleaning method in the above embodiments are implemented by the execution subject being a server or a client. For example, when the processor executes the computer program, the following steps are implemented:

[0152] Step 100: If the HBase cluster is currently in the file merging period, then search the HBase cluster for a target data table whose current data storage duration exceeds a timeout threshold, and obtain the current target batch identifier of the target data table.

[0153] In step 100, the file merge may refer to a major compacting process of the HBase database; if the HBase cluster is currently in the file merge period, searching the HBase cluster for a target data table whose current storage time exceeds a timeout threshold means: utilizing an automatic cleanup mechanism in the major compacting process of the HBase database to clean up overdue data during the periodic major compacting process of the HBase database.

[0154] It is understandable that the target data table refers to the data table in the HBase cluster whose current storage time exceeds the timeout threshold. Specifically, the critical value TTL (Time to Live, used to limit the timeout of data) based on the timestamp can be cited, that is, the critical value TTL based on the timestamp is used as a specific example of the timeout threshold mentioned in one or more embodiments of the present application.

[0155] Step 200: Cleaning the data with the target batch identifier from the target data table, wherein the target batch identifier is pre-added to the data in the target data table operated by the user.

[0156] In step 200, the internal management mechanism in the HBase cluster will clean up expired data during the periodic major compaction of files; tables that need to be frequently cleaned up and truncated need to be cleaned up according to the set TTL life cycle; at the same time, the user automatically adds a batch identifier when writing and reading the table, and only operates on the data of the corresponding batch each time, and no longer performs a full table cleanup and truncation operation. The data that needs to be cleaned up is deleted according to the TTL setting when the HBase database is merged by major compaction.

[0157] From the above description, it can be seen that the computer-readable storage medium provided in the embodiment of the present application can effectively reduce the operating pressure of the master node in the HBase cluster, and can improve the reliability and effectiveness of the HBase data cleaning process, and improve the overall performance and operating stability of the HBase cluster by automatically searching for the target data table whose current storage time exceeds the timeout threshold during the file merging of the HBase cluster for table clearing. By obtaining the current target batch identifier of the target data table and clearing the data with the target batch identifier from the target data table, the data to be cleaned can be processed in batches without searching and clearing the entire table, thereby further reducing the operating pressure of the master node in the HBase cluster and effectively improving the efficiency of HBase data cleaning and the utilization efficiency of HBase.

[0158] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (devices), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the functions in the process. Figure 1 A process or multiple processes and / or boxes Figure 1These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0159] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0160] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A HBase data cleaning method, characterized in that: include: Set the current expired file automatic cleanup parameter status of the HBase cluster to executable state; According to the correspondence between the pre-acquired data table identifier, the column family name and the timeout threshold, and the pre-acquired table creation configuration information, a configuration table with the timeout threshold added thereto is created in the HBase cluster; If the HBase cluster is currently in the process of merging files, then the target data table whose current storage time exceeds the timeout threshold is searched in the HBase cluster, and the current target batch identifier of the target data table is obtained; Cleaning data with the target batch identifier from the target data table, wherein the target batch identifier is pre-added to the data in the target data table operated by the user; A write request for a data table in the HBase cluster is obtained, wherein the write request includes a data table identifier, a data location identifier, write data, and a batch identifier.

2. The HBase data cleaning method according to claim 1, characterized in that: Also includes: According to the data position identifier and the write data, data writing processing is performed on the write position in the data table corresponding to the data table identifier, and a corresponding relationship between the batch identifier in the write request and the write data is added.

3. The HBase data cleaning method according to claim 1, characterized in that: Also includes: Obtain a read request for a data table in the HBase cluster, wherein the read request includes a data table identifier, a data location identifier, and a batch identifier; The data at the reading position in the data table corresponding to the data table identifier is retrieved according to the data position identifier for user reading, and a corresponding relationship between the batch identifier in the read request and the reading position is added.

4. The HBase data cleaning method according to claim 3, characterized in that: Also includes: Receive and store an HBase write configuration table, wherein the HBase write configuration table is used to store the corresponding relationship between the data table identifier, data location identifier, write data, and batch identifier input by the user; Correspondingly, obtaining a write request for a data table in the HBase cluster includes: The newly added tasks in the HBase write configuration table are scanned periodically to obtain write requests for the data tables in the HBase cluster.

5. The HBase data cleaning method according to claim 4, characterized in that: Also includes: Receive and store an HBase read configuration table, wherein the HBase read configuration table is used to store the corresponding relationship between the data table identifier, the data location identifier, and the batch identifier input by the user; Correspondingly, obtaining a read request for a data table in the HBase cluster includes: The newly added tasks in the HBase read configuration table are scanned periodically to obtain read requests for the data tables in the HBase cluster.

6. The HBase data cleaning method according to any one of claims 1 to 5, characterized in that: If the HBase cluster is currently in the file merging period, before searching the HBase cluster for a target data table whose current storage data duration exceeds the timeout threshold and obtaining the current target batch identifier of the target data table, the method further includes: Based on the preset file merging time and frequency, the HBase cluster is subjected to file merging processing.

7. An HBase data cleaning device, characterized in that: include: The timeout verification module is used to set the current expired file automatic cleanup parameter status of the HBase cluster to an executable state; According to the correspondence between the pre-acquired data table identifier, the column family name and the timeout threshold, and the pre-acquired table creation configuration information, a configuration table with the timeout threshold added thereto is created in the HBase cluster; If the HBase cluster is currently in the process of merging files, then the target data table whose current storage time exceeds the timeout threshold is searched in the HBase cluster, and the current target batch identifier of the target data table is obtained; A batch cleaning module, used for cleaning data with the target batch identifier from the target data table, wherein the target batch identifier is pre-added to the data in the target data table operated by the user; A write request for a data table in the HBase cluster is obtained, wherein the write request includes a data table identifier, a data location identifier, write data, and a batch identifier.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the HBase data cleaning method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the HBase data cleaning method according to any one of claims 1 to 6 is implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.