File processing method and device based on segmentation, equipment and storage medium
By generating segmentation and merge queues in HBase and monitoring the load status in real time, the problem of poor quality of HBase segmentation and merge tasks is solved, and reasonable resource allocation and task success rate are improved.
Patent Information
- Application Number
- CN202410089148.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-22
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, the splitting and merging tasks of HBase have poor quality and low efficiency, and are easy to compete for resources with production tasks, resulting in task failure, data loss and merging progress uncontrollable.
By scanning the HBase partition storage unit directory, a queue is generated based on pre-set splitting and merging thresholds, and the load status of the HBase cluster is monitored in real time, triggering splitting and merging processing in a timely manner to avoid resource competition.
Improves the success rate and efficiency of slicing and merging tasks, avoids task failure and data loss caused by resource competition, and is suitable for Minor and Major Compaction.
Smart Images

Figure CN120353659A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular, to a file processing method, device, equipment, and storage medium based on splitting. Background Art
[0002] HBase is a distributed, column-oriented database, which is implemented based on HDFS (Hadoop Distributed File System) at the bottom layer and managed based on ZooKeeper (a reliable coordination system for distributed systems). Due to its good architecture design, HBase has demonstrated excellent performance in aspects such as rapid storage and random query of massive data. Data splitting (Spilt) and data compaction are the two most core functions of HBase. When the partition storage unit (Region) is too large, the partition storage unit is split; HFile is the file organization form for storing data in HBase. When the number of HFile files is too large, it will affect the efficiency of data query and also occupy too much metadata storage space, so the HFile files are merged. The main methods include presetting a merge period, judging the number of HFile files, judging the priority of HFile file sets, and manual merging, etc.
[0003] In the prior art, for the method of presetting a merge period, the merge period is configured in advance before merging. When the merge period arrives, the merge process is triggered and executed; for the method of judging the number of HFile files, after each flush is executed, the current number of HFile files is checked. If the total number exceeds the threshold, the merge process is executed; for the method of judging the priority of HFile file sets, the priority of each HFile subset is identified according to the priority algorithm. After the merge period arrives, the high-priority HFile subsets are merged first; for the method of manual merging, the merge process is manually executed by the operator according to the actual situation.
[0004] Therefore, the existing splitting or merging tasks have poor quality and low efficiency. Summary of the Invention
[0005] This application provides a file processing method, device, equipment, and storage medium based on splitting, so as to solve the technical problems of poor quality and low efficiency of splitting or merging tasks based on HBase.
[0006] In a first aspect, this application provides a file processing method based on splitting, including:
[0007] Scanning the partition storage unit directory of HBase to obtain the file metrics of the partition storage unit data in the partition storage unit directory, where the file metrics include file size and file quantity;
[0008] Compare the file metrics according to a preset splitting threshold to obtain a splitting queue, where the splitting queue is a queue of partition storage unit data waiting to be split;
[0009] Monitor the HBase cluster to determine the load status in the current HBase cluster, where the HBase cluster is used to maintain data of at least one partition storage unit;
[0010] Split the partition storage unit data in the splitting queue according to the load status in the HBase cluster to obtain initial target partition storage unit data.
[0011] In the embodiment of the present application, comparing the file metrics according to a preset splitting threshold to obtain a splitting queue includes:
[0012] Determine the partition storage unit data with a file size greater than the splitting threshold as the first target data;
[0013] Generate a splitting task according to the first target data;
[0014] Place the splitting task in the splitting queue.
[0015] In the embodiment of the present application, after splitting the partition storage unit data in the splitting queue according to the load status in the HBase cluster to obtain initial target partition storage unit data, the method further includes:
[0016] Determine the file metrics of the initial target partition storage unit data;
[0017] Compare the file metrics according to a preset merging threshold to obtain a merging queue, where the merging queue is a queue of initial target partition storage unit data waiting to be merged;
[0018] Determine the load status in the current HBase cluster according to the monitoring of the HBase cluster;
[0019] Merge the initial target partition storage unit data in the merging queue according to the load status in the HBase cluster to obtain target partition storage unit data.
[0020] In the embodiment of the present application, comparing the file metrics according to a preset merging threshold to obtain a merging queue includes:
[0021] Determine the partition storage unit data with a file quantity greater than the merging threshold as the second target data;
[0022] Generate a merging task according to the second target data;
[0023] Place the merge task in the merge queue.
[0024] In an embodiment of the present application, monitor the HBase cluster to determine the load status in the current HBase cluster, including:
[0025] Determine the load in the current HBase cluster;
[0026] When the load is less than the load threshold and the duration for which the load is less than the load threshold is greater than or equal to the time threshold, determine that the load status is the idle state; otherwise, determine that the load status is the busy state.
[0027] In an embodiment of the present application, after monitoring the HBase cluster to determine the load status in the current HBase cluster, the method further includes:
[0028] Continuously monitor the HBase cluster according to the load status in the HBase cluster being the busy state.
[0029] In an embodiment of the present application, during the process of splitting the partition storage unit data in the split queue, or during the process of merging the initial target partition storage unit data in the merge queue, the method further includes:
[0030] When the load status in the HBase cluster changes from the idle state to the busy state, pause the splitting process or the merging process and send an alarm prompt message.
[0031] In a second aspect, the present application provides a file processing device based on splitting, including:
[0032] A scanning module, configured to scan the partition storage unit directory to obtain the file size and the number of files of the partition storage unit data in the partition storage unit directory;
[0033] A comparison module, configured to compare the file metrics according to a preset splitting threshold to obtain a split queue, where the split queue is a queue of partition storage unit data waiting to be split;
[0034] A monitoring module, configured to monitor the HBase cluster to determine the load status in the current HBase cluster, where the HBase cluster is used to maintain at least one piece of partition storage unit data, and the load status is the idle state and the busy state;
[0035] A processing module, configured to split the partition storage unit data in the split queue according to the load status in the HBase cluster being the idle state to obtain the initial target partition storage unit data.
[0036] In a third aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0037] The memory stores computer-executable instructions;
[0038] The processor executes the computer-executable instructions stored in the memory to execute the file processing method based on segmentation of the present application.
[0039] In a fourth aspect, the present application provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the file processing method based on segmentation of the present application when executed by a processor.
[0040] The file processing method based on segmentation provided by the present application determines a segmentation queue according to the file metrics of Region data before the merging process, performs segmentation processing on the Region data first and then performs merging processing, solving the merging problem of HFile files generated for offline processing tasks. Moreover, during the merging process and the segmentation process, the load status in the HBase cluster is monitored in real time. When the load continuously remains below the load threshold for a period of time and the load status during this period is determined to be an idle status, the Region data in the segmentation queue or the merging queue is further segmented or merged, realizing timely and accurate judgment of the load condition of the HBase cluster, avoiding risks such as the failure of the merging process or the segmentation process, data loss, and uncontrollable progress caused by the competition for cluster resources between production tasks and operation tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0042] Figure 1 It is a schematic diagram of a scenario of the file processing method based on segmentation provided by an embodiment of the present application;
[0043] Figure 2 It is a schematic flowchart of a file processing method based on segmentation provided by an embodiment of the present application;
[0044] Figure 3 It is a schematic flowchart of another file processing method based on segmentation provided by an embodiment of the present application;
[0045] Figure 4 It is a schematic flowchart of a method for obtaining the load status in the HBase cluster provided by an embodiment of the present application;
[0046] Figure 5Schematic structural diagram of a file processing device based on segmentation provided by an embodiment of the present application;
[0047] Figure 6 Schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0048] Through the above drawings, specific embodiments of the present application have been shown, and more detailed descriptions will be given later. These drawings and written descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0049] Here, exemplary embodiments will be described in detail, and examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0050] In the prior art, a partition storage unit (Region) consists of several Stores (storages), and each Store contains a MemStore and multiple HFiles. When user data is written, it will first be written into the MemStore in the memory. When the size of the MemStore exceeds a set threshold, the system will write the data in the MemStore to disk to form a file in the HFile format.
[0051] The existing segmentation method only depends on directly segmenting based on the judgment of the size of the partition storage unit, which will result in the situation where the segmentation task and the production task compete for resources; moreover, the object processed by the existing segmentation method is the data in the MemStore written into the memory, and it cannot effectively process the HFile files in HDFS. The existing merging method is to preset a merging period in advance, including the method of judging the priority of the HFile file set, which can only relieve the competition for cluster resources and cannot fundamentally solve the problem that the merging task competes for cluster resources. Similarly, there will be a situation where the merging task and the production task of the cluster compete for resources, resulting in risks such as a large number of task failures, data loss, and uncontrollable merging progress; secondly, the object processed by the existing merging method is usually the HFile files flushed to HDFS, which is prone to the phenomenon of file accumulation; finally, the existing file merging method is suitable for selecting some small and adjacent files in location to merge into larger files, and is more suitable for Minor Compaction, but not suitable for Major Compaction.
[0052] To solve the above problems, the file processing method based on splitting provided by the embodiments of the present application monitors the HBase cluster, enabling the splitting and merging processes to be triggered in a timely manner according to the load conditions of the HBase cluster, achieving more reasonable allocation of cluster resources, avoiding problems such as a large number of task failures, data loss, and service unavailability caused by resource competition among splitting, merging, and normal production tasks, improving the utilization rate of cluster resources, the success rate of task execution, and the efficiency of merge task execution; moreover, first determining a splitting queue based on the file sizes of the partition storage unit data, splitting the partition storage unit data in the splitting queue, then determining a merging queue based on the number of files of the partition storage unit data, and merging the partition storage unit data in the merging queue, solving the problem of difficult merging due to large and numerous HFile files when bulk loading data into HDFS after offline processing, and avoiding the phenomenon of file accumulation.
[0053] The file processing method based on splitting provided by the present application aims to solve the above technical problems in the prior art.
[0054] The following uses specific embodiments to elaborate in detail on the technical solution of the present application and how the technical solution of the present application solves the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0055] Figure 1 It is a schematic diagram of the scenario of the file processing method based on splitting provided by the embodiments of the present application. As Figure 1 shown, the execution subject of the file processing method based on splitting provided by the embodiments of the present application is a server, such as a computer, etc., and the file processing method based on splitting is implemented using HBase technology.
[0056] Figure 2 It is a schematic flowchart of the file processing method based on splitting provided by the embodiments of the present application. As Figure 2 shown, it includes:
[0057] S201. Scan the partition storage unit directory of HBase to obtain the file metrics of the partition storage unit data in the partition storage unit directory, where the file metrics include file size and number of files.
[0058] Among them, the Region (partition storage unit) is the smallest unit for distributed storage and load balancing in HBase; the Region directory is part of the directory of HBase on HDFS (Hadoop Distributed File System); the Region data is the data contained in several Stores; the file metrics are parameters related to the basic configuration information of the Region data.
[0059] Among them, the file size refers to the actual file volume of the file. For example, if a file itself is 6 bytes in size, then these 6 bytes are called the file size of the file; the number of files refers to the number of files contained in the Region data.
[0060] S202. Compare the file metrics according to a preset splitting threshold to obtain a splitting queue, which is a queue of Region data waiting to be split.
[0061] Among them, the splitting threshold is a preset threshold for comparing file sizes; the splitting queue is a queue composed of Region data that meets the splitting threshold conditions.
[0062] It should be noted that comparing the file metrics according to a preset splitting threshold to obtain a splitting queue includes:
[0063] Determine the Region data with a file size greater than the splitting threshold as the first target data;
[0064] Generate a splitting task according to the first target data;
[0065] Place the splitting task in the splitting queue.
[0066] Among them, the first target data is Region data that meets the splitting threshold conditions.
[0067] Among them, the splitting task is to split the Region data.
[0068] After the task starts, by scanning the Region directory of HBase, collect the file metrics of the Region data, and add the splitting tasks generated by the Region data that meet the conditions to the splitting queue for queuing, waiting for subsequent splitting processing. At this time, to avoid conflicts between splitting processing and merging processing, during the scan, it will judge the Region data in the splitting queue or the merging queue. The same Region data can only appear in one of the splitting queue and the merging queue at the same time.
[0069] S203. Monitor the HBase cluster to determine the load status in the current HBase cluster. The HBase cluster is used to maintain the data of at least one partition storage unit.
[0070] Among them, an HBase cluster is generally composed of a Master (primary server) and multiple RegionServers (partition storage unit servers). Among them, the Master mainly assigns Regions to each RegionServer and is responsible for the load balancing of RegionServers; the RegionServer is mainly responsible for reading and writing Region data, responding to user I / O (Input / Output) requests, and is responsible for splitting Regions that become too large during operation.
[0071] Among them, the load status includes the idle state and the busy state. The idle state is the state when the load of the HBase cluster is lower than the load threshold condition and meets the time threshold condition; the busy state is the state when the load of the HBase cluster exceeds the load threshold condition or does not meet the time threshold. Among them, both the load threshold and the time threshold are pre-set thresholds.
[0072] Monitoring the HBase cluster can be achieved through prometheus, grafana, zabbix, Java development, etc. When the load of the HBase cluster is high, performing the splitting process will cause resource contention with normal production tasks. Therefore, in order to avoid a large number of production task failures and data loss due to resource contention, it is necessary to monitor the load of the HBase cluster in real time, including monitoring the load conditions of the CPU, memory, network IO, and disk of each host in the cluster. After determining the load status of the HBase cluster, corresponding processing is then performed according to the load status.
[0073] S204. Perform a splitting process on the partition storage unit data in the splitting queue according to the load status in the HBase cluster to obtain the initial target partition storage unit data.
[0074] Among them, the initial target partition storage unit data is the partition storage unit data after the splitting process in the splitting queue.
[0075] From the fact that the load status is the idle state, it can be seen that the current HBase cluster has sufficient resources. Therefore, performing a splitting process on the partition storage unit data in the splitting queue can avoid the risk of resource contention between the splitting task and the production task, making the operating conditions controllable, the process controllable, and the accuracy controllable.
[0076] It should be noted that when the load status in the HBase cluster is busy, the HBase cluster is continuously monitored.
[0077] Among them, when the load status in the HBase cluster is busy, no processing is performed on the data of the partition storage unit, and only the load in the HBase cluster is continuously monitored.
[0078] It should be noted that when the load status in the HBase cluster is idle and the data of the partition storage unit in the splitting queue is being split, if the load status in the HBase cluster changes from idle to busy, the splitting process is paused and an alarm prompt message is sent.
[0079] Among them, during the process of splitting the data of the partition storage unit, the load in the HBase cluster is in a changing state. Therefore, when it is monitored that the load status changes from idle to busy, the splitting task of the partition storage unit data and the subsequent processing flow are paused, an alarm prompt message is sent, and manual intervention is performed.
[0080] In summary, the file processing method based on splitting provided by the embodiments of the present application monitors the load conditions such as CPU, disk, and network IO of each host in the HBase cluster in real time. When it is found that the cluster load is low and lasts for a certain period of time, the splitting process is executed, avoiding the risk of the splitting task competing with the production task for cluster resources, making the operating conditions controllable, the process controllable, and the accuracy controllable, and improving the success rate of subsequent merging processing of the partition storage unit data.
[0081] On the basis of the above embodiments, this embodiment further provides an implementation manner of merging the initial target partition storage unit data after obtaining the initial target partition storage unit data.
[0082] Figure 3 It is a schematic flowchart of another file processing method based on splitting provided by the embodiments of the present application. As Figure 3 shown, the method includes the following steps:
[0083] S301. Determine the file metrics of the initial target partition storage unit data.
[0084] Among them, according to the file metrics obtained by scanning the partition storage unit directory of HBase in the above embodiments, the file metrics of the initial target partition storage unit data can be directly determined.
[0085] S302. Compare the file metrics according to the preset merging threshold to obtain a merging queue, where the merging queue is a queue of the initial target partition storage unit data waiting to be merged.
[0086] Among them, the merging threshold is a threshold preset for comparing the number of files; the merging queue is a queue composed of the data of the initial target partition storage units that meet the merging threshold condition.
[0087] It should be noted that by comparing the number of files of the data of the initial target partition storage units according to the preset merging threshold, a merging queue is obtained, including:
[0088] Determine the data of the partition storage units with the number of files greater than the merging threshold as the second target data;
[0089] Generate a merging task according to the second target data;
[0090] Place the merging task in the merging queue.
[0091] Among them, the second target data is the data of the initial target partition storage units that meet the merging threshold condition.
[0092] Among them, the merging task is to perform a merging process on the data of the initial target partition storage units.
[0093] After the splitting task is executed, the merging tasks generated from the data of the initial target partition storage units that meet the conditions are added to the merging queue for queuing, waiting for subsequent merging processing.
[0094] S303. According to the monitoring of the HBase cluster, determine the load status in the current HBase cluster.
[0095] By performing real-time monitoring on the HBase cluster, the merging tasks can be triggered in a timely manner according to the load status of the HBase cluster, which is beneficial to more reasonable allocation of cluster resources.
[0096] S304. According to the load status in the HBase cluster, perform a merging process on the data of the initial target partition storage units in the merging queue to obtain the data of the target partition storage units.
[0097] Among them, the data of the target partition storage units is the data of the initial target partition storage units after the merging process in the merging queue.
[0098] According to the fact that the load status in the HBase cluster is in an idle state, it can be known that the current HBase cluster has sufficient resources. Then, execute the merging tasks in the merging queue to avoid the risk of the merging tasks competing for resources with the production tasks of the HBase cluster, and ensure the completion quality and success efficiency of file merging.
[0099] It should be noted that during the process of merging the initial target Region data in the merging queue according to the idle load status in the HBase cluster, when the load status in the HBase cluster changes from the idle status to the busy status, the merging process is paused and an alarm prompt message is issued.
[0100] Among them, after the splitting task and the merging task are processed, the entire process enters a dormant state for a period of time to prepare for the next execution of splitting and merging.
[0101] In summary, the file processing method based on splitting provided by the embodiments of the present application can timely judge the load status of the current HBase cluster through the scanning and monitoring processes, so that the splitting task and the merging task can be triggered in a timely manner according to the current load status, avoiding risks such as task failure and data loss caused by resource contention with production tasks. Moreover, by introducing the splitting task before the merging task, it can solve the problem of difficult merging caused by the large scale and large quantity of HFile files imported into HDFS in batches for offline processing tasks, avoiding the phenomenon of file accumulation. At the same time, the present application is applicable not only to Minor Compaction but also to Major Compaction.
[0102] In the embodiments of the present application, there is also provided an implementation manner for determining the load status in the current HBase cluster by monitoring the HBase cluster.
[0103] Figure 4 It is a schematic flowchart of a process for obtaining the load status in the HBase cluster provided by the embodiments of the present application. As Figure 4 shown, the method includes the following steps:
[0104] S401. Determine the load in the current HBase cluster.
[0105] Among them, the load in the HBase cluster includes the load of CPU, disk, and network IO of each host.
[0106] S402. When the load is less than the load threshold and the duration for which the load is less than the load threshold is greater than or equal to the time threshold, determine that the load status is the idle status; otherwise, determine that the load status is the busy status.
[0107] Among them, both the load threshold and the time threshold are pre-set thresholds.
[0108] Figure 5 It is a schematic structural diagram of a file processing device based on splitting provided by the embodiments of the present application. As Figure 5As shown, the segmentation-based file processing device 500 includes: a scanning module 501, a comparison module 502, a monitoring module 503, and a processing module 504. The scanning module 501 is configured to scan the partition storage unit directory to obtain the file size and the number of files of the partition storage unit data in the partition storage unit directory.
[0109] The comparison module 502 is configured to compare file metrics according to a preset segmentation threshold to obtain a segmentation queue, and the segmentation queue is a queue of partition storage unit data waiting to be segmented.
[0110] The monitoring module 503 is configured to monitor the HBase cluster to determine the load status in the current HBase cluster. The HBase cluster is used to maintain at least one partition storage unit data, and the load status includes an idle state and a busy state.
[0111] The processing module 504 is configured to segment the partition storage unit data in the segmentation queue according to the idle load status in the HBase cluster to obtain initial target partition storage unit data.
[0112] Wherein, in the embodiment of the present application, the comparison module 502 further includes:
[0113] Determine the partition storage unit data with a file size greater than the segmentation threshold as the first target data;
[0114] Generate a segmentation task according to the first target data;
[0115] Place the segmentation task in the segmentation queue.
[0116] Wherein, in the embodiment of the present application, the segmentation-based file processing device 500 further includes:
[0117] Determine the file metrics of the initial target partition storage unit data;
[0118] Compare the file metrics according to a preset merging threshold to obtain a merging queue, and the merging queue is a queue of initial target partition storage unit data waiting to be merged;
[0119] Determine the load status in the current HBase cluster according to the monitoring of the HBase cluster;
[0120] Merge the initial target partition storage unit data in the merging queue according to the load status in the HBase cluster to obtain the target partition storage unit data.
[0121] Wherein, in the embodiment of the present application, the segmentation-based file processing device 500 further includes:
[0122] Determine the data of the partition storage unit with the number of files greater than the merging threshold as the second target data;
[0123] Generate a merging task according to the second target data;
[0124] Place the merging task in the merging queue.
[0125] Among them, in the embodiment of the present application, the monitoring module 503 further includes:
[0126] Determine the load in the current HBase cluster;
[0127] When the load is less than the load threshold and the duration of the load being less than the load threshold is greater than or equal to the time threshold, determine that the load status is the idle state; otherwise, determine that the load status is the busy state.
[0128] Among them, in the embodiment of the present application, the monitoring module 503 further includes:
[0129] Continuously monitor the HBase cluster according to the load status in the HBase cluster being the busy state.
[0130] Among them, in the embodiment of the present application, the file processing device 500 based on splitting further includes:
[0131] When the load status in the HBase cluster changes from the idle state to the busy state, pause the splitting process or the merging process and send an alarm prompt message.
[0132] In summary, the file processing device based on splitting provided by the embodiment of the present application, through the scanning module 501 and the monitoring module 503, enables timely and accurate judgment of the current load status of the HBase cluster, and then according to the load status, through the comparison module 502 and the processing module 504, the data of the partition storage unit is first split and then merged, solving the current situation where the splitting task or the merging task competes with the normal production task for cluster resources, effectively handling the problem of ignoring the splitting of HFile files falling on HDFS. At the same time, it solves the problem that the HFile files falling on HDFS after offline processing are large and numerous, resulting in difficult merging, avoiding the phenomenon of file accumulation, and improving the success efficiency of the splitting and merging tasks.
[0133] Figure 6 It is a schematic structural diagram of the electronic device provided by the embodiment of the present application. As Figure 6 shown, the electronic device 600 includes:
[0134] The electronic device 600 may include components such as a processor 601 with one or more processing cores, a memory 602 with one or more computer-readable storage media, a communication component 603, etc. Among them, the processor 601, the memory 602, and the communication component 603 are connected through a bus 604.
[0135] In a specific implementation process, at least one processor 601 executes the computer-executable instructions stored in the memory 602, so that at least one processor 601 executes the message processing method as described above.
[0136] For the specific implementation process of the processor 601, reference may be made to the above method embodiments. Their implementation principles and technical effects are similar, and will not be elaborated here in this embodiment.
[0137] In the above Figure 6 In the shown embodiment, it should be understood that the processor may be a central processing unit (English: Central Processing Unit, abbreviated: CPU), or other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated: DSP), application-specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with this application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0138] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0139] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.
[0140] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0141] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A file processing method based on segmentation, characterized in that, Including: Scanning the partition storage unit directory of HBase to obtain file metrics of the partition storage unit data in the partition storage unit directory, where the file metrics include file size and file quantity; Comparing the file metrics according to a preset splitting threshold to obtain a splitting queue, where the splitting queue is a queue of partition storage unit data waiting to be split; Monitoring the HBase cluster to determine the current load status in the HBase cluster, where the HBase cluster is used to maintain at least one piece of the partition storage unit data; Splitting the partition storage unit data in the splitting queue according to the load status in the HBase cluster to obtain initial target partition storage unit data.
2. The method according to claim 1, wherein The step of comparing the file metrics according to a preset splitting threshold to obtain a splitting queue includes: Determining the partition storage unit data with a file size greater than the splitting threshold as first target data; Generating a splitting task according to the first target data; Placing the splitting task in the splitting queue.
3. The method according to claim 1, wherein After splitting the partition storage unit data in the splitting queue according to the load status in the HBase cluster to obtain initial target partition storage unit data, the method further includes: Determining the file metrics of the initial target partition storage unit data; Comparing the file metrics according to a preset merging threshold to obtain a merging queue, where the merging queue is a queue of initial target partition storage unit data waiting to be merged; Determining the current load status in the HBase cluster according to the monitoring of the HBase cluster; Merging the initial target partition storage unit data in the merging queue according to the load status in the HBase cluster to obtain target partition storage unit data.
4. The method according to claim 3, characterized in that The step of comparing the file metrics according to a preset merging threshold to obtain a merging queue includes: Determining the partition storage unit data with a file quantity greater than the merging threshold as second target data; Generating a merging task according to the second target data; Placing the merging task in the merging queue.
5. The method according to claim 1, wherein The step of monitoring the HBase cluster to determine the current load status in the HBase cluster includes: Determining the current load in the HBase cluster; When the load is less than the load threshold and the duration for which the load is less than the load threshold is greater than or equal to the time threshold, determining the load status as the idle state, otherwise, determining the load status as the busy state.
6. The method according to claim 1, wherein After monitoring the HBase cluster to determine the current load status in the HBase cluster, the method further includes: Continuously monitoring the HBase cluster according to the load status in the HBase cluster being the busy state.
7. The method according to claim 1 or 3, characterized in that, During the process of splitting the partition storage unit data in the splitting queue or during the process of merging the initial target partition storage unit data in the merging queue, the method further includes: When the load status in the HBase cluster changes from the idle state to the busy state, suspend the splitting process or the merging process, and send an alarm prompt message.
8. A file processing device based on segmentation, characterized in that Including: A scanning module, configured to scan the partition storage unit directory to obtain the file size and the number of files of the partition storage unit data in the partition storage unit directory; A comparison module, configured to compare the file metrics according to a preset splitting threshold to obtain a splitting queue, where the splitting queue is a queue of partition storage unit data waiting to be split; A monitoring module, configured to monitor the HBase cluster to determine the current load status in the HBase cluster, where the HBase cluster is used to maintain at least one piece of the partition storage unit data, and the load status is the idle state and the busy state; A processing module, configured to split the partition storage unit data in the splitting queue according to the idle state of the load status in the HBase cluster to obtain initial target partition storage unit data.
9. An electronic device, characterized in that, Including: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 7.