Data processing method and device based on distributed file model and electronic equipment
Through the data processing method based on the distributed file model, combined with the characteristics of the memory computing engine and HBase, efficient processing and storage of big data is achieved, solving the problem of low system throughput in the existing technology, resulting in high response delay, and improving the speed and accuracy of data processing.
Patent Information
- Application Number
- CN202510385146.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology has low system throughput in the case of large data volume, resulting in high response delay. The traditional HBase data deletion method uses large network resources, low throughput and high response delay. The table deletion and reconstruction method can easily affect the normal operation of the cluster.
The data processing method based on the distributed file model is adopted, and the data is filtered and repartitioned through the memory computing engine. Combined with the storage unit interval and marking mechanism of HBase, the data is accurately written and deleted, and the interaction with the server is reduced. BulkLoad and HFile incremental file loading interfaces are used for batch operations.
It improves the efficiency and accuracy of data processing, reduces the consumption of network resources and computing resources, significantly improves the speed and accuracy of big data processing, and provides strong support for data analysis and mining.
Smart Images

Figure CN120277045A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing based on a distributed file model. Specifically, it relates to a data processing method, device, computer-readable storage medium, and electronic device based on a distributed file model. Background Art
[0002] The system's big data cluster needs to obtain transaction data from the banking business system through EDB every day, complete the calculation of management report data in multiple dimensions through the job scheduling engine and Spark, and finally synchronize and store the data in the HBase cluster for subsequent analysis of the transaction link by management personnel. However, there are scenarios of upstream system data error correction and various special correction data within the system. To ensure data consistency, considering the multi-version and data append storage characteristics of HBase, it is necessary to perform a one-time large batch and custom deletion of the existing data, and then rewrite the data.
[0003] The native Delete method of HBase requires the client to initiate an RPC connection and send a single data deletion request to the server. The server will search and delete based on the RowKey of the single data. This method consumes a large amount of network resources, takes a long time on the server side, and has a low throughput. The existing batch Delete method, although it can utilize concurrency and batch send deletion requests, still needs to send a large number of RPC requests in the interaction with the server, and still uses the Delete mechanism of single data processing for deletion. In the case of a large amount of data, there are still problems of high response latency and low throughput. The method of deleting the table and reconstructing involves multiple Regions on multiple machines, which is very likely to cause the cluster to freeze and affect normal read and write requests. Summary of the Invention
[0004] The main purpose of this application is to provide a data processing method, device, computer-readable storage medium, and electronic device based on a distributed file model to at least solve the problem of high response latency caused by low system throughput when the data volume is large in the prior art.
[0005] To achieve the above object, according to one aspect of the present application, a data processing method based on a distributed file model is provided, including: obtaining source data, determining the current job type, and configuring the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine; screening the source data through the in-memory computing engine to obtain data fields, and establishing a mapping relationship between the HBase data table and the data fields; scanning the data files in the Hive data table according to the mapping relationship to form an original distributed data set, and obtaining the storage unit interval in the HBase data table at the same time, and repartitioning the partitions of the original distributed data set according to the storage unit interval to obtain a partitioned distributed data set; mapping each key-value pair in the partitioned distributed data set to the basic data storage unit of HBase, inputting the source data into the basic data storage unit, marking the basic data storage unit according to the current job type to obtain a marked storage basic unit, and performing write or delete operations on the source data according to the marked storage basic unit.
[0006] Optionally, before configuring the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine, the method further includes: configuring a function switch by using an XML configuration file, where the function switch is a write function switch or a delete function switch; submitting the function switch to the in-memory computing engine job.
[0007] Optionally, after determining the current job type, the method further includes: in the case where the current job type is writing data to the HBase data table, querying whether the HBase data table exists by using the HBase client interface; in the case where it is determined that the HBase data table does not exist, creating the HBase data table by using the HBase client interface according to the preset configuration information in the XML configuration file, where the configuration information includes the data table name, column family name, data expiration time, and partition interval.
[0008] Optionally, performing write or delete operations on the source data according to the marked storage basic unit and the current job type includes: writing the HFile file to the temporary directory of the distributed file system; using the BulkLoad of HBase and the HFile incremental file loading interface to load the HFile file in the temporary directory into the Region partition directory of HBase.
[0009] Optionally, repartition the original distributed data set according to the storage unit range to obtain a partitioned distributed data set, including: repartitioning the data of the original distributed data set according to the range of the storage unit range of the values of the Rowkey column in the Hive data table, and inputting the data into the corresponding partitioned distributed data set.
[0010] Optionally, after inputting the data into the corresponding partitioned distributed data set, the method further includes: sorting and deduplicating the data of the partitioned distributed data set according to the values of the Rowkey column, and generating a key-value pair sequence in the partitioned distributed data set.
[0011] Optionally, configure the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine, including: setting the job memory and the number of cores of the in-memory computing engine according to the data volume of the source data.
[0012] According to another aspect of the present application, there is provided a data processing device based on a distributed file model, including: a first configuration unit, configured to obtain source data, determine the current job type, and configure the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine; a screening unit, configured to screen the source data through the in-memory computing engine to obtain data fields, and establish a mapping relationship between the HBase data table and the data fields; a repartition processing unit, configured to scan the data files in the Hive data table according to the mapping relationship to form an original distributed data set, and obtain the storage unit range in the HBase data table, and repartition the original distributed data set partition according to the storage unit range to obtain a partitioned distributed data set; a marking unit, configured to map each key-value pair in the partitioned distributed data set to a basic data storage unit of HBase, input the source data into the basic data storage unit, and mark the basic data storage unit according to the current job type to obtain a marked storage basic unit, and perform a write or delete operation on the source data according to the marked storage basic unit.
[0013] According to still another aspect of the present application, there is provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the data processing methods based on the distributed file model.
[0014] According to another aspect of the present application, an electronic device is provided, including: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the data processing methods based on the distributed file model.
[0015] Applying the technical solution of the present application, source data is obtained, the current job type is determined, and the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine are configured; the source data is screened by the in-memory computing engine to obtain data fields, and at the same time, a mapping relationship between the HBase data table and the data fields is established; the data files in the Hive data table are scanned according to the mapping relationship to form an original distributed data set, and at the same time, the storage unit range in the HBase data table is obtained, and the original distributed data set partition is repartitioned according to the storage unit range to obtain a partitioned distributed data set; each key-value pair in the partitioned distributed data set is mapped to the basic data storage unit of HBase, the source data is input into the basic data storage unit, and the basic data storage unit is marked according to the current job type to obtain a marked basic data storage unit, and write or delete operations are performed on the source data according to the marked basic data storage unit. By combining the characteristics of the in-memory computing engine and HBase, efficient processing and storage of big data are realized. Through the screening process of the in-memory computing engine, unnecessary data processing is reduced, and the data processing efficiency is improved. At the same time, through the storage unit range repartitioning and marking mechanism of HBase, accurate writing and deletion of data in the case of a large amount of data are realized, overcoming the defects in the prior art that a large number of RPC requests need to be sent in the interaction between the system and the server, and the deletion and writing speeds of the Delete mechanism based on single data processing are slow. Data redundancy and errors are avoided, and the accuracy and efficiency of data processing are further improved. In practical applications, this solution can significantly improve the speed and accuracy of big data processing, provide strong support for data analysis and mining, have broad application prospects and significant economic benefits. It solves the problem of high response latency caused by low system throughput when the data volume is large in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The specification drawings forming a part of the present application are used to provide a further understanding of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0017] Figure 1 The hardware structure block diagram of a mobile terminal for executing a data processing method based on a distributed file model provided in an embodiment of the present application is shown;
[0018] Figure 2 The flowchart shows a data processing method based on a distributed file model provided according to an embodiment of the present application;
[0019] Figure 3 The flowchart shows the schematic process of data deletion principle provided according to an embodiment of the present application;
[0020] Figure 4 The flowchart shows the schematic process of a specific data processing method based on a distributed file model provided according to an embodiment of the present application;
[0021] Figure 5 The block diagram shows a data processing device based on a distributed file model provided according to an embodiment of the present application.
[0022] Among them, the above-mentioned drawings include the following reference numerals:
[0023] 102, processor; 104, memory; 106, transmission device; 108, input / output device; 51, first configuration unit; 52, screening unit; 53, re-partitioning processing unit; 54, marking unit. Detailed implementation manners
[0024] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0025] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of the present application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0027] For ease of description, some nouns or terms related to the embodiments of the present application are described below:
[0028] Hadoop: Hadoop is a big data distributed system basic framework provided by the Apache Foundation. This framework mainly includes HDFS (a distributed file system that can provide high-throughput access to application data) and YARN (a framework for job scheduling and cluster resource management). This framework allows for distributed processing of large datasets across computer clusters using a simple programming model. It aims to scale from a single server to thousands of machines, with each machine providing local computing and storage.
[0029] HBase: HBase is a distributed column storage system built on HDFS, a NoSQL database with high reliability, high performance, column-oriented, and scalable characteristics.
[0030] HFile: The underlying data file storage model of HBase stored on HDFS.
[0031] BulkLoad: A common way in HBase to write a large number of data files into a data table.
[0032] Rowkey: The row key of the HBase table, commonly used to index column values.
[0033] Cell: The basic data structure for storing data in HBase.
[0034] RegionServer: The data storage node of HBase.
[0035] Region: The partition storage unit of the data table distributed on the HBase data storage node.
[0036] Spark: A fast and general-purpose engine based on in-memory computing designed specifically for large-scale data processing.
[0037] SparkSQL: Spark SQL is the Spark module for processing structured data in Spark.
[0038] RDD: Resilient Distributed Dataset, an abstract structure in which Spark data is stored in memory.
[0039] Hive: Hive is a data warehouse tool based on Hadoop, used for data extraction, transformation, and loading. It is a mechanism for storing, querying, and analyzing large-scale data stored in Hadoop; it can map structured data files into a database table and provide SQL-like query functions.
[0040] EDB: Data Transmission Bus, a channel for transmitting a large amount of data between systems within an enterprise.
[0041] RPC: Remote Procedure Call.
[0042] As introduced in the background art, the existing technology has a low system throughput and high response latency in the case of a large amount of data. To solve the problem of low system throughput and high response latency when the amount of data in the existing technology is large, embodiments of the present application provide a data processing method, apparatus, computer-readable storage medium, and electronic device based on a distributed file model.
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention.
[0044] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking the operation on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal that executes a data processing method based on a distributed file model in an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0045] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the data processing method based on the distributed file model in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the mobile terminal through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0046] In this embodiment, a data processing method based on a distributed file model running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0047] Figure 2 It is a flowchart of the data processing method based on the distributed file model according to the embodiments of the present application. As Figure 2 shown, the method includes the following steps:
[0048] Step S201, obtain source data, determine the current job type, and configure the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine;
[0049] Among them, the current job type refers to the purpose of data processing, including but not limited to writing a large amount of data into the HBase data table and deleting HBase data;
[0050] Specifically, for both writing and deleting HBase data, it is necessary to specify target parameters such as the database, data table, and partition in the data warehouse Hive, and to pre-set resources such as the memory and number of cores of the Spark job according to the data volume of the source data, or to use the adaptive resource adjustment strategy of Spark.
[0051] Among them, the adaptive resource adjustment strategy of Spark refers to the technology of dynamically adjusting resource allocation during runtime to optimize performance and cluster resource utilization. This strategy can help address resource waste or performance degradation caused by issues such as data skew and task imbalance.
[0052] In the above step S201, by obtaining the source data, determining the current job type based on the obtained source data, and then configuring the target parameters of the data warehouse Hive and the parameters of the memory computing engine according to the current job type, through this parameter configuration mechanism combined with the memory computing characteristics of the memory computing engine Spark, the resources of the input data can be dynamically adjusted, and it can dynamically adapt to deleting, cleaning, or writing data of different scales, improving the robustness of job operation.
[0053] Step S202: Screen and process the above source data through the above memory computing engine to obtain data fields, and at the same time establish a mapping relationship between the HBase data table and the above data fields;
[0054] Specifically, through the SQL framework interface of the memory computing engine Spark, standard SQL statements can be used to customize the screening and processing of the Hive data source. Among them, SparkSQL can utilize the memory computing characteristics of Spark to quickly read and process the data in Hive. When dealing with a large amount of data, the data can be processed in the memory of Spark, reducing disk I / O and improving the data processing speed. At the same time, establish a mapping relationship between the columns of the HBase data table and the data fields after screening the Hive data source. The mapping relationship ensures data consistency.
[0055] Through the above step S202, when the data is screened and processed, Spark can directly convert the result data into the format required by HBase and efficiently write it into HBase through BulkLoad, avoiding the performance bottleneck of writing one by one. In addition, establishing a mapping relationship between the HBase data table and the above data fields ensures that the fields after screening the Hive data source can accurately correspond to the column families and columns of HBase. This mapping relationship enables the information to remain consistent during the data conversion process, avoiding data loss or confusion that may occur during the conversion between different systems and ensuring the integrity and accuracy of the data.
[0056] Step S203: Scan and process the data files in the Hive data table according to the above mapping relationship to form an original distributed dataset. At the same time, obtain the storage unit range in the above HBase data table, and perform re-partitioning processing on the partitions of the original distributed dataset according to the above storage unit range to obtain a partitioned distributed dataset;
[0057] Specifically, inside the Spark job, according to the information such as the Hive data table and the custom processing SQL configured in the previous steps, parallelly scan the data files of the Hive data table to form an original RDD (original distributed dataset). At the same time, obtain the Region partition range of the HBase data table, re-partition the original RDD, re-partition the data of the value corresponding to the Rowkey column in the Hive table according to the range of the Region interval into the corresponding Spark partition RDD. At the same time, sort and deduplicate the data of the partition RDD according to the value of the Rowkey, and form a key-value pair sequence of [(Rowkey + column family name + column name, column value)] inside each partition, so as to obtain a partitioned distributed dataset.
[0058] Rowkey is the key used to uniquely identify a row of data in HBase. The first step to ensure accurate data writing is to design a reasonable Rowkey to ensure that it can reflect the uniqueness of the data and the hot spot distribution of access. For example, the Rowkey may consist of a date, a business type, a user ID, etc., so as to quickly locate and access specific data. During the data processing, Spark sorts the data in the RDD according to the Rowkey, which ensures that the data can be effectively located according to the distribution of the Region when writing to HBase, reduces the overhead of cross-Region writing, and thus realizes more accurate data writing.
[0059] Applying the above step S203, through the distributed processing ability of Spark, the original distributed dataset can be scanned and processed in parallel, which means that the data processing tasks can be executed simultaneously on multiple nodes in the cluster, greatly improving the efficiency and speed of data processing. The data storage in HBase is based on the storage unit Region, and each Region contains a continuous Rowkey range. By obtaining the storage unit range of the data table in HBase, the distributed dataset in Spark can be repartitioned according to the distribution of the storage units, ensuring that the data in each partitioned distributed dataset can correspond to a specific Region in HBase. This optimized partitioning strategy reduces the data transfer between networks because the data can be written more directly into the corresponding HBase storage unit, thereby reducing the data write latency and improving the data write efficiency. In addition, using the HBase Region range for repartitioning processing can ensure that the data is written into HBase in the order of Rowkey, guaranteeing data consistency.
[0060] Step S204, map each key-value pair in the above partitioned distributed dataset to the basic data storage unit of HBase, input the above source data into the above basic data storage unit, mark the above basic data storage unit according to the above current job type to obtain a marked storage basic unit, and perform a write or delete operation on the above source data according to the above marked storage basic unit.
[0061] Before writing data into HBase, the data needs to be converted into the Cell data structure (basic data storage unit). A Cell contains information such as Rowkey, column family name, column name, column value, and timestamp. During the mapping process, the present invention marks each Cell with specific marks, such as write marks and delete marks. This not only ensures the correct storage of the data but also provides a basis for subsequent data deletion because when deleting data, the marks can be used to filter out the Cells that need to be retained or deleted.
[0062] Specifically, when the current job type is the data writing HBase data table type, the basic data storage unit is marked as written to obtain the written storage basic unit. According to the key-value pair, the source data is input into the written storage basic unit, and the written storage basic unit is concurrently written into the temporary directory of HDFS. Using the BulkLoad of HBase and the HFile incremental file loading interface, the written storage basic unit in the HDFS temporary directory is loaded into the corresponding Region partition directory of HBase, and the system automatically performs a write operation on the data according to the Region partition directory. Similarly, when the current job type is the HBase data deletion type, the basic data storage unit is marked as deleted to obtain the deleted storage basic unit. According to the key-value pair, the source data is input into the deleted storage basic unit, and the deleted storage basic unit is concurrently written into the temporary directory of HDFS. Using the BulkLoad of HBase and the HFile incremental file loading interface, the deleted storage basic unit in the HDFS temporary directory is loaded into the corresponding Region partition directory of HBase, and the data is deleted according to the Region partition directory. It realizes accurate writing and deletion of data in the case of a large amount of data, avoids the need to send a large number of RPC requests in the interaction between the system and the server, and still uses the Delete mechanism for single data processing for deletion and writing.
[0063] In the above step S204, by mapping each key-value pair in the partitioned distributed data set to the basic data storage unit of HBase, the unity of the data storage format is realized, so that the data can be seamlessly compatible with the data model of HBase. Then, the source data is input into the basic data storage unit, and the basic data storage unit is marked according to the current job type to obtain the marked storage basic unit, which reduces the interaction with the HBase server during the data processing process, thereby reducing the consumption of network resources and computing resources. The batch writing and deletion operations use the BulkLoad and tombstone marking mechanisms, making the resource use more intensive and efficient, and reducing the common resource bottleneck problems in big data operations.
[0064] In this embodiment, source data is first obtained, the current job type is determined, and the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine are configured. The source data is filtered by the in-memory computing engine to obtain data fields, and at the same time, a mapping relationship between the HBase data table and the data fields is established. The data files in the Hive data table are scanned according to the mapping relationship to form an original distributed data set. At the same time, the storage unit range in the HBase data table is obtained, and the original distributed data set partition is repartitioned according to the storage unit range to obtain a partitioned distributed data set. Each key-value pair in the partitioned distributed data set is mapped to the basic data storage unit of HBase, the source data is input into the basic data storage unit, and the basic data storage unit is marked according to the current job type to obtain a marked storage basic unit, and write or delete operations are performed on the source data according to the marked storage basic unit.
[0065] By combining the characteristics of the in-memory computing engine and HBase, the efficient processing and storage of big data are realized. Through the filtering process of the in-memory computing engine, unnecessary data processing is reduced, and the data processing efficiency is improved. At the same time, through the repartitioning and marking mechanism of the storage unit range of HBase, batch write operations and batch delete operations are performed on the data according to the value of the Rowkey, realizing accurate writing and deletion of data in the case of a large amount of data, avoiding the problem that a large number of RPC requests need to be sent in the interaction between the existing technology system and the server, and still using the single-data processing Delete mechanism for deletion and writing. It avoids data redundancy and errors, further improves the accuracy and efficiency of data processing. In practical applications, this solution can significantly improve the speed and accuracy of big data processing, provide strong support for data analysis and mining, have broad application prospects and significant economic benefits. It solves the problem of high response latency caused by low system throughput when the data volume is large in the existing technology.
[0066] In the specific implementation process, before configuring the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine, the above method further includes: configuring a function switch using an XML configuration file, where the function switch is a write function switch or a delete function switch; submitting the function switch to the in-memory computing engine job.
[0067] Among them, the XML configuration file is a file format for storing and transmitting data, which can clearly describe the data structure and data content. The function switch is used to control the type of data processing;
[0068] Specifically, when the function switch is set to write, the in-memory computing engine will perform a data write operation; when the function switch is set to delete, the in-memory computing engine will perform a data delete operation;
[0069] By applying the switch for the write or delete function configured in the XML file, the configuration file is submitted to the in-memory computing engine Spark job. By switching the configuration to reuse the data write channel, the data deletion or write channel can be seamlessly switched, improving the overall system performance and meeting the requirements for data correction and deletion. This method can flexibly adapt to different data processing requirements and improve the flexibility and efficiency of data processing.
[0070] Specifically, after determining the current job type, the above method further includes: when the current job type is writing data to an HBase data table, using the HBase client interface to query whether the above HBase data table exists; when it is determined that the above HBase data table does not exist, according to the preset configuration information in the XML configuration file, using the above HBase client interface to create the above HBase data table, where the above configuration information includes the data table name, column family name, data expiration time, and partition range.
[0071] Specifically, to use the HBase client interface to query whether the above HBase data table exists, specifically, the HBase client interface can be used to query whether an HBase data table exists. The Java API of HBase can be used: first, an HBase configuration object is created, then "ConnectionFactory" is used to create a connection, and then the "tableExists" method of the "Admin" object is used to check whether the specified table exists, and finally the corresponding information is output according to the result.
[0072] Among them, HBase is a distributed, versioned columnar storage database that can handle random reads and writes of large-scale data.
[0073] In the data write job, this solution first checks whether the target HBase data table exists. If it does not exist, the data table is created according to the configuration information in the XML configuration file, which can ensure the correctness and integrity of data writing and avoid data writing errors or losses.
[0074] More specifically, performing a write or delete operation on the above source data according to the above marker storage basic unit and the above current job type includes: writing the HFile file to the temporary directory of the distributed file system; using the BulkLoad of HBase and the HFile incremental file loading interface to load the HFile file in the above temporary directory into the Region partition directory of the above HBase.
[0075] The HFile is the data storage format of HBase, which optimizes the read and write performance of data. After the data processing is completed, the Spark job will convert the data into the HFile format. When generating the HFile file, the present invention utilizes the multi-version storage feature of HFile to write the Cell data with the normal write mark to ensure the accuracy of the data. The generation of the HFile file also considers the compression and encoding of the data to improve the storage efficiency and read and write speed.
[0076] In addition, the traditional HBase writing is to write one by one through the put API of the HBase client, with low efficiency. The present invention utilizes the BulkLoad mechanism to directly load a large number of HFile files into HBase, avoiding frequent interactions with the server side and greatly improving the writing speed and efficiency. Through BulkLoad, the data can be accurately written into the corresponding Region, reducing the latency and resource consumption during data writing and achieving accurate writing.
[0077] This method writes the HFile file to the temporary directory of the distributed file system, and then uses the BulkLoad of HBase and the HFile incremental file loading interface to load the HFile file in the temporary directory into the Region partition directory of HBase, and performs deletion or writing operations on the source data according to the mark of the basic unit stored in the Region partition directory. Being able to perform data writing or deletion operations efficiently according to the mark of the basic unit stored in the Region partition directory reduces the data processing time and improves the efficiency and accuracy of data processing.
[0078] Furthermore, repartition the above-mentioned original distributed dataset according to the above-mentioned storage unit range to obtain a partitioned distributed dataset, including: repartition the data of the above-mentioned original distributed dataset according to the range of the above-mentioned storage unit range based on the value of the column corresponding to the Rowkey of the above-mentioned HBase data table in the above-mentioned Hive data table, and input it into the corresponding above-mentioned partitioned distributed dataset.
[0079] Rowkey is a key concept in HBase. It is the primary key of the data and is used to uniquely identify a row of data. By repartitioning the data of the original distributed dataset according to the range of the storage unit range based on the value of the column corresponding to the Rowkey of the HBase data table in the Hive data table, the accuracy and integrity of data writing or data deletion can be ensured, avoiding data redundancy and errors.
[0080] Furthermore, after inputting into the corresponding partitioned distributed dataset as described above, the above method further includes: sorting and deduplicating the data in the partitioned distributed dataset according to the value of the Rowkey column, and generating a key-value pair sequence in the partitioned distributed dataset. By generating a key-value pair sequence in the partitioned distributed dataset, it is more efficient to perform data writing or deletion operations according to the key-value pair sequence, reducing data processing time and improving the efficiency and accuracy of data processing.
[0081] Specifically, configure the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine, including: setting the job memory and the number of cores of the in-memory computing engine according to the data volume of the above source data. The setting of job memory and the number of cores is crucial for the performance of the in-memory computing engine, which directly affects the speed and efficiency of data processing. By setting the job memory and the number of cores of the in-memory computing engine according to the data volume of the data source, it is possible to ensure the efficiency and accuracy of the in-memory computing engine in data processing, and avoid resource waste and errors during data processing.
[0082] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the data processing method based on the distributed file model of the present application will be described in detail below in conjunction with specific embodiments.
[0083] This embodiment relates to a specific data processing method based on a distributed file model. Based on the data structure distribution in the file storage model HFile, this embodiment proposes a big data fast deletion mechanism that can dynamically adjust read and write resources, customize operation data configuration, and have low latency. By configuring the data writing channel through switching, seamlessly switching the deletion or writing channel of the data, improving the overall performance of the system, and meeting the requirements of data correction and deletion. The schematic flow diagram of the data deletion principle is as Figure 3 shown.
[0084] The schematic flow diagram of the specific data processing method based on the distributed file model in this embodiment is as Figure 4 shown, and specifically includes the following steps:
[0085] Step S001: Configure the switch for the write or delete function using an XML file, and submit the configuration file to the Spark job to provide a dependency for subsequent processing logic. If this is a large amount of data writing to the HBase data table this time, jump to step S002; if it is to delete HBase data, jump to step S004.
[0086] Step S002: If this job is to write to the HBase cluster, it is necessary to use the HBase client interface to query whether the HBase data exists. If it does not exist, jump to step S003; if the data table already exists, jump to step S004.
[0087] Step S003: For the case where there is no data table in the HBase cluster, it is necessary to create the target data table according to the configurations such as the preset data table name, column family name, data expiration time, partition range, etc. in the XML configuration file. After completion, proceed to Step S004.
[0088] Step S004: For both writing and deleting HBase data, it is necessary to specify target parameters such as the database, data table, and partition in the data warehouse Hive, and pre-set resources such as the memory and number of cores of the Spark job according to the data volume of the source data, or use the adaptive resource adjustment strategy of Spark. After completion, proceed to Step S005;
[0089] Step S005: Through the SQL framework interface of Spark, standard SQL statements can be used to perform custom filtering and processing on the Hive data source, and at the same time, a mapping relationship between the columns of the HBase data table and the data fields after filtering the Hive data source is established. After completion, proceed to Step S006.
[0090] Step S006: Inside the Spark job, according to the information such as the Hive source data table and the custom processing SQL configured in the previous steps, the data files of the Hive data table are scanned in parallel to form the original RDD. At the same time, the Region partition range of the HBase data table is obtained, and the original RDD is repartitioned. The values of the Rowkey column corresponding to the HBase table in the Hive table are repartitioned into the corresponding Spark partition RDDs according to the range of the Region interval. At the same time, the data in the partition RDD is sorted and de-duplicated according to the value of the Rowkey, and a key-value pair sequence of [(Rowkey + column family name + column name, column value)] is formed inside each partition. After completion, proceed to Step S007;
[0091] Step S007: After obtaining the Spark partition RDD in Step S006, according to the switch configuration in Step S001, each key-value pair in the partition RDD needs to be mapped to the basic data storage unit Cell of HBase. If the delete switch is turned on, then proceed to Step S009; if the write switch is turned on, then proceed to Step S008.
[0092] Step S008: If the write switch is turned on, then when the key-value pair data in the partition RDD is mapped to the Cell data of HBase, a normal creation mark is written to the Cell metadata structure. After completion, proceed to Step S010.
[0093] Step S009: If the deletion switch is turned on, then when mapping the key-value pair data in the partitioned RDD to the Cell data of HBase, write a deletion tombstone marker to the Cell metadata structure. After completion, proceed to Step S010.
[0094] Step S010: After mapping the key-value pairs in the partitioned RDD to the Cells of HBase, concurrently write the partitioned data to a temporary directory in HDFS. After completion, proceed to Step S011.
[0095] Step S011: Utilize the BulkLoad of HBase and the HFile incremental file loading interface to load the HFile files in the HDFS temporary directory into the corresponding Region partition directory of HBase.
[0096] Regarding the data structure distribution in the file storage model HFile of HBase in this embodiment, a big data fast deletion mechanism that can be custom-configured, out-of-the-box, and share the data writing and deletion links is proposed. By reusing the data writing channel through switch-based configuration, the development efficiency can be improved, the complexity of the system link can be reduced. At the same time, using Spark's dynamic resource adjustment and SparkSql's custom filtering of data items increases the flexibility of massive data deletion, reduces the latency of massive data deletion, and improves the performance of the entire system.
[0097] The embodiment of the present application also provides a data processing device based on a distributed file model. It should be noted that the data processing device based on the distributed file model in the embodiment of the present application can be used to execute the data processing method based on the distributed file model provided by the embodiment of the present application. The device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0098] The following introduces the data processing device based on the distributed file model provided by the embodiment of the present application.
[0099] Figure 5 is a schematic diagram of the data processing device based on the distributed file model according to the embodiment of the present application. As Figure 5 shown, the device includes:
[0100] A first configuration unit 51, configured to obtain source data, determine the current job type, and configure the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine;
[0101] Specifically, by obtaining source data, determining the current job type based on the obtained source data, and then configuring the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine according to the current job type, the mechanism of this parameter configuration combined with the in-memory computing characteristics of the in-memory computing engine Spark can dynamically adjust the resources of the input data, and can dynamically adapt to the deletion, cleaning, or writing of different scales of data volumes, improving the robustness of job operation.
[0102] The screening unit 52 is used to screen and process the above source data through the above in-memory computing engine to obtain data fields, and at the same time establish a mapping relationship between the HBase data table and the above data fields;
[0103] Specifically, through the SQL framework interface of the in-memory computing engine Spark, standard SQL statements can be used to customize the screening and processing of the Hive data source. Among them, SparkSQL can utilize the in-memory computing characteristics of Spark to quickly read and process the data in Hive. When dealing with a large amount of data, the data can be processed in the memory of Spark, reducing disk I / O and improving the data processing speed. At the same time, a mapping relationship between the columns of the HBase data table and the data fields after screening the Hive data source is established, and the mapping relationship ensures data consistency.
[0104] The repartition processing unit 53 is used to scan and process the data files in the Hive data table according to the above mapping relationship to form an original distributed dataset, and at the same time obtain the storage unit interval in the above HBase data table, and repartition the above original distributed dataset according to the above storage unit interval to obtain a partitioned distributed dataset;
[0105] Specifically, through the distributed processing ability of Spark, the original distributed dataset can be scanned and processed in parallel, which means that the data processing tasks can be executed simultaneously on multiple nodes in the cluster, greatly improving the efficiency and speed of data processing. The data storage in HBase is based on the storage unit Region, and each Region contains a continuous Rowkey interval. By obtaining the storage unit interval of the HBase data table, the distributed dataset in Spark can be repartitioned according to the distribution of the storage units, ensuring that the data in each partitioned distributed dataset can correspond to a specific Region in HBase. This optimized partitioning strategy reduces the data transmission between networks because the data can be written more directly to the corresponding HBase storage unit, thereby reducing the data write latency and improving the data write efficiency. In addition, using the HBase Region interval for repartition processing can ensure that the data is written to HBase in the order of Rowkey, ensuring data consistency.
[0106] A tagging unit 54 is configured to map each key-value pair in the above partitioned distributed dataset into a basic data storage unit of HBase, input the above source data into the basic data storage unit, tag the basic data storage unit according to the above current job type to obtain a tagged basic storage unit, and perform a write or delete operation on the above source data according to the tagged basic storage unit.
[0107] Specifically, by mapping each key-value pair in the partitioned distributed dataset into a basic data storage unit of HBase, the uniformity of the data storage format is achieved, enabling the data to be seamlessly compatible with the data model of HBase. Then, the source data is input into the basic data storage unit, and the basic data storage unit is tagged according to the current job type to obtain a tagged basic storage unit, reducing the interaction with the HBase server during the data processing process, thereby reducing the consumption of network resources and computing resources. The bulk write and delete operations use the BulkLoad and tombstone marking mechanisms, making resource usage more intensive and efficient, and reducing the common resource bottleneck problems during big data operations.
[0108] In this embodiment, a first configuration unit is used to obtain source data, determine the current job type, and configure the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine; a screening unit is used to screen the source data through the in-memory computing engine to obtain data fields, and at the same time establish a mapping relationship between the HBase data table and the data fields; a re-partitioning processing unit is used to scan the data files in the Hive data table according to the mapping relationship to form an original distributed dataset, and at the same time obtain the storage unit range in the HBase data table, and perform re-partitioning processing on the partitions of the original distributed dataset according to the storage unit range to obtain a partitioned distributed dataset; a tagging unit is used to map each key-value pair in the partitioned distributed dataset into a basic data storage unit of HBase, input the source data into the basic data storage unit, tag the basic data storage unit according to the current job type to obtain a tagged basic storage unit, and perform a write or delete operation on the source data according to the tagged basic storage unit.
[0109] By combining the characteristics of in-memory computing engines and HBase, efficient processing and storage of big data are achieved. Through the filtering process of the in-memory computing engine, unnecessary data processing is reduced, and the efficiency of data processing is improved. At the same time, through the re-partitioning and marking mechanism of the storage unit intervals in HBase, batch write operations and batch delete operations are performed on the data according to the value of the Rowkey, realizing accurate writing and deletion of data in the case of a large amount of data, avoiding the problem that a large number of RPC requests need to be sent in the interaction between the existing technology system and the server, and still using the Delete mechanism for single data processing for deletion and writing. Data redundancy and errors are avoided, further improving the accuracy and efficiency of data processing. In practical applications, this solution can significantly improve the speed and accuracy of big data processing, provide strong support for data analysis and mining, have broad application prospects and significant economic benefits. It solves the problem of high response latency caused by low system throughput when the amount of data in the existing technology is large.
[0110] In some embodiments of the present application, the device further includes a second configuration unit and a submission unit; the second configuration unit is used to configure a function switch using an XML configuration file before configuring the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine, where the above function switch is a write function switch or a delete function switch; the submission unit is used to submit the above function switch to the above in-memory computing engine job.
[0111] By applying the switch in the XML file to configure the write or delete function and submitting the configuration file to the in-memory computing engine Spark job. By using switch-based configuration to reuse the data write channel and seamlessly switch the data deletion or write channel, the overall performance of the system is improved to meet the requirements of data correction and deletion. In this way, it can flexibly adapt to different data processing requirements and improve the flexibility and efficiency of data processing.
[0112] In some embodiments of the present application, the device further includes a query unit and a creation unit; the query unit is used to query whether the above HBase data table exists using the HBase client interface when determining the current job type and when the current job type is writing data to the HBase data table; the creation unit is used to create the above HBase data table using the above HBase client interface according to the preset configuration information in the XML configuration file when determining that the above HBase data table does not exist, where the above configuration information includes the data table name, column family name, data expiration time, and partition interval.
[0113] Specifically, in the data writing operation, first check whether the target HBase data table exists. If it does not exist, create the data table according to the configuration information in the XML configuration file, which can ensure the correctness and integrity of data writing and avoid data writing errors or losses.
[0114] In some embodiments of the present application, the marking unit includes a writing module and a loading module; the writing module is used to write the HFile file into the temporary directory of the distributed file system; the loading module is used to load the HFile file in the above temporary directory into the Region partition directory of the above HBase by using the BulkLoad of HBase and the HFile incremental file loading interface.
[0115] By writing the HFile file into the temporary directory of the distributed file system, and then using the BulkLoad of HBase and the HFile incremental file loading interface to load the HFile file in the temporary directory into the Region partition directory of HBase, delete or write the source data according to the mark of the basic unit of mark storage in the Region partition directory. Being able to perform data writing or deletion operations efficiently according to the mark of the basic unit of mark storage in the Region partition directory can reduce the data processing time and improve the efficiency and accuracy of data processing.
[0116] In some embodiments of the present application, the repartition processing unit includes an input module, and the input module is used to repartition the data of the original distributed data set according to the values of the columns corresponding to the Rowkey of the above HBase data table in the above Hive data table within the range of the above storage unit interval and input it into the corresponding above partitioned distributed data set.
[0117] Specifically, Rowkey is a key concept in HBase. It is the primary key of data and is used to uniquely identify a row of data. By repartitioning the data of the original distributed data set according to the values of the columns corresponding to the Rowkey of the HBase data table in the Hive data table within the range of the storage unit interval, the accuracy and integrity of data writing or data deletion can be ensured, and data redundancy and errors can be avoided.
[0118] In some embodiments of the present application, the device further includes a deduplication operation unit, and the deduplication operation unit is used to sort and deduplicate the data of the above partitioned distributed data set according to the values of the Rowkey column after inputting it into the corresponding above partitioned distributed data set, and at the same time generate a key-value pair sequence in the above partitioned distributed data set.
[0119] In some embodiments of the present application, the first configuration unit includes a configuration module, which is used to set the job memory and the number of cores of the above in-memory computing engine according to the data volume of the above source data.
[0120] Specifically, by generating a key-value pair sequence in a partitioned distributed data set, performing data writing or deletion operations based on the key-value pair sequence is more efficient, reducing data processing time and improving the efficiency and accuracy of data processing.
[0121] The above-mentioned data processing device based on a distributed file model includes a processor and a memory. The above-mentioned first configuration unit, screening unit, re-partitioning processing unit, marking unit, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement corresponding functions. The above-mentioned modules are all located in the same processor; or, the above-mentioned each module is separately located in different processors in any combination form.
[0122] The processor contains a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem of high response latency caused by low system throughput when the data volume is large in the prior art can be solved.
[0123] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0124] An embodiment of the present invention provides a computer-readable storage medium. The above-mentioned computer-readable storage medium includes a stored program. Wherein, when the above-mentioned program runs, it controls the device where the above-mentioned computer-readable storage medium is located to execute the above-mentioned data processing method based on a distributed file model.
[0125] Specifically, the data processing method based on a distributed file model includes:
[0126] Step S201, obtain source data, determine the current job type, and configure the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine;
[0127] Step S202, perform screening processing on the above-mentioned source data through the above-mentioned in-memory computing engine to obtain data fields, and at the same time establish a mapping relationship between the HBase data table and the above-mentioned data fields;
[0128] Step S203, perform scanning processing on the data files in the Hive data table according to the above-mentioned mapping relationship to form an original distributed data set, and at the same time obtain the storage unit interval in the above-mentioned HBase data table, and perform re-partitioning processing on the partitions of the above-mentioned original distributed data set according to the above-mentioned storage unit interval to obtain a partitioned distributed data set;
[0129] Step S204: Map each key-value pair in the above partitioned distributed dataset to the basic data storage unit of HBase, input the above source data into the above basic data storage unit, mark the above basic data storage unit according to the above current job type to obtain a marked storage basic unit, and perform a write or delete operation on the above source data according to the above marked storage basic unit.
[0130] An embodiment of the present invention provides a processor, which is used to run a program. When the program runs, it executes the above data processing method based on a distributed file model.
[0131] Specifically, the data processing method based on a distributed file model includes:
[0132] Step S201: Obtain source data, determine the current job type, and configure the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine;
[0133] Step S202: Screen the above source data through the above in-memory computing engine to obtain data fields, and at the same time establish a mapping relationship between the HBase data table and the above data fields;
[0134] Step S203: Scan the data files in the Hive data table according to the above mapping relationship to form an original distributed dataset, and at the same time obtain the storage unit range in the above HBase data table, and re-partition the above original distributed dataset partition according to the above storage unit range to obtain a partitioned distributed dataset;
[0135] Step S204: Map each key-value pair in the above partitioned distributed dataset to the basic data storage unit of HBase, input the above source data into the above basic data storage unit, mark the above basic data storage unit according to the above current job type to obtain a marked storage basic unit, and perform a write or delete operation on the above source data according to the above marked storage basic unit.
[0136] An embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements at least the following steps:
[0137] Step S201: Obtain source data, determine the current job type, and configure the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine;
[0138] Step S202: Screen the above source data through the above in-memory computing engine to obtain data fields, and at the same time establish a mapping relationship between the HBase data table and the above data fields;
[0139] Step S203: Scan and process the data files in the Hive data table according to the above mapping relationship to form an original distributed data set. Meanwhile, obtain the storage unit range in the above HBase data table, and re-partition the partitions of the above original distributed data set according to the above storage unit range to obtain a partitioned distributed data set;
[0140] Step S204: Map each key-value pair in the above partitioned distributed data set to the basic data storage unit of HBase, input the above source data into the above basic data storage unit, mark the above basic data storage unit according to the above current job type to obtain a marked storage basic unit, and perform write or delete operations on the above source data according to the above marked storage basic unit.
[0141] The devices in this article can be servers, PCs, PADs, mobile phones, etc.
[0142] This application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program initialized with at least the following method steps:
[0143] Step S201: Obtain source data, determine the current job type, and configure the target parameters of the data warehouse Hive and the parameters of the memory computing engine;
[0144] Step S202: Screen and process the above source data through the above memory computing engine to obtain data fields, and establish a mapping relationship between the HBase data table and the above data fields;
[0145] Step S203: Scan and process the data files in the Hive data table according to the above mapping relationship to form an original distributed data set. Meanwhile, obtain the storage unit range in the above HBase data table, and re-partition the partitions of the above original distributed data set according to the above storage unit range to obtain a partitioned distributed data set;
[0146] Step S204: Map each key-value pair in the above partitioned distributed data set to the basic data storage unit of HBase, input the above source data into the above basic data storage unit, mark the above basic data storage unit according to the above current job type to obtain a marked storage basic unit, and perform write or delete operations on the above source data according to the above marked storage basic unit.
[0147] Obviously, those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0148] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0149] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0150] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0151] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing in the processFigure 1 One or more processes and / or blocks Figure 1 Steps of the functions specified in one block or more blocks.
[0152] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0153] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0154] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0155] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0156] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0157] 1) A data processing method based on a distributed file model in this application. The method includes obtaining source data, determining the current job type, and configuring the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine; screening the source data through the in-memory computing engine to obtain data fields, and establishing a mapping relationship between the HBase data table and the data fields; scanning the data files in the Hive data table according to the mapping relationship to form an original distributed data set, and at the same time obtaining the storage unit range in the HBase data table, and performing re-partitioning processing on the partitions of the original distributed data set according to the storage unit range to obtain a partitioned distributed data set; mapping each key-value pair in the partitioned distributed data set to the basic data storage unit of HBase, inputting the source data into the basic data storage unit, and marking the basic data storage unit according to the current job type to obtain a marked storage basic unit, and performing write or delete operations on the source data according to the marked storage basic unit. By combining the characteristics of the in-memory computing engine and HBase, efficient processing and storage of big data are achieved. Through the screening process of the in-memory computing engine, unnecessary data processing is reduced, and the data processing efficiency is improved. At the same time, through the re-partitioning and marking mechanisms of the storage units in HBase, accurate writing and deletion of data are realized, avoiding data redundancy and errors, and further improving the accuracy and efficiency of data processing. In practical applications, this solution can significantly improve the speed and accuracy of big data processing, provide strong support for data analysis and mining, and has broad application prospects and significant economic benefits. It solves the problem that the system throughput is low and the response delay is high when the data volume is large in the prior art.
[0158] 2) A data processing device based on a distributed file model according to the present application applies a first configuration unit to obtain source data, determine the current job type, and configure the target parameters of the data warehouse Hive and the parameters of the in-memory computing engine; applies a screening unit to screen the source data through the in-memory computing engine to obtain data fields, and at the same time establish a mapping relationship between the HBase data table and the data fields; a re-partitioning processing unit is used to scan the data files in the Hive data table according to the mapping relationship to form an original distributed data set, and at the same time obtain the storage unit interval in the HBase data table, and re-partition the original distributed data set according to the storage unit interval to obtain a partitioned distributed data set; applies a marking unit to map each key-value pair in the partitioned distributed data set to the basic data storage unit of HBase, input the source data into the basic data storage unit, and mark the basic data storage unit according to the current job type to obtain a marked storage basic unit, and perform write or delete operations on the source data according to the marked storage basic unit. By combining the characteristics of the in-memory computing engine and HBase, efficient processing and storage of big data are achieved. Through the screening process of the in-memory computing engine, unnecessary data processing is reduced, and the data processing efficiency is improved. At the same time, through the re-partitioning and marking mechanism of the storage unit interval of HBase, accurate writing and deletion of data are achieved, avoiding data redundancy and errors, and further improving the accuracy and efficiency of data processing. In practical applications, this solution can significantly improve the speed and accuracy of big data processing, provide strong support for data analysis and mining, and has broad application prospects and significant economic benefits. It solves the problem that the system throughput is low and the response delay is high when the data volume is large in the prior art.
[0159] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A data processing method based on a distributed file model, characterized in that, Including: Obtain source data, determine the current job type, and configure the target parameters of the data warehouse Hive and the parameters of the memory computing engine; Filter the source data through the memory computing engine to obtain data fields, and at the same time establish a mapping relationship between the HBase data table and the data fields; Scan and process the data files in the Hive data table according to the mapping relationship to form an original distributed data set, and at the same time obtain the storage unit interval in the HBase data table, and re-partition the partitions of the original distributed data set according to the storage unit interval to obtain a partitioned distributed data set; Map each key-value pair in the partitioned distributed data set to the basic data storage unit of HBase, input the source data into the basic data storage unit, and mark the basic data storage unit according to the current job type to obtain a marked storage basic unit, and perform a write or delete operation on the source data according to the marked storage basic unit.
2. The method according to claim 1, wherein Before configuring the target parameters of the data warehouse Hive and the parameters of the memory computing engine, the method further includes: Configure a function switch using an XML configuration file, where the function switch is a write function switch or a delete function switch; Submit the function switch to the memory computing engine job.
3. The method according to claim 1, wherein After determining the current job type, the method further includes: When the current job type is to write data to the HBase data table, use the HBase client interface to query whether the HBase data table exists; When it is determined that the HBase data table does not exist, create the HBase data table using the HBase client interface according to the preset configuration information in the XML configuration file, where the configuration information includes the data table name, column family name, data expiration time, and partition interval.
4. The method according to claim 1, wherein Performing a write or delete operation on the source data according to the marked storage basic unit and the current job type includes: Write the HFile file to the temporary directory of the distributed file system; Use the BulkLoad of HBase and the HFile incremental file loading interface to load the HFile file in the temporary directory into the Region partition directory of HBase.
5. The method according to claim 1, wherein Re-partitioning the partitions of the original distributed data set according to the storage unit interval to obtain a partitioned distributed data set includes: Re-partition the data of the original distributed data set according to the range of the storage unit interval for the value of the Rowkey column corresponding to the HBase data table in the Hive data table, and input it into the corresponding partitioned distributed data set.
6. The method according to claim 1, wherein After inputting into the corresponding partitioned distributed data set, the method further includes: Sort and deduplicate the data of the partitioned distributed data set according to the value of the Rowkey column, and at the same time generate a key-value pair sequence in the partitioned distributed data set.
7. The method according to claim 1, characterized in that Configuring the target parameters of the data warehouse Hive and the parameters of the memory computing engine includes: Set the job memory and the number of cores of the in-memory computing engine according to the data volume of the source data.
8. A data processing device based on a distributed file model, characterized in that, Including: A first configuration unit, configured to obtain source data, determine the current job type, and configure target parameters of the data warehouse Hive and parameters of the in-memory computing engine; A screening unit, configured to screen the source data through the in-memory computing engine to obtain data fields, and simultaneously establish a mapping relationship between the HBase data table and the data fields; A re-partitioning processing unit, configured to scan the data files in the Hive data table according to the mapping relationship to form an original distributed data set, and simultaneously obtain the storage unit range in the HBase data table, and perform re-partitioning processing on the partitions of the original distributed data set according to the storage unit range to obtain a partitioned distributed data set; A marking unit, configured to map each key-value pair in the partitioned distributed data set to a basic data storage unit of HBase, input the source data into the basic data storage unit, and mark the basic data storage unit according to the current job type to obtain a marked storage basic unit, and perform a write or delete operation on the source data according to the marked storage basic unit.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the data processing method based on the distributed file model according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Including: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a data processing method based on the distributed file model according to any one of claims 1 to 7.
Citation Information
Cited By
HBase off-line data import optimization method based on key value pre-partitioning
CN120804106A
A HBase offline data import optimization method based on key value pre-partitioning
CN120804106B