System, device and method for writing data warehouse data into HBase
Patent Information
- Application Number
- CN202510171265.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-07-18
AI Technical Summary
[0004]1、离线推送数据期间,在线应用jrHBase将数据写入slave会存在积压情况(1天出现4次波峰,最久积压4小时,峰值为3000万),若积压期间切换slave,部分数据有读取不到的风险
[0044]1. Directly convert the offline data in the data warehouse into the Hfile format, so that each node in HBase can batch load the Hfile files into HBase, avoiding the operation of writing each record of the data one by one, improving the data writing speed and at the same time enhancing the freshness of the data in HBase.
Smart Images

Figure CN120336423A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage, and in particular, to a system, device, and method for writing data warehouse data into HBase. Background Art
[0002] With the digital transformation of enterprises, enterprises generally establish their own data warehouse systems (abbreviation: data warehouse) internally. Along with the increase in online services, the combination of the data warehouse and online service system applications (abbreviation: online system applications) has become increasingly close. It is often necessary to write hundreds or thousands of data warehouse result table data into Hbase online storage through the online system application consuming data warehouse kafka messages for candidate data processing (such as business decision-making, label usage, customer group screening, etc.).
[0003] Currently, these result table data are mainly sent from the data warehouse to the online service system for consumption and warehousing by sending kafka. During this process, HBase, kafka, HBase master-slave synchronization rocketMq, and the online system all face significant stability challenges, which are mainly manifested in:
[0004] 1. During the offline data push period, there will be a backlog situation when the online application jrHBase writes data to the slave (4 peaks appear in 1 day, the longest backlog is 4 hours, and the peak value is 30 million). If the slave is switched during the backlog period, there is a risk that some data cannot be read.
[0005] 2. When writing HBase data, the number of requests per second (Queries Per Seconds, qps) for writing will soar sharply, posing a stability hazard to the nodes;
[0006] 3. As the number of tables gradually increases, the node restart recovery is slow. Summary of the Invention
[0007] In view of this, the main purpose of the present invention is to propose a system, device, and method for writing data warehouse data into HBase, in order to at least partially solve at least one of the above technical problems.
[0008] To solve the above technical problems, the first aspect of the present invention proposes a system for writing data warehouse data into HBase, including:
[0009] A new data warehouse push node, configured to convert offline data into the Hfile format to obtain an Hfile file, batch copy the Hfile file to a specified hdfs directory of Hbase; after the batch copy is successful, generate a status message to notify the online system application;
[0010] An online system application is used to concurrently control the batch loading of Hfile files to each node of Hbase when receiving status messages sent by newly pushed nodes in the data warehouse.
[0011] To solve the above technical problems, a second aspect of the present invention provides a device for writing data warehouse data into HBase, and the device includes:
[0012] A batch replication module is used to convert offline data into the Hfile format to obtain Hfile files, and batch copy the Hfile files to a specified hdfs directory of Hbase.
[0013] A generation module is used to generate a status message indicating the successful batch replication of the Hfile file after the batch replication is successful.
[0014] A notification module is used to notify the online system application of the status message through kafka, so that the online system application can concurrently control the batch loading of Hfile files to each node of Hbase.
[0015] According to a preferred embodiment of the present invention, the batch replication module uses the distcp tool to batch copy the Hfile files to a specified hdfs directory of the slave nodes of HBase.
[0016] According to a preferred embodiment of the present invention, the device further includes:
[0017] A sending module is used to send the Hfile files to the online system application, so that the online system application can query the Hbase data after import based on the Hfile files and process the online feature data.
[0018] According to a preferred embodiment of the present invention, the device further includes:
[0019] A switching module is used to perform AB table switching after the loading is completed.
[0020] According to a preferred embodiment of the present invention, the device further includes:
[0021] A verification module is used to compare and verify the case of key values for the offline data.
[0022] To solve the above technical problems, a third aspect of the present invention provides an online system application, including:
[0023] A receiving module is used to receive the status message sent by the data warehouse through kafka, and the status message is used to indicate that the data warehouse has successfully batch copied the converted Hfile files to a specified hdfs directory of Hbase.
[0024] A concurrency control module for concurrently controlling the bulk loading of Hfile files to each node of Hbase.
[0025] According to a preferred embodiment of the present invention, the concurrency control module, through concurrency control, respectively calls the bulkLoad operation of the slave nodes in HBase to bulk load the specified files in the hdfs directory of the slave nodes to the slave nodes; calls the distcp operation and the bulkLoad operation of the master node in HBase, and after bulk copying the specified files in the hdfs directory of the slave nodes to the specified hdfs directory of the master node of HBase, then calls the bulkLoad operation of the master node of HBase to bulk load the files in the specified hdfs directory of the master node to the master node.
[0026] To solve the above technical problems, a fourth aspect of the present invention provides a method for writing data warehouse data into HBase, the method comprising:
[0027] Convert the offline data into the Hfile format to obtain an Hfile file, and bulk copy the Hfile file to the specified hdfs directory of Hbase;
[0028] Generate a status message indicating the successful bulk copy of the Hfile file after the successful bulk copy;
[0029] Notify the online system application of the status message through kafka so that the online system application can concurrently control the bulk loading of the Hfile file to each node of Hbase.
[0030] According to a preferred embodiment of the present invention, use the distcp tool to bulk copy the Hfile file to the specified hdfs directory of the slave nodes of HBase.
[0031] According to a preferred embodiment of the present invention, the online system application, through concurrency control, respectively calls the bulkLoad operation of the slave nodes in HBase to bulk load the specified files in the hdfs directory of the slave nodes to the slave nodes; calls the distcp operation and the bulkLoad operation of the master node in HBase, and after bulk copying the specified files in the hdfs directory of the slave nodes to the specified hdfs directory of the master node of HBase, then calls the bulkLoad operation of the master node of HBase to bulk load the files in the specified hdfs directory of the master node to the master node.
[0032] According to a preferred embodiment of the present invention, after obtaining the Hfile file, the method further includes:
[0033] Sending the Hfile file to an online system application so that the online system application queries the imported Hbase data based on the Hfile file and processes the online feature data.
[0034] According to a preferred embodiment of the present invention, the method further includes:
[0035] Performing an AB table switch after the loading is completed.
[0036] According to a preferred embodiment of the present invention, before converting the offline data into the Hfile format, the method further includes:
[0037] Comparing and verifying the case of the key value of the offline data.
[0038] To solve the above technical problems, a fifth aspect of the present invention provides an electronic device, including:
[0039] A processor; and
[0040] A memory storing computer-executable instructions, which when executed cause the processor to execute the method according to any one of the above.
[0041] To solve the above technical problems, a sixth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by a processor, the method according to any one of the above is implemented.
[0042] To solve the above technical problems, a seventh aspect of the present invention provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, the method according to any one of the above is implemented.
[0043] In summary, in the present invention, the data warehouse converts the offline data into the Hfile format, obtains the Hfile file, and batch copies the Hfile file to the specified hdfs directory of Hbase; and generates a status message to notify the online system application after the batch copy is successful; so that the online system application can concurrently control the batch loading of the Hfile file by each node of HBase through the way of batch loading of Hfile data (the function that data in Hfile format in HBase can be batch loaded), thereby improving the speed of writing data from the data warehouse to Hbase and reducing the load of the online system application. Compared with the prior art, the present invention has at least the following beneficial effects:
[0044] 1. Directly convert the offline data in the data warehouse into the Hfile format, so that each node in HBase can batch load the Hfile files into HBase, avoiding the operation of writing each record of the data one by one, improving the data writing speed and at the same time enhancing the freshness of the data in HBase.
[0045] 2. Use the method of batch loading Hfile data to concurrently control the batch loading of each node in HBase, reducing the number of connections for data transmission between the online system and the data warehouse, and alleviating the writing pressure on HBase.
[0046] 3. Implement the writing operation of the online system through the method of batch loading Hfile data, reducing the queries per second (QPS), saving the consumption of online resources, and thus reducing the operating cost. Description of the Drawings
[0047] In order to make the technical problems solved by the present invention, the technical means adopted and the technical effects achieved clearer, the specific embodiments of the present invention will be described in detail below with reference to the drawings. However, it should be stated that the drawings described below are only the drawings of the exemplary embodiments of the present invention, and those skilled in the art can obtain the drawings of other embodiments based on these drawings without creative efforts.
[0048] Figure 1 It is a schematic flow chart of a method for writing data warehouse data into HBase provided by an embodiment of the present invention;
[0049] Figure 2 It is a schematic flow chart of another method for writing data warehouse data into HBase provided by an embodiment of the present invention;
[0050] Figure 3 It is a schematic structural framework diagram of a device for writing data warehouse data into HBase provided by an embodiment of the present invention;
[0051] Figure 4 It is a schematic diagram of the structural framework of an online system application provided by an embodiment of the present invention;
[0052] Figure 5 It is a structural block diagram of an exemplary embodiment of an electronic device according to the present invention;
[0053] Figure 6 It is a schematic diagram of an embodiment of a computer-readable medium of the present invention;
[0054] Figure 7 It is a schematic structural framework diagram of a system for writing data warehouse data into HBase provided by an embodiment of the present invention. Detailed Embodiments
[0055] On the premise of conforming to the technical concept of the present invention, the structures, performances, effects or other features described in a certain specific embodiment can be combined into one or more other embodiments in any suitable manner.
[0056] In the process of introducing specific embodiments, the detailed descriptions of structures, performances, effects or other features are to enable those skilled in the art to fully understand the embodiments. However, it does not exclude that those skilled in the art can implement the present invention with technical solutions that do not contain the above-mentioned structures, performances, effects or other features under specific circumstances. The figures in the drawings are only exemplary demonstrations and do not represent that all the contents, operations and steps in the figures must be included in the solutions of the present invention, nor does it represent that the execution order must be in accordance with the order shown in the figures.
[0057] Reference Figure 1 , Figure 1 is a schematic diagram of a method for writing data warehouse data into HBase provided by an embodiment of the present invention. This method can batch write the offline data warehouse data into HBase through the uds-app online system application. As Figure 1 , the method for writing data warehouse data into HBase includes:
[0058] S1. Convert the offline data into the Hfile format to obtain an Hfile file, and batch copy the Hfile file to the specified hdfs directory of Hbase;
[0059] Since the bulk load operation in HBase can only process data in the Hfile format, this step converts the offline data warehouse data into the Hfile format to obtain an Hfile file, and batch copies the Hfile file to the specified hdfs directory of Hbase; then the uds-app online system application can directly call the bulkLoad operation to batch load the Hfile files in the specified hdfs directory into each node of HBase. This avoids writing data record by record, improves the data writing speed, and enhances the freshness of the data in HBase.
[0060] Among them: The offline data can be the data obtained by processing the workflow data in the data warehouse through cleaning, transformation, integration, analysis, etc. These offline data can be stored in HBase or other types of databases, such as the document database mongodb. Then, before this step, it is also possible to: determine whether the storage category of the offline data is HBase; if the storage category is HBase, execute this step; if the storage category is not HBase, store the offline data whose storage category is not HBase in mongodb. Among them: The storage category can be judged according to the data category. For example, the storage category of document data is mongodb, and the storage category of non-document data is HBase. Or, the storage category can be judged according to the value of the category field. For example, the storage category with the category field being 0 is mongodb, and the storage category with the category field being 1 is HBase.
[0061] In addition, in order to ensure the data quality of the data written into HBase, before this step, it is also possible to compare and perform case-insensitive key value verification on the offline data.
[0062] Furthermore, this step can also send the converted Hfile file to the uds-app online system application. In this way, the uds-app online system application can query the imported Hbase data based on the Hfile file and process the online feature data.
[0063] Exemplarily, as Figure 2 shown, this step can use the distributed replication tool (DistributedCopy, distcp) to batch copy the Hfile file to the specified hdfs directory under the slave node in Hbase.
[0064] S2. After the batch copy is successful, generate a status message indicating that the batch copy of the Hfile file is successful;
[0065] In this step, it is possible to monitor in real time whether the batch copy is successful. If the current Hfile file is completely copied to the specified hdfs directory under the slave node in Hbase, a first status message indicating that the batch copy of the Hfile file is successful can be generated; otherwise, a second status message indicating that the batch copy of the Hfile file fails can be generated.
[0066] Furthermore, when the second status message is generated, the distcp tool can be used to re-batch copy the current Hfile file to the specified hdfs directory under the slave node in Hbase until the first status message is generated. In addition, the number of times n of the re-batch copy can also be recorded. When n is greater than the threshold and the first status message is not generated, a prompt indicating that the batch copy has an error is issued.
[0067] S3. Notify the online system application of the status message through Kafka, so that the online system application can concurrently control the bulk loading of Hfile files to each node of Hbase.
[0068] Exemplarily, as Figure 2 In, the first status message can be input into Kafka (distributed publish-subscribe message system), and the first status message is notified to the uds-app online system application through Kafka. Then, when the online system application (uds-app) receives the first status message, it indicates that the Hfile files in the data warehouse have been successfully bulk copied to the specified hdfs directory of Hbase. The online system application can use the method of bulk loading Hfile data to concurrently control each node of HBase to perform bulk loading through the bulkLoad operation. On the one hand, it reduces the number of connections between the online system and Kafka, and on the other hand, it also reduces the write pressure on HBase, shortening the data write timeliness from one day to within 10 minutes. At the same time, notifying the online system through Kafka can achieve low-latency data synchronization.
[0069] In HBase, there are two types of nodes, master and slave, where: the master is the main node and the slave is the secondary node. As Figure 2 , the online system application (uds-write-app) can, through concurrent control, separately call the bulkLoad operation of the slave nodes in HBase to bulk load the specified files in the hdfs directory of the slave nodes to the slave nodes; call the distcp operation and the bulkLoad operation of the master node in HBase to copy the specified files in the hdfs directory of the slave nodes to the specified hdfs directory of the master node of HBase, and then call the bulkLoad operation of the master node of HBase to bulk load the files in the specified hdfs directory of the master node to the master node. Thus, the bulk loading of data for both slave nodes and master nodes is completed. In this way, the bulkLoad operations on the slave nodes and master nodes in HBase are managed and executed through the concurrent control mechanism, optimizing the bulk data loading process and avoiding resource contention.
[0070] Furthermore, during the loading process, the loading status of the bulkLoad operation can be queried in real time (query Status). When the loading is completed and the data is successfully imported into HBase, the AB table can be switched to complete the data update, ensuring business continuity during the data update process.
[0071] Furthermore, it is also possible to query the switched AB table in the ODS table to obtain the name of the AB table to be written. Among them: The AB table stores the same offline data in two tables, such as Jr_User_A and jr_user_B. The A table is used to store data on odd days, and the B table is used to store data on even days. If there is a problem with the data of the day in the data warehouse offline processing, affecting the online feature data, then the switched AB table can be queried in the ODS table, and the data of the previous day can be reverted in time through the read-write switch for downgrading, and the name of the correct AB table to be written can be obtained.
[0072] Table 1 shows the comparison results of the method of writing data warehouse data into HBase Figure 2 shown in the prior art. It can be seen from Table 1 that Figure 2 the method of writing data warehouse data into HBase shown in the figure far exceeds the prior art in terms of performance and efficiency, stability, and cost control.
[0073]
[0074] Table 1 shows the comparison results of the method of writing data warehouse data into HBase Figure 2 shown in the figure with the prior art
[0075] Based on Figure 1 the method of writing data warehouse data into HBase described above, the embodiment of the present invention further provides a device for writing data warehouse data into HBase, as Figure 3 shown in the figure. The device includes:
[0076] A batch replication module 31, configured to convert offline data into the Hfile format to obtain an Hfile file, and batch copy the Hfile file to a specified hdfs directory of Hbase;
[0077] A generation module 32, configured to generate a status message indicating the successful batch replication of the Hfile file after the batch replication is successful;
[0078] A notification module 33, configured to notify the online system application of the status message through kafka, so that the online system application can concurrently control the batch loading of the Hfile file to each node of Hbase.
[0079] In a specific embodiment, the batch replication module 31 uses the distcp tool to batch copy the Hfile file to a specified hdfs directory of the slave nodes of HBase.
[0080] Furthermore, the device further includes:
[0081] A sending module, configured to send an Hfile file to an online system application, so that the online system application queries the imported Hbase data based on the Hfile file and processes the online feature data.
[0082] A switching module, configured to perform AB table switching after the loading is completed.
[0083] A verification module, configured to perform comparison on the offline data and verify the case of the key value.
[0084] Based on Figure 1 the method for writing data warehouse data into HBase described above, an embodiment of the present invention further provides an online system application, as Figure 4 shown, the online system includes:
[0085] A receiving module 41, configured to receive a status message sent by the data warehouse through kafka, where the status message is used to identify that the Hfile files to be converted by the data warehouse are successfully batch-copied to a specified hdfs directory of Hbase;
[0086] A concurrency control module 42, configured to perform concurrency control on the batch loading of Hfile files to each node of Hbase.
[0087] In a specific embodiment, the concurrency control module 43, through concurrency control, respectively calls the bulkLoad operation of the slave nodes in HBase to batch-load the specified files in the hdfs directory of the slave nodes to the slave nodes; calls the distcp operation and the bulkLoad operation of the master node in HBase to batch-copy the specified files in the hdfs directory of the slave nodes to the specified hdfs directory of the master node of HBase, and then calls the bulkLoad operation of the master node of HBase to batch-load the files in the specified hdfs directory of the master node to the master node.
[0088] Those skilled in the art can understand that the modules in the above device embodiments can be distributed in the device according to the description, or can be correspondingly changed and distributed in one or more devices different from the above embodiments. The modules in the above embodiments can be combined into one module, or further split into multiple sub-modules.
[0089] Next, an embodiment of the electronic device of the present invention is described. This electronic device can be regarded as an implementation form of the entity for the above method and device embodiments of the present invention. For the details described in the embodiment of the electronic device of the present invention, it should be regarded as a supplement to the above method or device embodiments; for the details not disclosed in the embodiment of the electronic device of the present invention, reference can be made to the above method or device embodiments for implementation.
[0090] Figure 5 It is a structural block diagram of an exemplary embodiment of an electronic device according to the present invention. Figure 5 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0091] As Figure 5 shown, the electronic device 500 of this exemplary embodiment is presented in the form of a general-purpose data processing device. The components of the electronic device 500 may include, but are not limited to: at least one processing unit 510, at least one storage unit 520, a bus 530 connecting different electronic device components (including the storage unit 520 and the processing unit 510), a display unit 540, etc.
[0092] Among them, the storage unit 520 stores a computer-readable program, which may be the source program or the code of a read-only program. The program can be executed by the processing unit 510, so that the processing unit 510 executes the steps of various embodiments of the present invention. For example, the processing unit 510 can execute as Figure 1 shown in the steps.
[0093] The bus 530 can represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.
[0094] The electronic device 500 can also communicate with one or more external devices 100 (such as a keyboard, a display, a network device, a Bluetooth device, etc.), so that the user can interact with the electronic device 500 via these external devices 100, and / or enable the electronic device 500 to communicate with one or more other data processing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 550, and can also be through the network adapter 560 with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network). The network adapter 560 can communicate with other modules of the electronic device 500 through the bus 530.
[0095] Figure 6 It is a schematic diagram of an embodiment of a computer-readable medium of the present invention. As Figure 6 shown, the computer program can be stored on one or more computer-readable media. The computer-readable medium can be a readable signal medium or a readable storage medium. When the computer program is executed by one or more data processing devices, the computer-readable medium can implement the Figure 1 method of the present invention.
[0096] The present invention also provides a computer program product, including a computer program which, when executed by a processor, implements Figure 1 the method described above.
[0097] Based on Figure 1 the method for writing data warehouse data into HBase, an embodiment of the present invention also provides a system for writing data warehouse data into HBase, as Figure 7 shown. The system includes:
[0098] A new data warehouse push node 71, configured to convert offline data into the Hfile format to obtain an Hfile file, batch copy the Hfile file to a specified hdfs directory of Hbase; after the batch copy is successful, generate a status message to notify the online system application 72;
[0099] An online system application 72, configured to, when receiving the status message sent by the new data warehouse push node 71, perform concurrent control to batch load the Hfile file to each node of Hbase.
[0100] In a specific embodiment, the new data warehouse push node 71 is further configured to send the Hfile file to the online system application; correspondingly,
[0101] the online system application 72 queries the Hbase data after import based on the Hfile file, and processes the online feature data.
[0102] In the present invention, the data warehouse converts offline data into the Hfile format to obtain an Hfile file, and batch copies the Hfile file to a specified hdfs directory of Hbase; and generates a status message to notify the online system application after the batch copy is successful; thus, the online system application can perform concurrent control on each node of HBase to batch load the Hfile file by means of batch loading of Hfile data (the function that data in the Hfile format in HBase can be batch loaded), thereby improving the speed of writing data warehouse data into Hbase and reducing the load of the online system application. Compared with the prior art, the present invention has at least the following
[0103] beneficial effects:
[0104] 1. Directly convert the data warehouse offline data into the Hfile format, so that each node in HBase can batch load the Hfile file into HBase, avoiding the operation of writing each record of the data one by one, improving the data writing speed and at the same time enhancing the freshness of the data in HBase.
[0105] 2. By using the method of batch loading of Hfile data to control the batch loading of each node of HBase concurrently, the number of connections for data transmission between the online system and the data warehouse is reduced, and the writing pressure on HBase is alleviated.
[0106] 3. By implementing the writing operation of the online system through the method of batch loading of Hfile data, the queries per second (QPS) are reduced, the consumption of online resources is saved, and thus the operation cost is reduced.
[0107] In summary, the present invention can be implemented by a method, apparatus, system, electronic device or computer-readable medium that can execute computer programs. Some or all functions of the present invention can be implemented by using general data processing devices such as microprocessors or digital signal processors (DSPs) in practice.
[0108] The specific embodiments described above have further elaborated on the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the present invention is not inherently related to any specific computer, virtual device or electronic device, and various general devices can also implement the present invention. The above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A system for writing data warehouse data into HBase, comprising: A data warehouse new push node, configured to convert offline data into the Hfile format to obtain an Hfile file, and batch copy the Hfile file to a specified hdfs directory of Hbase; Generate a status message notification to the online system application after successful batch copy; An online system application, configured to, when receiving the status message sent by the data warehouse new push node, perform concurrent control to batch load the Hfile file into each node of Hbase.
2. A device for writing data warehouse data into HBase, characterized in that, The device includes: A batch copy module, configured to convert offline data into the Hfile format to obtain an Hfile file, and batch copy the Hfile file to a specified hdfs directory of Hbase; A generation module, configured to generate a status message indicating successful batch copy of the Hfile file after successful batch copy; A notification module, configured to notify the online system application of the status message through kafka, so that the online system application performs concurrent control to batch load the Hfile file into each node of Hbase.
3. The device according to claim 2, characterized in that, The batch copy module uses the distcp tool to batch copy the Hfile file to a specified hdfs directory of the slave nodes of HBase.
4. The device according to claim 2, characterized in that, The device further includes: A sending module, configured to send the Hfile file to the online system application, so that the online system application queries the Hbase data after import based on the Hfile file and processes the online feature data.
5. The device according to claim 2, characterized in that The device further includes: A switching module, configured to perform AB table switching after loading is completed.
6. The device according to claim 2, wherein The device further includes: A verification module, configured to perform comparison and key value case sensitivity verification on the offline data.
7. An online system application, characterized in that, Includes: A receiving module, configured to receive the status message sent by the data warehouse through kafka, where the status message is used to indicate that the converted Hfile file of the data warehouse is successfully batch copied to a specified hdfs directory of Hbase; A concurrent control module, configured to perform concurrent control to batch load the Hfile file into each node of Hbase.
8. The online system application according to claim 7, wherein The concurrent control module, through concurrent control, respectively calls the bulkLoad operation of the slave nodes in HBase to batch load the specified files in the hdfs directory of the slave nodes to the slave nodes; calls the distcp operation and the bulkLoad operation of the master node in HBase to batch copy the specified files in the hdfs directory of the slave nodes to the specified hdfs directory of the master node of HBase, and then calls the bulkLoad operation of the master node of HBase to batch load the files in the specified hdfs directory of the master node to the master node.
9. A method for writing data warehouse data into HBase, characterized in that, The method includes: Convert offline data into the Hfile format to obtain an Hfile file, and batch copy the Hfile file to a specified hdfs directory of Hbase; Generate a status message indicating successful batch copy of the Hfile file after successful batch copy; Notify the online system application of the status message through Kafka so that the online system application can perform concurrent control to batch load the Hfile files into each node of HBase.
10. The method according to claim 9, wherein Use the distcp tool to batch copy the Hfile files to the specified hdfs directory of the slave nodes of HBase.
11. The method according to claim 9, wherein Through concurrent control, the online system application separately calls the bulkLoad operation of the slave nodes in HBase to batch load the specified files in the hdfs directory of the slave nodes into the slave nodes; calls the distcp operation and the bulkLoad operation of the master node in HBase to batch copy the specified files in the hdfs directory of the slave nodes to the specified hdfs directory of the master node of HBase, and then calls the bulkLoad operation of the master node of HBase to batch load the files in the specified hdfs directory of the master node into the master node.
12. The method according to claim 9, wherein After obtaining the Hfile files, the method further includes: Send the Hfile files to the online system application so that the online system application can query the imported HBase data based on the Hfile files and process the online feature data.
13. The method according to claim 9, wherein The method further includes: Perform an AB table switch after the loading is completed.
14. The method according to claim 9, wherein Before converting the offline data into the Hfile format, the method further includes: Compare the offline data and perform a case-insensitive check on the key values.
15. An electronic device, comprising: A processor; And A memory storing computer-executable instructions, which when executed cause the processor to execute the method according to any one of claims 9 to 14.
16. A computer-readable storage medium, wherein, The computer-readable storage medium stores one or more programs, which when executed by the processor, implement the method according to any one of claims 9 to 14.
17. A computer program product comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the method according to any one of claims 9 to 14.
Citation Information
Cited By
Stream batch collaborative super-large-scale user risk real-time assessment system and method, electronic equipment and computer program product
CN121788158A