A Method, System, Storage Medium and Device for Triggering Alluxio Acceleration by Hive Events

By intercepting SQL updates in Hive and temporarily disabling Alluxio acceleration, the method ensures data consistency and efficient caching, addressing data inconsistency and metadata pressure issues in Alluxio synchronization with Hive tables.

CN115982294BActive Publication Date: 2025-07-15SHANGHAI ZHONGTONGJI NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310031100.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-10
Publication Date
2025-07-15
Estimated Expiration
2043-01-10

Smart Images

  • Figure CN115982294B_ABST
    Figure CN115982294B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, storage medium and device for triggering Alluxio acceleration by Hive events. The method includes: the data warehouse Hive receives and parses an SQL statement input by a user. When the SQL statement is an SQL statement for updating data in an accelerated table, the following steps are executed. The accelerated table is a Hive data table that has been synchronized to Alluxio and can be directly accelerated for query through Alluxio; intercept the SQL statement and trigger a pre-event; according to the pre-event, disable the acceleration attribute of the accelerated table so that the accelerated table is switched to query through the distributed file system Hdfs of Hive; execute the SQL statement to update the data in the accelerated table; intercept the SQL statement after the data is updated and trigger a post-event; according to the post-event, update the data corresponding to the accelerated table in Alluxio; enable the acceleration attribute of the accelerated table so that the accelerated table can be accelerated for query again through Alluxio. The present invention uses Hive events to trigger Alluxio updates, effectively achieving accelerated query while ensuring data consistency between Hdfs and Alluxio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data query, and in particular, to a method, system, storage medium and device for accelerating Alluxio by Hive event triggering. Background Art

[0002] Alluxio is the world's first memory-centric virtual distributed storage system. It unifies the data access method and builds a bridge for the upper-layer computing framework and the underlying storage system. An application only needs to connect to Alluxio to access the data stored in any underlying storage system. In addition, the memory-centric architecture of Alluxio enables the data access speed to be several orders of magnitude faster than existing conventional solutions.

[0003] Hive is a data warehouse software that uses SQL language to assist in reading, writing, and managing large datasets stored on the distributed file system HDFS. The biggest feature is to analyze big data through SQL-like statements. The data is stored on HDFS, and it does not provide data storage function itself.

[0004] In the existing Alluxio technology solution for synchronizing Hive table data, the data on Hdfs is synchronously transferred to Alluxio actively and periodically. The specific process is as follows:

[0005] (1) A timing task initiates polling to detect whether the data of the source table has changed. The data change usually takes the table partition as a unit.

[0006] (2) Initiate a data synchronization task. If it is a new partition, first add the partition to the Alluxio table, and then mount the partition data to Alluxio; for historical partitions, remount the partitions with data changes to Alluxio.

[0007] However, the above solution has the following disadvantages:

[0008] 1. There is data inconsistency. When the data on Hdfs is updated, it cannot be reflected in Alluxio in a timely manner. At this time, when a user queries Alluxio, old data is still pulled.

[0009] 2. Cold data in Alluxio cannot be unmounted, resulting in wasted cache.

[0010] 3. There are many updates to historical data on Hdfs, resulting in a large amount of metadata. Then, the Alluxio metadata synchronization requests will increase, which will put pressure on the Alluxio master node and cause query jitter. Summary of the Invention

[0011] Based on this, it is necessary to propose a method, system, storage medium and device for accelerating Alluxio triggered by Hive events to address the above problems.

[0012] The present invention discloses a method for accelerating Alluxio triggered by Hive events, and the method includes:

[0013] The data warehouse Hive receives and parses the SQL statement input by the user. When the SQL statement is an SQL statement for updating data in the acceleration table, the following steps are executed. The acceleration table is a Hive data table that has been synchronized to the distributed storage system Alluxio and can be directly accelerated for query through the distributed storage system Alluxio:

[0014] Intercept the SQL statement and trigger a pre-event;

[0015] According to the pre-event, disable the acceleration attribute of the acceleration table so that the acceleration table is switched to query through the distributed file system Hdfs of the data warehouse Hive;

[0016] Execute the SQL statement to update the data in the acceleration table;

[0017] Intercept the SQL statement after the data is updated and trigger a post-event;

[0018] According to the post-event, update the data corresponding to the acceleration table in the distributed storage system Alluxio;

[0019] Enable the acceleration attribute of the acceleration table so that the acceleration table can be accelerated for query again through the distributed storage system Alluxio.

[0020] Further, updating the data corresponding to the acceleration table in the distributed storage system Alluxio specifically includes:

[0021] Delete the old data table corresponding to the acceleration table in the distributed storage system Alluxio;

[0022] Disassociate the data path on the distributed storage system Alluxio corresponding to the old data table from the distributed file system Hdfs;

[0023] Re-create a new data table corresponding to the acceleration table in the distributed storage system Alluxio according to the table creation statement of the acceleration table;

[0024] Modify the data path of the new data table from the distributed file system Hdfs to the distributed storage system Alluxio;

[0025] Reassign the data permissions of the accelerometer table to the new data table;

[0026] Mount the new data table and re-add the partitions of the accelerometer table to the new data table.

[0027] Furthermore, the data warehouse Hive includes two threads: a pre-thread and a post-thread. The pre-thread processes the pre-events, and the post-thread processes the post-events.

[0028] Furthermore, after the step of the data warehouse Hive receiving and parsing the SQL statement input by the user, it further includes:

[0029] Filter the SQL statement input by the user to filter out the SQL statements for updating the data in the accelerometer table. Among them, the filtered SQL statements include at least one of the following:

[0030] SQL statements for implementing the data table deletion function;

[0031] SQL statements for implementing the data table creation function;

[0032] SQL statements for implementing the data table insertion function;

[0033] SQL statements for implementing the function of adding, deleting, or changing the data table structure or table name;

[0034] SQL statements for implementing the permission assignment function;

[0035] SQL statements for implementing the permission recycling function.

[0036] On the other hand, the present invention also discloses a system for triggering Alluxio acceleration by Hive events. The system includes:

[0037] A receiving and parsing module for the data warehouse Hive to receive and parse the SQL statement input by the user. When the SQL statement is an SQL statement for updating the data in the accelerometer table, the following modules are used to implement the triggering of Alluxio acceleration by Hive events. The accelerometer table is a hive data table that has been synchronized to the distributed storage system Alluxio and can be directly accelerated for query through the distributed storage system Alluxio;

[0038] A pre-event triggering module for intercepting the SQL statement and triggering a pre-event;

[0039] An acceleration disabling module for disabling the acceleration attribute of the accelerometer table according to the pre-event, so that the accelerometer table is switched to query through the distributed file system Hdfs of the data warehouse Hive;

[0040] An acceleration table update module for executing the SQL statement to update the data in the acceleration table;

[0041] A post-event trigger module for intercepting the SQL statement after data update and triggering a post-event;

[0042] An Alluxio update module for updating the data corresponding to the acceleration table in the distributed storage system Alluxio according to the post-event;

[0043] An acceleration restart module for enabling the acceleration attribute of the acceleration table so that the acceleration table can perform accelerated queries again through the distributed storage system Alluxio.

[0044] Furthermore, the Alluxio update module specifically includes:

[0045] A deletion module for deleting the old data table corresponding to the acceleration table in the distributed storage system Alluxio;

[0046] An unmounting module for disassociating the data path of the distributed storage system Alluxio corresponding to the old data table from the distributed file system Hdfs;

[0047] A reconstruction module for recreating a new data table corresponding to the acceleration table in the distributed storage system Alluxio according to the table creation statement of the acceleration table;

[0048] A path modification module for modifying the data path of the new data table from the distributed file system Hdfs to the distributed storage system Alluxio;

[0049] An authorization module for reassigning the data permissions that the acceleration table has to the new data table;

[0050] A mounting module for mounting the new data table and adding the partitions that the acceleration table has back to the new data table.

[0051] Furthermore, the data warehouse Hive includes two threads: a pre-thread and a post-thread, and the pre-event is processed through the pre-thread, and the post-event is processed through the post-thread.

[0052] Furthermore, the system further includes:

[0053] An SQL filtering module for filtering the SQL statement input by the user and filtering out the SQL statement for updating the data in the acceleration table, where the filtered SQL statement includes at least one of the following:

[0054] SQL statements for implementing the function of deleting data tables;

[0055] SQL statements for implementing the function of creating data tables;

[0056] SQL statements for implementing the function of inserting data into data tables;

[0057] SQL statements for implementing the function of adding, deleting, or changing the structure or name of data tables;

[0058] SQL statements for implementing the function of assigning permissions;

[0059] SQL statements for implementing the function of reclaiming permissions.

[0060] On the other hand, the present invention also discloses a computer device, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs the following steps:

[0061] The data warehouse Hive receives and parses the SQL statement input by the user. When the SQL statement is an SQL statement for updating data in the acceleration table, the following steps are executed. The acceleration table is a hive data table that has been synchronized to the distributed storage system Alluxio and can be directly accelerated for query through the distributed storage system Alluxio:

[0062] Intercept the SQL statement and trigger a pre-event;

[0063] According to the pre-event, disable the acceleration attribute of the acceleration table so that the acceleration table is converted to query through the distributed file system Hdfs of the data warehouse Hive;

[0064] Execute the SQL statement to update the data in the acceleration table;

[0065] Intercept the SQL statement after the data is updated and trigger a post-event;

[0066] According to the post-event, update the data corresponding to the acceleration table in the distributed storage system Alluxio;

[0067] Enable the acceleration attribute of the acceleration table so that the acceleration table can be accelerated for query again through the distributed storage system Alluxio.

[0068] On the other hand, the present invention also discloses a computer-readable storage medium, storing a computer program. When the computer program is executed by a processor, the processor performs the following steps:

[0069] The data warehouse Hive receives and parses the SQL statements input by users. When the SQL statement is an SQL statement for updating data in the acceleration table, the following steps are executed. The acceleration table is a Hive data table that has been synchronized to the distributed storage system Alluxio and can be directly accelerated for query through the distributed storage system Alluxio:

[0070] Intercept the SQL statement and trigger a pre-event;

[0071] According to the pre-event, disable the acceleration attribute of the acceleration table so that the acceleration table is switched to query through the distributed file system Hdfs of the data warehouse Hive;

[0072] Execute the SQL statement to update the data in the acceleration table;

[0073] Intercept the SQL statement after the data is updated and trigger a post-event;

[0074] According to the post-event, update the data corresponding to the acceleration table in the distributed storage system Alluxio;

[0075] Enable the acceleration attribute of the acceleration table so that the acceleration table can be accelerated for query again through the distributed storage system Alluxio.

[0076] Adopting the embodiment of the present invention has the following beneficial effects:

[0077] The present invention intercepts the SQL statement for changing the data in the acceleration table, triggers the Hive pre-event and post-event, and according to the pre-event, disables the acceleration attribute of the acceleration table and switches to query by pulling data through Hdfs. After the SQL statement for updating the data is executed, according to the post-event, the data in Alluxio is updated and the data table is remounted, so that the acceleration query can be realized again through Alluxio, effectively ensuring the data consistency between Hdfs and Alluxio while realizing the acceleration query. Description of the Drawings

[0078] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0079] Among them:

[0080] Figure 1 It is a flowchart of a method for triggering Alluxio acceleration by Hive events in an embodiment;

[0081] Figure 2 It is a structural block diagram of a system for triggering Alluxio acceleration by Hive events in an embodiment;

[0082] Figure 3 It is a structural block diagram of a computer device in an embodiment.

[0083] Explanation of reference numerals: receiving and parsing module 100, pre-event trigger module 200, acceleration disabling module 300, acceleration table update module 400, post-event trigger module 500, Alluxio update module 600, acceleration restart module 700. Specific implementation mode

[0084] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0085] As Figure 1 shown, in an embodiment, a method for triggering Alluxio acceleration by Hive events is provided. The method specifically includes the following steps:

[0086] S1. The data warehouse Hive receives and parses the SQL statement input by the user. When the SQL statement is an SQL statement for updating data in the acceleration table, the following steps are executed. The acceleration table is a hive data table that has been synchronized to the distributed storage system Alluxio and can be directly accelerated for query through the distributed storage system Alluxio:

[0087] S2. Intercept the SQL statement and trigger a pre-event;

[0088] S3. According to the pre-event, disable the acceleration attribute of the acceleration table so that the acceleration table is converted to query through the distributed file system Hdfs of the data warehouse Hive;

[0089] S4. Execute the SQL statement to update the data in the acceleration table;

[0090] S5. Intercept the SQL statement after the data is updated and trigger a post-event;

[0091] S6. According to the post-event, update the data corresponding to the acceleration table in the distributed storage system Alluxio;

[0092] S7. Enable the acceleration attribute of the acceleration table so that the acceleration table can perform accelerated queries through the distributed storage system Alluxio again.

[0093] The data in the data warehouse Hive is stored in the distributed file system Hdfs. Actively and periodically pulling the data in the distributed file system Hdfs to the distributed storage system Alluxio can enable users to directly perform data query access through the distributed storage system Alluxio, greatly improving the data access speed and effectively achieving accelerated queries. However, in the prior art, during the process of actively and periodically pulling the data in the distributed file system Hdfs to the distributed storage system Alluxio, there are certain operation gaps, such that during this gap time, the requests of users query the old data in Alluxio, resulting in data inconsistency between Hdfs and Alluxio. In addition, in the existing solution, Alluxio caches all the data of the source table. Since there are hot and cold data in the table, it is bound to cause waste of caching. And if there are many data changes and updates on Hdfs, a large amount of metadata is generated, which will inevitably put pressure on the Alluxio master node during data synchronization, thus causing query jitter.

[0094] Therefore, to solve the above problems existing in the prior art, in this embodiment, the Hive data tables that can be already synchronized to the distributed storage system Alluxio and can directly perform accelerated queries through the distributed storage system Alluxio are used as acceleration tables. When Hive receives an SQL statement for updating the data of the acceleration table, it first intercepts the statement, triggers a pre-event, and disables the acceleration attribute of the acceleration table, so that the acceleration table switches to query through the distributed file system Hdfs. Because at this time, the data of the acceleration table in Hdfs starts to change, and Alluxio has not been synchronized and updated yet, so there is no latest data in Alluxio. This step can effectively avoid the occurrence of data inconsistency between Hdfs and Alluxio, so that even at this gap moment, what the user queries is not the old data from Alluxio, but the new data to be updated in Hdfs. Then, update the data of the acceleration table according to the specific SQL statement. After the update, intercept again, trigger a post-event, and synchronize and update the accelerated table after the data update to Alluxio. Finally, re-enable the acceleration attribute of the acceleration table so that the user can perform accelerated queries through Alluxio again. At this time, since Alluxio is re-synchronized according to the updated acceleration table, the data in it is already the latest data, so what the user queries through Alluxio is also the updated data, effectively solving the problem of data inconsistency between Hdfs and Alluxio.

[0095] Specifically, in this embodiment, the interception of SQL trigger pre - and post - events in steps S2 and S4 can be implemented by Hive hooks in Hive. By inheriting the Hive interface ExecuteWithHookContext and implementing the run() method, while the processing of pre - and post - events in steps S3, S6, and S7 can be achieved through an additional deployed service Hook handle. After the service starts, two threads are resident. One is a pre - thread for processing pre - events, and the other is a post - thread for processing post - events.

[0096] Further, in one embodiment, step S6 specifically includes:

[0097] S61. Delete the old data table corresponding to the acceleration table in the distributed storage system Alluxio;

[0098] S62. Disassociate the data path on the distributed storage system Alluxio corresponding to the old data table from the distributed file system Hdfs;

[0099] S63. Re - create a new data table corresponding to the acceleration table in the distributed storage system Alluxio according to the table - creation statement of the acceleration table;

[0100] S64. Modify the data path of the new data table from the distributed file system Hdfs to the distributed storage system Alluxio;

[0101] S65. Re - assign the data permissions of the acceleration table to the new data table;

[0102] S66. Mount the new data table and re - add the partitions that the acceleration table has to the new data table.

[0103] In this embodiment, to update the data corresponding to the acceleration table in Alluxio, first delete the old data table corresponding to the acceleration table contained in the original Alluxio, and disassociate the data path on Alluxio from Hdfs. Then re - create a new data table and re - modify the path for mounting. After the above process, the hive acceleration table and the data table in Alluxio are re - associated, so as to open the table acceleration to achieve the fast query function. It should be noted that in this embodiment, the entire data table is re - mounted instead of mounting a certain partition of a certain table. After re - mounting the table, the corresponding partition data can be queried and pulled as needed, which is significantly different from the existing technical solutions.

[0104] In this embodiment, when the accelerometer data changes, all the data cached in the previous Alluxio can be deleted and unmounted in a timely manner, and then remounted after the data is updated, so that the data cached in the current Alluxio are all hot data that may be queried in real time, effectively avoiding the problem of Alluxio cache waste caused by the inability to unmount cold data.

[0105] Further, in one embodiment, the data warehouse Hive includes two threads: a pre-thread and a post-thread. The pre-event is processed by the pre-thread, and the post-event is processed by the post-thread.

[0106] In this embodiment, the triggered pre-events and post-events can be sent to different message queues for queuing processing during the specific operation process, and different threads are used to listen to and process different events respectively, thereby effectively alleviating and reducing the pressure on the Alluxio master node during data processing and improving the data processing and query efficiency.

[0107] Further, in one embodiment, after step S1, it further includes:

[0108] Filter the SQL statements input by the user, and filter out the SQL statements for updating the data in the acceleration table. Among them, the filtered SQL statements include at least one of the following:

[0109] SQL statements for implementing the function of deleting a data table, such as: drop table such SQL statements, triggering pre-events and post-events of the drop type;

[0110] SQL statements for implementing the function of creating or inserting a data table, such as: create table as and insert such SQL statements, triggering pre-events and post-events of the insert type;

[0111] SQL statements for implementing the function of modifying the data table name, such as: alter tablename such SQL statements, triggering pre-events and post-events of the rename type;

[0112] SQL statements for implementing the function of modifying the data table, such as: alter table such SQL statements, triggering pre-events and post-events of the ddl type;

[0113] SQL statements for implementing the function of allocating or revoking permissions, such as: grant and revoke such SQL statements, triggering pre-events and post-events of the grant type;

[0114] Not all SQL statements input by users can trigger Hive events. Therefore, in this embodiment, it is necessary to first perform category screening and filtering on the SQL statements to intercept the SQL statements that affect the change of the accelerated table data, and then the Hive event and subsequent steps such as Alluxio data update can be triggered to continue running.

[0115] In addition, as Figure 2 shown, in one embodiment, a system for triggering Alluxio acceleration by Hive events is further provided. The system includes:

[0116] A receiving and parsing module 100, configured to receive and parse SQL statements input by users in the data warehouse Hive. When the SQL statement is an SQL statement for updating data in the accelerated table, the following modules are used to implement triggering Alluxio acceleration by Hive events. The accelerated table is a Hive data table that has been synchronized to the distributed storage system Alluxio and can be directly accelerated for query through the distributed storage system Alluxio;

[0117] A pre-event triggering module 200, configured to intercept the SQL statement and trigger a pre-event;

[0118] An acceleration disabling module 300, configured to disable the acceleration attribute of the accelerated table according to the pre-event, so that the accelerated table is converted to be queried through the distributed file system Hdfs of the data warehouse Hive;

[0119] An accelerated table update module 400, configured to execute the SQL statement and update the data in the accelerated table;

[0120] A post-event triggering module 500, configured to intercept the SQL statement after the data is updated and trigger a post-event;

[0121] An Alluxio update module 600, configured to update the data corresponding to the accelerated table in the distributed storage system Alluxio according to the post-event;

[0122] An acceleration restart module 700, configured to enable the acceleration attribute of the accelerated table, so that the accelerated table can be accelerated for query again through the distributed storage system Alluxio.

[0123] Further, in one embodiment, the Alluxio update module 600 specifically includes:

[0124] A deletion module 610, configured to delete the old data table corresponding to the accelerated table in the distributed storage system Alluxio;

[0125] An uninstallation module 620, configured to disassociate the Alluxio data path corresponding to the old data table from the Hadoop Distributed File System (HDFS);

[0126] A reconstruction module 630, configured to recreate a new data table corresponding to the acceleration table in the Alluxio distributed storage system according to the table creation statement of the acceleration table;

[0127] A path modification module 640, configured to modify the data path of the new data table from the HDFS to the Alluxio distributed storage system;

[0128] An authorization module 650, configured to re - assign the data permissions of the acceleration table to the new data table;

[0129] A mounting module 660, configured to mount the new data table and re - add the partitions of the acceleration table to the new data table.

[0130] Further, in one embodiment, the data warehouse Hive includes two threads: a pre - thread and a post - thread. The pre - thread is used to process the pre - events, and the post - thread is used to process the post - events.

[0131] Further, in one embodiment, the system further includes:

[0132] An SQL filtering module 800, configured to filter the SQL statements input by the user, and filter out the SQL statements for updating the data in the acceleration table. Among them, the filtered SQL statements include at least one of the following:

[0133] SQL statements for implementing the data table deletion function;

[0134] SQL statements for implementing the data table creation function;

[0135] SQL statements for implementing the data table insertion function;

[0136] SQL statements for implementing functions of adding, deleting, or changing the data table structure or table name;

[0137] SQL statements for implementing the permission assignment function;

[0138] SQL statements for implementing the permission recycling function.

[0139] Figure 3 The internal structure diagram of a computer device in one embodiment is shown. This computer device can specifically be a terminal or a server. For example, Figure 3As shown in the figure, the computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and can also store a computer program. When the computer program is executed by the processor, the processor can implement the method for triggering Alluxio acceleration by Hive events. The internal memory can also store a computer program. When the computer program is executed by the processor, the processor can execute the method for triggering Alluxio acceleration by Hive events. Those skilled in the art can understand that Figure 3 The structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0140] In one embodiment, a computer device is proposed, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs the following steps:

[0141] The data warehouse Hive receives and parses the SQL statement input by the user. When the SQL statement is an SQL statement for updating data in the acceleration table, the following steps are executed. The acceleration table is a hive data table that has been synchronized to the distributed storage system Alluxio and can be directly accelerated for query through the distributed storage system Alluxio:

[0142] Intercept the SQL statement and trigger a pre-event;

[0143] According to the pre-event, disable the acceleration attribute of the acceleration table so that the acceleration table is switched to query through the distributed file system Hdfs of the data warehouse Hive;

[0144] Execute the SQL statement to update the data in the acceleration table;

[0145] Intercept the SQL statement after the data is updated and trigger a post-event;

[0146] According to the post-event, update the data corresponding to the acceleration table in the distributed storage system Alluxio;

[0147] Enable the acceleration attribute of the acceleration table so that the acceleration table can be accelerated for query again through the distributed storage system Alluxio.

[0148] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which when executed by a processor, causes the processor to perform the following steps:

[0149] The data warehouse Hive receives and parses an SQL statement input by a user. When the SQL statement is an SQL statement for updating data in an acceleration table, the following steps are performed. The acceleration table is a Hive data table that has been synchronized to the distributed storage system Alluxio and can be directly accelerated for query through the distributed storage system Alluxio:

[0150] Intercept the SQL statement and trigger a pre-event;

[0151] According to the pre-event, disable the acceleration attribute of the acceleration table so that the acceleration table is switched to query through the distributed file system Hdfs of the data warehouse Hive;

[0152] Execute the SQL statement to update the data in the acceleration table;

[0153] Intercept the SQL statement after the data is updated and trigger a post-event;

[0154] According to the post-event, update the data corresponding to the acceleration table in the distributed storage system Alluxio;

[0155] Enable the acceleration attribute of the acceleration table so that the acceleration table can be accelerated for query again through the distributed storage system Alluxio.

[0156] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0157] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0158] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for triggering Alluxio acceleration by Hive events, characterized in that The method includes: The data warehouse Hive receives and parses the SQL statement input by the user. When the SQL statement is an SQL statement for updating data in the acceleration table, the following steps are executed. The acceleration table is a Hive data table that has been synchronized to the distributed storage system Alluxio and can be directly accelerated for query through the distributed storage system Alluxio: Intercept the SQL statement and trigger a pre-event; According to the pre-event, disable the acceleration attribute of the acceleration table so that the acceleration table is changed to query through the distributed file system Hdfs of the data warehouse Hive; Execute the SQL statement to update the data in the acceleration table; Intercept the SQL statement after the data is updated and trigger a post-event; According to the post-event, update the data corresponding to the acceleration table in the distributed storage system Alluxio; Enable the acceleration attribute of the acceleration table so that the acceleration table can be accelerated for query again through the distributed storage system Alluxio; Updating the data corresponding to the acceleration table in the distributed storage system Alluxio specifically includes: Delete the old data table corresponding to the acceleration table in the distributed storage system Alluxio; Disassociate the data path on the distributed storage system Alluxio corresponding to the old data table from the distributed file system Hdfs; Re-create a new data table corresponding to the acceleration table in the distributed storage system Alluxio according to the table creation statement of the acceleration table; Modify the data path of the new data table from the distributed file system Hdfs to the distributed storage system Alluxio; Re-assign the data permissions that the acceleration table has to the new data table; Mount the new data table and re-add the partitions that the acceleration table has to the new data table.

2. The method for triggering Alluxio acceleration by Hive events according to claim 1, wherein The data warehouse Hive includes two threads: a pre-thread and a post-thread, and the pre-event is processed through the pre-thread and the post-event is processed through the post-thread.

3. A method for triggering Alluxio acceleration by Hive events according to claim 1, characterized in that, After the step where the data warehouse Hive receives and parses the SQL statement input by the user, it further includes: Filter the SQL statement input by the user to filter out the SQL statement for updating the data in the acceleration table. Among them, the filtered SQL statement includes at least one of the following: An SQL statement for implementing the function of deleting a data table; An SQL statement for implementing the function of creating a data table; An SQL statement for implementing the function of inserting into a data table; An SQL statement for implementing the function of adding, deleting, or changing the data table structure or table name; An SQL statement for implementing the function of assigning permissions; An SQL statement for implementing the function of revoking permissions.

4. A system for triggering Alluxio acceleration by Hive events, characterized in that, The system includes: A receiving and parsing module for receiving and parsing SQL statements input by users in the data warehouse Hive. When the SQL statement is an SQL statement for updating data in the acceleration table, the following modules are used to implement Hive event-triggered Alluxio acceleration. The acceleration table is a hive data table that has been synchronized to the distributed storage system Alluxio and can be directly accelerated for query through the distributed storage system Alluxio; A pre-event trigger module for intercepting the SQL statement and triggering a pre-event; An acceleration disabling module for disabling the acceleration attribute of the acceleration table according to the pre-event, so that the acceleration table is converted to query through the distributed file system Hdfs of the data warehouse Hive; An acceleration table update module for executing the SQL statement to update the data in the acceleration table; A post-event trigger module for intercepting the SQL statement after the data is updated and triggering a post-event; An Alluxio update module for updating the data corresponding to the acceleration table in the distributed storage system Alluxio according to the post-event; An acceleration restart module for enabling the acceleration attribute of the acceleration table so that the acceleration table can be accelerated for query again through the distributed storage system Alluxio; The Alluxio update module specifically includes: A deletion module for deleting the old data table corresponding to the acceleration table in the distributed storage system Alluxio; An unmounting module for disassociating the distributed storage system Alluxio data path corresponding to the old data table from the distributed file system Hdfs; A reconstruction module for re-creating a new data table corresponding to the acceleration table in the distributed storage system Alluxio according to the table creation statement of the acceleration table; A path modification module for modifying the data path of the new data table from the distributed file system Hdfs to the distributed storage system Alluxio; A permission granting module for re-granting the data permissions that the acceleration table has to the new data table; A mounting module for mounting the new data table and re-adding the partitions that the acceleration table has to the new data table.

5. The system for triggering Alluxio acceleration by Hive events according to claim 4, wherein The data warehouse Hive includes two threads: a pre-thread and a post-thread, and the pre-event is processed through the pre-thread, and the post-event is processed through the post-thread.

6. The system for triggering Alluxio acceleration by Hive events according to claim 4, wherein The system further includes: An SQL filtering module for filtering the SQL statements input by users to filter out the SQL statements for updating data in the acceleration table. Among them, the filtered SQL statements include at least one of the following: An SQL statement for implementing the function of deleting a data table; An SQL statement for implementing the function of creating a data table; An SQL statement for implementing the function of inserting into a data table; An SQL statement for implementing the function of adding, deleting, or changing the data table structure or table name; An SQL statement for implementing the function of allocating permissions; An SQL statement for implementing the function of reclaiming permissions.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 3.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the computer program is executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Method and device used for improving stability of distributed database middleware

    CN107943870A

  • Distributed transaction processing method and framework

    CN111008202A