Data processing method, device, storage medium and program product for data lake
Patent Information
- Application Number
- CN202610953760.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-18
AI Technical Summary
[0003]然而,在数据湖数据重构或优化完成后,如何将成百上千的下游数据应用任务平滑、低风险地迁移至新的数据链路,是一个普遍存在的技术挑战
[0018] This paper presents a data lake data processing method, device, storage medium, and program product. It receives a first query request to a first data table in the data lake. The first data table includes at least one virtual partition, which is pre-created and associated with a first query statement. This first query statement is associated with at least one second data table. The method determines the partition filtering conditions included in the first query request, queries matching partitions in the first data table based on these conditions, and determines the type of matching partition. If the matching partition is a virtual partition, the method executes the first query statement associated with the virtual partition, obtaining the corresponding first query result. By adding virtual partitions to the first data table and associating them with the first query statement, data from other data sources can be obtained through the first query statement. This allows for small-scale reconstruction or optimization of the first data table. Data consumers only need to obtain the reconstructed or optimized data from the first data table through the virtual partitions, without requiring a complete data migration of the first data table. It also enables partition-level canary releases, greatly reducing the risk of data migration, ensuring business continuity, and providing complete transparency to downstream data consumers. No modification to the data consumer's query code is required, thus reducing the impact of reconstruction or optimization on downstream businesses.
Smart Images

Figure CN122777583A_ABST
Abstract
Description
Technical Field
[0001] This article relates to the field of computer technology, and in particular to a data processing method, device, storage medium and program product for a data lake. Background Technology
[0002] In the process of data lake evolution and maintenance, data reconstruction and optimization is a routine and necessary activity.
[0003] However, after the data lake has been reconstructed or optimized, how to smoothly and with low risk migrate hundreds or thousands of downstream data application tasks to the new data link is a common technical challenge. Summary of the Invention
[0004] This article provides a data processing method, device, storage medium, and program product for data lakes to ensure the stability of data lake reconstruction and optimization.
[0005] Firstly, this paper provides a data processing method for data lakes, including:
[0006] Receive a first query request for a first data table in the data lake; wherein the partition of the first data table includes at least one virtual partition, the virtual partition is pre-created and associated with a first query statement, the first query statement being a query statement associated with at least one second data table;
[0007] Determine the partition filtering conditions included in the first query request, query the matching partitions in the first data table according to the partition filtering conditions, and determine the type of the matching partitions;
[0008] If the matched partition is a virtual partition, the first query statement associated with the virtual partition is executed to obtain the corresponding first query result.
[0009] Secondly, this paper provides a data processing device for a data lake, including:
[0010] A receiving unit is configured to receive a first query request for a first data table in a data lake; wherein the partition of the first data table includes at least one virtual partition, the virtual partition is pre-created and associated with a first query statement, the first query statement being a query statement associated with at least one second data table;
[0011] A partition management unit is used to determine the partition filtering conditions included in the first query request, query the matching partitions in the first data table according to the partition filtering conditions, and determine the type of the matching partitions.
[0012] An execution unit is configured to, in response to the matching partition being a virtual partition, execute a first query statement associated with the virtual partition to obtain the corresponding first query result.
[0013] Thirdly, this article provides an electronic device, including: a processor and a memory;
[0014] The memory stores computer-executed instructions;
[0015] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the data processing method for the data lake as described in the first aspect and various possible designs of the first aspect.
[0016] Fourthly, this document provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the data processing method for the data lake as described in the first aspect and various possible designs of the first aspect.
[0017] Fifthly, this document provides a computer program product, including a computer program that, when executed by a processor, implements the data processing method for the data lake as described in the first aspect and various possible designs of the first aspect.
[0018] This paper presents a data lake data processing method, device, storage medium, and program product. It receives a first query request to a first data table in the data lake. The first data table includes at least one virtual partition, which is pre-created and associated with a first query statement. This first query statement is associated with at least one second data table. The method determines the partition filtering conditions included in the first query request, queries matching partitions in the first data table based on these conditions, and determines the type of matching partition. If the matching partition is a virtual partition, the method executes the first query statement associated with the virtual partition, obtaining the corresponding first query result. By adding virtual partitions to the first data table and associating them with the first query statement, data from other data sources can be obtained through the first query statement. This allows for small-scale reconstruction or optimization of the first data table. Data consumers only need to obtain the reconstructed or optimized data from the first data table through the virtual partitions, without requiring a complete data migration of the first data table. It also enables partition-level canary releases, greatly reducing the risk of data migration, ensuring business continuity, and providing complete transparency to downstream data consumers. No modification to the data consumer's query code is required, thus reducing the impact of reconstruction or optimization on downstream businesses. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this document or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this document. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A schematic diagram illustrating a data lake data processing method provided in one embodiment of this paper;
[0021] Figure 2 This is a schematic diagram of a data lake data processing method provided in one embodiment of the present invention;
[0022] Figure 3 A schematic diagram of a data lake data processing method provided in another embodiment of this article;
[0023] Figure 4 This is a structural block diagram of a data processing device for a data lake provided in one embodiment of the present invention;
[0024] Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in one embodiment of this article. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this document clearer, the technical solutions described below will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments described herein, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments described herein without inventive effort are within the scope of protection of this document.
[0026] In the evolution and maintenance of data lakes, data reconstruction and optimization are routine and necessary activities, including but not limited to data migration, logical reconstruction, and storage optimization. A data lake is a centralized data storage architecture or technology system used to store raw data of any size. It supports direct storage of structured, semi-structured, and unstructured data in its raw format without the need for a predefined model (i.e., "read-time mode").
[0027] However, after the data lake has been reconstructed or optimized, how to smoothly and with low risk migrate hundreds or thousands of downstream data application tasks to the new data link is a common technical challenge.
[0028] In existing technologies, after the reconstruction or optimization of data lake data, migrating the old data link to the new data link typically relies on a one-time replacement of the data link. This means that all downstream applications will switch from the old data link to the new data link at the same time. This "one-size-fits-all" approach lacks an effective gray-scale and verification period. If the data logic of the new data link has unforeseen defects or performance bottlenecks, it will cause large-scale data quality problems or task failures, severely impacting business operations.
[0029] To address the aforementioned technical challenges, this paper presents a data lake processing method that introduces the concept of virtual partitions. By adding virtual partitions to a first data table and associating them with a first query statement, data from other data sources can be retrieved through this first query statement. This allows for small-scale reconstruction or optimization of the first data table. Data consumers only need to obtain the reconstructed or optimized data from the first data table through the virtual partitions, without requiring a complete data migration of the entire first data table. Furthermore, it enables partition-level canary releases, significantly reducing the risk of data migration, ensuring business continuity, and remaining completely transparent to downstream data consumers. No modification to the data consumer's query code is required, thus minimizing the impact of reconstruction or optimization on downstream businesses.
[0030] The application scenarios of the data lake data processing methods described in this article are as follows: Figure 1 As shown, this can be applied to the query engine of a data lake. A data consumer can send a first query request to a first data table in the data lake, wherein the partitions of the first data table include at least one virtual partition. The virtual partition is pre-created and associated with a first query statement, which is a query statement associated with at least one second data table. The partition filtering conditions included in the first query request are determined, and matching partitions in the first data table are queried according to the partition filtering conditions, and the type of matching partition is determined. In response to the matching partition being a virtual partition, the first query statement associated with the virtual partition is executed, the corresponding first query result is obtained, and returned to the data consumer.
[0031] The data processing method for data lakes described in this paper will be described in detail below with reference to specific embodiments.
[0032] refer to Figure 2 , Figure 2 This is a schematic flowchart of a data lake data processing method provided in one embodiment of this paper. This data lake data processing method can be applied to the query engine of a data lake, and includes:
[0033] S201. Receive a first query request for a first data table in the data lake; wherein the partition of the first data table includes at least one virtual partition, the virtual partition is pre-created and associated with a first query statement, the first query statement is a query statement associated with at least one second data table.
[0034] To avoid a one-time full replacement of data lake data after reconstruction or optimization, this paper supports small-scale reconstruction or optimization of data lake data. Through a gradual strategy, the granularity of reconstruction or optimization can be finely controlled, reducing the risk of data migration, ensuring business continuity, and facilitating problem location and canary release.
[0035] For the first data table in a data lake, after performing data migration, logical restructuring, and storage optimization on some data in the first data table, the results of the restructuring or optimization are usually written to at least one second data table. To ensure that data consumers, such as downstream users or applications, are unaware of the changes in the data chain in the first data table and do not need to modify any query code when querying data, a virtual partition can be created in the first data table. A first query statement can be associated with the virtual partition. The first query statement is a query statement for the results of the restructuring or optimization, that is, a query statement for at least one second data table. It can be preset or generated in real time. In this way, if the data consumer needs to query the restructured or optimized data in the first data table, it can still continue to query the first data table. The query engine will locate the virtual partition, execute the first query statement associated with the virtual partition, query at least one second data table, and then return the results to the data consumer, making the data consumer unaware of the changes. In addition, the restructured or optimized data can still be stored in the physical partition of the first data table, which can support gray-scale testing and rollback (the virtual partition can be deleted during rollback).
[0036] Based on this, the first data table may specifically include physical partitions and virtual partitions, such as... Figure 3 As shown, the content in the physical partition is the physical file or a static pointer to the physical file. The physical file can be data that has not been reconstructed or optimized. The virtual partition, on the other hand, does not include the physical file or the static pointer to the physical file. Instead, it is associated with the first query statement. This is equivalent to turning the virtual partition into a dynamic execution unit that encapsulates the computing logic. Its core idea is to extend the concept of data partitioning from physical storage to logical definition, thus realizing a partition-level view.
[0037] Of course, the first data table can also consist entirely of virtual partitions. For example, if all data in the first data table has been reconstructed or optimized and does not require rollback, then physical partitions can be deleted. In addition, if the first data table has not been reconstructed or optimized, the first data table may not include virtual partitions and may only include physical partitions.
[0038] S202. Determine the partition filtering conditions included in the first query request, query the matching partitions in the first data table according to the partition filtering conditions, and determine the type of matching partitions.
[0039] The query engine can parse the partition filtering conditions included in the first query request. Optionally, the query engine can parse the partition filtering conditions in the WHERE clause of the first query request through the optimizer rule. Of course, the partition filtering conditions included in the first query request can be determined in other ways, which are not limited here.
[0040] Furthermore, the query engine can query partitions in the first data table that match the partition filter conditions; these are referred to here as matching partitions. Specifically, it can request matching partitions from the data lake's metadata service unit based on the partition filter conditions. The metadata service unit then performs matching based on the assigned filter conditions and the metadata of each partition, returning the matching partitions. There can be one or more matching partitions; a matching partition can be at least one physical partition, at least one virtual partition, or both. Therefore, it is necessary to determine whether the matching partition is a physical partition or a virtual partition.
[0041] Optionally, when creating a virtual partition, a virtual partition identifier can be added to the virtual partition's metadata. Therefore, when determining the type of a matching partition, the type can be determined by reading the matching partition's metadata and checking if a virtual partition identifier exists. If it does, the matching partition is determined to be a virtual partition. Of course, determining the type of a matching partition is not limited to the above method. For example, a partition type identifier can be added to the metadata of each partition—that is, a physical partition identifier can be added to the metadata of physical partitions, and a virtual partition identifier can be added to the metadata of virtual partitions. The type of the matching partition can then be determined based on the partition type identifier in the matching partition's metadata. Alternatively, other possible methods can be used, which are not limited here.
[0042] S203. In response to the matching partition being a virtual partition, the first query statement associated with the virtual partition is executed to obtain the corresponding first query result.
[0043] When the query engine determines that the matching partition is a virtual partition, it can execute a first query statement associated with the virtual partition. Specifically, the first query statement can be executed in at least one of the aforementioned second data tables to obtain the corresponding first query result, which can then be returned to the data consumer. The specific execution process is not limited here.
[0044] Of course, since the matching partition may also be a physical partition, in response to the matching partition being a physical partition, a standard query statement (such as a standard Select query statement) is constructed for the physical partition and executed. That is, the standard query statement is executed in the physical partition. The process is the same as querying data in a traditional data table, which will not be elaborated here.
[0045] Furthermore, a matching partition may include both physical and virtual partitions. Since the matching partition includes both physical and virtual partitions, a standard query statement can be constructed for the physical partition. The virtual partition itself is associated with a pre-defined standard query statement. Therefore, the first query statement associated with the virtual partition can be joined with the standard query statement of the physical partition, for example, using the UNION ALL operator, to obtain a complete query statement, denoted here as the first union query statement, which is then executed. More specifically, a unified, executable query plan can be generated based on the first union query statement, and then that query plan can be executed.
[0046] For example, the first query request for the first data table is: SELECT * FROM db.table WHERE date BETWEEN '2025-01-01' AND '2025-01-03'; assuming date='2025-01-01' belongs to the physical partition, while date='2025-01-02' and date='2025-01-03' belong to the virtual partition, then the first join query statement can be as follows:
[0047] (SELECT * FROM db.table WHERE date = '2025-01-01')
[0048] UNION ALL
[0049] (SELECT ... FROM source_A WHERE ...) -- Definition of '2025-01-02'
[0050] UNION ALL
[0051] (SELECT ... FROM source_B WHERE ...) -- Definition of '2025-01-03'
[0052] Where db.table is the first data table, and source_A and source_B are the data sources for the virtual partitions, which are at least one of the second data tables mentioned above.
[0053] It should be noted that in the data processing method of the data lake described above, when a data consumer needs to query data in the first data table, it can send a first query request to the first data table in the data lake. The physical partitions and virtual partitions in the first data table are completely transparent to the data consumer. The data consumer does not need to care whether the underlying partitions are physical partitions or virtual partitions, and can still query data as if querying a regular table. This seamless experience greatly reduces the impact of the reconstruction or optimization of the first data table on downstream business and saves a lot of communication, coordination and code modification work.
[0054] The data processing method for the aforementioned data lake involves receiving a first query request for a first data table in the data lake. The first data table includes at least one virtual partition, which is pre-created and associated with a first query statement. This first query statement is linked to at least one second data table. The method then determines the partition filtering conditions included in the first query request, queries matching partitions in the first data table based on these conditions, and determines the type of matching partition. If the matching partition is a virtual partition, the method executes the first query statement associated with the virtual partition to obtain the corresponding first query result. By adding virtual partitions to the first data table and associating them with the first query statement, data from other data sources can be retrieved through the first query statement. This allows for small-scale reconstruction or optimization of the first data table. Data consumers only need to obtain the reconstructed or optimized data from the first data table through the virtual partitions, without requiring a complete data migration of the entire first data table. Furthermore, it enables partition-level canary releases, significantly reducing the risk of data migration, ensuring business continuity, and providing complete transparency to downstream data consumers. No modification to the data consumer's query code is required, thus minimizing the impact of reconstruction or optimization on downstream businesses.
[0055] In one alternative example, the matching partition may contain multiple virtual partitions. If the number of virtual partitions in the matching partition is very large, directly using the above-mentioned join (UNION ALL) method may result in an excessively large query plan. To solve this problem, a merge optimization strategy is introduced into the query engine. Specifically, in response to the matching partition being a virtual partition, S203 executes the first query statement related to the virtual partitions, which may include:
[0056] In response to the fact that the matching partition includes multiple virtual partitions and the first query statements associated with the multiple virtual partitions have the same or similar structure, the first query statements associated with the multiple virtual partitions are merged to obtain a second join query statement, and the second join query statement is executed.
[0057] If the first query statements associated with multiple virtual partitions in a matching partition have the same or similar structure—for example, if the first query statement only contains simple operators such as TableScan, Project, and Filter, and the source tables are the same, or if other structures are the same or similar—then the first query statements associated with multiple virtual partitions can be merged into a single query statement, referred to here as the second join query statement. This can be achieved by combining the filtering conditions in the first query statements associated with multiple virtual partitions using expression operators such as OR or IN. The query engine then executes this second join query statement; the execution process is not detailed here.
[0058] In another alternative example, the first query statement associated with one or more virtual partitions may be complex, for example, including a large number of subqueries, making simple merging unsuitable. The first query statement can be split and executed in stages to improve execution stability. Specifically, in response to the matched partition being a virtual partition, S203 executes the first query statement associated with the virtual partition, including:
[0059] In response to the matching partition being a virtual partition, and the number of subqueries in the first query statement associated with the virtual partition exceeding a preset threshold, each subquery is executed in stages, and the results of each subquery are stored in a temporary storage unit; after all subqueries have been executed, the results of each subquery in the temporary storage unit are merged.
[0060] When executing subqueries in stages, they can be grouped, meaning each stage can execute a group of subqueries, improving execution efficiency. After each stage is completed, the results of the subqueries in that stage can be stored in a temporary storage unit. Once all subqueries have been executed, the results in the temporary storage unit are merged to obtain a complete query result, which is then returned to the data consumer. This process breaks down large, complex query tasks into multiple smaller, controllable query tasks, improving execution stability.
[0061] Based on any of the examples above, this article also provides a method for creating virtual partitions, as follows.
[0062] In one optional example, when performing a change operation such as reconstruction or optimization on any partition in the first data table, where the specific change operation includes but is not limited to data migration, logical reconstruction, storage optimization, etc., in response to the change operation instruction on the partition in the first data table, the change operation instruction is executed on the partition, and the execution result is written into the second data table (at least one); then a virtual partition can be created in the first data table, and a first query statement of the second data table is generated, and the first query statement is associated with the virtual partition.
[0063] The first query statement of the second data table is a query statement that queries the execution results of the reconstructed or optimized second data table. The first query statement of the second data table can be generated by a query engine, or it can be generated by other feasible methods, or it can be written manually.
[0064] Optionally, the first query statement can be associated with the virtual partition. This can be done by storing the first query statement directly in the preset partition or in the metadata of the virtual partition; or by storing it in a separate data table (such as a relational data table) and associating it with the virtual partition through a primary key. When the first query statement associated with the virtual partition needs to be executed, the first query statement associated with the virtual partition can be retrieved from the data table through the primary key.
[0065] In addition, after creating a virtual partition in the first data table, you can also create metadata for the virtual partition in the metadata service unit of the data lake, and add a virtual partition identifier to the metadata of the virtual partition.
[0066] Optionally, the metadata service unit can maintain the metadata of each partition, including physical and virtual partitions, through data tables. Furthermore, a virtual partition identifier can be added to the metadata of a virtual partition, specifically in the parameters section of the virtual partition's metadata. Using this virtual partition identifier in the metadata, the type of partition can be quickly identified.
[0067] Optionally, to ensure the accuracy and security of metadata, the addition and deletion operations of this flag are strictly controlled by the internal interface (API) of the metadata service. Users cannot directly modify it through the regular ALTER TABLE command. That is, when a user's modification command for the flag is detected, the modification command will not be executed, and a response message indicating that it cannot be executed will be returned.
[0068] In another alternative example, considering that some restructuring or optimization scenarios derive new data from existing data through computation, and with the application of virtual partitions, there is no need to physically persist the derived new data to disk; it is only necessary to associate the derivation logic (i.e., computation logic) of the derivation process with the virtual partition. Specifically, it can be as follows:
[0069] In response to a derivation instruction on any partition in the first data table, a virtual partition is created in the first data table, and a first query statement is generated according to the derivation logic in the derivation instruction and associated with the virtual partition.
[0070] Here, the first query statement can be obtained by transforming the derived logic into the operation logic in the query statement, or the derived logic can be directly used as the first query statement and then associated with the virtual partition. Then, when the data consumer needs to query the derived data, the query engine can locate the virtual partition, execute the first query statement associated with the virtual partition, generate the required derived data, and return it to the data consumer. Through the above process, the storage overhead of derived data and copies can be eliminated, effectively reducing operating costs.
[0071] In one alternative example, after data reconstruction or optimization, a rollback may be necessary, i.e., returning to the state before data reconstruction or optimization. This can be achieved by deleting the corresponding virtual partitions. Specifically, in response to a rollback command, the virtual partition, its metadata, and the first query statement associated with the virtual partition can be deleted. Of course, the physical partition before data reconstruction or optimization must not have been deleted from the first data table. This way, when data consumers query again subsequently, they can directly retrieve the required data from the physical partition before reconstruction or optimization. This process effectively avoids the need for a one-time full replacement during rollback, reducing latency and uncertainty in the rollback process and improving stability.
[0072] In one optional example, to facilitate user access to and management of virtual partitions, corresponding query statements (e.g., SQL) syntax and a metadata service interface (API) can be designed. The query statement syntax can extend existing DML (Data Manipulation Language) syntax to provide intuitive operations on virtual partitions, including but not limited to syntax for creating / updating virtual partitions, syntax for deleting virtual partitions, and syntax for querying virtual partition information, among others. The metadata service interface can be used to support some atomic operations on virtual partitions, such as performing one or more of the following operations in response to a specific call instruction to the metadata service interface: creating virtual partition metadata, querying single or multiple virtual partition metadata, and deleting virtual partition metadata. The specific syntax and metadata service interface are not limited in this embodiment and can be implemented in any feasible manner.
[0073] The data processing method for the data lake corresponding to the above embodiment, Figure 4 This is a block diagram of the data processing equipment for the data lake presented in this paper. For ease of illustration, only the parts relevant to this paper are shown. (See also...) Figure 4 The data processing equipment 400 of the data lake includes: a receiving unit 401, a partition management unit 402, and an execution unit 403.
[0074] The receiving unit is configured to receive a first query request for a first data table in the data lake; wherein the partition of the first data table includes at least one virtual partition, the virtual partition is pre-created and associated with a first query statement, the first query statement being a query statement associated with at least one second data table;
[0075] The partition management unit is used to determine the partition filtering conditions included in the first query request, query the matching partitions in the first data table according to the partition filtering conditions, and determine the type of matching partitions.
[0076] The execution unit is used to execute the first query statement associated with the virtual partition in response to the matching partition being a virtual partition, and obtain the corresponding first query result.
[0077] In one or more embodiments, the partition management unit, when determining the type of matching partition, is used to:
[0078] Read the metadata of the matching partition;
[0079] The type of the matching partition is determined by whether a virtual partition identifier exists in the metadata. The partition type includes physical partitions or virtual partitions. The virtual partition identifier is added to the metadata of the virtual partition when it is created.
[0080] In one or more embodiments, the execution unit is further configured to:
[0081] In response to the matching partitions including virtual and physical partitions, a standard query statement for the physical partitions is constructed, and a first query statement associated with the virtual partitions is joined with the standard query statement to obtain a first union query statement, which is then executed; or
[0082] In response to the matching partition being a physical partition, a standard query statement is constructed for the physical partition and executed.
[0083] In one or more embodiments, when the execution unit executes a first query statement associated with a virtual partition in response to the matching partition being a virtual partition, it is configured to:
[0084] In response to the fact that the matching partition includes multiple virtual partitions and the first query statements associated with the multiple virtual partitions have the same or similar structure, the first query statements associated with the multiple virtual partitions are merged to obtain a second join query statement, and the second join query statement is executed.
[0085] In one or more embodiments, when the execution unit executes a first query statement associated with a virtual partition in response to the matching partition being a virtual partition, it is configured to:
[0086] In response to the matching partition being a virtual partition, and the number of subqueries in the first query statement associated with the virtual partition exceeding a preset threshold, each subquery is executed in stages, and the results of each subquery are stored in a temporary storage unit.
[0087] After all subqueries have been executed, the results of each subquery in the temporary storage unit are merged.
[0088] In one or more embodiments, the execution unit is further configured to, in response to a change operation instruction for any partition in the first data table, execute a change operation instruction on any partition and write the execution result into the second data table;
[0089] The partition management unit is also used to create virtual partitions in the first data table, generate a first query statement in the second data table, and associate the first query statement with the virtual partition.
[0090] In one or more embodiments, after creating a virtual partition in the first data table, the partition management unit is further configured to:
[0091] In the metadata service unit of the data lake, create the metadata for the virtual partition and add the virtual partition identifier to the metadata of the virtual partition.
[0092] In one or more embodiments, the partition management unit is further configured to:
[0093] In response to a derivation instruction on any partition in the first data table, a virtual partition is created in the first data table, and a first query statement is generated according to the derivation logic in the derivation instruction and associated with the virtual partition.
[0094] In one or more embodiments, the partition management unit is further configured to:
[0095] In response to the rollback command, the virtual partition, its metadata, and the first query statement associated with the virtual partition are deleted.
[0096] The data processing equipment of the aforementioned data lake can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.
[0097] To implement the above embodiments, this document also provides an electronic device.
[0098] refer to Figure 5The diagram illustrates the structure of an electronic device 500 suitable for implementing this document, which can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not impose any limitations on the functionality and scope of this article.
[0099] like Figure 5 As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, the ROM 502, and the RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0100] Typically, the following devices can be connected to the input / output interface 505: input devices 506 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 507 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 508 including, for example, magnetic tape, hard disk, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0101] In particular, according to embodiments herein, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments herein include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 509, or installed from storage device 508, or installed from read-only memory 502. When the computer program is executed by processing device 501, it performs the functions defined in the methods herein.
[0102] It should be noted that the computer-readable storage medium described herein can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this document, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0103] The aforementioned computer-readable storage medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0104] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.
[0105] Computer program code for performing the operations described herein can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments herein. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0107] The units described herein can be implemented in software or hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0108] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0109] In a first aspect, according to one or more embodiments herein, a data processing method for a data lake is provided, comprising:
[0110] According to one or more embodiments herein,
[0111] Receive a first query request for a first data table in the data lake; wherein the partition of the first data table includes at least one virtual partition, the virtual partition is pre-created and associated with a first query statement, the first query statement is a query statement associated with at least one second data table;
[0112] Determine the partition filtering conditions included in the first query request, query the matching partitions in the first data table according to the partition filtering conditions, and determine the type of matching partitions;
[0113] If the matched partition is a virtual partition, the first query statement associated with the virtual partition is executed, and the corresponding first query result is obtained.
[0114] In one or more embodiments, determining the type of matching partition includes:
[0115] Read the metadata of the matching partition;
[0116] The type of the matching partition is determined by whether a virtual partition identifier exists in the metadata. The partition type includes physical partitions or virtual partitions. The virtual partition identifier is added to the metadata of the virtual partition when it is created.
[0117] In one or more embodiments, the method further includes:
[0118] In response to the matching partitions including virtual and physical partitions, a standard query statement for the physical partitions is constructed, and a first query statement associated with the virtual partitions is joined with the standard query statement to obtain a first union query statement, which is then executed; or
[0119] In response to the matching partition being a physical partition, a standard query statement is constructed for the physical partition and executed.
[0120] In one or more embodiments, in response to the matching partition being a virtual partition, a first query statement associated with the virtual partition is executed, including:
[0121] In response to the fact that the matching partition includes multiple virtual partitions and the first query statements associated with the multiple virtual partitions have the same or similar structure, the first query statements associated with the multiple virtual partitions are merged to obtain a second join query statement, and the second join query statement is executed.
[0122] In one or more embodiments, in response to the matching partition being a virtual partition, a first query statement associated with the virtual partition is executed, including:
[0123] In response to the matching partition being a virtual partition, and the number of subqueries in the first query statement associated with the virtual partition exceeding a preset threshold, each subquery is executed in stages, and the results of each subquery are stored in a temporary storage unit.
[0124] After all subqueries have been executed, the results of each subquery in the temporary storage unit are merged.
[0125] In one or more embodiments, the method further includes:
[0126] In response to a change operation instruction on any partition in the first data table, the change operation instruction is executed on any partition, and the execution result is written to the second data table;
[0127] Create a virtual partition in the first data table, generate the first query statement in the second data table, and associate the first query statement with the virtual partition.
[0128] In one or more embodiments, after creating a virtual partition in the first data table, the method further includes:
[0129] In the metadata service unit of the data lake, create the metadata for the virtual partition and add the virtual partition identifier to the metadata of the virtual partition.
[0130] In one or more embodiments, the method further includes:
[0131] In response to a derivation instruction on any partition in the first data table, a virtual partition is created in the first data table, and a first query statement is generated according to the derivation logic in the derivation instruction and associated with the virtual partition.
[0132] In one or more embodiments, the method further includes:
[0133] In response to the rollback command, the virtual partition, its metadata, and the first query statement associated with the virtual partition are deleted.
[0134] Secondly, according to one or more embodiments herein, a data processing apparatus for a data lake is provided, comprising:
[0135] The receiving unit is configured to receive a first query request for a first data table in the data lake; wherein the partition of the first data table includes at least one virtual partition, the virtual partition is pre-created and associated with a first query statement, the first query statement being a query statement associated with at least one second data table;
[0136] The partition management unit is used to determine the partition filtering conditions included in the first query request, query the matching partitions in the first data table according to the partition filtering conditions, and determine the type of matching partitions.
[0137] The execution unit is used to execute the first query statement associated with the virtual partition in response to the matching partition being a virtual partition, and obtain the corresponding first query result.
[0138] In one or more embodiments, the partition management unit, when determining the type of matching partition, is used to:
[0139] Read the metadata of the matching partition;
[0140] The type of the matching partition is determined by whether a virtual partition identifier exists in the metadata. The partition type includes physical partitions or virtual partitions. The virtual partition identifier is added to the metadata of the virtual partition when it is created.
[0141] In one or more embodiments, the execution unit is further configured to:
[0142] In response to the matching partitions including virtual and physical partitions, a standard query statement for the physical partitions is constructed, and a first query statement associated with the virtual partitions is joined with the standard query statement to obtain a first union query statement, which is then executed; or
[0143] In response to the matching partition being a physical partition, a standard query statement is constructed for the physical partition and executed.
[0144] In one or more embodiments, when the execution unit executes a first query statement associated with a virtual partition in response to the matching partition being a virtual partition, it is configured to:
[0145] In response to the fact that the matching partition includes multiple virtual partitions and the first query statements associated with the multiple virtual partitions have the same or similar structure, the first query statements associated with the multiple virtual partitions are merged to obtain a second join query statement, and the second join query statement is executed.
[0146] In one or more embodiments, when the execution unit executes a first query statement associated with a virtual partition in response to the matching partition being a virtual partition, it is configured to:
[0147] In response to the matching partition being a virtual partition, and the number of subqueries in the first query statement associated with the virtual partition exceeding a preset threshold, each subquery is executed in stages, and the results of each subquery are stored in a temporary storage unit.
[0148] After all subqueries have been executed, the results of each subquery in the temporary storage unit are merged.
[0149] In one or more embodiments, the execution unit is further configured to, in response to a change operation instruction for any partition in the first data table, execute a change operation instruction on any partition and write the execution result into the second data table;
[0150] The partition management unit is also used to create virtual partitions in the first data table, generate a first query statement in the second data table, and associate the first query statement with the virtual partition.
[0151] In one or more embodiments, after creating a virtual partition in the first data table, the partition management unit is further configured to:
[0152] In the metadata service unit of the data lake, create the metadata for the virtual partition and add the virtual partition identifier to the metadata of the virtual partition.
[0153] In one or more embodiments, the partition management unit is further configured to:
[0154] In response to a derivation instruction on any partition in the first data table, a virtual partition is created in the first data table, and a first query statement is generated according to the derivation logic in the derivation instruction and associated with the virtual partition.
[0155] In one or more embodiments, the partition management unit is further configured to:
[0156] In response to the rollback command, the virtual partition, its metadata, and the first query statement associated with the virtual partition are deleted.
[0157] Thirdly, according to one or more embodiments herein, an electronic device is provided, comprising: at least one processor and a memory;
[0158] The memory stores the instructions that the computer executes;
[0159] At least one processor executes computer execution instructions stored in memory, causing at least one processor to perform the data processing method of the data lake as described in the first aspect above and various possible designs of the first aspect.
[0160] Fourthly, according to one or more embodiments herein, a computer-readable storage medium is provided, which stores computer-executable instructions that, when executed by a processor, implement the data processing method of the data lake as described in the first aspect and various possible designs of the first aspect.
[0161] Fifthly, according to one or more embodiments herein, a computer program product is provided, including a computer program that, when executed by a processor, implements the data processing method for a data lake as described in the first aspect above and various possible designs of the first aspect.
[0162] The above description is merely a preferred embodiment and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure herein is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed herein that have similar functions.
[0163] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this document. Certain features described in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0164] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A data processing method for a data lake, comprising: Receive the first query request for the first data table in the data lake; The partitions of the first data table include at least one virtual partition, which is pre-created and associated with a first query statement, which is a query statement associated with at least one second data table. Determine the partition filtering conditions included in the first query request, query the matching partitions in the first data table according to the partition filtering conditions, and determine the type of the matching partitions; If the matched partition is a virtual partition, the first query statement associated with the virtual partition is executed to obtain the corresponding first query result.
2. The method according to claim 1, wherein determining the type of the matching partition includes: Read the metadata of the matched partition; The type of the matching partition is determined based on whether a virtual partition identifier exists in the metadata; The partition types include physical partitions or virtual partitions, and the virtual partition identifier is added to the metadata of the virtual partition when it is created.
3. The method according to claim 1, further comprising: In response to the matching partition including a virtual partition and a physical partition, a standard query statement for the physical partition is constructed, and a first query statement associated with the virtual partition is concatenated with the standard query statement to obtain a first combined query statement, and the first combined query statement is executed. or In response to the matching partition being a physical partition, a standard query statement is constructed for the physical partition and executed.
4. The method according to claim 1, wherein in response to the matching partition being a virtual partition, executing the first query statement associated with the virtual partition includes: In response to the fact that the matching partition includes multiple virtual partitions and the first query statements associated with the multiple virtual partitions have the same or similar structure, the first query statements associated with the multiple virtual partitions are merged to obtain a second combined query statement, and the second combined query statement is executed.
5. The method according to claim 1, wherein in response to the matching partition being a virtual partition, executing the first query statement associated with the virtual partition includes: In response to the fact that the matching partition is a virtual partition and the number of subqueries in the first query statement associated with the virtual partition exceeds a preset threshold, each subquery is executed in stages and the results of each subquery are stored in a temporary storage unit. After all subqueries have been executed, the results of each subquery in the temporary storage unit are merged.
6. The method according to any one of claims 1-5, further comprising: In response to a change operation instruction for any partition in the first data table, the change operation instruction is executed on the partition, and the execution result is written into the second data table; Create a virtual partition in the first data table, generate a first query statement for the second data table, and associate the first query statement with the virtual partition.
7. The method according to claim 6, further comprising, after creating a virtual partition in the first data table: In the metadata service unit of the data lake, the metadata of the virtual partition is created, and a virtual partition identifier is added to the metadata of the virtual partition.
8. The method according to any one of claims 1-5, further comprising: In response to a derivation instruction for any partition in the first data table, a virtual partition is created in the first data table, and a first query statement is generated according to the derivation logic in the derivation instruction and associated with the virtual partition.
9. The method according to any one of claims 1-5, further comprising: In response to the rollback command, the virtual partition, the metadata of the virtual partition, and the first query statement associated with the virtual partition are deleted.
10. A data processing device for a data lake, characterized in that, include: The receiving unit is used to receive the first query request for the first data table in the data lake. The partitions of the first data table include at least one virtual partition, which is pre-created and associated with a first query statement, which is a query statement associated with at least one second data table. A partition management unit is used to determine the partition filtering conditions included in the first query request, query the matching partitions in the first data table according to the partition filtering conditions, and determine the type of the matching partitions. An execution unit is configured to, in response to the matching partition being a virtual partition, execute a first query statement associated with the virtual partition to obtain the corresponding first query result.
11. An electronic device, characterized in that, include: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-9.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-9.