Metadata query method, data access method, and related device
Patent Information
- Application Number
- PCT/CN2026/079890
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2026-02-25
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026079890_01102026_PF_FP_ABST
Abstract
Description
Metadata query methods, data access methods and related equipment
[0001] This disclosure claims priority to Chinese Patent Application No. 202510369868.7, filed with the China Patent Office on March 26, 2025, entitled “Metadata Query Method, Data Access Method and Related Equipment”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the field of data processing technology, and in particular to a metadata query method, a data access method, and related equipment. Background Technology
[0003] In a data lake system, metadata management is crucial for ensuring efficient querying. To reduce the overhead of recording metadata changes, metadata is typically stored as records of those changes. For example, metadata records operations such as adding or deleting a data file, or changes to an appendage of a data file. These changes are also periodically merged, leading to unordered metadata changes. Thus, querying metadata in a data lake system may require traversing all metadata to determine if a particular data file is matched, resulting in significant latency, especially as the amount of metadata increases. Summary of the Invention
[0004] This disclosure provides a metadata query method, a data access method, and related equipment to accelerate metadata query.
[0005] This disclosure provides a metadata query method, comprising: periodically acquiring a snapshot list of business tables in a data lake system, and determining snapshot change information of the snapshot list acquired in the current period relative to the snapshot list acquired in the previous period, wherein the snapshot list includes snapshots with multiple version numbers corresponding to the business tables; determining a first existence and a second existence, wherein the first existence indicates whether a first metadata image corresponding to the business table exists; the second existence indicates whether a snapshot with a first version number exists in the snapshot list acquired in the current period, the first version number being the version number corresponding to the snapshot used to construct the first metadata image; constructing a second metadata image corresponding to the business table based on the snapshot change information, the first existence, and the second existence; and caching the second metadata image corresponding to the business table in a database system so that the database system can perform metadata queries on the second metadata image.
[0006] This disclosure also provides a data access method, comprising: receiving an access request sent by a client, the access request including a snapshot ID to be accessed and query conditions; using a streaming query interface to query metadata in a metadata mirror cached locally in the database system based on the snapshot ID to be accessed and the query conditions; responding to each batch of data files retrieved by the streaming query interface, returning the retrieved file metadata of the current batch of data files to the computing layer in the database system; and using the computing layer to perform data processing operations corresponding to the access request on the current batch of data files in the data lake system based on the file metadata of the current batch of data files, obtaining data processing results and returning them to the client.
[0007] This disclosure also provides an electronic device, including: a memory and a processor; the memory for storing a computer program; and the processor coupled to the memory for executing the computer program to perform steps in a metadata query method or a data query method.
[0008] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, enables the processor to implement steps in a metadata query method or a data access method.
[0009] This disclosure also provides a computer program product, including a computer program or instructions that, when executed by a processor, enable the processor to implement the steps in a metadata query method or a data access method.
[0010] In this embodiment, snapshot change information of business tables in the data lake system is periodically acquired. Considering multiple factors, including snapshot change information, the existence of a metadata image corresponding to the business table, and whether the snapshot list acquired in the current period contains a snapshot used to build the metadata image, a decision is made on whether to update the metadata image incrementally or rebuild it entirely. Once a new metadata image corresponding to the business table is obtained, it is cached in the database system. This allows the database system to perform metadata queries based on the locally cached new metadata image, eliminating the need to query the metadata in the data lake system, significantly reducing metadata query latency and improving metadata query efficiency. Attached Figure Description
[0011] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this disclosure, illustrate exemplary embodiments of the present disclosure and are used to explain the disclosure, but do not constitute an undue limitation of the disclosure. In the drawings:
[0012] Figure 1 is a flowchart of a metadata query method provided in an embodiment of this disclosure;
[0013] Figure 2 is a flowchart of a data access method provided in an embodiment of this disclosure;
[0014] Figure 3 shows an exemplary application scenario.
[0015] Figure 4 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0017] In the embodiments of this disclosure, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the access relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone, where A and B can be singular or plural. In the textual description of this disclosure, the character " / " generally indicates that the preceding and following associated objects have an "or" relationship. Furthermore, in the embodiments of this disclosure, "first," "second," "third," etc., are only used to distinguish the content of different objects and have no other special meaning.
[0018] It should be noted that, in the cases involving user information in the embodiments of this disclosure, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this disclosure (including but not limited to language models or large models) comply with relevant laws and standards.
[0019] The technical solutions of this disclosure and how they solve the aforementioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The technical solutions provided by each embodiment of this disclosure are described in detail below with reference to the accompanying drawings.
[0020] The following explains some terms used in the embodiments of this disclosure:
[0021] A data lake system is a platform for centrally storing large amounts of raw data, which can be structured, semi-structured, or unstructured. Data lakes allow users to store massive amounts of data at low cost and support various types of processing and analysis tasks.
[0022] Business tables refer to tables in a data lake system used to store data related to specific business operations. These tables typically have a well-defined schema and contain data relevant to a particular business domain. For example, on an e-commerce platform, business tables might include order tables, product tables, customer tables, etc.
[0023] System tables are used to manage and maintain the state and metadata of the data lake, helping the data lake system to effectively organize and manage business tables and other objects. System tables record all metadata information about business tables, including but not limited to the business table's structure definition, partition information, snapshot version, etc.
[0024] Figure 1 is a flowchart of a metadata query method provided in an embodiment of this disclosure. Referring to Figure 1, the method may include the following steps:
[0025] 101. Periodically obtain a snapshot list of business tables in the data lake system, and determine the snapshot change information of the snapshot list obtained in the current period relative to the snapshot list obtained in the previous period. The snapshot list includes snapshots with multiple version numbers corresponding to the business tables.
[0026] Specifically, in a data lake system, when changes such as inserting, deleting, or updating records are performed on a business table, the data lake system generates a snapshot to reflect these changes. In simpler terms, a snapshot of a business table is a complete or partial record of the table's data state at a specific point in time; it provides a complete view of the table's data and its metadata at that particular point in time, recording all data and related metadata information. Each snapshot has a unique identifier (such as a version number), allowing users to query and restore to the data state at any historical point in time. Over time, the data lake system generates snapshots of different version numbers for the business tables, forming a snapshot list of the business tables.
[0027] The data lake system itself provides metadata related to business tables. To accelerate metadata queries within the data lake system, a metadata image corresponding to the business table metadata also needs to be generated. A metadata image can be understood as a reorganization of the business table metadata provided by the data lake system itself; it can be understood as a complete record or snapshot of the business table metadata at a specific point in time. It contains all information describing the structure and state of the business tables in the data lake, such as table schema, partition information, and file locations.
[0028] To generate a metadata mirror of the latest state of the business tables, it is necessary to periodically retrieve a snapshot list of the business tables in the data lake system and determine the snapshot change information of the snapshot list retrieved in the current period compared to the snapshot list retrieved in the previous period. For example, a scheduled task can be set up to perform snapshot list retrieval and comparison operations at fixed intervals (e.g., every half hour, hour, or day). By comparing the snapshot list retrieved in the current period with the snapshot list retrieved in the previous period, newly added, deleted, or modified snapshots can be identified. New snapshots: Snapshots that appear in the snapshot list retrieved in the current period but are not in the snapshot list retrieved in the previous period. Deleted snapshots: Snapshots that appear in the snapshot list retrieved in the previous period but are not found in the snapshot list retrieved in the current period. Modified snapshots: If snapshot content change detection is supported, it can be further identified whether the content of a snapshot has changed. In this embodiment, the snapshot change information can also reflect whether the maximum version number of the snapshots in the snapshot list retrieved in the current period is higher than or equal to the maximum version number of the snapshots in the snapshot list retrieved in the previous period.
[0029] 102. Determine the first existence and the second existence, wherein the first existence indicates whether there is a first metadata image corresponding to the business table. The first metadata image is obtained by reorganizing the metadata of the business table stored in the data lake system. The second existence indicates whether there is a snapshot with the first version number in the snapshot list obtained in the current period. The first version number is the version number corresponding to the snapshot used to build the first metadata image.
[0030] Specifically, in order to generate a metadata mirror of the latest state of a business table, it is also necessary to determine whether a first metadata mirror corresponding to the business table exists. Here, the first existence is defined to characterize whether a first metadata mirror corresponding to the business table exists. The first metadata mirror can be understood as a currently existing metadata mirror (also called an existing metadata mirror) or a metadata mirror before the update.
[0031] In addition, assuming a first metadata image corresponding to the business table exists, it is also necessary to determine whether a snapshot with the first version number exists in the snapshot list obtained in the current period. The first version number is the version number corresponding to the snapshot used to build the first metadata image. Here, a second existence is defined to characterize whether a snapshot with the first version number exists in the snapshot list obtained in the current period.
[0032] 103. Based on the snapshot change information, the first existence and the second existence, update the first metadata image to obtain the second metadata image.
[0033] 104. Cache the second metadata mirror in the database system so that the database system can query metadata from the second metadata mirror.
[0034] Specifically, the second metadata image can be understood as the updated metadata image obtained by updating the first metadata image (also known as the new metadata image).
[0035] In practical applications, considering snapshot change information, first existence, and second existence, a decision is made as to whether to update the metadata image incrementally or rebuild it entirely. Once the new metadata image corresponding to the business table is obtained, it is cached in the database system. This allows the database system to perform metadata queries based on the locally cached image, eliminating the need to query the metadata in the data lake system. This significantly reduces metadata query latency and improves query efficiency.
[0036] In scenarios involving incremental updates to the metadata image, the goal is to make minimal updates to the existing metadata image based on incremental snapshot information, rather than completely rebuilding the entire metadata image. This reduces processing time, saves resources, and ensures that the metadata image always reflects the latest data state.
[0037] In the scenario of creating a full metadata mirror, the goal is to build a complete metadata mirror from scratch, which will comprehensively reflect all metadata information of the business tables. This typically occurs when creating a metadata mirror for the first time, or when an existing metadata mirror is deemed unreliable and needs to be regenerated.
[0038] In some optional embodiments, the implementation of updating the first metadata image to obtain the second metadata image based on snapshot change information, first existence, and second existence is as follows: Determine that the snapshot change information reflects that the maximum version number of the snapshots in the snapshot list obtained in the current period is higher than or equal to the maximum version number of the snapshots in the snapshot list obtained in the previous period; if the first existence indicates that there is no first metadata image corresponding to the business table, or if the second existence indicates that there is no snapshot with the first version number in the snapshot list obtained in the current period, then the snapshot with the maximum version number in the snapshot list obtained in the current period is taken as the snapshot with the latest version number; and construct the second metadata image corresponding to the business table in a full-scale construction manner based on the snapshot with the latest version number.
[0039] In some alternative embodiments, the implementation of updating the first metadata image to obtain the second metadata image based on snapshot change information, first existence, and second existence is as follows: determine that the snapshot change information reflects that the maximum version number of the snapshots in the snapshot list obtained in the current period is higher than the maximum version number of the snapshots in the snapshot list obtained in the previous period; if the first existence indicates that there is a first metadata image corresponding to the business table, and the second existence indicates that there is a first version number in the snapshot list obtained in the current period, then based on the incremental snapshot information between the snapshot with the maximum version number in the snapshot list obtained in the current period and the snapshot with the first version number, update the first metadata image corresponding to the business table in an incremental construction manner to obtain the second metadata image.
[0040] For example, sorted by version number from low to high, the snapshot list obtained in the current period includes snapshots of version V2, version V3, and version V4, while the snapshot list obtained in the previous period includes snapshots of version V1, version V2, and version V3.
[0041] By comparing the snapshot lists, it was found that the largest version number of the snapshots in the current period's snapshot list (V4 version) is greater than the largest version number of the snapshots in the previous period's snapshot list (V3 version). In this case, it is necessary to determine whether there is a first metadata image corresponding to the business table. The first metadata image corresponding to the business table can be understood as the metadata image built before the current time. If there is no first metadata image corresponding to the business table, it means that a second metadata image corresponding to the business table needs to be built in a full build mode based on the V4 version snapshot.
[0042] If a first metadata image exists for a business table, and this first metadata image is built based on a snapshot of version V1, but the snapshot of version V1 has been deleted from the snapshot list obtained in the current period (the snapshot of version V1 can be understood as the snapshot with the first version number used to build the first metadata image), then in this case, a second metadata image for the business table also needs to be built using a full build method based on the snapshot of version V4.
[0043] If a first metadata image exists for a business table, and this first metadata image is built based on a snapshot of version V1, then if a snapshot of version V1 still exists in the snapshot list obtained in the current period, then the first metadata image for the business table needs to be updated incrementally based on the incremental snapshot information between the snapshots of version V4 and version V3 to obtain the second metadata image.
[0044] In practical applications, metadata mirroring can be constructed as multiple system tables. For example, metadata mirroring may include a first system table (denoted as the snapshot table), a second system table (denoted as the partition_version table), a third system table (denoted as the file_version table), a fourth system table (denoted as the file_meta table), and a fifth system table (denoted as the sub_files table).
[0045] The first system table may include a business table ID field to record the uniqueness of the business table, a snapshot ID field to record the uniqueness of the snapshot, and a snapshot metadata field to record the snapshot metadata. The business table ID field and the snapshot ID field are used as primary key fields. The first system table stores the metadata of each snapshot, which includes, but is not limited to, snapshot creation time and version number. Optionally, the snapshot ID can represent the snapshot version number.
[0046] Assume the business table ID field is named table_id, the snapshot ID field is named snapshot_id, the snapshot metadata field is named snapshot_meta, the snapshot version number is named vesion, and the timestamp corresponding to the snapshot creation time is named timestamp. The first system table is shown in Table 1. One record stored in the first system table is "business table ID is 1, snapshot ID is 1, the snapshot metadata records the snapshot version number (vesion) as 1, and the snapshot creation time is 160000000".
[0047] Table 1
[0048] The second system table includes a business table ID field, a snapshot ID field, and a partition version number field. The business table ID field and the snapshot ID field are the primary key fields, and the value of the partition version number field includes the partition name of each partition and its corresponding partition version number.
[0049] Assuming the field name for the partition version number field is denoted as map_partition_version, and the second system table is shown in Table 2, one record stored in the second system table is "Business table ID is 1; snapshot ID is 2; the partition version number corresponding to partition name p1 is 1, and the partition version number corresponding to partition name p2 is 3".
[0050] Table 2
[0051] The third system table includes a business table ID field, a snapshot ID field, a partition name field, and a file version number field. The business table ID field, the snapshot ID field, and the partition name field are the primary key fields, and the file version number field includes the file name and file version number of each data file.
[0052] Assuming the field name of the partition name field is denoted as partition_name, the third system table is shown in Table 3. One record stored in the third system table is "the partition with business table ID of 1, snapshot ID of 2, and partition name of p1; the file version number corresponding to data file f1 in partition p1 is 1; the file version number corresponding to data file f2 in partition p1 is 2".
[0053] Table 3
[0054] The fourth system table includes a business table ID field, a file name field, and a file metadata field, with the business table ID field and the file name field as primary key fields.
[0055] Assume the file name field is named `file_id` and the file metadata field is named `file_meta`. The fourth system table is shown in Table 4. One record in the fourth system table is: "Data file with business table ID 1, snapshot ID 2, and file name f1; file metadata of data file f1". The file metadata of data file f1 can be recorded in key-value (KV) pairs, for example, in Table 4, the file metadata is {"key1":"value1"}.
[0056] Table 4
[0057] The fifth system table includes a business table ID field, a snapshot ID field, a file name field, and an auxiliary file information field, with the business table ID field, snapshot ID field, and file name field as primary key fields.
[0058] Assuming the field name for the supplementary file information field is denoted as `sub_files`, and the fifth system table is shown in Table 5, one record stored in the fifth system table is "Data file with business table ID 1, snapshot ID 2, and filename f1; supplementary file information of data file f1". The supplementary file information of data file f1 can be recorded in the form of key-value pairs (KV). For example, in Table 5, the supplementary file information of data file f1 is {"subfile1": "subfilemeta"}, which represents the file metadata `subfilemeta` of the supplementary file `subfile1` under data file f1.
[0059] Table 5
[0060] In some optional embodiments, the implementation of building the second metadata image corresponding to the business table in a full-scale build based on the snapshot of the latest version number is as follows: The business table ID, first snapshot ID, and snapshot metadata corresponding to the snapshot of the latest version number are inserted as a record into a new first system table; the partition version number of each partition associated with the snapshot of the latest version number is determined as the first snapshot ID, and the business table ID, first snapshot ID, partition name of each partition, and its corresponding partition version number corresponding to the snapshot of the latest version number are inserted as a record into a new second system table; the file version number of each data file in each partition is determined as the first snapshot ID, and the business table ID, first snapshot ID, partition name of each partition, and all data files under each partition are added as a record. The file name and file version number of the data file are inserted as a record into a new third system table; the file name and file metadata of each data file associated with the snapshot of the latest version number are obtained; the business table ID corresponding to the snapshot of the latest version number, the file name and file metadata of each data file are inserted as a record into a new fourth system table; the information of the auxiliary files corresponding to each data file associated with the snapshot of the latest version number is obtained, and the business table ID corresponding to the snapshot of the latest version number, the first snapshot ID, the file name of each data file, and the information of the auxiliary files corresponding to each data file are inserted as a record into a new fifth system table; based on the new first system table, the new second system table, the new third system table, the new fourth system table, and the new fifth system table, the second metadata mirror corresponding to the business table is constructed.
[0061] Specifically, when building a second metadata image of a business table in a full manner, the snapshot based on the latest version number can determine the various partitions of the business table, the various data files under each partition, and the various subsidiary files under each data file. Partitioning refers to dividing data into different subsets according to the values of one or more fields. Partitioning can speed up querying and facilitate the management of large-scale datasets.
[0062] A data file is a file that stores the actual data. Each data file may correspond to a part of an entire business table, such as data from a specific partition or simply a fragment of the table. Supporting files may include index files, statistics files, etc. Data file metadata includes, but is not limited to: file name, file path, file size, file format, table schema information, partition information, number of rows, column-level statistics (e.g., maximum, minimum, and number of null values per column), supporting file information, etc. Supporting file information includes, but is not limited to: supporting file name, file path, file size, file format, number of rows, column-level statistics, etc.
[0063] Specifically, the first system table mainly stores snapshot metadata corresponding to the latest version number snapshot, such as the snapshot version number; the second system table mainly stores the partition version number of each partition determined based on the latest version number snapshot; the third system table mainly stores the file version number of each data file under each partition determined based on the latest version number snapshot; the fourth system table mainly stores the file metadata of each data file under each partition determined based on the latest version number snapshot; and the fifth system table mainly stores the auxiliary file information of each data file under each partition determined based on the latest version number snapshot. In this way, each system table stores different types or aspects of metadata information, and the metadata mirror composed of the first, second, third, fourth, and fifth system tables can completely reproduce the metadata of the business table.
[0064] In some optional embodiments, the first metadata mirror includes: an existing first system table, an existing second system table, an existing third system table, an existing fourth system table, and an existing fifth system table; correspondingly, the implementation of updating the first metadata mirror corresponding to the business table in an incremental construction manner based on the incremental snapshot information between the snapshot with the largest version number in the snapshot list obtained in the current period and the snapshot with the first version number, to obtain the second metadata mirror, is as follows: based on the incremental snapshot information between the snapshot with the largest version number in the snapshot list obtained in the current period and the snapshot with the first version number, including: determining the target incremental snapshot information of the snapshot with the next version number relative to the previous version number, wherein the next version number is greater than the first version number and less than or equal to the largest version number corresponding to the snapshot in the snapshot list obtained in the current period, and the previous version number is greater than or equal to the first version number and less than the largest version number corresponding to the snapshot in the snapshot list obtained in the current period; and setting the next version number as the target incremental snapshot information between the snapshot with the next version number and the snapshot with the first version number. The business table ID, snapshot ID, and snapshot metadata corresponding to the snapshot of the version number are inserted as a record into the existing first system table to obtain a new first system table. The changed partition is determined based on the target incremental snapshot information. The business table ID, snapshot ID, partition name, and partition version number corresponding to the snapshot of the next version number are inserted as a record into the existing second system table to obtain a new second system table. The partition version number of the changed partition is the snapshot ID of the snapshot of the next version number. The existing third, fourth, and fifth system tables are updated according to the change type of the changed files in the changed partition to obtain new third, fourth, and fifth system tables, respectively. Based on the new first, second, third, fourth, and fifth system tables, a second metadata mirror corresponding to the business table is constructed.
[0065] In this embodiment, during incremental construction, the target incremental snapshot information of the next version number relative to the previous version number is determined. The next version number is greater than the first version number and less than or equal to the maximum version number corresponding to the snapshot in the snapshot list obtained in the current period. The previous version number is greater than or equal to the first version number and less than the maximum version number corresponding to the snapshot in the snapshot list obtained in the current period.
[0066] Assume the maximum version number corresponding to the snapshot with the highest version number in the snapshot list acquired in the current period is denoted as Vx, and the first version number is denoted as Vy. When determining the incremental snapshot information between the snapshot with the highest version number and the snapshot with the first version number in the snapshot list acquired in the current period, it is necessary to consider all incremental snapshot information from Vy to Vx. Incremental snapshot information can be generated based on the snapshot with the previous version number, recording the changes since the snapshot with the previous version number; that is, incremental snapshot information reflects all changes (additions, modifications, deletions, etc.) of the next version number snapshot relative to the previous version number snapshot.
[0067] For all version numbers from Vy to Vx, determine the incremental snapshot information between snapshots of two adjacent version numbers. For example, if Vx is V8, the first version number is V5, and the version numbers are sorted in ascending order as V5, V6, V7, and V8, then it is necessary to determine the incremental snapshot information between the snapshot of version V6 and the snapshot of version V5, the incremental snapshot information between the snapshot of version V7 and the snapshot of version V6, and the incremental snapshot information between the snapshot of version V8 and the snapshot of version V7.
[0068] Specifically, when updating an existing first system table, the business table ID, snapshot ID, and snapshot metadata corresponding to the snapshot of the next version number are inserted as a record into the existing first system table to obtain a new first system table. Using the example above, the snapshot table needs to add 3 records, which respectively record the business table ID, snapshot ID, and snapshot metadata corresponding to V6, V7, and V8.
[0069] In this embodiment, when updating the existing second system table, the changed partition can be determined based on the target incremental snapshot information of the next version number snapshot relative to the previous version number. The changed partition can be understood as the partition in the business table corresponding to the next version number snapshot that has been changed (added, modified, or deleted) relative to the previous version number snapshot. Then, the business table ID, snapshot ID, partition name of the changed partition, and partition version number corresponding to the next version number snapshot are inserted as a new record into the existing first system table.
[0070] For example, the snapshot ID of the previous version is 6, and the snapshot ID of the next version is 7. Partition p1 is not a changed partition, while partition p2 is a changed partition. There is an existing record in the second system table with "table_id = 1, snapshot_id = 6, map_partition_version = {"p1":1, "p2":2}". Compared to the existing second system table, the new second system table adds a new record with "table_id = 1, snapshot_id = 7, map_partition_version = {"p1":1, "p2":7}".
[0071] In this embodiment, when updating an existing third system table, the changed files in the changed partition are determined based on the target incremental snapshot information of the next version number relative to the previous version number. If the changed file is a deleted file, the file name and file version number of the changed file are deleted from the existing third system table to obtain a new third system table. If the changed file is a newly added file, the snapshot ID of the next version number is used as the file version number of the changed file, and the file name and file version number of the changed file are added to the existing third system table to obtain a new third system table.
[0072] Specifically, based on the target incremental snapshot information of the next version number relative to the previous version number, all changed partitions and their changed files can be identified. For example, by comparing the file lists of changed partitions between two snapshot versions, newly added or deleted files corresponding to the latest version number snapshot can be found. For newly added files, the corresponding file version number needs to be added to the existing third-party system table; for deleted files, the corresponding file version number needs to be deleted from the existing third-party system table.
[0073] For example, the snapshot ID of the previous version is 6, and the snapshot ID of the next version is 7. Partition p1 is a changed partition, and partition p2 is a changed partition. There are already two records in the third system table. One record is "table_id is 1, snapshot_id is 6, partition_name is p1, map_file_version is {"f1":6, "f2":6}". The other record is "table_id is 1, snapshot_id is 6, partition_name is p2, map_file_version is {"f3":6}".
[0074] Compared to the existing third system table, the new third system table adds two records. One of the new records is "table_id is 1, snapshot_id is 7, partition_name is p1, map_file_version is {"p1":{"f1":6}". Because the file f2 under the p1 partition was deleted, the map_file_version does not contain information about the file f2.
[0075] Another newly added record is "table_id is 1, snapshot_id is 7, partition_name is p2, map_file_version is {"p2":{"f3":6,"f4":7}", because a file f4 has been added under the p2 partition, so the map_file_version record contains information about file f4.
[0076] In this embodiment, when updating an existing fourth system table, if the changed file is a newly added file, the business table ID corresponding to the snapshot of the next version number, the file name of the newly added file, and the file metadata are inserted as a record into the existing fourth system table to obtain a new fourth system table.
[0077] When updating an existing fifth system table, if the changed file is a new file, insert the business table ID corresponding to the snapshot of the next version number, the snapshot ID, the file name of the new file, and the information of the corresponding subsidiary file of the new file as a record into the existing fifth system table.
[0078] In some optional embodiments, in order to ensure the validity of the metadata mirror, after constructing the second metadata mirror corresponding to the business table, in response to the deletion of the second version number snapshot in the data lake system, the metadata related to the second version number snapshot in the second metadata mirror can be deleted.
[0079] For example, when a snapshot with a certain version number expires and is deleted in the data lake system, its associated metadata information also needs to be synchronously deleted in the metadata mirror. First, the `map_partition_version` of the expired snapshot is compared with that of the next snapshot in the `partition_version` table to check which partitions were deleted. For deleted partitions, their data file list is retrieved, and these data file records are deleted from the `file_meta` and `sub_files` tables. Then, the record for the deleted partition in the `file_version` table is deleted. For modified partitions, the `file_version` table is read to check which data files were deleted or modified. For deleted data files, they are deleted from the `file_meta` and `sub_files` tables. For modified data files, the version number of the data file in the expired snapshot is retrieved, and the corresponding record in the `sub_files` table is deleted. Finally, the record for the expired snapshot in the `partition_version` table is deleted.
[0080] In practical applications, a metadata mirror can be constructed as multiple system tables. For example, a metadata mirror may include a base system table (denoted as the base table) and a difference system table (denoted as the delta table).
[0081] The basic system table can include a business table ID field (table_id), a snapshot ID field (snapshot_id), a partition name field (partition_name), a file name field (file_id), and a file metadata field (file_meta). As shown in Table 6, a record stored in the basic system table is "file metadata with business table ID of 1, snapshot ID of 1, partition name of p1, and file name of f2".
[0082] Table 6
[0083] The difference system table can include a business table ID field (table_id), a snapshot ID field (snapshot_id), a partition name field (partition_name), a file name field (file_id), a change type field (change_kind), and a file metadata field (changed_file_meta). The file metadata field in the difference system table records the file metadata of the changed files. The change type field records the file change type (change_kind), which includes: added (ADD) file, modified (MODIFY) file, or deleted (DELETE) file.
[0084] Table 7
[0085] In one optional embodiment, the implementation of constructing the second metadata image corresponding to the business table in a full-scale construction mode based on the snapshot of the latest version number is as follows: insert the business table ID corresponding to the snapshot of the latest version number, the first snapshot ID, the partition name of each partition, the file name of each data file under each partition and its corresponding file metadata as a record into the new basic system table; construct the second metadata image corresponding to the business table based on the new basic system table.
[0086] As time progresses, the version number of the snapshot used by the base system tables changes. For example, at different points in time, the base system tables may use snapshot version number V1 at some points, V2 at others, and so on. The base system tables record all file information of the business tables under a specific snapshot version. Of course, during a full build, the second metadata image may include not only the new base system tables but also the difference system tables. When building a new base system table, it is not necessary to update the difference system tables.
[0087] In one optional embodiment, updating the first metadata image corresponding to the business table in an incremental construction manner to obtain the second metadata image, based on the incremental snapshot information between the snapshot with the largest version number in the snapshot list obtained in the current period and the snapshot with the first version number, includes: determining the target incremental snapshot information of the next version number relative to the previous version number, wherein the next version number is greater than the first version number and less than or equal to the largest version number corresponding to the snapshot in the snapshot list obtained in the current period, and the previous version number is greater than or equal to the first version number and less than the largest version number corresponding to the snapshot in the snapshot list obtained in the current period; determining the changed partition and the change type of the changed files in the changed partition based on the target incremental snapshot information; inserting the business table ID, snapshot ID, partition name of the changed partition, file name of the changed file, change type, and file metadata corresponding to the snapshot with the next version number as a record into the existing difference system table to obtain a new difference system table; and constructing the second metadata image corresponding to the business table based on the new difference system table.
[0088] In this embodiment, the snapshot of the first version number can be understood as the snapshot used when building the basic system table in full. For example, the basic system table records all file information under the snapshot of version V3; when performing an incremental build, if the snapshot of the latest version number (i.e., the snapshot of the highest version number in the snapshot list obtained in the current period) is the snapshot of version V6, then it is necessary to determine the incremental snapshot information between the snapshot of version V6 and the snapshot of version V5, the incremental snapshot information between the snapshot of version V5 and the snapshot of version V4, and the incremental snapshot information between the snapshot of version V4 and the snapshot of version V3. The difference system table is updated based on the incremental snapshot information between snapshots of adjacent version numbers.
[0089] For example, based on the incremental snapshot information between the V4 version snapshot (assuming snapshot_id is 4) and the V3 version snapshot (assuming snapshot_id is 3), we know that partition p1 of the business table table_id is 1 and a new data file f3 has been added to this partition. Therefore, a new record will be inserted into the new difference system table as "table_id is 1, snapshot_id is 4, partition_name is p1, file_id is f3, change_kind is ADD, changed_file_meta is {f3's file metadata}".
[0090] For example, based on the incremental snapshot information between the V5 version snapshot (assuming snapshot_id is 5) and the V4 version snapshot, we know that partition p1 of the business table table_id is 1 and is a changed partition. The data file f3 has been modified in this partition. Then, a record is inserted into the new difference system table as "table_id is 1, snapshot_id is 5, partition_name is p1, file_id is f3, change_kind is MODIFY, changed_file_meta is {f3's file metadata}".
[0091] For example, based on the incremental snapshot information between the V6 version snapshot (assuming snapshot_id is 6) and the V5 version snapshot, we know that partition p2 of the business table table_id is 1 and is a changed partition. If data file f4 is deleted from this partition, then a record will be inserted into the new difference system table as "table_id is 1, snapshot_id is 6, partition_name is p2, file_id is f4, change_kind is DELETE, changed_file_meta is {f4's file metadata}".
[0092] Of course, during incremental builds, the second metadata image includes not only the new difference system tables but also the base system tables. When incrementally building new difference system tables, it is not necessary to update the base system tables.
[0093] In practical applications, to prevent the difference system table from becoming too large, the base system table can be updated periodically, and records corresponding to snapshots of older version numbers in the difference system table can be deleted. For example, the base system table used snapshot version number V5 before the update, and the updated base system table uses snapshot version number V8. The records in the difference system table correspond to incremental snapshot information. For example, the difference system table has records corresponding to snapshots of version numbers V6, V7, and V8. The record corresponding to the snapshot of version number V7 is related to the changed files determined by the incremental snapshot information between the V7 snapshot and the V6 snapshot. For example, the V7 snapshot is equivalent to the V6 snapshot having added, modified, or deleted data files. The record corresponding to the V7 snapshot in the difference system table is related to adding, modifying, or deleting data files. When updating the base system table, records corresponding to the V6, V7, and V8 version snapshots can be retrieved sequentially from the difference system table. These records are then applied to the base system table containing the file metadata of all data files under the V5 version snapshot, resulting in the file metadata of all data files under the V8 version snapshot. This file metadata is then inserted into the base system table to update it. The updated base system table includes the full data corresponding to both the V5 and V8 version snapshots. The full data corresponding to the V5 version snapshot is the file metadata of all data files under the V5 version snapshot, and vice versa. After advancing the base system table, you can delete the records corresponding to the snapshots of version V6, V7, and V8 in the difference system table to reduce the data volume of the difference system table. Of course, the full data corresponding to the snapshot of version V5 in the base system table can also be deleted.
[0094] It is worth noting that if the version updates of the basic system tables are promoted regularly, when an update of the existing metadata image is triggered, the first version number can be understood as the latest version number of the metadata corresponding to the snapshot of the latest version number in the basic system table.
[0095] In some optional embodiments, if there is an ongoing query request using the base table and delta table, to ensure query success rate, some snapshot-related records of older version numbers in the base table and delta table can be deleted after the query request has been responded to. That is, the execution of some snapshot-related records of older version numbers in the base table and delta table should be asynchronous and delayed.
[0096] The technical solution provided in this disclosure periodically acquires snapshot change information of business tables in the data lake system. It considers multiple factors, including snapshot change information, the existence of a metadata image corresponding to the business table, and whether the snapshot list acquired in the current period contains snapshots used to build the metadata image, to decide whether to update the metadata image incrementally or rebuild it completely. Once a new metadata image corresponding to the business table is obtained, it is cached in the database system. This allows the database system to perform metadata queries based on the locally cached new metadata image, eliminating the need to query the metadata in the data lake system, significantly reducing metadata query latency and improving metadata query efficiency.
[0097] Figure 2 is a flowchart of a data access method provided in an embodiment of this disclosure. Referring to Figure 2, the method may include the following steps:
[0098] 201. Receive the access request sent by the client. The access request includes the snapshot ID to be accessed and the query conditions.
[0099] 202. Use the streaming query interface to query metadata in the local cached metadata mirror of the database system based on the snapshot ID to be accessed and the query conditions.
[0100] 203. In response to each batch of data files retrieved by the streaming query interface, the number of file metadata files in the current batch is returned to the computing layer of the database system.
[0101] 204. The computing layer performs data processing operations corresponding to the access requests of the data files in the current batch in the data lake system based on the file metadata of the current batch of data files, obtains the data processing results, and returns them to the client.
[0102] Specifically, after the data lake system generates a new metadata image, this image can be cached in the database system to accelerate metadata queries. When a client initiates an access request, the request is parsed to obtain the snapshot ID to be accessed and the query conditions. The snapshot ID is the snapshot ID the client wants to access, and the query conditions specify which partitions, data files, and auxiliary files to query. Then, a streaming query interface is used to perform metadata queries on the locally cached metadata image in the database system based on the snapshot ID and the query conditions.
[0103] In practical applications, multiple metadata images may be built over time, and version management can be performed on these multiple metadata images. The later the metadata image is built, the larger its version number will be. The database system can perform metadata queries based on the metadata image with the largest version number cached locally.
[0104] Taking a metadata mirror that includes the first system table (denoted as the snapshot table), the second system table (denoted as the partition_version table), the third system table (denoted as the file_version table), the fourth system table (denoted as the file_meta table), and the fifth system table (denoted as the sub_files table) as an example, the partition_version table is checked to see if there is a record for the snapshot ID to be accessed. If not, it means that the metadata mirror does not have the file metadata of the data file under the snapshot corresponding to the snapshot ID to be accessed. At this time, it is necessary to access the metadata in the data lake to obtain the file metadata of the data file under the snapshot corresponding to the snapshot ID to be accessed. If a record for the snapshot ID to be accessed is found in the `partition_version` table, the partition name and partition version number under the snapshot ID to be accessed are obtained from the `map_partition_version` field of the `partition_version` table according to the query conditions. For each partition to be accessed, the file name and its corresponding file version number of the data file to be accessed under that partition are obtained from the `file_version` table. Based on the file name and its corresponding file version number of the data file to be accessed, the file metadata of the data file to be accessed is obtained from the `file_meta` and `sub_files` tables. The file metadata of the data file to be accessed can be filtered using query conditions. The streaming query interface returns the file metadata of the found data file to be accessed.
[0105] Taking the metadata mirroring base system table (denoted as the base table) and the difference system table (denoted as the delta table) as examples, based on the snapshot ID to be accessed, in order from the newest version snapshot to the oldest version snapshot, all records in the delta table whose snapshot IDs are less than or equal to the snapshot ID to be accessed are searched. If a record of an add-on (ADD) type data file corresponding to the snapshot ID to be accessed is found, the file metadata of the add-on (ADD) type data file found in the delta table is directly returned. If a record of a delete-on (DELETE) type data file corresponding to the snapshot ID to be accessed is found, the data file can be marked as deleted in memory; if a record of the add-on type of the data file is subsequently found again in the delta table or the base table, the deletion mark of the data file in memory can be removed first. If a record of a modified (MODIFY) type data file corresponding to the snapshot ID to be accessed is found, the file metadata of the data file can be recorded in memory; if a record of the add-on type of the data file is subsequently found again in the delta table or the base table, the file metadata in the add-on type record of the data file and the file metadata of the data file in memory are merged, and the merged file metadata is returned. After searching the delta table, the base table is searched based on the snapshot ID to be accessed to obtain the file metadata of each data file.
[0106] For example, the `base` table records the full data corresponding to the V8 version snapshot (i.e., the file metadata of all data files under the V8 version snapshot); the `delta` table records the data files (including data files of the add, delete, or modify types) corresponding to the snapshots of various version numbers such as V9, V10, V11, and V12. When a client needs to query the file metadata of the data file corresponding to the V11 version snapshot, it first searches the `delta` table sequentially for the file metadata of the data file corresponding to the snapshot of each version number, in the order of V11, V10, and V9. If the file metadata of a data file of the add type is found in the `delta` table, it is directly returned. If the file metadata of a data file of the delete type is found in the `delta` table, the data file is marked as deleted in memory. If the file metadata of a data file of the modify type is found in the `delta` table, the file metadata of that data file can be recorded in memory. After searching for the file metadata of the data files corresponding to the snapshots of version numbers V11, V10, and V9 in the delta table, search for the file metadata of the data files corresponding to the snapshot of version number V8 in the base table. Merge the file metadata of the data files in the base table with the file metadata of the data files with deletion marks and modification types in memory, and return the file metadata of the merged data file.
[0107] A streaming query interface is an interface that allows a database system to return query results sequentially and progressively, rather than waiting for the entire query operation to complete before returning all results at once. This interface is particularly suitable for scenarios involving large amounts of data or requiring real-time responses. In this disclosure, the streaming query interface supports batch querying and returning file metadata of data files. The streaming query interface retrieves the file metadata of one batch of data files at a time and returns the retrieved file metadata to the computation layer. The streaming query interface can then continue querying the file metadata of the next batch of data files. After the computation layer uses the file metadata returned by the streaming query interface, it searches for the corresponding data file in the data lake system. The computation layer can then perform various data processing operations on the found data file, such as SELECT, INSERT, UPDATE, and DELETE, and return the data processing operations to the client. It is understood that parsing the access request sent by the client can determine the type of data processing operation. For example, access requests sent by clients can contain various forms of SQL (Structured Query Language) statements, which can indicate the type of data processing operation such as SELECT, INSERT, UPDATE, and DELETE.
[0108] Optionally, a streaming query interface can be used to sequentially search for the file metadata of the target data files in the local cached metadata mirror of the database system, according to the snapshot ID to be accessed and the query conditions, following the order of each partition. Understandably, the streaming query interface returns the file metadata of a portion of the data files in a partition as soon as it retrieves it, allowing subsequent data processing operations to be performed on these files in the data lake system. This eliminates the need to wait until the file metadata of all data files in a partition has been retrieved before performing data processing operations on all data files in that partition in the data lake, significantly improving data access performance. After querying one partition, the streaming query interface switches to the next partition to continue querying. For example, if partition 1 has 10,000 data files, the file metadata of 10 data files retrieved can be returned to the computation layer first, allowing the computation layer to read the data of those 10 data files based on their file metadata and return the data to the client. After querying partition 1, the query can switch to partition 2 to continue querying, with the streaming query interface returning the metadata of 10 data files to the computation layer for processing. When the computing layer processes files in partition 1, the streaming query interface has already prefetched some metadata from partition 2. This model ensures efficient system operation, maintaining low overall latency regardless of the data size.
[0109] In some optional embodiments, in order to ensure the reliability of metadata query, if the file metadata of the target data file corresponding to the snapshot ID to be accessed is not found in the metadata mirror cached locally in the database system, the streaming query interface is used to query the metadata stored locally in the data lake system according to the snapshot ID to be accessed and the query conditions.
[0110] The technical solution provided in this disclosure reorganizes the metadata storage method of the data lake system to obtain a new metadata image corresponding to the business tables in the data lake system. This new metadata image is then cached in the database system. The database system can then perform metadata queries based on the locally cached new metadata image, eliminating the need to query metadata within the data lake system. This significantly reduces metadata query latency and improves query efficiency. Furthermore, a streaming query interface can be used for metadata queries. This interface supports batch querying and returning file metadata for data files. The streaming query interface retrieves the file metadata for one batch of data files at a time and returns the retrieved file metadata to the computation layer. The interface can then continue querying the file metadata for the next batch of data files. This allows the computation layer to process a certain number of data files in batches with each access request, effectively minimizing the overall latency of metadata queries, achieving a near-zero latency experience, and improving data access performance.
[0111] To facilitate understanding, a specific scenario implementation is described below with reference to Figure 3. Referring to Figure 3, in this application scenario, the data lake system can store business tables, a list of snapshots of the business tables, and metadata of the business tables. Over time, a metadata mirror can be built based on the metadata of the business tables in the data lake system and cached in the database system. After a client initiates a query request, the database system can use a streaming query interface to first query the file metadata of the required data file in the metadata mirror. Then, through the computation layer, it accesses the data lake system based on the file metadata of the required data file to find the file data and returns the search results to the client. Of course, to ensure query reliability, if the file metadata of the required data file is not found in the metadata mirror, the streaming query interface can also be used to query the metadata stored locally in the data lake. Then, through the computation layer, it queries the corresponding data file in the data lake based on the file metadata found in the data lake and returns the query results to the client.
[0112] In this implementation, the metadata storage method of the data lake system is reorganized, and the query interface for retrieving file metadata of data files is optimized to return query results in a streaming manner. This improvement allows the computing layer to process a certain number of data files in batches with each access request, enabling batch processing operations on these data files. Simultaneously, while processing the current batch of data files, the system can asynchronously pre-load the file metadata of subsequent batches of data files. Therefore, except for the initial synchronous wait during the first call to the streaming query interface, subsequent operations are performed asynchronously. Regardless of the file size, the latency of the entire process is mainly concentrated during the initial call to the streaming query interface, while subsequent operations can achieve almost instantaneous response, effectively minimizing the overall latency of metadata queries and approaching a zero-latency experience.
[0113] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 101 to 103 can be device A; or the execution subject of steps 101 and 102 can be device A, and the execution subject of step 103 can be device B; and so on.
[0114] Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations appearing in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0115] Figure 4 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. As shown in Figure 4, the electronic device includes: a memory 41 and a processor 42;
[0116] Memory 41 is used to store computer programs and can be configured to store various other data to support operation on the computing platform. Examples of this data include instructions for any application or method operating on the computing platform, data structures, contact data, phone book data, messages, pictures, videos, etc.
[0117] Processor 42, coupled to memory 41, is used to execute computer programs in memory 41 for steps in: metadata query methods or data access methods.
[0118] Optionally, as shown in Figure 4, the electronic device may also include other components such as a communication component 43, a display 44, a power supply component 45, and an audio component 46. Figure 4 only schematically shows some components and does not imply that the electronic device only includes the components shown in Figure 4. Furthermore, the components within the dashed boxes in Figure 4 are optional, not mandatory, and their specific inclusion depends on the product form of the electronic device. The electronic device of this embodiment can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT (Internet of Things) device, or a server-side device such as a conventional server, cloud server, or server array. If the electronic device of this embodiment is a terminal device such as a desktop computer, laptop computer, or smartphone, it may include the components within the dashed boxes in Figure 4; if the electronic device of this embodiment is a server-side device such as a conventional server, cloud server, or server array, it may not include the components within the dashed boxes in Figure 4.
[0119] The aforementioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0120] The aforementioned communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as 2G (2nd Generation), 3G (3rd Generation), 4G (4th Generation) / LTE (long Term Evolution), 5G (5th Generation), or combinations thereof. In an exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.
[0121] The aforementioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.
[0122] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.
[0123] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0124] Accordingly, this disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium may be volatile, non-volatile, or a combination thereof, and may be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random-access memory (SRAM), dynamic random-access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, digital video disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium.
[0125] Accordingly, this disclosure also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, cause the processor to implement the steps in the above-described method embodiments. It should be understood that each step or combination of steps in the above-described method flow can be implemented by the computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above-described method embodiments.
[0126] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0127] The above are merely embodiments of this disclosure and are not intended to limit the scope of this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of the claims of this disclosure.
Claims
1. A metadata query method, wherein, include: The snapshot list of business tables in the data lake system is periodically retrieved, and the snapshot change information of the snapshot list retrieved in the current period is determined relative to the snapshot list retrieved in the previous period. The snapshot list includes snapshots of multiple version numbers corresponding to the business tables. Determine a first existence and a second existence, wherein the first existence indicates whether there is a first metadata image corresponding to the business table, the first metadata image being obtained by reorganizing the metadata of the business table stored in the data lake system; the second existence indicates whether there is a snapshot with a first version number in the snapshot list obtained in the current period, the first version number being the version number corresponding to the snapshot used to build the first metadata image. Based on the snapshot change information, the first existence and the second existence, update the first metadata image to obtain the second metadata image; The second metadata image is cached in the database system so that the database system can perform metadata queries on the second metadata image.
2. The method according to claim 1, wherein, Based on the snapshot change information, the first existence, and the second existence, the first metadata image is updated to obtain the second metadata image, including: The snapshot change information is determined to reflect that the maximum version number of the snapshots in the snapshot list acquired in the current period is higher than or equal to the maximum version number of the snapshots in the snapshot list acquired in the previous period. If the first existence representation does not contain the first metadata image corresponding to the business table, or if the second existence representation does not contain a snapshot with the first version number in the snapshot list obtained in the current period, then the snapshot with the largest version number in the snapshot list obtained in the current period shall be taken as the snapshot with the latest version number. The second metadata image corresponding to the business table is built using a full build method based on the snapshot of the latest version number.
3. The method according to claim 1, wherein, Based on the snapshot change information, the first existence, and the second existence, the first metadata image is updated to obtain the second metadata image, including: The snapshot change information is determined to reflect that the maximum version number of the snapshots in the snapshot list acquired in the current period is higher than the maximum version number of the snapshots in the snapshot list acquired in the previous period. If the first existence representation has a first metadata image corresponding to the business table, and the second existence representation has a first version number in the snapshot list obtained in the current period, then based on the incremental snapshot information between the snapshot with the largest version number in the snapshot list obtained in the current period and the snapshot with the first version number, the first metadata image corresponding to the business table is updated in an incremental construction manner to obtain the second metadata image.
4. The method according to claim 2, wherein, Based on the latest version number snapshot, a second metadata image corresponding to the business table is built using a full build method, including: Insert the business table ID, first snapshot ID, and snapshot metadata corresponding to the latest version snapshot into a new first system table as a record; The partition version number of each partition associated with the snapshot of the latest version number is determined as the first snapshot ID. The business table ID corresponding to the snapshot of the latest version number, the first snapshot ID, the partition name of each partition and its corresponding partition version number are inserted as a record into the new second system table. The file version number of each data file in each partition is determined as the first snapshot ID, and the business table ID corresponding to the latest version snapshot number, the first snapshot ID, the partition name of each partition, the file name and file version number of all data files under each partition are inserted as a record into a new third system table; Get the file name and file metadata of each data file associated with the latest version snapshot; insert the business table ID corresponding to the latest version snapshot, the file name and file metadata of each data file as a record into a new fourth system table; Obtain the information of the auxiliary files corresponding to each data file associated with the latest version number snapshot, and insert the business table ID corresponding to the latest version number snapshot, the first snapshot ID, the file name of each data file, and the information of the auxiliary files corresponding to each data file as a record into the new fifth system table; Based on the new first system table, the new second system table, the new third system table, the new fourth system table, and the new fifth system table, construct the second metadata mirror corresponding to the business table.
5. The method according to claim 3, wherein, The first metadata mirror includes: an existing first system table, an existing second system table, an existing third system table, an existing fourth system table, and an existing fifth system table; correspondingly, the incremental snapshot information between the snapshot with the largest version number in the snapshot list obtained in the current period and the snapshot with the first version number includes: Determine the target incremental snapshot information of the next version number relative to the previous version number, wherein the next version number is greater than the first version number and less than or equal to the maximum version number corresponding to the snapshot in the snapshot list obtained in the current period, and the previous version number is greater than or equal to the first version number and less than the maximum version number corresponding to the snapshot in the snapshot list obtained in the current period. Insert the business table ID, snapshot ID, and snapshot metadata corresponding to the next version number snapshot into the existing first system table as a record to obtain the new first system table; Based on the target incremental snapshot information, the changed partition is determined. The business table ID, snapshot ID, partition name and partition version number of the changed partition corresponding to the snapshot of the next version number are inserted as a record into the existing second system table to obtain a new second system table. The partition version number of the changed partition is the snapshot ID of the snapshot of the next version number. Update the existing third system table, the existing fourth system table, and the existing fifth system table according to the change type of the changed files in the changed partition, so as to obtain the new third system table, the new fourth system table, and the new fifth system table respectively. Based on the new first system table, the new second system table, the new third system table, the new fourth system table, and the new fifth system table, construct the second metadata mirror corresponding to the business table.
6. The method according to claim 4 or 5, characterized in that, The first system table uses the business table ID field and the snapshot ID field as primary key fields. The fields of the first system table also include a snapshot metadata field, which is used to record at least one of the snapshot version number and the snapshot creation timestamp.
7. The method according to any one of claims 4 to 6, characterized in that, The second system table uses the business table ID field and the snapshot ID field as primary key fields. The second system table also includes a partition version number field, which is used to record key-value pairs of each partition name and its corresponding partition version number.
8. The method according to any one of claims 4 to 7, characterized in that, The third system table uses the business table ID field, snapshot ID field, and partition name field as a composite primary key field. The fields of the third system table also include a file version number field, which is used to record the key-value pairs of the names of each data file under the partition and their corresponding file version numbers.
9. The method according to claim 2, wherein, The second metadata image corresponding to the business table is constructed using a full build method based on the latest version snapshot, including: Insert the business table ID corresponding to the latest version snapshot number, the first snapshot ID, the partition name of each partition, the file name of each data file under each partition and its corresponding file metadata as a record into the new basic system table; The second metadata mirror corresponding to the business table is constructed based on the new basic system table.
10. The method according to claim 3, wherein, Based on the incremental snapshot information between the snapshot with the largest version number in the snapshot list obtained in the current period and the snapshot with the first version number, the first metadata image corresponding to the business table is updated in an incremental construction manner to obtain the second metadata image, including: Determine the target incremental snapshot information of the next version number relative to the previous version number, wherein the next version number is greater than the first version number and less than or equal to the maximum version number corresponding to the snapshot in the snapshot list obtained in the current period, and the previous version number is greater than or equal to the first version number and less than the maximum version number corresponding to the snapshot in the snapshot list obtained in the current period. The changed partition and the change type of the changed files in the changed partition are determined based on the target incremental snapshot information. Insert the business table ID, snapshot ID, partition name of the changed partition, file name of the changed file, change type, and file metadata corresponding to the snapshot of the next version number into the existing difference system table as a record to obtain the new difference system table; Construct a second metadata mirror corresponding to the business table based on the new difference system table.
11. The method according to claim 9, characterized in that, The fields of the basic system table include the business table ID field, snapshot ID field, partition name field, file name field, and file metadata field.
12. The method according to claim 10, characterized in that, The fields of the difference system table include the business table ID field, snapshot ID field, partition name field, file name field, change type field, and change file metadata field. The change type field can take at least one of the following values: add, modify, and delete.
13. The method according to any one of claims 1 to 12, characterized in that, The snapshot change information includes at least one of the following: newly added snapshots, deleted snapshots, and modified snapshots in the current period snapshot list relative to the previous period snapshot list. It also includes the relationship between the maximum version number of the current period snapshot list and the maximum version number of the previous period snapshot list.
14. The method according to any one of claims 1 to 13, wherein, After constructing the second metadata mirror corresponding to the business table, the following is also included: In response to the deletion of a snapshot with the second version number in the data lake system, the metadata associated with the snapshot with the second version number in the second metadata image is deleted.
15. A data access method, wherein, include: Receive an access request sent by a client, the access request including the snapshot ID to be accessed and query conditions; The streaming query interface is used to query metadata in the local cached metadata mirror of the database system based on the snapshot ID to be accessed and the query conditions; In response to each batch of data files retrieved by the streaming query interface, the number of file metadata files in the current batch is returned to the computing layer of the database system. The computing layer uses the file metadata of the data files in the current batch to perform the data processing operation corresponding to the access request on the data files in the current batch in the data lake system, obtains the data processing result, and returns it to the client.
16. The method according to claim 15, wherein, The streaming query interface is used to perform metadata queries in the local cached metadata mirror of the database system based on the snapshot ID to be accessed and query conditions, including: Using the streaming query interface, based on the snapshot ID to be accessed and the query conditions, the file metadata of the data files under each partition is searched sequentially in the local cache metadata mirror of the database system according to the order of each partition.
17. The method according to claim 15 or 16, wherein, Also includes: If the file metadata corresponding to the snapshot ID to be accessed is not found, the streaming query interface is used to query the metadata stored locally in the data lake system based on the snapshot ID to be accessed and the query conditions.
18. An electronic device, wherein, include: Memory and processor; The memory is used to store computer programs; The processor is coupled to the memory for executing the computer program to perform the steps of the method according to any one of claims 1-14 or 15-17.
19. A computer-readable storage medium storing a computer program, wherein, When the computer program is executed by a processor, it causes the processor to perform the steps of the method according to any one of claims 1-14 or 15-17.
20. A computer program product, wherein, Includes a computer program or instructions that, when executed by a processor, cause the processor to perform the steps of the method according to any one of claims 1-14 or 15-17.