A Method for Incremental Consumption of Data Lake Data Based on Two-Layer Time Identification
By using two-layer time identification method in the data lake, the problem that data in the data lake does not support incremental consumption is solved, the efficiency of incremental query is achieved, and the needs of real-time computing is met.
Patent Information
- Application Number
- CN202211070114.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-02
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-09-02
AI Technical Summary
The current data lake architecture has an important functional flaw: the data in the data lake does not support incremental consumption, resulting in the time complexity of incremental queries is O(n), which cannot meet the business needs of real-time computing.
The incremental consumption method of data lake data based on two-layer time identification is adopted. When data is written to the data lake, time identification is added to different batches of data, and time identification is added to the file name to quickly locate the storage path of data and support incremental query.
Through the two-layer time identification method, the time complexity of incremental query is reduced, from O(n) to O(1), which meets the business needs of real-time computing and solves the problem that data cannot be consumed incrementally after entering the data lake.
Smart Images

Figure CN115470223B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data lakes, and specifically, to a method for incrementally consuming data in a data lake based on two-layer time identifiers. Background Art
[0002] The development of data management technology has mainly gone through three stages: data warehouses, data lakes, and lakehouse integration proposed at the current stage.
[0003] Data warehouses mainly rely on traditional databases to achieve data storage, calculation, and access, and are mainly used for functions such as BI (Business Intelligence, also known as business wisdom or business intelligence, which refers to using modern data warehouse technology, online analytical processing technology, data mining, and data presentation technology for data analysis to achieve business value) and reporting. The main characteristics of data warehouses are: strict data systems, standard formats, relatively easy data governance, and easy to obtain high optimization for specific engines. The disadvantages are that they can only support structured data and have poor cluster scalability.
[0004] The development of data lakes has been approximately less than 10 years so far. Currently, it mainly relies on the Hadoop ecosystem and is used to build data storage for structured, semi-structured, and unstructured data, and can be used for scientific exploration and value mining of heterogeneous data. However, the data is relatively flexible, and the difficulty of data governance is relatively large, resulting in a low degree of data utilization.
[0005] To integrate the advantages of both, a lakehouse integration big data architecture has been proposed at the current stage. The data lake absorbs the advantages of the data warehouse, breaks through the data barriers between the two, provides support for data ingestion, storage, calculation, governance, data services, machine learning, etc., forms a complete closed-loop system, and truly achieves a big data solution of "once ingested, used multiple times".
[0006] Currently, one of the most important features of the lakehouse platform is its support for the computing scenario of stream-batch integration. To enable the application of the stream computing scenario, the platform needs to support incremental data writing and consumption. Hadoop is the mainstream big data ecosystem framework, and excellent frameworks for big data storage and computing such as Hive, HBase, and Spark have been derived based on Hadoop. Building a data lake based on the Hadoop ecosystem is the current mainstream big data technology trend. However, there is an important functional defect in the current data lake architecture: the data in the data lake does not support incremental consumption. For example, when querying incremental data in Hadoop through Hive, Spark, etc., it is necessary to filter partitions and full tables, which is actually not much different from full-table query calculation, and the time complexity is both O(n). This means that the larger the data, the longer the query time. For incremental data, the rate of data writing into the data lake is basically constant. Therefore, to achieve incremental consumption, the target of its time complexity is O(1). Summary of the Invention
[0007] In view of the requirements and deficiencies of the current technological development, the present invention provides a method for incremental consumption of data lake data based on two-layer time identifiers, which solves the defect that data cannot be incrementally consumed after entering the data lake.
[0008] The method for incremental consumption of data lake data based on two-layer time identifiers of the present invention adopts the following technical solutions to solve the above technical problems:
[0009] A method for incremental consumption of data lake data based on two-layer time identifiers, the method includes two stages: writing data into the data lake and querying data in the data lake;
[0010] (1) In the stage of writing data into the data lake,
[0011] (1.1) Create an incremental table in the "metastore" according to the table structure information of the data. The "metastore" is a metadata service.
[0012] (1.2) Obtain this batch of data, start a thread as a time server, and the client generates a timestamp T by passing through the local time of the time server operating system. i , the timestamp T i is used as the time identifier for writing this batch of data into the data lake.
[0013] (1.3) Estimate the data volume included in this batch of data and create Y files.
[0014] (1.4) Divide the data of this batch according to the number of files and write them into Y files correspondingly. During the process of writing data into files, write data statistical information at the footer of the file. The data statistical information includes the amount of data contained in the file, the maximum value information and the minimum value information of column storage. Write a Bloom index at the header of the file.
[0015] (1.5) After all the data of this batch is written into the data lake, record the writing of this batch of data as a Log in the commit file.
[0016] (2) In the stage of querying data in the data lake,
[0017] (2.1) Specify the incremental table to be consumed, the starting consumption timestamp T0, and the time range between_time for each consumption by executing the set method.
[0018] (2.2) Determine whether the incremental table specified in step (2.1) supports incremental query. If it supports, continue to execute step (2.3).
[0019] (2.3) Parse the SQL statement to generate a Job. In the Job, obtain the value of the timestamp field “_commit_time_”, that is, the starting consumption timestamp T0.
[0020] (2.4) Filter the current incremental table by the timestamp T0 to obtain the storage paths of the files that meet the condition of being greater than the timestamp T0. The storage paths of multiple files form an array of file lists[]. Return the array of file lists[] to the Job to generate the task to be executed.
[0021] Optionally, the incremental table created in step (1.1) includes the name of the table, the fields of the table, the storage format of the table, and the actual storage location of the table.
[0022] Optionally, when creating the incremental table in step (1.1), a timestamp field “_commit_time_” needs to be added. The storage format of the data to be executed is the Parquet format. A unique field needs to be provided as the primary key information of the table, and the UUID default mode is supported.
[0023] Optionally, after using the generated timestamp Ti as the time identifier for writing this batch of data into the data lake in step (1.2),
[0024] The client first calls the API interface to obtain the timestamp Ti-1 of the previous batch of data written into the data lake.
[0025] Subsequently, compare the timestamp Ti of this batch of data with the timestamp Ti-1 of the previous batch of data.
[0026] (a) If the timestamp Ti is less than the timestamp Ti-1, it indicates that there is an abnormality in the time server, or there is a time conflict caused by concurrent data writing. At this time, the client will write the data of this batch into the failure queue, and then throw an exception to the foreground, prompting the client to handle the exception and then continue to write the data of this batch.
[0027] (b) If the timestamp Ti is greater than the timestamp Ti-1, then directly use the timestamp Ti as the time identifier for writing the data of this batch into the data lake.
[0028] Further optionally, when performing step (1.3), the specific process of creating Y files is as follows:
[0029] Estimate that the amount of data to be written into the data lake in this batch is X, the storage space occupied by each piece of data is m, and set the threshold threshold for each file. Then, the number of files to be created is Y = mX / threshold.
[0030] Preferably, the generated files are in Parquet format, and the naming rule is: random string + timestamp + sequence of the number of files written this time.
[0031] Further optionally, when performing step (1.4), the specific operation of writing the Bloom index in the file header is as follows:
[0032] First, based on the amount of data written in the file, obtain the actual threshold of the file.
[0033] Then, determine how many bit positions are needed to store the Bloom index according to the actual threshold of the file.
[0034] Then, calculate the result flags at multiple positions for each row of UUID through multiple hash algorithms, and write the flags into the bit storage according to the bit flag positions.
[0035] Finally, when writing each row of data, add a timestamp field "_commit_time_" to this record data and assign it the value Ti.
[0036] Further optionally, the content of the Log includes: how much data is written in this batch, which new files are created, which files are merged resulting in the invalidation of old files, and the timestamp Ti written in this batch.
[0037] Preferably, the format of the timestamp is yyyymmddhhmmss.
[0038] Further optionally, when performing step (2.3), obtain the value of the timestamp field "_commit_time_" in the Job, that is, the starting consumption timestamp T0. The specific process is as follows:
[0039] (2.3.1) Parse the SQL statement to generate a Job. Obtain Conditions through the syntax analyzer and determine whether the syntax conforms to the format of incremental query. If it conforms, continue to execute (2.3.2).
[0040] (2.3.2) Obtain the timestamp field _commit_time_, and obtain the starting time identifier T0 of the incremental query from the hash table through keywords.
[0041] (2.3.3) Determine whether the time identifier T0 conforms to the format of the timestamp. If it conforms, return the starting consumption timestamp T0.
[0042] (2.3.4) Obtain the time range between_time in the configuration parameters when executing set. Based on the starting consumption timestamp T0, generate the ending timestamp Tend.
[0043] A method for incremental consumption of data lake data based on two-layer time identifiers according to the present invention has the beneficial effects compared with the prior art as follows:
[0044] When writing data into the data lake in the present invention, time identifiers are added to different batches of data, and time identifiers are also added to the file names when writing the data of the same batch into the files of the data lake. The addition of these two time identifiers can quickly locate the storage path of the data, realize incremental query of the data, meet the business requirements of real-time computing, solve the defect that incremental consumption cannot be performed after the data enters the data lake, and reduce the time complexity from O(n) to O(1). Brief Description of the Drawings
[0045] Attached Figure 1 is a schematic flow chart of the data writing stage into the data lake of the present invention;
[0046] Attached Figure 2 is a schematic flow chart of querying data in the data lake of the present invention. Detailed Embodiments
[0047] To make the technical solutions, technical problems to be solved, and technical effects of the present invention clearer and more understandable, the following describes the technical solutions of the present invention clearly and completely in combination with specific embodiments.
[0048] Embodiment 1:
[0049] This embodiment proposes a method for incremental consumption of data lake data based on two-layer time identifiers. This method includes two stages: writing data into the data lake and querying data in the data lake.
[0050] (1) Combining with Attached Figure 1 , in the stage of writing data into the data lake:
[0051] (1.1) Create an incremental table in the "metastore" based on the table structure information of the data. The "metastore" is a metadata service.
[0052] The created incremental table includes the table name, table fields, table storage format, and the actual storage location of the table.
[0053] When creating the incremental table, a timestamp field "_commit_time_" needs to be added. The storage format of the data to be executed is Parquet format. A unique field needs to be provided as the primary key information of the table, and the UUID default mode is supported.
[0054] (1.2) Obtain the data of this batch. Start a thread as a time server. The client obtains the local time of the operating system through the time server and generates a timestamp T in the format of yyyymmddhhmmss. i , the timestamp T i is used as the time identifier for writing this batch of data into the data lake. Subsequently, the client calls the API interface to obtain the timestamp Ti-1 of the previous batch of data written into the data lake, and further compares the timestamp Ti of this batch of data with the timestamp Ti-1 of the previous batch of data.
[0055] (a) If the timestamp Ti is less than the timestamp Ti-1, it means that there is an exception in the time server, or concurrent writing of data causes a time conflict. At this time, the client will write this batch of data into the failure queue and then throw an exception to the foreground, prompting the client to handle the exception and then continue writing this batch of data.
[0056] (b) If the timestamp Ti is greater than the timestamp Ti-1, then directly use the timestamp Ti as the time identifier for writing this batch of data into the data lake.
[0057] (1.3) Estimate the data volume included in this batch of data and generate Y Parquet format files. The specific process is as follows:
[0058] Estimate that the data volume to be written into the data lake for this batch is X, and the storage space occupied by each piece of data is m.
[0059] Set the threshold threshold for each file.
[0060] Then, the number of files to be created is Y = mX / threshold.
[0061] Set the naming rule of the file as: random string + timestamp + sequence of the number of files written this time. For example, 123e4567-e89b-12d3-a456-426655440000_20211102171312789_2.parquet, which is the file name that meets the file naming rule.
[0062] (1.4) Divide the data of this batch according to the number of files and write them into Y files correspondingly.
[0063] During the process of writing data into the file,
[0064] Write data statistical information at the footer of the file. The data statistical information includes the amount of data contained in the file, the maximum value information and the minimum value information of column storage;
[0065] Write a Bloom index at the header of the file. The specific operation is as follows:
[0066] First, based on the amount of data written in the file, obtain the actual threshold of the file.
[0067] Then, determine how many bit positions are needed to store the Bloom index according to the actual threshold of the file.
[0068] Then, calculate the result flags at multiple positions for each row of UUID through multiple hash algorithms, and write the flags into the bit storage according to the bit flag positions.
[0069] Finally, when writing each row of data, add a timestamp field "_commit_time_" to this record data and assign it the value Ti.
[0070] (1.5) After all the data of this batch is written into the data lake, record the writing of this batch of data as a Log in the commit file. The content of the Log includes: how much data is written in this batch, which new files are created, which files are merged resulting in the invalidation of old files, and the timestamp Ti of this batch of writing.
[0071] (2) Combine the appendix Figure 2 , in the stage of querying data in the data lake:
[0072] (2.1) Specify the incremental table to be consumed, the starting consumption timestamp T0, and the time range between_time for each consumption by executing the set method. The format for executing the set method is: support.increment.table = database.tablename.
[0073] (2.2) Determine whether the specified incremental table in step (2.1) supports incremental query. If it does, continue to execute step (2.3).
[0074] (2.3) Parse the SQL statement to generate a Job, and obtain the value of the timestamp field "_commit_time_", i.e., the starting consumption timestamp T0, in the Job. The specific process is as follows:
[0075] (2.3.1) Parse the SQL statement to generate a Job, obtain Conditions through the syntax analyzer, and determine whether the syntax conforms to the format of incremental query. If it does, continue to execute (2.3.2);
[0076] (2.3.2) Obtain the timestamp field _commit_time_, and obtain the starting time identifier T0 of the incremental query from the hash table through keywords;
[0077] (2.3.3) Determine whether the time identifier T0 conforms to the timestamp format. If it does, return the starting consumption timestamp T0;
[0078] (2.3.4) Obtain the time range between_time in the configuration parameters when executing set, and generate the ending timestamp Tend based on the starting consumption timestamp T0.
[0079] (2.4) Filter the current incremental table through the timestamp T0, obtain the storage paths of the files that satisfy being greater than the timestamp T0. The storage paths of multiple files form an array of file lists[], and return the array of file lists[] to the Job to generate the executed task.
[0080] In summary, by adopting a data lake data incremental consumption method based on two - layer time identifiers of the present invention, the storage path of data can be quickly located, incremental query of data can be realized, the business requirements of real - time calculation can be met, the defect that incremental consumption cannot be performed after data enters the data lake can be solved, and the time complexity is reduced from O(n) to O(1).
[0081] The above applications use specific examples to elaborate in detail the principle and implementation manner of the present invention. These embodiments are only used to help understand the core technical content of the present invention. Based on the above - mentioned specific embodiments of the present invention, those skilled in the art of this technology, without departing from the principle of the present invention, any improvements and modifications made to the present invention shall fall within the scope of patent protection of the present invention.
Claims
1. A method for incremental consumption of data lake data based on two - layer time identifiers, characterized in that, The method includes two stages: writing data to the data lake and querying data in the data lake; (1) In the stage of writing data to the data lake, (1.1) Create an incremental table in the "metastore" according to the table structure information of the data, (1.2) Get this batch of data, start a thread as a time server, and the client generates a timestamp T through the local time of the time server operating system. i , timestamp T i As the time stamp of the data batch written into the data lake, (1.3) Estimate the amount of data contained in this batch of data, create Y files, and the generated files are in Parquet format. The naming rule is: random string + timestamp + sequence number of the number of files written this time; (1.4) Divide this batch of data according to the number of files and write it to Y files correspondingly. During the process of writing data to the files, write data statistics information at the footer of the file. The data statistics information includes the amount of data contained in the file, the maximum value information and the minimum value information of column storage, and write a Bloom index at the header of the file, (1.5) After all the data in this batch is written to the data lake, record the writing of this batch of data as a Log in the commit file; (2) In the stage of querying data in the data lake, (2.1) Specify the incremental table to be consumed, the starting consumption timestamp T0, and the time range between_time for each consumption by executing the set method, (2.2) Judge whether the specified incremental table in step (2.1) supports incremental query. If it supports, continue to execute step (2.3), (2.3) Parse the SQL statement to generate a Job, and obtain the value of the timestamp field "_commit_time_" in the Job, that is, the starting consumption timestamp T0, (2.4) Filter the current incremental table by the timestamp T0 to obtain the storage paths of the files that meet the condition of being greater than the timestamp T0. The storage paths of multiple files form an array of file lists[], and return the file lists[] array to the Job to generate an executable task.
2. The method for incrementally consuming data in a data lake based on two - layer time identifiers according to claim 1, wherein, The incremental table created by executing step (1.1) includes the name of the table, the fields of the table, the storage format of the table, and the actual storage location of the table.
3. The method for incrementally consuming data in a data lake based on two - layer time identifiers according to claim 2, wherein, When creating an incremental table by executing step (1.1), a timestamp field "_commit_time_" needs to be added. The storage format of the data to be executed is Parquet format, and a unique field needs to be provided as the primary key information of the table, supporting the UUID default mode.
4. A method for incrementally consuming data in a data lake based on two - layer time identifiers according to claim 1 or 3, characterized in that, After executing step (1.2) and using the generated timestamp Ti as the time identifier for writing this batch of data to the data lake, The client first calls the API interface to obtain the timestamp Ti-1 of the previous batch of data written to the data lake, Subsequently, compare the timestamp Ti of this batch of data with the timestamp Ti-1 of the previous batch of data, (a) If the timestamp Ti is less than the timestamp Ti-1, it means that there is an abnormality in the time server, or concurrent data writing causes a time conflict. At this time, the client will write this batch of data to the failure queue, and then throw an exception to the foreground, prompting the client to handle the exception and then continue writing this batch of data, (b) If the timestamp Ti is greater than the timestamp Ti-1, directly use the timestamp Ti as the time identifier for writing this batch of data to the data lake.
5. A method for incrementally consuming data in a data lake based on two-layer time identifiers according to claim 3, characterized in that, The specific process of creating Y files by executing step (1.3) is: It is estimated that the amount of data to be written into the data lake in this batch is X, the storage space occupied by each piece of data is m, and the threshold of each file is set to threshold. Then, the number of files to be created is Y = mX / threshold.
6. A method for incrementally consuming data in a data lake based on two - layer time identifiers according to claim 3, wherein, Execute step (1.4). The specific operation of writing the Bloom index in the header of the file is as follows: First, based on the amount of data written in the file, obtain the actual threshold of the file. Then, determine how many bit positions are needed to store the Bloom index according to the actual threshold of the file. Next, calculate the result flags at multiple positions for each line of UUID through multiple hash algorithms, and write the flags into the bit storage according to the bit flags. Finally, when writing each line of data, add a timestamp field "_commit_time_" to this record data and assign it the value Ti.
7. A method for incremental consumption of data lake data based on two - layer time identifiers according to claim 1, wherein, The content of the Log includes: how much data is written in this batch, which new files are created, which files are merged resulting in the invalidation of old files, and the timestamp Ti of this batch of writes.
8. A method for incrementally consuming data in a data lake based on two - layer time identifiers according to claim 1, characterized in that, The format of the timestamp is yyyymmddhhmmss.
9. A method for incrementally consuming data in a data lake based on two-layer time identifiers according to claim 8, characterized in that Execute step (2.3). The specific process of obtaining the value of the timestamp field "_commit_time_" in the Job, that is, the starting consumption timestamp T0, is as follows: (2.3.1) Parse the SQL statement to generate a Job. Use the syntax analyzer to obtain Conditions and determine whether the syntax conforms to the format of incremental query. If it conforms, continue to execute (2.3.2); (2.3.2) Obtain the timestamp field _commit_time_, and obtain the starting time identifier T0 of the incremental query from the hash table through keywords. (2.3.3) Determine whether the time identifier T0 conforms to the format of the timestamp. If it conforms, return the starting consumption timestamp T0. (2.3.4) Obtain the time range between_time in the configuration parameters when executing set. Based on the starting consumption timestamp T0, generate the end timestamp Tend.
Citation Information
Patent Citations
System for importing data into a data repository
CN109997125A
Data processing method and device, electronic equipment and storage medium
CN111209352A