A file merging method and apparatus

By directly using the initial data of multiple data lake files to be merged as storage data in the target data lake files and obtaining the target metadata based on the storage location, the inefficiency problem caused by the small file problem in the HDFS cluster is solved, and efficient file merging is achieved.

CN116932497BActive Publication Date: 2025-05-30BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310919449.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-25
Publication Date
2025-05-30
Estimated Expiration
2043-07-25

AI Technical Summary

Technical Problem

Due to small file problems in HDFS clusters, the number of metadata increases, waste of storage space, inefficient data replication, inefficient access efficiency and increased system management complexity, the existing technology has low efficiency in merging small files.

Method used

Provide a file merging method, by obtaining the initial data stored in multiple data lake files to be merged, and directly using it as storage data in the target data lake file, obtaining the target metadata according to the location of the stored data, and writing it to the target data lake file, thereby realizing file merging.

Benefits of technology

There is no need to decoding, decompressing, deserializing the data stored in the merged data lake file, and the original data is directly used as the stored data in the target data lake file, which improves the efficiency of file merging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116932497B_ABST
    Figure CN116932497B_ABST
Patent Text Reader

Abstract

The present application discloses a file merging method, including: obtaining initial data stored in a plurality of data lake files to be merged, where the initial data is data that has been processed by one or more of serialization, compression, and encoding. Directly use the initial data as the stored data in the target data lake file, and obtain the target metadata of the target data lake file according to the storage location of the stored data in the target data lake file. Write the stored data and the target metadata into the target data lake file, so as to obtain the target data lake file merged from the plurality of data lake files to be merged. By using the solution of the embodiment of the present application, when merging a plurality of data lake files to be merged into a target data lake file, it is not necessary to first process the data stored in the data lake files to be merged, but the original data can be directly used as the stored data in the target data lake file, thereby improving the efficiency of merging a plurality of data lake files to be merged into a target data lake file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and particularly to a file merging method and apparatus. Background Art

[0002] Lakehouse is a new data management mode that integrates the differences between a data warehouse and a data lake, and constructs a data warehouse on top of the data lake to form a lakehouse system. The lakehouse system can effectively simplify the data infrastructure, improve data storage elasticity and quality while reducing costs and data redundancy. The underlying data of the lakehouse system is generally stored in a Hadoop Distributed File System (HDFS) cluster in Parquet format.

[0003] As the data scale and application scenarios continue to expand, some problems will occur in the HDFS cluster. One of the more common problems is the small file problem. Among them: small files refer to files whose size is much smaller than the size of a data block. Storing a large number of small files may cause various problems, including an increase in the number of metadata, wasted storage space, low data replication efficiency, low access efficiency, and increased system management complexity.

[0004] To avoid the above problems caused by storing a large number of small files, small files can be merged. Currently, the efficiency of merging small files is low. Therefore, there is an urgent need for a solution to solve the above problems. Summary of the Invention

[0005] To solve or partially solve the above technical problems, embodiments of this application provide a file merging method and apparatus.

[0006] In a first aspect, embodiments of this application provide a file merging method, the method comprising:

[0007] Obtain initial data stored in multiple data lake files to be merged, where the initial data is data that has been processed by one or more of serialization, compression, and encoding;

[0008] Directly use the initial data as the stored data in the target data lake file;

[0009] Obtain target metadata of the target data lake file according to the storage location of the stored data in the target data lake file;

[0010] Write the stored data and the target metadata into the target data lake file.

[0011] Optionally, the writing the stored data and the target metadata into the target data lake file includes:

[0012] First, directly write the initial data into the target data lake file, and then write the target metadata into the target data lake file.

[0013] Optionally, obtaining the target metadata of the target data lake file according to the storage location of the stored data in the target data lake file includes:

[0014] After directly writing the initial data into the target data lake file, obtain the target metadata of the target data lake file according to the storage location of the stored data in the target data lake file.

[0015] Optionally, before obtaining the initial data stored in multiple data lake files, the method further includes:

[0016] Receive a target Structured Query Language (SQL) statement, where the target SQL statement is used to specify the data lake files to be merged;

[0017] Determine the multiple data lake files to be merged according to the target SQL statement.

[0018] Optionally, the target SQL statement further includes a data lake file merge statement, and the data lake file merge statement is used to indicate merging the multiple data lake files to be merged.

[0019] Optionally, the target SQL statement further includes a file size setting statement, and the file size setting statement is used to indicate that the size of the target data lake file is a target size, and the size of the generated target data lake file is the target size.

[0020] Optionally, obtaining the initial data stored in multiple data lake files to be merged includes:

[0021] Traverse the multiple data lake files to be merged, and for each data lake file to be merged, perform the following operations:

[0022] Determine the offset and size of the row groups in the data lake file to be merged according to the metadata of the data lake file to be merged;

[0023] Determine the offset and size of the column chunks in the data lake file to be merged according to the metadata of the row groups;

[0024] Obtain the data in the column chunks as the initial data.

[0025] In a second aspect, an embodiment of the present application provides a file merging device, and the device includes:

[0026] An acquisition unit, configured to acquire initial data stored in a plurality of data lake files to be merged, where the initial data is data that has been processed by one or more of serialization, compression, and encoding;

[0027] A first determination unit, configured to directly use the initial data as the stored data in the target data lake file;

[0028] A second determination unit, configured to obtain target metadata of the target data lake file according to the storage location of the stored data in the target data lake file;

[0029] A writing unit, configured to write the stored data and the target metadata into the target data lake file.

[0030] Optionally, the writing unit is configured to:

[0031] First, directly write the initial data into the target data lake file, and then write the target metadata into the target data lake file.

[0032] Optionally, the second determination unit is configured to:

[0033] After directly writing the initial data into the target data lake file, obtain the target metadata of the target data lake file according to the storage location of the stored data in the target data lake file.

[0034] Optionally, the apparatus further includes:

[0035] A receiving unit, configured to receive a target structured query language statement before acquiring the initial data stored in a plurality of data lake files, where the target structured query language statement is used to specify the data lake files to be merged;

[0036] A third determination unit, configured to determine the plurality of data lake files to be merged according to the target structured query language statement.

[0037] Optionally, the target structured query language statement further includes a data lake file merging statement, where the data lake file merging statement is used to indicate merging of the plurality of data lake files to be merged.

[0038] Optionally, the target structured query language statement further includes a file size setting statement, where the file size setting statement is used to indicate that the size of the target data lake file is a target size, and the size of the generated target data lake file is the target size.

[0039] Optionally, the acquisition unit is configured to:

[0040] Traverse the multiple data lake files to be merged, and for each data lake file to be merged, perform the following operations:

[0041] According to the metadata of the data lake file to be merged, determine the offset and size of the row groups in the data lake file to be merged within the data lake file to be merged;

[0042] According to the metadata of the row groups, determine the offset and size of the column chunks in the data lake file to be merged within the data lake file to be merged;

[0043] Obtain the data in the column chunks as the initial data.

[0044] In a third aspect, an embodiment of the present application provides a file merging device, which includes a processor and a memory;

[0045] The processor is used to execute the instructions stored in the memory, so that the device executes the method described in any one of the above first aspects.

[0046] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions, where the instructions direct the device to execute the method described in any one of the above first aspects.

[0047] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when running on a computer, causes the computer to execute the method described in any one of the above first aspects.

[0048] Compared with the prior art, the embodiments of the present application have the following advantages:

[0049] An embodiment of the present application provides a file merging method, which includes: obtaining initial data stored in multiple data lake files to be merged, where the initial data is data that has been processed by one or more of serialization, compression, and encoding. After obtaining the initial data, the initial data can be directly used as the stored data in the target data lake file, and according to the storage location of the stored data in the target data lake file, the target metadata of the target data lake file can be obtained. After obtaining the stored data and the target metadata, the stored data and the target metadata can be written into the target data lake file, thereby obtaining the target data lake file merged from the multiple data lake files to be merged. It can be seen that by using the solution of the embodiment of the present application, when merging multiple data lake files to be merged into a target data lake file, it is not necessary to first process the data stored in the data lake files to be merged, but the original data can be directly used as the stored data in the target data lake file, thereby improving the efficiency of merging multiple data lake files to be merged into a target data lake file. Description of the Drawings

[0050] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0051] Figure 1 It is a schematic flowchart of the file merging method in the traditional technology;

[0052] Figure 2 It is a schematic structural diagram of the content stored in a data lake file provided by an embodiment of the present application;

[0053] Figure 3 It is a schematic diagram of storing an original data table as a data lake file provided by an embodiment of the present application;

[0054] Figure 4 It is a schematic flowchart of a file merging method provided by an embodiment of the present application;

[0055] Figure 5 It is a schematic diagram of the scenario of a file merging method provided by an embodiment of the present application;

[0056] Figure 6 It is a schematic structural diagram of a file merging device provided by an embodiment of the present application. Detailed implementation manners

[0057] To enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0058] The inventors of the present application have found through research that the efficiency of merging data lake files in the traditional technology is relatively low. Specifically, reference can be made to Figure 1 for understanding, Figure 1 which shows a schematic flowchart of the file merging method in the traditional technology.

[0059] As Figure 1As shown, when merging multiple data lake files to be merged (such as small Parquet files), the data in the multiple data lake files to be merged can first be decoded, decompressed, and deserialized to obtain the data to be merged. Then, when merging the data to be merged, the data to be merged needs to be serialized, compressed, and encoded to obtain the merged data lake file.

[0060] However, the processes of decoding, decompressing, deserializing, serializing, compressing, and encoding the data mentioned above are rather cumbersome, resulting in low efficiency of file merging. Therefore, there is an urgent need for a solution to solve the above problems and improve the efficiency of file merging.

[0061] In view of this, the embodiments of the present application provide a file merging method and apparatus.

[0062] Before introducing the file merging method and apparatus provided by the embodiments of the present application, the data lake file will be introduced first.

[0063] See Figure 2 , Figure 2 which is a schematic structural diagram of the content stored in a data lake file provided by the embodiments of the present application. As Figure 2 shown:

[0064] The data lake file includes three parts: a header, data, and a footer, where:

[0065] The content of the Header is very small, only 4 bytes, which is a fixed magic number used to indicate that this is a Parquet file.

[0066] The data is used to store data. When the original data is stored in the Parquet file, it will first be split into multiple row groups by rows. For example, in Figure 2 , 3 RowGroup0s may store the data of the 1st to 100th rows in the original data table.

[0067] Within each RowGroup, it will be further split into multiple column chunks by columns, corresponding to each column in the original data table. For example, in Figure 2 , ColumnChunk0 may store the data of the 1st column in RowGroup1.

[0068] The ColumnChunk can be further divided into multiple pages internally. Each page includes a page header and page data, and the Page Header is used to describe the encoding information of the page.

[0069] The Footer is mainly used to record metadata, which includes the storage locations of each RowGroup and each ColumnChunk in the data lake file. The storage locations can be reflected by information such as offset and size. Through the Footer, the location of the data in the file can be quickly located and read as needed.

[0070] It should be noted that Figure 2 it is shown only for the purpose of understanding the content structure of the data lake file method, and it does not constitute a limitation to the embodiments of the present application. A data lake file is not limited to including 3 RowGroups, a RowGroup is not limited to being split into 3 ColumnChunks, and the number of Pages included in a ColumnChunk is not limited to Figure 2 the 3 shown.

[0071] See Figure 3 , which is a schematic diagram of storing an original data table as a data lake file provided by an embodiment of the present application. Figure 3 It shows an original data table including 2 rows and 3 columns and a data lake file generated therefrom. For the content included in the data lake file, reference can be made to the relevant descriptions above, and no repeated description will be given here.

[0072] Next, various non-limiting embodiments of the present application will be described in detail with reference to the accompanying drawings.

[0073] Exemplary method

[0074] See Figure 4 , which is a schematic flowchart of a file merging method provided by an embodiment of the present application. In this embodiment, the method can be executed by a server, for example. As an example, the method can include the following steps: S101 - S104.

[0075] S101: Obtain the initial data stored in multiple data lake files to be merged, where the initial data is data that has been processed by one or more of serialization, compression, and encoding.

[0076] The initial data mentioned here refers to the data in the data lake files to be merged. In the embodiments of the present application, the initial data stored in the multiple data lake files to be merged can be directly read from the multiple data lake files.

[0077] In one example, before executing S101, the multiple data lake files to be merged can also be determined. In a specific example, the multiple data lake files to be merged are pre-specified. In yet another specific example, the multiple data lake files to be merged are determined by the server based on a target Structured Query Language (SQL) statement. Specifically, the server can receive a target SQL statement for specifying the data lake files to be merged. For example, the target SQL statement includes the conditions that the multiple data lake files need to meet. After receiving the target SQL statement, the server can perform a query operation according to the target SQL statement to obtain the multiple data lake files to be merged. The embodiments of the present application do not specifically limit the conditions that the data lake files to be merged need to meet. For example, the condition can be the date when the data lake file is generated, or the size of the data lake file, and so on.

[0078] Illustrative example: The target SQL statement can be: ALTER TABLE $tableName PARTITION(date='***'). This target SQL statement specifies the date of the multiple data lake files to be merged. In this case, the server can determine the data lake files with the date of "***" as the data lake files to be merged.

[0079] In one example, the target SQL statement can also include a data lake file merge statement, which is used to indicate the merging of the multiple data lake files to be merged. In this case, the target SQL statement can not only be used to specify the data lake files to be merged, but also be used to specify the merge operation to be performed on the data lake files to be merged. In this case, after receiving the target SQL statement, the server can, on the one hand, determine the multiple data lake files to be merged, and on the other hand, execute S101 - S104 to implement the merging of the multiple data lake files to be merged.

[0080] Illustrative example: The data lake file merge statement can be: COMPACT 'FastCompactionStrategy', where: the Fast Compaction Strategy indicates that the file merge method provided by the embodiments of the present application is used to merge the multiple data lake files to be merged.

[0081] In yet another example, the target SQL statement may further include a file size setting statement for setting the size of the target data lake file obtained after merging to a target size. In this case, when the server merges the multiple data lake files to be merged, the size of the generated target data lake file is the target size. For this case, the target SQL statement can be used not only to specify the data lake files to be merged, but also to specify the size of the target data lake file obtained by merging the data lake files to be merged.

[0082] In one example, the server may determine the storage locations of the data in the multiple data lake files to be merged in the data lake files to be merged based on the metadata of the multiple data lake files to be merged, and further read the initial data from the corresponding storage locations. In a specific example, when specifically implemented, S101 may traverse each of the multiple data lake files to be merged and perform the following steps A1 - A3 on each data lake file to be merged respectively.

[0083] A1: Determine the offset and size of the row groups in the data lake file to be merged in the data lake file to be merged according to the metadata of the data lake file to be merged.

[0084] In one example, the metadata of the data lake file to be merged may be read to determine the offset and size of each row group in the data lake file to be merged in the data lake file to be merged.

[0085] A2: Determine the offset and size of the column chunks in the data lake file to be merged in the data lake file to be merged according to the metadata of the row groups.

[0086] In one example, for each row group, the offset and size of the column chunks included in the row group in the data lake file to be merged may be determined based on the metadata of the row group, so as to obtain the offset and size of each column chunk in the data lake file to be merged in the data lake file to be merged.

[0087] A3: Obtain the data in the column chunks as the initial data.

[0088] After obtaining the offset and size of the column chunks in the data lake file to be merged, the storage location of the column chunks in the data lake file to be merged may be determined based on the offset and size. Correspondingly, the data in the column chunks may be read from the determined storage location as the initial data. Among them, the data in the column chunks may be binary fragments.

[0089] In one example, the server may first create a target data lake file and write a header (i.e., magic number) into the target data lake file. Further, S102 and subsequent steps are executed.

[0090] S102: Use the initial data directly as the stored data in the target data lake file.

[0091] S103: Obtain the target metadata of the target data lake file according to the storage location of the stored data in the target data lake file.

[0092] S104: Write the stored data and the target metadata into the target data lake file.

[0093] In the embodiments of the present application, in order to improve the efficiency of merging data lake files, after obtaining the initial data, in the embodiments of the present application, the initial data is no longer subjected to one or more processes such as decoding, decompression, deserialization, serialization, compression, and encoding, but the initial data is directly used as the stored data in the target data lake file, thereby reducing the processing of data and improving the file merging efficiency.

[0094] As can be seen from the previous description of the data lake file, the data lake file includes a header, data, and a footer, and the footer is used to store metadata. Therefore, in the embodiments of the present application, the target metadata of the target data lake file can also be obtained based on the storage location of the stored data in the target data lake file. In a specific example, the location of the stored data in the target data lake file may include the offsets and sizes of each row group in the aforementioned initial data in the target data lake file. After determining the stored data and the target metadata, the stored data and the target metadata can be written into the target data lake file, thereby obtaining the target data lake file merged from multiple data lake files to be merged.

[0095] In the embodiments of the present application, when specifically implemented, S103 may have multiple implementation manners.

[0096] In one example, the storage location of the stored data in the target data lake file can be pre-planned according to the size of the stored data, and then, based on the planned storage location, the target metadata is obtained. For this case, when specifically implemented, S104 may, for example, synchronously write the stored data and the target metadata into the target data lake file.

[0097] In yet another example, after executing S102, the initial data may be directly written into the target data lake file first, and then, according to the storage location of the initial data in the target data lake file, the target metadata of the target data lake file is obtained. For this case, after writing the initial data into the target data lake file, the storage location of the initial data in the target data lake file can be obtained, and there is no need to pre-plan the storage location of the initial data in the target data lake file. Correspondingly, for this case, after obtaining the target metadata of the target data lake file according to the storage location of the initial data in the target data lake file, the target metadata can be further written into the target data lake file, so as to obtain the target data lake file merged from multiple data lake files to be merged.

[0098] As can be seen from the above description, by using the solution of the embodiments of the present application, when merging multiple data lake files to be merged into a target data lake file, there is no need to perform one or more processes such as decoding, decompressing, deserializing, serializing, compressing, and encoding on the data stored in the data lake files to be merged. Instead, the original data can be directly used as the stored data in the target data lake file, thereby improving the efficiency of merging multiple data lake files to be merged into a target data lake file.

[0099] Next, a file merging method provided by the embodiments of the present application will be introduced in combination with specific examples.

[0100] See Figure 5 , which is a schematic diagram of a scenario of a file merging method provided by the embodiments of the present application. As Figure 5 shown, the Parquet file 1 to be merged and the Parquet file 2 to be merged are merged into a target Parquet file. The data 510 and 520 in the Parquet file 1 to be merged are directly copied into the target Parquet file as the stored data of the target Parquet file, and the data 530 in the Parquet file 2 to be merged is directly copied into the target Parquet file as the stored data of the target Parquet file.

[0101] Exemplary device

[0102] Based on the method provided in the above embodiments, the embodiments of the present application further provide a device, which will be introduced below with reference to the accompanying drawings.

[0103] See Figure 6 , which is a schematic structural diagram of a file merging device provided by the embodiments of the present application. The device 600 may specifically include, for example: an acquisition unit 601, a first determination unit 602, a second determination unit 603, and a writing unit 604.

[0104] An acquisition unit 601, configured to acquire initial data stored in a plurality of data lake files to be merged, where the initial data is data that has been processed by one or more of serialization, compression, and encoding;

[0105] A first determination unit 602, configured to directly use the initial data as the stored data in the target data lake file;

[0106] A second determination unit 603, configured to obtain target metadata of the target data lake file according to the storage location of the stored data in the target data lake file;

[0107] A writing unit 604, configured to write the stored data and the target metadata into the target data lake file.

[0108] Optionally, the writing unit 604 is configured to:

[0109] First, directly write the initial data into the target data lake file, and then write the target metadata into the target data lake file.

[0110] Optionally, the second determination unit 603 is configured to:

[0111] After directly writing the initial data into the target data lake file, obtain the target metadata of the target data lake file according to the storage location of the stored data in the target data lake file.

[0112] Optionally, the apparatus further includes:

[0113] A receiving unit, configured to receive a target structured query language statement before acquiring the initial data stored in a plurality of data lake files, where the target structured query language statement is used to specify the data lake files to be merged;

[0114] A third determination unit, configured to determine the plurality of data lake files to be merged according to the target structured query language statement.

[0115] Optionally, the target structured query language statement further includes a data lake file merging statement, where the data lake file merging statement is used to indicate merging of the plurality of data lake files to be merged.

[0116] Optionally, the target structured query language statement further includes a file size setting statement, where the file size setting statement is used to indicate that the size of the target data lake file is a target size, and the size of the generated target data lake file is the target size.

[0117] Optionally, the acquisition unit 601 is configured to:

[0118] Traverse the multiple data lake files to be merged, and for each data lake file to be merged, perform the following operations:

[0119] Determine the offset and size of the row groups in the data lake file to be merged in the data lake file to be merged according to the metadata of the data lake file to be merged;

[0120] Determine the offset and size of the column chunks in the data lake file to be merged in the data lake file to be merged according to the metadata of the row groups;

[0121] Obtain the data in the column chunks as the initial data.

[0122] Since the device 600 corresponds to the method provided in the above method embodiment, the specific implementation of each unit of the device 600 is based on the same concept as the above method embodiment. Therefore, for the specific implementation of each unit of the device 600, reference may be made to the description part of the above method embodiment, which will not be elaborated here.

[0123] An embodiment of the present application also provides a file merging device, which includes a processor and a memory;

[0124] The processor is configured to execute the instructions stored in the memory to cause the device to execute the file merging method provided in the above method embodiment.

[0125] An embodiment of the present application provides a computer-readable storage medium, including instructions, where the instructions direct a device to execute the file merging method provided in the above method embodiment.

[0126] An embodiment of the present application also provides a computer program product, which, when running on a computer, causes the computer to execute the file merging method provided in the above method embodiment.

[0127] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0128] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

[0129] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.

Claims

1. A file merging method, characterized in that, the method includes: Obtaining the initial data stored in multiple data lake files to be merged, where the initial data is data processed by one or more of serialization, compression, and encoding, and the data lake files to be merged are Parquet small files with a file size smaller than the size of a data block; Directly using the initial data as the stored data in the target data lake file; Obtaining the target metadata of the target data lake file according to the storage location of the stored data in the target data lake file; Writing the stored data and the target metadata into the target data lake file; The writing the stored data and the target metadata into the target data lake file includes: First, directly writing the initial data into the target data lake file, and then writing the target metadata into the target data lake file; The obtaining the target metadata of the target data lake file according to the storage location of the stored data in the target data lake file includes: After directly writing the initial data into the target data lake file, obtaining the target metadata of the target data lake file according to the storage location of the stored data in the target data lake file; The obtaining the initial data stored in multiple data lake files to be merged includes: Traversing the multiple data lake files to be merged, and for each data lake file to be merged, performing the following operations: Determining the offset and size of the row group in the data lake file to be merged according to the metadata of the data lake file to be merged; Determining the offset and size of the column block in the data lake file to be merged according to the metadata of the row group; Obtaining the data in the column block as the initial data.

2. The method according to claim 1, characterized in that, Before obtaining the initial data stored in multiple data lake files, the method further includes: Receiving a target Structured Query Language (SQL) statement, where the target SQL statement is used to specify the data lake files to be merged; Determining the multiple data lake files to be merged according to the target SQL statement.

3. The method according to claim 2, characterized in that, The target SQL statement further includes a data lake file merging statement, and the data lake file merging statement is used to indicate the merging of the multiple data lake files to be merged.

4. The method according to claim 2 or 3, characterized in that, The target SQL statement further includes a file size setting statement, and the file size setting statement is used to indicate that the size of the target data lake file is a target size, and the size of the generated target data lake file is the target size.

5. A file merging device, characterized in that, the device includes: An acquisition unit, configured to acquire initial data stored in a plurality of data lake files to be merged, where the initial data is data that has been processed by one or more of serialization, compression, and encoding, and the data lake files to be merged are Parquet small files with a file size smaller than a data block size; A first determination unit, configured to directly use the initial data as the stored data in the target data lake file; A second determination unit, configured to obtain target metadata of the target data lake file according to the storage location of the stored data in the target data lake file; A writing unit, configured to write the stored data and the target metadata into the target data lake file; The writing unit is specifically configured to: First, directly write the initial data into the target data lake file, and then write the target metadata into the target data lake file; The second determination unit is specifically configured to: After directly writing the initial data into the target data lake file, obtain target metadata of the target data lake file according to the storage location of the stored data in the target data lake file; The acquisition unit is specifically configured to: Traverse the plurality of data lake files to be merged, and for each data lake file to be merged, perform the following operations: Determine the offset and size of a row group in the data lake file to be merged according to the metadata of the data lake file to be merged; Determine the offset and size of a column block in the data lake file to be merged according to the metadata of the row group; Acquire the data in the column block as the initial data.

6. A file merging device, characterized in that the device includes a processor and a memory; the processor is configured to execute instructions stored in the memory, so that the device executes the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that it includes instructions that instruct a device to execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Data processing method and device

    CN114911809A

  • Data lake-based batch flow integrated data processing method, device and equipment

    CN115878642A