A data processing method and apparatus based on a data lake

The method integrates data lake files with the same primary key to construct a wide table in real-time, addressing the challenge of varying data lifecycles and eliminating the need for computation engines in data processing.

CN114968938BActive Publication Date: 2025-07-15BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210609515.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-07-15
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

How to build a large wide table in real time based on multiple data sources with the same primary key, especially how to deal with the life cycle setting and maintenance problems of frequently updated metric data and low-frequency updated dimensional data.

Method used

By receiving data read requests, the initial data lake file with the same primary key is integrated, the integrated data lake file is generated, and the target file is determined based on the data read request, avoiding the introduction of the computing engine and directly building a large-width table in the data read stage.

Benefits of technology

It realizes that without maintaining the data state and life cycle during the data reading stage, and directly builds large-wide tables based on the primary key integration of the data lake file, improving the efficiency and flexibility of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114968938B_ABST
    Figure CN114968938B_ABST
Patent Text Reader

Abstract

The present application discloses a data processing method based on a data lake, including: receiving a data reading request and reading at least one initial data lake file. Then, integrating at least one initial data lake file based on the primary key of at least one initial data lake file to obtain multiple integrated data lake files. When integrating at least one initial data lake file, multiple initial data lake files with the same primary key are integrated into one integrated data lake file. Then, based on the multiple integrated data lake files, a target file that meets the data reading request is obtained. Thus, it can be seen that in the data reading stage of this solution, based on the primary key of the data lake file, data lake files with the same key are integrated into one integrated data lake file. By adopting this method, there is no need to introduce a computing engine. Therefore, there is no need to maintain the state of data and set the life cycle of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular, to a data processing method and apparatus based on a data lake. Background Art

[0002] In some scenarios, it may be necessary to construct a large wide table in real time based on multiple data sources with the same primary key (key). The data sources mentioned here may include metric data, which can be data obtained based on summary calculations and is generally updated relatively frequently. The data sources can also include dimension data, which can also be understood as attribute data and has a low update frequency.

[0003] Currently, how to construct a large wide table in real time based on multiple data sources with the same primary key is a problem yet to be solved. Summary of the Invention

[0004] This application provides a data processing method and apparatus based on a data lake.

[0005] In a first aspect, an embodiment of this application provides a data processing method based on a data lake, and the method includes:

[0006] Receiving a data reading request and reading at least one initial data lake file;

[0007] Integrating the at least one initial data lake file based on the primary key key of the at least one initial data lake file to obtain multiple integrated data lake files, where: multiple initial data lake files with the same primary key are integrated into one integrated data lake file;

[0008] Based on the multiple integrated data lake files, obtaining a target file that meets the data reading request.

[0009] Optionally, the obtaining a target file that meets the data reading request based on the multiple integrated data lake files includes:

[0010] Obtaining multiple basic files;

[0011] Using the multiple integrated data lake files to update the multiple basic files to obtain at least one updated file corresponding to the basic file;

[0012] Determining a target file that meets the data reading request from the updated files.

[0013] Optionally, the multiple basic files include a first basic file, and using the multiple integrated data lake files to update the first basic file to obtain the updated file corresponding to the first basic file includes:

[0014] Obtain a second record from the multiple integrated data lake files that has the same primary key as the first record in the first base file;

[0015] Use the second record to update the first record to obtain an updated first record. The updated file corresponding to the first base file includes the updated first record.

[0016] Optionally, the base file is obtained in the following manner:

[0017] Read at least one historical data lake file;

[0018] Integrate the at least one historical data lake file to obtain multiple base files, where multiple historical data lake files with the same primary key are integrated into one base file.

[0019] Optionally, the at least one initial data lake file includes a first initial data lake file and a second initial data lake file. The first initial data lake file and the second initial data lake file are written in the following manner:

[0020] Receive a first data write request and a second data write request within a certain time period. The first data write request is used to request writing first stream data, and the second data write request is used to request writing second stream data. The first stream data and the second stream data have the same primary key;

[0021] Determine that the first stream data and the second stream data do not include the same attributes other than the primary key;

[0022] Write the first stream data into the first initial data lake file and write the second stream data into the second initial data lake file.

[0023] Optionally, the at least one initial data lake file includes a third initial data lake file. The third initial data lake file is written in the following manner:

[0024] Receive a third data write request and a fourth data write request within a certain time period. The third data write request is used to request writing third stream data, and the fourth data write request is used to request writing fourth stream data. The time of receiving the third data write request is earlier than the time of receiving the fourth data write request. The third stream data and the fourth stream data have the same primary key;

[0025] Determine that the third stream data and the fourth stream data include the same attributes other than the primary key;

[0026] Write the third stream of data into the third initial data lake file and reject writing the fourth stream of data.

[0027] In a second aspect, an embodiment of the present application provides a data processing device based on a data lake. The device includes:

[0028] A receiving unit, configured to receive a data reading request and read at least one initial data lake file;

[0029] An integrating unit, configured to integrate the at least one initial data lake file based on the primary key key of the at least one initial data lake file to obtain a plurality of integrated data lake files, where: a plurality of initial data lake files with the same primary key are integrated into one integrated data lake file;

[0030] A determining unit, configured to obtain a target file that meets the data reading request based on the plurality of integrated data lake files.

[0031] Optionally, the determining unit is configured to:

[0032] Obtain a plurality of base files;

[0033] Use the plurality of integrated data lake files to update the plurality of base files to obtain at least one updated file corresponding to the base files;

[0034] Determine a target file that meets the data reading request from the updated files.

[0035] Optionally, the plurality of base files include a first base file. Using the plurality of integrated data lake files to update the first base file to obtain the updated file corresponding to the first base file includes:

[0036] Obtain a second record with the same primary key as the first record in the first base file from the plurality of integrated data lake files;

[0037] Use the second record to update the first record to obtain an updated first record. The updated file corresponding to the first base file includes the updated first record.

[0038] Optionally, the base files are obtained by the following method:

[0039] Read at least one historical data lake file;

[0040] Integrate the at least one historical data lake file to obtain a plurality of base files, where: a plurality of historical data lake files with the same primary key are integrated into one base file.

[0041] Optionally, the at least one initial data lake file includes a first initial data lake file and a second initial data lake file, and the first initial data lake file and the second initial data lake file are written in the following manner:

[0042] Receive a first data write request and a second data write request within a certain time period. The first data write request is used to request writing of first stream data, and the second data write request is used to request writing of second stream data. The first stream data and the second stream data have the same primary key;

[0043] Determine that the first stream data and the second stream data do not include the same attributes other than the primary key;

[0044] Write the first stream data into the first initial data lake file, and write the second stream data into the second initial data lake file.

[0045] Optionally, the at least one initial data lake file includes a third initial data lake file, and the third initial data lake file is written in the following manner:

[0046] Receive a third data write request and a fourth data write request within a certain time period. The third data write request is used to request writing of third stream data, and the fourth data write request is used to request writing of fourth stream data. The time of receiving the third data write request is earlier than the time of receiving the fourth data write request. The third stream data and the fourth stream data have the same primary key;

[0047] Determine that the third stream data and the fourth stream data include the same attributes other than the primary key;

[0048] Write the third stream data into the third initial data lake file, and reject writing of the fourth stream data.

[0049] In a third aspect, an embodiment of the present application provides a device, and the device includes a processor and a memory;

[0050] The processor is configured to execute instructions stored in the memory to enable the device to execute the method according to any one of the above first aspects.

[0051] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions, and the instructions direct a device to execute the method according to any one of the above first aspects.

[0052] In a fifth aspect, an embodiment of the present application provides a computer program product, and when the computer program product runs on a computer, it causes the computer to execute the method according to any one of the above first aspects.

[0053] Compared with the prior art, the embodiments of the present application have the following advantages:

[0054] The embodiments of the present application provide a data processing method based on a data lake. In one example, a data reading request can be received, and at least one initial data lake file can be read. Then, based on the primary keys of the at least one initial data lake file, the at least one initial data lake file is integrated to obtain a plurality of integrated data lake files. Specifically, when integrating the at least one initial data lake file, multiple initial data lake files with the same primary key are integrated into one integrated data lake file. Then, based on the plurality of integrated data lake files, a target file that meets the data reading request is obtained. It can be seen that in the embodiments of the present application, in the data reading stage, based on the primary keys of the data lake files, the data lake files with the same key are integrated into one integrated data lake file. That is: in the data reading stage, a large wide table is constructed in real time based on multiple data sources with the same primary key. By adopting this method, there is no need to introduce a computing engine. Therefore, there is no need to maintain the state of the data and set the life cycle of the data. Description of the Drawings

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0056] Figure 1 It is a schematic flowchart of a data processing method based on a data lake provided by the embodiments of the present application;

[0057] Figure 2 It is a schematic flowchart of a method for determining a target file provided by the embodiments of the present application;

[0058] Figure 3 It is a schematic diagram of the data processing process during data reading provided by the embodiments of the present application;

[0059] Figure 4 It is a schematic structural diagram of a data processing device based on a data lake provided by the embodiments of the present application. Detailed Embodiments

[0060] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.

[0061] The inventors of this application have found through research that in traditional technologies, data sources include metric data and dimension data. The content involved in metric data is relatively large, and the differences in the life cycles between different data may be significant. Currently, a large wide table can be constructed in real time by a computing engine based on multiple data sources with the same primary key key. However, whether it is to construct a large wide table based on multiple metric data with the same primary key key or to construct a large wide table based on metric data and dimension data with the same primary key, setting and maintaining the life cycle of metric data are both major difficulties.

[0062] To solve the above problems, the embodiments of this application provide a data processing method and device based on a data lake.

[0063] The following will describe in detail various non-limiting embodiments of this application with reference to the accompanying drawings.

[0064] Exemplary method

[0065] See Figure 1 , which is a schematic flowchart of a data processing method based on a data lake provided by an embodiment of this application. The data processing method provided by the embodiment of this application can be executed by a server, for example.

[0066] In this embodiment, the method may include the following steps: S101 - S103, for example.

[0067] S101: Receive a data reading request and read at least one initial data lake file.

[0068] The data reading request is used to read data. In one example, the data reading request is a structured query language (SQL) statement.

[0069] In one example, the data lake file may be a file based on hudi, where hudi is a storage format of a data lake. The lake file may be stored on a disk in the form of a log file.

[0070] In the embodiments of the present application, the at least one initial data lake file may be all the data lake files written to the disk within a certain time period. Since the data lake does not support data screening during the data reading phase, the initial data lake files read in S101 are all the data lake files written to the disk within a certain time period.

[0071] In one example, the historical data lake files written to the disk before the certain time period may be integrated into corresponding base files.

[0072] S102: Integrate the at least one data lake file based on the primary key of the at least one initial data lake file to obtain a plurality of integrated data lake files, where: a plurality of initial data lake files with the same primary key are integrated into one integrated data lake file.

[0073] After reading the at least one data lake file, the primary keys of the at least one data lake file can be compared. When two or several (3 or more) data lake files all include the same primary key, then these two or several data lake files can be integrated to obtain one integrated data lake file. When integrating a plurality of data lake files, the integration can be performed separately for the plurality of primary keys in the plurality of data lake files. Taking key1 as an example, the records corresponding to key1 in the plurality of data lake files can be integrated. Among them, integrating the records corresponding to key1 in the plurality of data lake files may be to splice the records corresponding to key1 in the plurality of data lake files. For example: log file1 includes (key1, b0_new, c0_new), log file 2 includes (key1, d0_new), and after splicing, it gets (key1, b0_new, c0_new, d0_new).

[0074] In one example, after obtaining a plurality of integrated data lake files, the plurality of integrated data lake files can be stored in a map.

[0075] When integrating the initial data lake files, the integrated data lake files obtained may include content that is not included in the initial data lake files. For this situation, the value of this content can be the default value. For specific details, please refer to the description part for Figure 3 below, and no separate examples are given here.

[0076] S103: Obtain a target file that meets the data reading request based on the plurality of integrated data lake files.

[0077] After obtaining the multiple integrated data lake files, a target file that meets the data reading request can be obtained based on the multiple integrated data lake files.

[0078] In one example, if the at least one initial data lake file is a data lake file initially written to the disk, that is, there is no base file, then in specific implementation of S103, a target file that meets the data reading request can be filtered out from the multiple integrated data lake files. Among them, the filtering condition for filtering the integrated data lake files can be determined based on the data reading request. For example, if the data reading request is used to request the data corresponding to key1, then the key1 can be used as the retrieval condition to retrieve the target file from the multiple integrated data lake files. Another example is that if the data reading request is used to request the data corresponding to attribute A, then the attribute A can be used as the retrieval condition to retrieve the target file from the multiple integrated data lake files.

[0079] It should be noted that in one example, the "attribute" mentioned in the embodiments of the present application can correspond to the "column" in the table.

[0080] In yet another example, if there is a base file, then in specific implementation of S103, it can be implemented through Figure 2 S201 - S203 shown. Figure 2 It is a schematic flowchart of a method for determining a target file provided by an embodiment of the present application.

[0081] S201: Obtain multiple base files.

[0082] Similar to the at least one initial data lake file, the base file can also be stored on the disk. Therefore, the multiple base files can be read from the disk.

[0083] As described above, it can be known that the base file is obtained by integrating historical data lake files. That is: at least one historical data lake file can be read; the at least one historical data lake file is integrated to obtain multiple base files, where: multiple historical data lake files with the same main key are integrated into one base file. Regarding the specific implementation manner of obtaining the base file by integrating historical data lake files, it is the same as the manner of obtaining the integrated data lake file by integrating initial data lake files. The relevant content can refer to the relevant description part of S102 above and will not be repeated here.

[0084] In one example, the historical data lake files on the disk can be integrated at a certain period to obtain the corresponding base file. After obtaining the base file, the historical data lake files used to integrate the base file can be deleted from the disk to save disk space.

[0085] S202: Update the multiple basic files by using the multiple integrated data lake files to obtain update files corresponding to at least one of the basic files.

[0086] In one example, for each integrated data lake file, the corresponding basic file can be searched by using the primary key of the integrated data lake file, and then the corresponding basic file is updated by using the integrated data lake file.

[0087] In another example, for each basic file, the corresponding integrated data lake file can be searched by using the primary key of the basic file, and then the corresponding basic file is updated by using the found integrated data lake file.

[0088] The methods for updating each of the at least one basic file are similar. Next, taking the first basic file among the at least one basic file as an example, the implementation method for obtaining the update file corresponding to the first basic file is introduced. In one example, the update file corresponding to the first basic file can be obtained through the following steps A1 - A2.

[0089] A1: Obtain a second record having the same primary key as the first record in the first basic file from the multiple integrated data lake files.

[0090] In the specific implementation of step A1, for example, the integrated data lake file having the primary key can be searched by using the primary key in the first basic file, and the record corresponding to the primary key in the integrated data lake file is extracted. The primary key in the first basic file mentioned here is the aforementioned first record, and the record corresponding to the primary key in the integrated data lake file mentioned here is the aforementioned second record.

[0091] Illustrative example: The first basic file includes key1. Using key1 as the index, the integrated data lake file including key1 is searched, and the record corresponding to key1 in the integrated data lake file (i.e., the second record) is extracted. For example, the record corresponding to key1 is: (key1, b0_new, c0_new, d0_new).

[0092] A2: Update the first record by using the second record to obtain the updated first record. The update file corresponding to the first basic file includes the updated first record.

[0093] In the embodiments of the present application, compared with the first record, the second record has the same main key. In addition to the same main key, there may be the same attributes, or there may be no same attributes, or some attributes are the same and some attributes are different.

[0094] When using the second record to update the first record, the specific implementation may include any one or more of the following:

[0095] For the attributes that are the same in the second record and the first record, use the values in the second record to replace the values of the corresponding attributes in the first record. For example, if the first record is: (key1, b0, c0, d0, e0) and the second record is: (key1, b0_new, c0_new, d0_new), then replace (b0, c0, d0) in the first record with (b0_new, c0_new, d0_new).

[0096] For the attributes included in the second record but not included in the first record, add the attributes to the first record.

[0097] For the attributes included in the first record but not included in the second record, keep the attributes in the first record.

[0098] In addition, after updating the base file, the updated file may include content that is not included in either the base file or the integrated data lake file. For such a situation, the value of this content can be the default value. For specific details, refer to the description part for Figure 3 here, and no separate examples are given here.

[0099] S203: Determine the target file that meets the data reading request from the updated file.

[0100] When S203 is specifically implemented, the target file that meets the data reading request can be screened from the multiple updated files. Regarding the screening conditions for screening the updated files, refer to the description part of the screening conditions in S103, and no repeated description is given here.

[0101] As described above, the foregoing at least one initial data lake file is pre-written to the disk. Next, the implementation method of writing the initial data lake file to the disk will be introduced.

[0102] In a possible implementation, for concurrent data writing tasks, if the data requested to be written by the concurrent data writing tasks includes the same primary key, but does not include other identical attributes, then for the concurrent data writing tasks, the data can be written into the Hudi table in an upsert manner and stored on disk in the form of a log file. Among them, the concurrent data writing tasks can be the data writing tasks received within a certain time period.

[0103] In other words, if the first data writing request and the second data writing request are received within a certain time period, the first data writing request is used to request writing the first stream of data, the second data writing request is used to request writing the second stream of data, and the first stream of data and the second stream of data have the same primary key; then it can be further determined whether the first stream of data and the second stream of data do not include the same attributes other than the primary key; if it is determined that the first stream of data and the second stream of data do not include the same attributes other than the primary key; then write the first stream of data into the first initial data lake file, and write the second stream of data into the second initial data lake file. For example, the first stream of data and the second stream of data include the same primary key (it can be that all primary keys are the same, or some primary keys are the same), the first stream of data includes attributes (B, C), and the second stream of data includes attribute (D), then the first stream of data and the second stream of data can be written separately.

[0104] In another possible implementation, for concurrent data writing tasks, if the data requested to be written by the concurrent data writing tasks includes the same primary key and also includes other identical attributes. For this situation, in order to avoid data conflicts, for the concurrent data writing tasks, the successfully written task can be determined according to the time when the task is initiated. In one example, the data requested to be written by the data writing task with an earlier task initiation time can be written.

[0105] In other words, if a third data write request and a fourth data write request are received within a certain time period, the third data write request is used to request writing third stream data, the fourth data write request is used to request writing fourth stream data, the time when the third data write request is received is earlier than the time when the fourth data write request is received, and the third stream data and the fourth stream data have the same primary key; then it can be further determined whether the third stream data and the fourth stream data include the same attributes other than the primary key; if it is determined that the third stream data and the fourth stream data include the same attributes other than the primary key; then write the third stream data into the third initial data lake file and reject writing the fourth stream data. For example, if the third stream data and the fourth stream data have the same primary key, the third stream data includes attributes (B, C), and the fourth stream data includes attributes (B, D), then the third stream data can be written and writing the fourth stream data can be rejected.

[0106] As can be seen from the above description, in the embodiment of the present application, in the data reading stage, based on the primary key of the data lake file, the data lake files with the same key are integrated into an integrated data lake file. That is: in the data reading stage, a large wide table is constructed in real time based on multiple data sources with the same primary key. In this way, there is no need to introduce a computing engine, and therefore, there is no need to maintain the state of the data and set the life cycle of the data.

[0107] To facilitate understanding of the solution of the embodiment of the present application, next, in combination with Figure 3 , the data processing method provided by the embodiment of the present application will be introduced. Figure 3 It is a schematic diagram of the data processing process during the data reading process provided by the embodiment of the present application.

[0108] As Figure 3 shown, log file1 and log file2 include the same primary keys (key1 and key2), log file1 and log file2 are integrated to obtain an integrated file 301. Then, the base file is updated using the integrated file 301 to obtain an updated file 302.

[0109] When integrating log file1 and log file2, different columns are concatenated, and for the extra values in the concatenated table (as shown in 301), for example, column B and column C of key3, default values can be filled.

[0110] When using the integrated file 301 to update the base file:

[0111] The overlapping part between the two ( Figure 3(the shaded part in the stitched table shown), and its value is the value in the integrated file 301.

[0112] For the content that exists in the base file but does not exist in the integrated file 301, keep its value unchanged as the original value in the base file. For example, e0 and e1 in the update file 302.

[0113] For the content that does not exist in the base file but exists in the integrated file 301, keep its value unchanged as the original value in the integrated file 301. For example, d2_new in the integrated file 301.

[0114] For the extra values in the update file 302, for example, column E of key3 in the update file 302, default values can be filled.

[0115] It should be noted that Figure 3 is shown for ease of understanding. In Figure 3 , for ease of understanding, the tables shown are all in the "view" form of the storage file.

[0116] Exemplary device

[0117] Based on the method provided in the above embodiments, the embodiments of the present application also provide a device. The device is introduced below with reference to the accompanying drawings.

[0118] See Figure 4 , Figure 4 which is a schematic structural diagram of a data processing device based on a data lake provided by an embodiment of the present application. Figure 4 The device 400 shown may specifically include, for example: a receiving unit 401, an integrating unit 402, and a determining unit 403.

[0119] The receiving unit 401 is configured to receive a data reading request and read at least one initial data lake file;

[0120] The integrating unit 402 is configured to integrate the at least one initial data lake file based on the primary key key of the at least one initial data lake file to obtain a plurality of integrated data lake files, where: a plurality of initial data lake files with the same primary key are integrated into one integrated data lake file;

[0121] The determining unit 403 is configured to obtain a target file that meets the data reading request based on the plurality of integrated data lake files.

[0122] Optionally, the determining unit 403 is configured to:

[0123] Obtain a plurality of base files;

[0124] Update the multiple basic files by using the multiple integrated data lake files to obtain at least one updated file corresponding to each of the basic files;

[0125] Determine a target file that meets the data reading request from the updated files.

[0126] Optionally, the multiple basic files include a first basic file. Updating the first basic file by using the multiple integrated data lake files to obtain an updated file corresponding to the first basic file includes:

[0127] Obtain a second record having the same primary key as a first record in the first basic file from the multiple integrated data lake files;

[0128] Update the first record by using the second record to obtain an updated first record. The updated file corresponding to the first basic file includes the updated first record.

[0129] Optionally, the basic files are obtained by the following method:

[0130] Read at least one historical data lake file;

[0131] Integrate the at least one historical data lake file to obtain multiple basic files, where: multiple historical data lake files having the same primary key are integrated into one basic file.

[0132] Optionally, the at least one initial data lake file includes a first initial data lake file and a second initial data lake file. The first initial data lake file and the second initial data lake file are written by the following method:

[0133] Receive a first data writing request and a second data writing request within a certain time period. The first data writing request is used to request writing of first stream data, and the second data writing request is used to request writing of second stream data. The first stream data and the second stream data have the same primary key;

[0134] Determine that the first stream data and the second stream data do not include the same attributes except the primary key;

[0135] Write the first stream data into the first initial data lake file and write the second stream data into the second initial data lake file.

[0136] Optionally, the at least one initial data lake file includes a third initial data lake file. The third initial data lake file is written by the following method:

[0137] Receive a third data write request and a fourth data write request within a certain time period. The third data write request is used to request writing third stream data, and the fourth data write request is used to request writing fourth stream data. The time when the third data write request is received is earlier than the time when the fourth data write request is received. The third stream data and the fourth stream data have the same primary key.

[0138] Determine that the third stream data and the fourth stream data include the same attributes except the primary key.

[0139] Write the third stream data into the third initial data lake file and reject writing the fourth stream data.

[0140] Since the device 400 is the device corresponding to the method provided in the above method embodiment, the specific implementation of each unit of the device 400 is the same concept as that of the above method embodiment. Therefore, for the specific implementation of each unit of the device 400, reference can be made to the description part of the above method embodiment, which will not be elaborated here.

[0141] An embodiment of the present application further provides a device, which includes a processor and a memory.

[0142] The processor is used to execute the instructions stored in the memory so that the device executes the data processing method based on the data lake provided in the above method embodiment.

[0143] An embodiment of the present application provides a computer-readable storage medium, including instructions, and the instructions direct the device to execute the data processing method based on the data lake provided in the above method embodiment.

[0144] An embodiment of the present application further provides a computer program product, which, when running on a computer, causes the computer to execute the data processing method based on the data lake provided in the above method embodiment.

[0145] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other implementation schemes of the present application. The present application is intended to cover any variations, uses, or adaptive changes of the present application, and these variations, uses, or adaptive changes follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0146] It should be understood that the present application is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

[0147] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A data processing method based on a data lake, characterized in that, The method includes: Receiving a data reading request and reading at least one initial data lake file; Integrating the at least one initial data lake file based on the primary key (key) of the at least one initial data lake file to obtain a plurality of integrated data lake files, where: a plurality of initial data lake files with the same primary key are integrated into one integrated data lake file; Obtaining a plurality of base files; Using the plurality of integrated data lake files to update the plurality of base files to obtain at least one updated file corresponding to the base file; Determining a target file that meets the data reading request from the updated file; The base file is obtained by the following method: Reading at least one historical data lake file; Integrating the at least one historical data lake file to obtain a plurality of base files, where: a plurality of historical data lake files with the same primary key are integrated into one base file.

2. The method according to claim 1, wherein The plurality of base files include a first base file. Using the plurality of integrated data lake files to update the first base file to obtain the updated file corresponding to the first base file includes: Obtaining, from the plurality of integrated data lake files, a second record that has the same primary key as the first record in the first base file; Using the second record to update the first record to obtain an updated first record, and the updated file corresponding to the first base file includes the updated first record.

3. The method according to claim 1, characterized in that, The at least one initial data lake file includes a first initial data lake file and a second initial data lake file. The first initial data lake file and the second initial data lake file are written in the following manner: Receiving a first data writing request and a second data writing request within a certain time period. The first data writing request is used to request writing of first stream data, and the second data writing request is used to request writing of second stream data. The first stream data and the second stream data have the same primary key; Determining that the first stream data and the second stream data do not include the same attributes other than the primary key; Writing the first stream data into the first initial data lake file and writing the second stream data into the second initial data lake file.

4. The method according to claim 1, characterized in that, The at least one initial data lake file includes a third initial data lake file. The third initial data lake file is written in the following manner: Receiving a third data writing request and a fourth data writing request within a certain time period. The third data writing request is used to request writing of third stream data, and the fourth data writing request is used to request writing of fourth stream data. The time of receiving the third data writing request is earlier than the time of receiving the fourth data writing request. The third stream data and the fourth stream data have the same primary key; Determining that the third stream data and the fourth stream data include the same attributes other than the primary key; Writing the third stream data into the third initial data lake file and rejecting writing of the fourth stream data.

5. A data processing device based on a data lake, characterized in that, The apparatus includes: A receiving unit, configured to receive a data reading request and read at least one initial data lake file; An integration unit, configured to integrate the at least one initial data lake file based on the primary key of the at least one initial data lake file to obtain a plurality of integrated data lake files, where: multiple initial data lake files with the same primary key are integrated into one integrated data lake file; A determination unit, configured to obtain a plurality of base files; use the plurality of integrated data lake files to update the plurality of base files to obtain at least one updated file corresponding to the base files; and determine a target file that meets the data reading request from the updated files; The base files are obtained by the following method: Read at least one historical data lake file; Integrate the at least one historical data lake file to obtain a plurality of base files, where: multiple historical data lake files with the same primary key are integrated into one base file.

6. A device, characterized in that, The device includes a processor and a memory; The processor is configured to execute instructions stored in the memory, so that the device executes the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, Includes instructions that direct the device to execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Heterogeneous data source integration method and device based on data lake

    CN111966750A

  • Data processing method and device, storage medium and electronic equipment

    CN114528127A