A batch-stream integrated data processing method, device and equipment based on a data lake

By implementing the parallel operation of stream processing processes and batch merging processes in Hudi architecture, the data backlog and delay of stream processing processes is solved, and the throughput and efficiency of data processing is improved.

CN115878642BActive Publication Date: 2025-06-27BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211581037.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2025-06-27
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

In Hudi architecture, the stream processing process and the batch merging process are run in serial, resulting in a data backlog of the stream processing process, affecting the throughput, and the batch merging process needs to wait for the stream processing process to end, resulting in unnecessary delays.

Method used

Get the full file by storing data in a log file in the stream processing process and when the batch merge process is triggered, the log file is merged with the original full file. In the merge process, the stream processing process can still receive and process data, realizing the parallel operation of the stream processing process and the merge process.

Benefits of technology

Reduces data backlog of stream processing processes, improves throughput of stream processing processes, and reduces the delay of the entire batch-stream integrated data processing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115878642B_ABST
    Figure CN115878642B_ABST
Patent Text Reader

Abstract

The present application discloses a batch-stream integrated data processing method, apparatus and device based on a data lake, including: storing a first data set in a first log file by using a stream processing process. When a task of a batch processing merging process is triggered, the first log file is merged with an original full-volume file to obtain a first full-volume file. When in the process of merging the first log file with the original full-volume file, the stream processing process can also be used to store a second data set in a second log file. After the first log file and the original full-volume file are merged to obtain a first full-volume file, the second log file can also be merged with the first full-volume file to obtain a second full-volume file. When the data obtained by stream processing is merged with the full-volume file, the stream processing process can still receive data for processing, thereby reducing the data backlog of the stream processing process, improving the throughput of the stream processing process, and reducing the latency of the entire batch-stream integrated data processing process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly relates to a batch-stream integrated data processing method, apparatus, and device based on a data lake. Background Art

[0002] With the wide application of database technology, enterprise information systems have generated a large amount of business data. How to extract useful information for enterprise decision-making analysis from the massive business data has become an important problem faced by enterprises. A data lake is an evolving and scalable data processing architecture for big data storage, processing, and analysis, supporting the storage of multiple data structures, such as unstructured data and multi-structured data. It can quickly obtain the original data from different data sources and then dynamically prepare the data when users access it. By interacting and integrating with various external heterogeneous data sources, it supports various enterprise-level applications. Online Analytical Processing (OLAP), as a data processing technology based on a data lake, enables analysts to quickly, consistently, and interactively observe information from all aspects to achieve the purpose of in-depth understanding of the data. In the OLAP scenario, the batch-stream integrated solution, as an industry trend, is widely used in the business side and the engine side.

[0003] Currently, the relatively common batch-stream integrated architecture is Hudi, which adopts an incremental data processing method, that is, storing the received data to be processed in an incremental log, and then applying the incremental log to the full-scale file through a batch processing merge process (Compaction) to generate new full-scale data.

[0004] In the data processing of the Hudi architecture, the stream processing process and the batch processing merge process run in a serial manner, that is, the batch and stream processing block each other. The stream processing process must wait for the batch processing merge process to end before processing the next batch of data, resulting in a large backlog of data to be processed in the stream processing process, affecting the throughput of the stream processing. The batch processing merge process also needs to wait for the last group of data being processed by the stream processing process to end, and then merge with the full-scale data to generate new full-scale data, thus causing unnecessary delays. Summary of the Invention

[0005] In view of this, this application provides a batch-stream integrated data processing method, apparatus, and device based on a data lake to improve the throughput of the stream processing process and reduce the latency of batch-stream data processing.

[0006] To achieve the above object, the technical solutions provided in this application are as follows:

[0007] In the first aspect of this application, a batch-stream integrated data processing method based on a data lake is provided. The method includes:

[0008] Store the first data set in a first log file using a stream processing process;

[0009] In response to a task that triggers a batch processing and merging process, merge the first log file with an original full-volume file to obtain a first full-volume file, where the first log file has an association relationship with the original full-volume file;

[0010] When in the process of merging the first log file with the original full-volume file, store a second data set in a second log file using the stream processing process.

[0011] In a second aspect of the present application, there is provided a batch-stream integrated data processing device based on a data lake, and the device includes:

[0012] A first storage unit for storing a first data set in a first log file using a stream processing process;

[0013] A merging unit for, in response to a task that triggers a batch processing and merging process, merging the first log file with an original full-volume file to obtain a first full-volume file, where the first log file has an association relationship with the original full-volume file;

[0014] A second storage unit for storing a second data set in a second log file using the stream processing process when in the process of merging the first log file with the original full-volume file.

[0015] In a third aspect of the present application, there is provided an electronic device, and the device includes: a processor and a memory;

[0016] The memory is used for storing instructions or computer programs;

[0017] The processor is used for executing the instructions or computer programs in the memory so that the electronic device executes the method described in the first aspect above.

[0018] In a fourth aspect of the present application, there is provided a computer-readable storage medium, and instructions are stored in the computer-readable storage medium, and when the instructions run on a device, the device is caused to execute the method described in the first aspect above.

[0019] In a fifth aspect of the present application, there is provided a computer program product, and the computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the method described in the first aspect is implemented.

[0020] Thus, the present application has the following beneficial effects:

[0021] In the above implementation manner of the present application, when processing data using a batch-stream integrated architecture, first, a stream processing process is used to store the first data set in the first log file. When a task of the batch processing merging process is triggered, the first log file is merged with the original full-volume file to obtain the first full-volume file. Among them, the first log file and the original full-volume file have an associated relationship. When in the process of merging the first log file and the original full-volume file, the stream processing process can also be used to store the second data set in the second log file. After the first log file and the original full-volume file are merged to obtain the first full-volume file, the second log file can also be merged with the first full-volume file to obtain the second full-volume file. Through the batch-stream integrated data processing method based on the data lake provided by the present application, when the data obtained by stream processing is merged with the full-volume file, the stream processing process can still receive data for processing, that is, the stream processing process and the merging process can be carried out in parallel, thereby reducing the data backlog of the stream processing process, improving the throughput of the stream processing process, and reducing the latency of the entire batch-stream integrated data processing process. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 Schematic diagram of a batch-stream integrated architecture provided by an embodiment of the present application;

[0024] Figure 2 Flowchart of a batch-stream integrated data processing method based on a data lake provided by an embodiment of the present application;

[0025] Figure 3 Schematic diagram of a batch-stream integrated data processing device provided by an embodiment of the present application;

[0026] Figure 4 Schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.

[0028] To facilitate the understanding of the technical method provided in the embodiments of this application, the following first introduces the technical background involved in this application.

[0029] A data lake is an evolving and scalable data processing architecture for big data storage, processing, and analysis, supporting the storage of multiple data structures, such as unstructured data and multi-structured data. It can quickly obtain the original data from different data sources and then dynamically prepare the data when accessed by users. Through the interaction and integration with various external heterogeneous data sources, it supports various enterprise-level applications. Online Analytical Processing (OLAP), as a data processing technology based on the data lake, enables analysts to quickly, consistently, and interactively observe information from all aspects to achieve a deep understanding of the data. In the OLAP scenario, the batch-stream integrated solution, as an industry trend, is widely applied to the business side and the engine side.

[0030] Currently, the relatively common batch-stream integrated architecture is Hudi, which adopts an incremental data processing method, that is, the received data to be processed is stored in the incremental log, and then the incremental log is applied to the full-scale file through the batch processing merge process (Compaction) to generate new full-scale data. See Figure 1 , Figure 1 which is a schematic diagram of a batch-stream integrated architecture provided in the embodiments of this application.

[0031] After the stream processing process obtains the data, it writes the data into Log File 1. When the batch processing merge process is triggered, Log File 1 can be merged with Full-scale File 1 to obtain a new full-scale file. Among them, Log File 1 and Full-scale File 1 have an associated relationship.

[0032] In the batch-stream integrated data processing based on the Hudi architecture, the stream processing process and the batch processing merge process run in a serial manner, that is, the batch and stream processing block each other. The stream processing process must wait for the batch processing merge process to end before processing the next batch of data, resulting in a large backlog of data to be processed in the stream processing process, affecting the throughput of the stream processing. The batch processing merge process also needs to wait for the last set of data being processed by the stream processing process to end, and then merge with the full-scale data to generate new full-scale data, thus causing unnecessary delays.

[0033] Based on this, an embodiment of the present application provides a batch-stream integrated data processing method based on a data lake to improve the throughput of the stream processing process and reduce the latency of batch-stream data processing. Specifically, when processing the acquired data using a batch-stream integrated architecture, first, the stream processing process stores the first data set in the first log file. When the task of the batch processing merge process is triggered, the first log file is merged with the original full-volume file to obtain the first full-volume file. Among them, the first log file and the original full-volume file have an association relationship. When the first log file is merged with the original full-volume file, the stream processing process can continue to acquire new data and store the acquired second data set in the second log file. After the first log file and the original full-volume file are merged to obtain the first full-volume file, the second log file can also be merged with the first full-volume file to obtain the second full-volume file. Through the batch-stream integrated data processing method based on the data lake provided by the present application, when the data obtained by stream processing is merged with the full-volume file, the stream processing process can still receive data for processing. That is, the stream processing process and the merge process can be performed in parallel, thereby reducing the data backlog of the stream processing process, improving the throughput of the stream processing process, and reducing the latency of the entire batch-stream integrated data processing process.

[0034] To facilitate understanding of the technical solution provided by the embodiment of the present application, the following will be specifically introduced in conjunction with the accompanying drawings.

[0035] See Figure 2 , Figure 2 which is a flowchart of a batch-stream integrated data processing method provided by an embodiment of the present application.

[0036] This method can be executed by a data processing system that implements batch-stream integration. This method may include the following steps:

[0037] S201: Use the stream processing process to store the first data set in the first log file.

[0038] When the data processing system processes data, it can be completed by processes that implement different functions, and each process processes the data differently. Specifically, the data processing system can obtain data from the data source in near real-time through the stream processing process. For example, it can achieve data acquisition at the minute level, and then store the acquired first data set in the first log file. Optionally, the stream processing process can also process the acquired data, including data association, data aggregation, etc., to obtain a processed data set. That is, when the stream processing process obtains the initial data set from the data source, it can process the initial data set to obtain the first data set.

[0039] When storing the first data set in the first log file, the data processing system can also store the first metadata information corresponding to the first data set and the first log file. For example, it can be stored in the metadata storage unit of the data processing system. Among them, metadata is data that describes data, mainly information describing data attributes, and is used to support functions such as indicating storage locations, historical data, resource searches, file records, etc. In this embodiment, the first metadata information may include the storage location of the first data set, such as the file name of the first log file, the size and data type of the first data set stored in the first log file. When a user needs to search for target data in the data processing system, they can first determine the metadata information corresponding to the target data in the metadata storage unit, and then obtain the target data from the corresponding file according to the metadata information.

[0040] S202: In response to a task that triggers the batch processing and merging process, merge the first log file with the original full-volume file to obtain the first full-volume file.

[0041] After the batch processing and merging process is triggered, the data processing system can, through the batch processing and merging process, merge the first log file with the original full-volume file to obtain the merged first full-volume file. Among them, the first log file and the original full-volume file have an association relationship. For example, the data in the first log file and the data in the original full-volume file have the same primary key. Therefore, the batch processing and merging process of the data processing system can determine the corresponding initial full-volume file according to the primary key information of the first log file, and then merge the data in the first log file with the data in the initial full-volume file.

[0042] Taking an application scenario as an example, if the total monthly salary of user A on November 10th is stored as 1000 in the initial full-volume file, and after the first data set is written into the first log file through the stream processing process, the data in the first log file includes the total monthly salary of user A on November 12th as 1300. Then, after merging the first log file with the initial full-volume file, the data content in the first full-volume file is the total monthly salary of user A on November 12th as 1300.

[0043] Optionally, in the embodiments of the present application, it is possible to determine whether the condition for triggering the batch processing and merging process is met by calculating the number of times the data set is stored in the log file. That is, after the data set is stored in the log file through the stream processing process, the count can be incremented by 1. When the cumulative number of times the data set is stored in the log file by the stream processing process reaches a preset number, the batch processing and merging process can be triggered to merge the log file with the full volume file. Before the batch processing and merging process is triggered, since the stream processing process stores different data sets in their respective corresponding log files multiple times, optionally, when the batch processing and merging process is triggered, multiple merging processes can be executed simultaneously. That is, for each log file storing the data set, the log file can be merged with its associated full volume file to obtain the merged full volume file. For example, after the first data set is stored in the first log file using the stream processing process, when the number of times the data set is stored in the log file by the stream processing process reaches the preset number, the batch processing and merging process is triggered. The batch processing and merging process can merge multiple log files with their respective corresponding full volume files. Taking the first log file as an example, the first log file is one of the multiple log files processed by the same batch processing and merging process. The initial full volume file associated with the first log file is determined, and the data of the first log file is merged with the data of the initial full volume file to obtain the first full volume file.

[0044] According to the above embodiments, after the first data set is written into the first log file using the stream processing process, the data processing system can also send the first metadata information to the metadata storage unit for storage. Therefore, the number of times the data set is stored in the log file by the stream processing process can also be determined by the number of times the metadata information is sent to the metadata storage unit. Optionally, before the batch processing and merging process is triggered, the number of times the metadata information is sent can be recorded through the first metadata information. When the cumulative number of times the metadata information is sent to the metadata storage unit reaches the preset number, the data processing system triggers the batch processing and merging process and clears the recorded number of times the metadata information is sent to enter the next new count statistics.

[0045] Optionally, after obtaining the first full volume file, since the data in the original full volume file is already invalid and the latest data needs to be queried through the first full volume file later, the data processing system can also store the second metadata information corresponding to the first log file and the first full volume file, including the association relationship between the first log file and the first full volume file, the file name, etc., to indicate that when the user queries the data, the data associated with the first log file can be queried in the first full volume file. For example, the second metadata information can be stored in the metadata storage unit.

[0046] S203: When in the process of merging the first log file with the original full - volume file, use the stream - processing process to store the second data set in the second log file.

[0047] When using the batch - processing merge process to merge the first log file with the original full - volume file, at this time, the data - processing system can still obtain the data set in the data source through the stream - processing process, and store the processed data set in the log file after processing. That is to say, the batch - processing merge process and the stream - processing process can be parallel processes. Since the stream - processing process can continuously obtain data from the data source, when using the batch - processing process to merge the full - volume file, the stream - processing process can store the obtained second data set in the second log file at this time, thereby reducing the data backlog in the stream - processing process and improving the throughput of the stream - processing process, that is, the amount of data processed per unit time.

[0048] Similarly, it can be known that the data - processing system can store the third metadata information corresponding to the second data set and the second log file. Among them, the third metadata information can include the storage location of the second data set, such as the file name of the second log file, the size and data type of the second data set stored in the second log file, etc.

[0049] Since the data - processing system is using the batch - processing merge process to merge the log file with the corresponding full - volume file at this time, it is possible to re - count the number of times the stream - processing process stores the data set in the log file in order to determine the condition for triggering the next batch - processing merge process.

[0050] Through the above - mentioned batch - stream integrated data - processing method based on the data lake, when the data obtained by stream - processing is merged with the full - volume file, the stream - processing process can still receive data for processing. That is to say, the stream - processing process and the merge process can be parallel. Since the stream - processing process can continuously obtain data from the data source, the stream - processing process can be unaffected by the batch - processing merge process, process the data and store it in the log file, thereby reducing the data backlog in the stream - processing process, improving the throughput of the stream - processing process, and reducing the latency of the entire batch - stream integrated data - processing process.

[0051] Based on the above - mentioned method embodiments, the embodiments of the present application provide a batch - stream integrated data - processing device based on the data lake. See Figure 3 , Figure 3 which is a schematic diagram of a batch - stream integrated data - processing device based on the data lake provided by the embodiments of the present application.

[0052] The device 300 includes:

[0053] A first storage unit 301, configured to store the first data set in the first log file by using the stream - processing process;

[0054] The merging unit 302 is configured to, in response to a task that triggers a batch merging process, merge the first log file with the original full - volume file to obtain a first full - volume file, where the first log file and the original full - volume file have an associated relationship;

[0055] The second storage unit 303 is configured to, when in the process of merging the first log file with the original full - volume file, store a second data set in a second log file by using the stream processing process.

[0056] In a possible implementation manner, the triggering condition of the batch merging process includes:

[0057] After storing the first data set in the first log file by using the stream processing process, the number of times of storing a data set in a log file by using the stream processing process accumulates to a preset number of times.

[0058] In a possible implementation manner, after storing the first data set in the first log file by using the stream processing process, the apparatus 300 further includes: a third storage unit; the third storage unit is configured to store first metadata information, where the first metadata information is used to represent the relevant information of the first data set and the first log file.

[0059] In a possible implementation manner, after obtaining the first full - volume file, the third storage unit is further configured to store second metadata information, where the second metadata information is used to represent the relevant information of the first log file and the first full - volume file.

[0060] In a possible implementation manner, after storing the second data set in the second log file by using the stream processing process, the third storage unit is further configured to store third metadata information, where the third metadata information is used to represent the relevant information of the second data set and the second log file, and the storing of the second metadata information and the storing of the third metadata information are serial processes.

[0061] In a possible implementation manner, the obtaining process of the first data set includes:

[0062] Obtain an initial data set from a data source;

[0063] Perform data processing on the initial data set to obtain the first data set.

[0064] In a possible implementation manner, the associated relationship between the first log file and the original full - volume file includes: the data in the first log file and the data in the original full - volume file have the same primary key.

[0065] The beneficial effects of the batch-stream integrated data processing device based on the data lake provided by the embodiments of the present application can be referred to the above method embodiments, which will not be elaborated here.

[0066] It should be noted that the specific implementation of each unit in this embodiment can be referred to the relevant descriptions in the above method embodiments. The division of units in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation. Each functional unit in the embodiments of the present application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. For example, in the above embodiment, the processing unit and the sending unit can be the same unit or different units. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0067] See Figure 4 , which shows a schematic structural diagram of an electronic device 400 suitable for implementing the embodiments of the present application. The terminal devices in the embodiments of the present application may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0068] As Figure 4 shown, the electronic device 400 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage device 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.

[0069] Generally, the following devices can be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 can allow the electronic device 400 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4An electronic device 400 with various devices is shown, but it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0070] In particular, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 409, or installed from a storage device 408, or installed from a ROM 402. When the computer program is executed by a processing device 401, the above-mentioned functions defined in the methods of the embodiments of the present application are performed.

[0071] The electronic device provided by the embodiments of the present application and the method provided by the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be referred to the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0072] The embodiments of the present application provide a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the method provided by the above embodiments is implemented.

[0073] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0074] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0075] The above computer-readable medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device.

[0076] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device can execute the above method.

[0077] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0078] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0079] The units involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the unit / module does not constitute a limitation on the unit itself in some cases.

[0080] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0081] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0082] According to one or more embodiments of the present application, there is provided a batch-stream integrated data processing method based on a data lake, the method comprising:

[0083] Storing a first data set in a first log file using a stream processing process;

[0084] In response to a task that triggers a batch processing merge process, merging the first log file with an original full amount file to obtain a first full amount file, where the first log file and the original full amount file have an association relationship;

[0085] When in the process of merging the first log file with the original full amount file, storing a second data set in a second log file using the stream processing process.

[0086] According to one or more embodiments of the present application, the triggering conditions for the batch processing merge process include:

[0087] After storing the first data set in the first log file using the stream processing process, the number of times of storing a data set in a log file using the stream processing process is cumulatively reached a preset number of times.

[0088] According to one or more embodiments of the present application, after storing a first data set in a first log file using a stream processing process, the method further comprises:

[0089] Storing first metadata information, where the first metadata information is used to represent information related to the first data set and the first log file;

[0090] According to one or more embodiments of the present application, after obtaining the first full amount file, the method further comprises:

[0091] Store second metadata information, where the second metadata information is used to represent the relevant information of the first log file and the first full - volume file.

[0092] According to one or more embodiments of the present application, after storing the second data set in the second log file by using the stream processing process, the method further includes:

[0093] Store third metadata information, where the third metadata information is used to represent the relevant information of the second data set and the second log file, and the storing of the second metadata information and the storing of the third metadata information are serial processes.

[0094] According to one or more embodiments of the present application, the process of obtaining the first data set includes:

[0095] Obtain an initial data set from a data source;

[0096] Perform data processing on the initial data set to obtain the first data set.

[0097] According to one or more embodiments of the present application, the association relationship between the first log file and the original full - volume file includes: the data in the first log file and the data in the original full - volume file have the same primary key.

[0098] According to one or more embodiments of the present application, there is provided a batch - stream integrated data processing device based on a data lake, and the device includes:

[0099] A first storage unit, configured to store a first data set in a first log file by using a stream processing process;

[0100] A merging unit, configured to, in response to a task that triggers a batch - processing merging process, merge the first log file and an original full - volume file to obtain a first full - volume file, where the first log file and the original full - volume file have an association relationship;

[0101] A second storage unit, configured to, when in the process of merging the first log file and the original full - volume file, store a second data set in a second log file by using the stream processing process.

[0102] In one or more embodiments of the present application, the triggering conditions of the batch - processing merging process include:

[0103] When the first data set is stored in the first log file by using the stream processing process, the cumulative number of times of storing a data set in a log file by using the stream processing process reaches a preset number of times.

[0104] In one or more embodiments of the present application, after storing a first data set in a first log file by using a stream processing process, the apparatus 300 further includes: a third storage unit; the third storage unit is configured to store first metadata information, and the first metadata information is used to represent information related to the first data set and the first log file.

[0105] In one or more embodiments of the present application, after obtaining a first full-volume file, the third storage unit is further configured to store second metadata information, and the second metadata information is used to represent information related to the first log file and the first full-volume file.

[0106] In one or more embodiments of the present application, after storing a second data set in a second log file by using the stream processing process, the third storage unit is further configured to store third metadata information, and the third metadata information is used to represent information related to the second data set and the second log file, and the storing of the second metadata information and the storing of the third metadata information are serial processes.

[0107] In one or more embodiments of the present application, the obtaining process of the first data set includes:

[0108] Obtaining an initial data set from a data source;

[0109] Performing data processing on the initial data set to obtain the first data set.

[0110] In one or more embodiments of the present application, the association relationship between the first log file and the original full-volume file includes: the data in the first log file and the data in the original full-volume file have the same primary key.

[0111] According to one or more embodiments of the present application, there is provided an electronic device, and the device includes: a processor and a memory;

[0112] The memory is configured to store instructions or computer programs;

[0113] The processor is configured to execute the instructions or computer programs in the memory so that the electronic device executes the data processing method based on batch-stream integration of a data lake.

[0114] According to one or more embodiments of the present application, there is provided a computer-readable storage medium, and instructions are stored in the computer-readable storage medium, and when the instructions run on a device, the device is caused to execute the data processing method based on batch-stream integration of a data lake.

[0115] It should be noted that the embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0116] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0117] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0118] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0119] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A batch-stream integrated data processing method based on a data lake, characterized in that, The method includes: Storing a first data set in a first log file by using a stream processing process; Calculating the number of times of storing the data set in the log file, and determining whether a condition for triggering a batch processing and merging process is met based on the number of times; In response to a task for triggering the batch processing and merging process, merging the first log file with an original full-volume file to obtain a first full-volume file, where the first log file and the original full-volume file have an association relationship; When in the process of merging the first log file with the original full-volume file, storing a second data set in a second log file by using the stream processing process.

2. The method according to claim 1, characterized in that, The triggering condition of the batch processing and merging process includes: After storing the first data set in the first log file by using the stream processing process, the number of times of storing the data set in the log file accumulatively reaches a preset number of times.

3. The method according to claim 1, characterized in that, After storing the first data set in the first log file by using the stream processing process, the method further includes: Storing first metadata information, where the first metadata information is used to represent information related to the first data set and the first log file.

4. The method according to claim 1, wherein After obtaining the first full-volume file, the method further includes: Storing second metadata information, where the second metadata information is used to represent information related to the first log file and the first full-volume file.

5. The method according to claim 4, characterized in that, After storing the second data set in the second log file by using the stream processing process, the method further includes: Storing third metadata information, where the third metadata information is used to represent information related to the second data set and the second log file, and the storing of the second metadata information and the storing of the third metadata information are serial processes.

6. The method according to claim 1, wherein The obtaining process of the first data set includes: Obtaining an initial data set from a data source; Performing data processing on the initial data set to obtain the first data set.

7. The method according to any one of claims 1 to 6, characterized in that, The association relationship between the first log file and the original full-volume file includes: the data in the first log file and the data in the original full-volume file have the same primary key.

8. A batch-stream integrated data processing device based on a data lake, characterized in that, The device includes: A first storage unit for storing a first data set in a first log file by using a stream processing process; A merging unit for calculating the number of times of storing the data set in the log file, and determining whether a condition for triggering a batch processing and merging process is met based on the number of times; in response to a task for triggering the batch processing and merging process, merging the first log file with an original full-volume file to obtain a first full-volume file, where the first log file and the original full-volume file have an association relationship; A second storage unit for storing a second data set in a second log file by using the stream processing process when in the process of merging the first log file with the original full-volume file.

9. An electronic device, characterized in that, The device includes: a processor and a memory; The memory for storing instructions or computer programs; The processor for executing the instructions or computer programs in the memory so that the electronic device executes the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when the instructions are run on a device, the device is caused to perform the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Real-time service data merging method and device and electronic equipment

    CN112988741A

  • Data processing method and device, equipment and storage medium

    CN114676161A