Multi-data-stream parallel processing method and apparatus and nonvolatile storage medium
By associating in the pre-write batch stage of the data stream and sinking the association logic to the storage layer of the data lake, the performance bottleneck caused by real-time data flow association of the computing layer is solved, and a lower amount of data cached by computing tasks is achieved.
Patent Information
- Application Number
- PCT/CN2024/120767
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2024-09-24
- Publication Date
- 2025-06-05
AI Technical Summary
When data flow association, the prior art performs real-time data flow association at the computing layer, resulting in a huge amount of data cached by the computing task, forming a performance bottleneck.
By associating data flows in the pre-write batch stage and sinking the association logic to the storage layer of the data lake, real-time data flow associations are avoided at the computing layer.
This enables no need to execute all data association tasks in the computing layer, reducing the amount of cached data of the computing tasks and avoiding performance bottlenecks.
Smart Images

Figure CN2024120767_05062025_PF_FP_ABST
Abstract
Description
Multi-data stream parallel processing method, device and non-volatile storage medium
[0001] Related applications
[0002] This application claims priority to Chinese patent application number 202311597814.3, filed on November 27, 2023, entitled “Multiple Data Stream Parallel Processing Method, Device and Non-volatile Storage Medium,” the entire text of which is hereby incorporated by reference. Technical Field
[0003] The present application relates to the field of data processing, and in particular to a method and device for parallel processing of multiple data streams, and a non-volatile computer-readable storage medium. Background Art
[0004] In related technologies, when associating data streams, real-time data streams are usually associated at the computing layer, and an external storage system needs to be introduced during the association, which leads to high pressure on storage and network transmission, easily forming performance bottlenecks, and the amount of cached data for computing tasks is extremely large.
[0005] To address the above-mentioned problems, no effective solutions have been proposed so far.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide a method, device and non-volatile storage medium for parallel processing of multiple data streams, so as to at least solve the technical problem in the related art that the amount of data cached for computing tasks is extremely large when real-time data stream association is performed at the computing layer when associating data streams.
[0008] In a first aspect, the present application provides a method for parallel processing of multiple data streams, including: determining task configuration information of a target task stream, the target task stream includes at least one data stream, the task configuration information includes: the association method between the data streams in the target task stream, the data feature information of each data stream in the target task stream, and the data table type of the data stream; in the pre-write batching stage of the target task stream, performing a first type of operation on the data stream in the target task stream according to the task configuration information, the first type of operation including at least one of the following: data deduplication and data association; in the data write and merge stage of the target task stream, performing persistent disk processing on the data stream in the target task stream according to the task configuration information.
[0009] In some embodiments, the association mode includes one of the following: normal write mode, wide table join mode, and data stream JOIN mode.
[0010] In some embodiments, when the association mode is a wide table connection mode, the steps of performing the first type of operation on the data stream in the target task flow according to the task configuration information include: determining whether there is target data with the same key value in the target task flow; when it is determined that there is target data with the same key value, determining whether the target data with the same key value belongs to the same data stream; when it is determined that the target data with the same key value belongs to the same data stream, updating the data stream to which the target data belongs; when it is determined that the target data with the same key value does not belong to the same data stream, performing data splicing processing on the data stream to which the target data belongs.
[0011] In some embodiments, when the association mode is a data stream JOIN mode, the steps of performing the first type of operation on the data stream in the target task stream according to the task configuration information include: determining whether there is target data with the same key value in the target task stream; when it is determined that there is target data with the same key value, determining whether the target data with the same key value belongs to the same data table; when it is determined that the target data with the same key value corresponds to the same data table, retaining the target data with a later data generation time; when it is determined that the target data with the same key value does not correspond to the same data table, performing data stream JOIN processing on the data stream corresponding to the data.
[0012] In some embodiments, the step of performing data stream JOIN processing on the data stream corresponding to the data includes: determining the JOIN order between the data streams, and associating the data streams according to the JOIN order; when the data streams are associated again according to the JOIN order, the association method is Inner.Join or Right.Join, and there is a JOIN association failure, returning and marking the target data as hidden data; when the data streams are associated again according to the JOIN order, if the association method is Left.Join, returning the target data, and adding a hidden mark to the data in the data stream that does not participate in the association.
[0013] In some embodiments, when the association method is a data stream JOIN mode, the step of persistently writing the data stream in the target task stream to the disk based on the task configuration information includes: determining the write merge logic corresponding to the table type information of the data stream, wherein the write merge logic includes a write merge method for a first type of data stream and a second type of data stream, the first type of data stream is a data stream with hidden marks in the corresponding data, and the second type of data stream is a data stream without hidden marks in the corresponding data; and persistently writing the data stream to the disk based on the write merge logic corresponding to the data stream.
[0014] In some embodiments, when the association method is a wide table connection mode, the steps of persistently writing the data stream in the target task stream to the disk according to the task configuration information include: determining the target data with the same key value; determining the generation time of the target data, and determining the target column where the latest generated target data is located; using the data in the target column to overwrite the historical data, and persistently writing the data associated with the data stream to the disk after the overwriting.
[0015] In some embodiments, the task configuration information also includes parallelism; the step of determining the configuration information of the target task flow includes: obtaining a preset parallelism parameter; or, determining the number of data sources associated with the target task flow, and configuring the parallelism parameter of the target task flow based on the number of data sources.
[0016] In some embodiments, after the step of persisting the data stream in the target task stream to disk according to the task configuration information, the multi-data stream parallel processing method also includes: determining the query type of the data query instruction and the table type of the data table corresponding to the data query instruction; and processing the data in the data table according to the query type and the table type.
[0017] In a second aspect, the present application also provides a multi-data stream parallel processing device, including: a first processing module, used to determine the task configuration information of the target task stream, the target task stream includes at least one data stream, and the task configuration information includes: the association method between the data streams in the target task stream, the data feature information of each data stream in the target task stream, and the data table type of the data stream; a second processing module, used to perform a first type of operation on the data stream in the target task stream according to the task configuration information in the pre-write batching stage of the target task stream, the first type of operation includes at least one of the following: data deduplication, data association; a third processing module, used to perform persistent disk processing on the data stream in the target task stream according to the task configuration information in the data write merging stage of the target task stream.
[0018] In some embodiments, the association mode includes one of the following: normal write mode, wide table connection mode, and data flow join (JOIN) mode.
[0019] In some embodiments, the first processing module is further configured to: obtain a preset parallelism parameter; or determine the number of data sources associated with the target task flow, and configure the parallelism parameter of the target task flow according to the number of data sources.
[0020] In some embodiments, when the association mode is a wide table connection mode, the second processing module is further used to: determine whether there is target data with the same key value in the target task flow; when it is determined that the target data with the same key value exists, determine whether the target data with the same key value belongs to the same data flow; when it is determined that the target data with the same key value belongs to the same data flow, update the data flow to which the target data belongs; when it is determined that the target data with the same key value does not belong to the same data flow, perform data splicing processing on the data flow to which the target data belongs.
[0021] In some embodiments, the second processing module, when the association mode is the data stream JOIN mode, is further used to: determine whether there is target data with the same key value in the target task flow; when it is determined that the target data with the same key value exists, determine whether the target data with the same key value belongs to the same data table; when it is determined that the target data with the same key value corresponds to the same data table, retain the target data with a later data generation time; when it is determined that the target data with the same key value does not correspond to the same data table, perform data stream JOIN processing on the data stream corresponding to the data.
[0022] In some embodiments, the second processing module is further used to: determine the JOIN order between the data streams, and associate the data streams according to the JOIN order; when the data streams are associated again according to the JOIN order and the association method is Inner.Join or Right.Join, and there is a JoinJOIN association failure, return and mark the target data as hidden data; when the data streams are associated again according to the JOIN order and the association method is Left.Join, return the target data, and add a hidden mark to the data in the data stream that does not participate in the association.
[0023] In some embodiments, when the association method is a data stream JOIN mode, the third processing module is further used to: determine the write merge logic corresponding to the table type information of the data stream, the write merge logic including the write merge method for the first type of data stream and the second type of data stream, the first type of data stream is the data stream in which the hidden mark exists in the corresponding data, and the second type of data stream is the data stream in which the hidden mark does not exist in the corresponding data; and perform persistent disk processing on the data stream according to the write merge logic corresponding to the data stream.
[0024] In some embodiments, when the association method is a wide table join mode, the third processing module is further used to: determine target data with the same key value; determine the generation time of the target data, and determine the target column where the most recently generated target data is located; use the data of the target column to overwrite historical data, and after overwriting, persist the data associated with the data stream to disk.
[0025] In some embodiments, the multi-data stream parallel processing device also includes a fourth processing module, which is used to: determine the query type of the data query instruction and the table type of the data table corresponding to the data query instruction; and process the data in the data table based on the query type and the table type.
[0026] In a third aspect, the present application further provides a non-volatile computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the multi-data stream parallel processing method in any embodiment of the first aspect is executed.
[0027] In a fourth aspect, the present application further provides an electronic device comprising: a memory and a processor, wherein the processor executes the multi-data stream parallel processing method in any embodiment of the first aspect when running a computer program.
[0028] According to the multi-data stream parallel processing method, device and non-volatile computer-readable storage medium provided in the embodiment of the present application, task configuration information of the target task stream is determined, the target task stream includes at least one data stream, and the task configuration information includes: the association method between the data streams in the target task stream, the data feature information of each data stream in the target task stream, and the data table type of the data stream; in the pre-write batching stage of the target task stream, a first type of operation is performed on the data stream in the target task stream according to the task configuration information, and the first type of operation includes at least one of the following: data deduplication, data association; in the data write merging stage of the target task stream, the data stream in the target task stream is persistently stored on the disk according to the task configuration information. Thus, by associating the data streams in the pre-write batching stage, the purpose of sinking the association logic of the data streams to the storage layer of the data lake is achieved, thereby achieving the technical effect of not having to execute all data association tasks in the computing layer, and thus solving the technical problem of extremely large computing task cache data caused by real-time data stream association in the computing layer when associating data streams in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings described below are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be derived from these drawings without inventive effort.
[0030] FIG1 is a schematic structural diagram of a computer device (mobile device) provided according to an embodiment of the present application.
[0031] FIG2 is a flow chart of a method for parallel processing of multiple data streams provided according to an embodiment of the present application.
[0032] FIG3 is a flowchart of a multi-data stream processing process provided according to an embodiment of the present application.
[0033] FIG4 is a schematic structural diagram of a device for parallel processing of multiple data streams according to an embodiment of the present application. DETAILED DESCRIPTION
[0034] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts should fall within the scope of protection of this application.
[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0036] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:
[0037] Hudi: In this application, it refers to the data lake.
[0038] Payload: Hudi Payload is a scalable data processing mechanism that allows for customized data writing for complex scenarios, greatly increasing data processing flexibility. Payload is a utility class that performs operations such as deduplication, filtering, and merging when writing and reading Hudi tables.
[0039] State: A data structure stored in a state backend (e.g., a state storage module, used to store the state of various objects) to meet the historical data needs of operator calculations and ensure fault tolerance using a checkpoint mechanism. State is used to store intermediate node results or metadata during the calculation process.
[0040] Compaction: Coordinates the operation of the difference data structure in Hudi, converting the updates from the row-based log file to the columnar format file.
[0041] Base file: In Hudi, base files are stored in the form of columnar files.
[0042] Log file: In Hudi, incremental files such as log files are stored in the form of line files.
[0043] Thanks to the development of real-time computing frameworks like Flink and Spark Streaming, as well as the rise of technologies like Kafka and MPP, real-time computing technology is becoming increasingly sophisticated. Furthermore, with the proliferation of technologies like the Internet of Things and machine learning, real-time streaming computing has found widespread application in areas such as intelligent recommendations, real-time fraud detection, complex event processing, and real-time machine learning. With the support of diverse application scenarios, traditional statistical aggregation calculations based on single data streams are no longer sufficient for calculating metrics in complex business scenarios, and the need for correlation and connection of real-time data streams is urgent.
[0044] In related technologies, when associating data streams, the following data stream JOIN connection schemes are commonly used: 1. Multi-task parallel processing, in which the data streams that need to be associated are written to an external storage system, and external data storage points are checked in the key data stream to complete the data association. 2. Using the Flink streaming computing framework, using Interval Join, Window Join, or its own State to cache associated data. 3. Columnar storage, in which each streaming task corresponds to a single real-time data source and only certain columns related to the single data source are updated.
[0045] However, all of the above solutions have drawbacks. Method 1 requires external storage, increasing the operational burden. The large amount of data puts pressure on storage and network I / O, easily creating performance bottlenecks. Method 2, using Flink Interval Join or Window Join, requires caching data in memory, placing a high workload on tasks. Furthermore, excessive state can lead to prolonged checkpoint times, impacting the timeliness of business data and the stability of Flink streaming tasks.
[0046] Moreover, the above association method may cause write contention when multiple streams write to the table simultaneously. Write locks need to be acquired sequentially, which can easily lead to deadlock and performance loss.
[0047] Furthermore, when performing data association in related technologies, when loading hot and cold data, it's impossible to set a reasonable time to live (TTL) for the hot data because it's stored in memory. This can lead to data anomalies caused by cache data not being updated in a timely manner. Furthermore, performing real-time association of data streams at the computational layer in related technologies can lead to extremely large amounts of cached data for computational tasks, causing back pressure and other issues.
[0048] According to an embodiment of the present application, a method for parallel processing of multiple data streams is provided. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed by a set of computer-executable instructions, such as in a computer system, and that although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in a different order than shown.
[0049] The method provided in the embodiment of the present application can be executed in a mobile terminal, a computer device or a similar computing device. Figure 1 shows a hardware structure block diagram of a computer device (or mobile terminal) for implementing a multi-data stream parallel processing method. As shown in Figure 1, the computer device 10 (or mobile terminal) may include one or more (as shown in 102a, 102b, ..., 102n in Figure 1) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be used as a port of a bus (BUS)), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that the structure shown in Figure 1 is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer device 10 may also include more or fewer components than those shown in Figure 1, or have a configuration different from that shown in Figure 1.
[0050] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer device 10 (or mobile terminal). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0051] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the multi-data stream parallel processing method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned multi-data stream parallel processing method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer device 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0052] The transmission device 106 is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by the communications provider of the computer device 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0053] The display may be, for example, a touch screen liquid crystal display (LCD), which may enable a user to interact with a user interface of the computer device 10 (or mobile terminal).
[0054] In the above operating environment, an embodiment of the present application provides a method for parallel processing of multiple data streams. As shown in FIG2 , the method includes the following steps S202 to S206 .
[0055] Step S202: determine the task configuration information of the target task flow, wherein the target task flow includes at least one data flow; the task configuration information includes: the association method between the data flows in the target task flow, the data feature information of each data flow in the target task flow, and the data table type of the data flow.
[0056] In the technical solution provided in step S202 , the association mode includes one of the following: a normal write mode, a wide table connection mode, and a data stream join (JOIN) mode.
[0057] As an optional implementation, when configuring the target task flow, the configuration content may include:
[0058] The connection method between the data streams corresponding to the target task stream, such as data stream join mode (JOIN), common write mode (common), wide table join mode (wide-mode), etc.
[0059] The parsed features (schema) of a single stream are used for data parsing and pattern matching. For example, t1.schema = {"type":"record","fields":[{"name":"field1","type":"string"},{"name":"field2","type":"string"}]}. The schema includes the structure and type of each stream, such as field name, data type, length, and constraints.
[0060] JOIN logic for data streams connected using the JOIN mode, such as t1.left.join.t2, where t1 and t2 represent two connected data streams.
[0061] In some embodiments of the present application, the task configuration information also includes parallelism; the step of determining the configuration information of the target task flow includes: obtaining a preset parallelism parameter; or, determining the number of data sources associated with the target task flow; and configuring the parallelism parameter of the target task flow based on the number of data sources.
[0062] Specifically, the above-mentioned parallelism refers to the number of data streams that can be processed simultaneously in the target task flow. When setting parallelism, users can configure global parallelism in the streaming computing framework or configure the parallelism of a single process task. When a task is executed, if the task does not have a personalized parallelism configuration, the global parallelism configuration is used as the default configuration for the task. If the user does not set the global task parallelism and the parallelism for a single task, the parallelism can be inferred based on the specific number of data sources. For example, for a data source of the Kafka Source type, the task parallelism is configured based on the number of Topic partitions. For example, the parallelism can be set to be equal to the number of data sources or the number of Topic partitions.
[0063] After configuration is complete, you can connect all data streams to be associated, configure streaming tasks to read the data source topic list, and set up a partition discovery mechanism to dynamically obtain the number of source topic partitions. For example, when partitions are added, corresponding consumer registrations are performed for the added partitions, included in the data reading list, and dynamically connected to the data streams to be associated in real time.
[0064] After the task configuration settings for the target task flow are completed, the target task flow can load the configuration file and read the task configuration during the startup phase, then determine the connection method between the data flows and load the Payload data processing class corresponding to different modes.
[0065] Step S204: in the pre-writing and batching stage of the target task flow, performing a first type of operation on the data flow in the target task flow according to the task configuration information, wherein the first type of operation includes at least one of the following: data deduplication and data association;
[0066] During the pre-write batching phase, different data deduplication and association logic can be configured based on the connection between data streams. For example, for a data stream in normal write mode, deduplication will only be performed on that data stream, and data will be updated and deleted according to the preset configuration scheme, but this data stream will not be connected or associated with other data streams.
[0067] In the technical solution provided in step S204, when the association method is a wide table connection mode, the steps of performing the first type of operation on the data stream in the target task flow according to the task configuration information include: determining whether there is target data with the same key value in the target task flow; when it is determined that there is target data with the same key value, determining whether the target data with the same key value belongs to the same data stream; when the target data with the same key value belongs to the same data stream, updating the data stream to which the target data belongs; when the target data with the same key value does not belong to the same data stream, performing data splicing processing on the data stream to which the target data belongs.
[0068] Specifically, because the data being processed at this time contains data written by different data streams, that is, the columns contained in each piece of data may be different. Therefore, when merging and deduplicating data, it is necessary to determine whether the two sets of data (records) with the same key value (recordKey) come from the same data stream. If it is confirmed that they come from the same data stream, the data corresponding to the key value is updated, that is, the latest generated data is retained. If not, data splicing is performed. When performing data splicing, if the two data streams have overlapping columns, they are sorted according to the pre-merged field (precombine.field), and the overlapping columns are applied as the latest data, that is, replaced with the latest data.
[0069] As an optional implementation, when the association mode is a data stream JOIN mode, the step of performing the first type of operation on the data stream in the target task stream according to the task configuration information includes: determining whether there is target data with the same key value in the target task stream; when it is determined that there is target data with the same key value, determining whether the target data with the same key value belongs to the same data table; when it is determined that the target data with the same key value corresponds to the same data table, retaining the target data with a later data generation time; when it is determined that the data with the same key value does not correspond to the same data table, performing data stream JOIN processing on the data stream corresponding to the data.
[0070] In some embodiments of the present application, the step of performing data stream JOIN processing on the data stream corresponding to the data includes: determining the JOIN order between the data streams, and associating the data streams according to the JOIN order; when the data streams are associated again according to the JOIN order, the association method is Inner.Join or Right.Join, and if there is a JOIN association failure, returning and marking the target data as hidden data; when the data streams are associated again according to the JOIN order, if the association method is Left.Join, returning the target data, and adding a hidden mark to the data in the data stream that does not participate in the association.
[0071] Specifically, when using the JOIN mode to connect different data streams, the data that the target task flow needs to process will contain data written by different data streams. That is, the data may come from different data tables. Therefore, at this stage, for two sets of data (records) with the same key value (recordKey), it is necessary to further determine whether the corresponding schemas of the data are the same. If they are the same, it means that the two sets of data come from the same data table, and the latest data is used to overwrite the old data to complete the data update. If they are not the same, it means that the two sets of data come from different data tables. At this time, it is necessary to match and correspond according to the data stream schema and the configuration table schema, distinguish which data stream table each data comes from (the data stream table can be regarded as a data stream in this application), and perform data JOIN on the data stream table according to the configuration connection JOIN method, and perform stream association according to the JOIN order.
[0072] If the next sequential join between related data stream tables uses an Inner Join or Right Join, and any joins are not found, the previously connected intermediate data (that is, data with the same key value as previously discovered) will be returned and marked as hidden data. If the next sequential join between related data stream tables uses a Left Join, the data will be returned without the need for data hiding. Any other data not involved in the join will be marked hidden and returned, preparing for the next stage of data writing.
[0073] It should be noted that returning data in the pre-write batching stage means allowing the data to be written and merged.
[0074] Step S206 , in the data writing and merging phase of the target task flow, the data flow in the target task flow is persistently written to the disk according to the task configuration information.
[0075] In the technical solution provided in step S206, during the data write and merge phase, the Hudi table can be used to determine the corresponding data write logic according to different data table types. For example, for the Copy On Write table type, the basic file can be copied again, and during the re-write process, the historical data and incremental data can be updated to complete the persistent disk. For the Merge On Read table type, the incremental data can be written to the log file first, and then during the merge (Compaction), the data in the historical data basic file and the incremental log file can be merged and updated.
[0076] As an optional implementation, for data streams in normal write mode, two sets of records with the same recordKey can be compared according to the precombine.field field based on the configured Payload, and data that meets the Payload retention rules can be persisted to disk.
[0077] In some embodiments of the present application, when the association method is a wide table connection mode, the steps of persistently writing the data stream in the target task stream to the disk based on the task configuration information include: determining the target data with the same key value; determining the generation time of the target data, and determining the target column where the latest generated target data is located; using the data in the target column to overwrite the historical data, and persistently writing the data associated with the data stream to the disk after the overwriting.
[0078] Specifically, for data streams with wide table joins, for two sets of data with the same recordKey, all columns of the latest data will be used to overwrite the historical data, splicing them into the latest data before persisting them to disk. For example, for a record with primary key key1 in the base file, when the Spill Map finds data with the same key value as that record in columns B, C, and D, there will be no data in column A with primary key key1. In this case, the latest data will be used to update columns B, C, and D, and after the update is complete, the updated columns B, C, and D will be spliced back with column A.
[0079] As an optional implementation method, when the association method is the data stream JOIN mode, the step of persistently writing the data stream in the target task stream to the disk according to the task configuration information includes: determining the write merge logic corresponding to the table type information of the data stream, wherein the write merge logic includes a write merge method for the first type of data stream and the second type of data stream, the first type of data stream is a data stream with hidden marks in the corresponding data, and the second type of data stream is a data stream without hidden marks in the corresponding data; the data stream is persistently written to the disk according to the write merge logic corresponding to the data stream.
[0080] Specifically, for data streams connected using the JOIN mode, when performing persistent disk processing, it is necessary to first determine the type of data table to be processed and then determine the corresponding write merge logic. For example, for data of the Merge On Read table type, during the merging (Compaction) phase of the base file and the Log file, for data with the same recordKey, if there is data that does not contain hidden marks in the incremental log file (Log file), and there is no data with the corresponding recordKey in the existing data in the corresponding base file (base file), then the data with the same recordKey is determined to be new associated data and written to the base file file corresponding to the new timestamp (instant time). If there is existing data with the same recordKey in the base file file, the incremental data is used to overwrite and update the historical data and write it to the base file file corresponding to the new instant time.
[0081] If the corresponding log file contains data with a hidden mark and data with the same recordKey exists in the base file, the columns in the incremental data are updated to the historical data, and the complete data is spliced and persisted to disk. If the base file does not contain data with the above recordKey, the hidden data is written to the log file corresponding to the new instant time, which is used to associate the relevant data streams during the next compaction.
[0082] For data in Copy On Write tables, writing this data is equivalent to a compaction and only includes the base file. Therefore, the specific writing logic is as follows: For data with the same recordKey, the new and old data are merged, and overlapping columns are updated using an overwrite update. For data without the same recordKey, it is marked as hidden data and written to the base file corresponding to the new instant time.
[0083] In some embodiments of the present application, after the step of persisting the data stream in the target task stream to disk according to the task configuration information, the multi-data stream parallel processing method also includes: determining the query type of the data query instruction and the table type of the data table corresponding to the data query instruction; and processing the data in the data table based on the query type and the table type.
[0084] Specifically, during the data application phase, such as when performing data queries, data in the data table can be processed differently depending on the table type and query instruction. For example, when querying a Copy on Write table, there's no need to merge incremental data with historical data. Therefore, only data containing hidden tags needs to be filtered out to obtain the final, post-association, real-time streaming wide table data.
[0085] For Merge On Read tables, if the query type is Snapshot View, the historical base file and the incremental log file must be merged. The specific operation logic of the merge is the same as the compaction process during the write merge phase. Data containing hidden tags is then filtered out to form the final, linked real-time streaming wide table data.
[0086] In summary, the complete process of processing multiple data streams in this application is shown in FIG3 , including the following steps S302 to S308 .
[0087] Step S302: Load the task configuration file in the stream computing task and determine the correlation pattern between data streams and the data stream parsing features;
[0088] Step S304, loading a corresponding mode payload according to the association mode configuration between the data flows;
[0089] Step S306: performing data pre-deduplication and association operations in the data pre-writing and batching stage according to the association pattern of the data stream;
[0090] Step S308: Based on the association pattern of the data stream and the data table type, historical data and incremental data are merged during the data writing and storage phase.
[0091] It can be seen that in this application, the JOIN process of implementing the data stream is sunk to the Hudi storage of the data lake, and the Hudi payload technology is used for wide table connection. There is no need to associate data between cache states in the computing engine, and there is no need to introduce external storage for associated data caching. Therefore, there is no need to operate and maintain the external storage system, and no additional pressure will be placed on the storage and network input / output (I / O), and no performance bottleneck will be caused. Moreover, in this application, for a single streaming computing task, data can be written to multiple associated dimension tables, avoiding the concurrent write lock situation that may exist when multiple tasks write to a single table at the same time, improving task processing performance, and avoiding deadlock problems that may occur during concurrent writing.
[0092] In addition, by adopting the task configuration information for determining the target task flow, wherein the target task flow includes at least one data flow. The task configuration information includes: the association method between the data flows in the target task flow, the data feature information of each data flow in the target task flow, and the data table type of the data flow. In the pre-write batching stage of the target task flow, the first type of operation is performed on the data flow in the target task flow according to the task configuration information, wherein the first type of operation includes at least one of the following: data deduplication and data association. In the data writing and merging stage of the target task flow, the data flow in the target task flow is persistently stored on the disk according to the task configuration information. By associating the data flow in the pre-write batching stage, the purpose of sinking the association logic of the data flow to the storage layer of the data lake is achieved, thereby achieving the technical effect of not having to execute all data association tasks in the computing layer, thereby solving the technical problem of the huge amount of computing task cache data caused by real-time data flow association in the computing layer when associating data flow in the related technology.
[0093] The embodiment of the present application provides a device for parallel processing of multiple data streams, and FIG4 is a schematic diagram of the structure of the device. As shown in FIG4, the device includes:
[0094] The first processing module 40 is used to determine the task configuration information of the target task flow; wherein the target task flow includes at least one data flow; the task configuration information includes: the association method between the data flows in the target task flow, the data feature information of each data flow in the target task flow, and the data table type of the data flow.
[0095] The second processing module 42 is used to perform a first type of operation on the data stream in the target task flow according to the task configuration information during the pre-writing batching stage of the target task flow, wherein the first type of operation includes at least one of the following: data deduplication and data association.
[0096] The third processing module 44 is configured to perform persistent write-to-disk processing on the data stream in the target task stream according to the task configuration information during the data write-to-disk merging phase of the target task stream.
[0097] In some embodiments of the present application, the association mode includes one of the following: normal write mode, wide table connection mode, and data stream JOIN mode.
[0098] In some embodiments of the present application, the task configuration information also includes parallelism; the first processing module 40 is further used to: obtain a preset parallelism parameter; or, determine the number of data sources associated with the target task flow, and configure the parallelism parameter of the target task flow based on the number of data sources.
[0099] In some embodiments of the present application, when the association mode is a wide table connection mode, the second processing module 42 is further used to: determine whether there is target data with the same key value in the target task flow; when it is determined that there is target data with the same key value, determine whether the target data with the same key value belongs to the same data stream; when it is determined that the target data with the same key value belongs to the same data stream, update the data stream to which the target data belongs; when it is determined that the target data with the same key value does not belong to the same data stream, perform data splicing processing on the data stream to which the target data belongs.
[0100] In some embodiments of the present application, when the association mode is a data stream JOIN mode, the second processing module 42 is further used to: determine whether there is target data with the same key value in the target task flow; when it is determined that there is target data with the same key value, determine whether the target data with the same key value belongs to the same data table; when it is determined that the target data with the same key value corresponds to the same data table, retain the target data with a later data generation time; when it is determined that the data with the same key value does not correspond to the same data table, perform data stream JOIN processing on the data stream corresponding to the data.
[0101] In some embodiments of the present application, the second processing module 42 is further used to: determine the JOIN order between data streams, and associate the data streams according to the JOIN order; when the data streams are associated again according to the JOIN order, the association method is Inner.Join or Right.Join, and there is a JOIN association failure, return and mark the target data as hidden data; when the data streams are associated again according to the JOIN order, if the association method is Left.Join, return the target data, and add a hidden mark to the data in the data stream that does not participate in the association.
[0102] In some embodiments of the present application, when the association method is a data stream JOIN mode, the third processing module 44 is further used to: determine the write merge logic corresponding to the table type information of the data stream, wherein the write merge logic includes a write merge method for the first type of data stream and the second type of data stream, the first type of data stream is a data stream with hidden marks in the corresponding data, and the second type of data stream is a data stream with no hidden marks in the corresponding data; and perform persistent disk processing on the data stream according to the write merge logic corresponding to the data stream.
[0103] In some embodiments of the present application, when the association method is a wide table connection mode, the third processing module 44 is further used to: determine the target data with the same key value; determine the generation time of the target data, and determine the target column where the latest generated target data is located; use the data of the target column to overwrite the historical data, and after overwriting, persist the data associated with the data stream to the disk.
[0104] In some embodiments of the present application, after the third processing module performs persistent disk processing on the data stream in the target task stream based on the task configuration information, the multi-data stream parallel processing device may also include a fourth processing module for determining the query type of the data query instruction and the table type of the data table corresponding to the data query instruction; and processing the data in the data table based on the query type and the table type.
[0105] It should be noted that the various modules in the above-mentioned multi-data stream parallel processing device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.
[0106] The embodiment of the present application also provides a non-volatile storage medium. A computer program is stored in the non-volatile storage medium, and when the computer program is executed by the processor, the computer program performs the following multi-data stream parallel processing method: determining the task configuration information of the target task stream, wherein the target task stream includes at least one data stream, and the task configuration information includes: the association method between the data streams in the target task stream, the data feature information of each data stream in the target task stream, and the data table type of the data stream; in the pre-write batching stage of the target task stream, performing a first type of operation on the data stream in the target task stream according to the task configuration information, wherein the first type of operation includes at least one of the following: data deduplication and data association; and in the data write merging stage of the target task stream, performing persistent disk processing on the data stream in the target task stream according to the task configuration information.
[0107] An embodiment of the present application also provides an electronic device, which includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor runs the computer program, it executes the following multi-data stream parallel processing method: determining task configuration information of a target task stream, wherein the target task stream includes at least one data stream, and the task configuration information includes: the association method between the data streams in the target task stream, the data feature information of each data stream in the target task stream, and the data table type of the data stream; in the pre-write batching stage of the target task stream, performing a first type of operation on the data stream in the target task stream according to the task configuration information, wherein the first type of operation includes at least one of the following: data deduplication and data association; and in the data write merging stage of the target task stream, performing persistent disk processing on the data stream in the target task stream according to the task configuration information.
[0108] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0109] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0110] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0111] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0112] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0113] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.
[0114] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0115] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for parallel processing of multiple data streams, comprising: Determine task configuration information of a target task flow, wherein the target task flow includes at least one data flow, and the task configuration information includes: an association mode between the data flows in the target task flow, data feature information of each data flow in the target task flow, and a data table type of the data flow; In the pre-writing batching stage of the target task flow, performing a first type of operation on the data flow in the target task flow according to the task configuration information, the first type of operation including at least one of the following: data deduplication, data association; and During the data writing and merging phase of the target task stream, the data stream in the target task stream is persistently written to the disk according to the task configuration information.
2. The multi-data stream parallel processing method according to claim 1, wherein the association mode comprises one of the following: a normal write mode, a wide table connection mode, and a data stream join (JOIN) mode.
3. The method for parallel processing of multiple data streams according to claim 2, wherein when the association mode is the wide table connection mode, the step of performing the first type of operation on the data stream in the target task stream according to the task configuration information comprises: Determine whether there is target data with the same key value in the target task flow; In the case where it is determined that the target data with the same key value exists, determining whether the target data with the same key value belongs to the same data stream; When it is determined that the target data with the same key value belong to the same data stream, updating the data stream to which the target data belongs; When it is determined that the target data with the same key value do not belong to the same data stream, data splicing processing is performed on the data stream to which the target data belongs.
4. The method for parallel processing of multiple data streams according to claim 2, wherein when the association mode is the data stream JOIN mode, the step of performing the first type of operation on the data stream in the target task stream according to the task configuration information comprises: Determine whether there is target data with the same key value in the target task flow; In the case where it is determined that the target data with the same key value exists, determining whether the target data with the same key value belongs to the same data table; When it is determined that the target data with the same key value corresponds to the same data table, retaining the target data with a later data generation time; When it is determined that the target data with the same key value do not correspond to the same data table, data stream JOIN processing is performed on the data stream corresponding to the data.
5. The method for parallel processing of multiple data streams according to claim 4, wherein the step of performing data stream JOIN processing on the data stream corresponding to the data comprises: Determine a JOIN order between the data streams, and associate the data streams according to the JOIN order; When the data stream is associated again according to the JOIN sequence, the association mode is Inner.Join or Right.Join, and if there is a JOIN association failure, the target data is returned and marked as hidden data; When the data stream is associated again according to the JOIN sequence and the association mode is Left.Join, the target data is returned, and a hidden mark is added to the data in the data stream that does not participate in the association.
6. The method for parallel processing of multiple data streams according to claim 5, wherein when the association mode is the data stream JOIN mode, the step of persistently storing the data stream in the target task stream according to the task configuration information comprises: Determine a write merge logic corresponding to the table type information of the data stream, wherein the write merge logic includes a write merge mode for a first type of data stream and a second type of data stream, the first type of data stream is a data stream in which the hidden mark exists in the corresponding data, and the second type of data stream is a data stream in which the hidden mark does not exist in the corresponding data; The data stream is persistently written to disk according to the write-merge logic corresponding to the data stream.
7. The method for parallel processing of multiple data streams according to claim 2, wherein when the association mode is the wide table connection mode, the step of persistently storing the data stream in the target task stream according to the task configuration information comprises: Determine the target data with the same key value; Determine the generation time of the target data, and determine the target column where the most recently generated target data is located; The data in the target column is used to overwrite the historical data, and after the overwriting, the data associated with the data stream is persisted to the disk.
8. The method for parallel processing of multiple data streams according to claim 1, wherein the task configuration information further includes a degree of parallelism; and the step of determining the configuration information of the target task stream comprises: Get the preset parallelism parameters; or, The number of data sources associated with the target task flow is determined, and the parallelism parameter of the target task flow is configured according to the number of data sources.
9. The method for parallel processing of multiple data streams according to claim 1, wherein after the step of persisting the data stream in the target task stream according to the task configuration information, the method for parallel processing of multiple data streams further comprises: Determine the query type of the data query instruction and the table type of the data table corresponding to the data query instruction; The data in the data table is processed according to the query type and the table type.
10. A device for parallel processing of multiple data streams, comprising: A first processing module is used to determine task configuration information of a target task flow, wherein the target task flow includes at least one data flow, and the task configuration information includes: an association mode between the data flows in the target task flow, data feature information of each data flow in the target task flow, and a data table type of the data flow; A second processing module is used to perform a first type of operation on the data stream in the target task stream according to the task configuration information during the pre-writing batching stage of the target task stream, wherein the first type of operation includes at least one of the following: data deduplication and data association; The third processing module is used to perform persistent disk write processing on the data stream in the target task stream according to the task configuration information during the data write merging stage of the target task stream.
11. The multi-data stream parallel processing device according to claim 10, wherein the association mode comprises one of the following: a normal write mode, a wide table connection mode, and a data stream join (JOIN) mode.
12. The multi-data stream parallel processing device according to claim 10, wherein the first processing module is further used for: Get the preset parallelism parameters; or, The number of data sources associated with the target task flow is determined, and a parallelism parameter of the target task flow is configured according to the number of data sources.
13. The multi-data stream parallel processing device according to claim 11, wherein the second processing module is further configured to: Determine whether there is target data with the same key value in the target task flow; In the case where it is determined that the target data with the same key value exists, determining whether the target data with the same key value belongs to the same data stream; When it is determined that the target data with the same key value belong to the same data stream, updating the data stream to which the target data belongs; When it is determined that the target data with the same key value do not belong to the same data stream, data splicing processing is performed on the data stream to which the target data belongs.
14. The device for parallel processing of multiple data streams according to claim 11, wherein when the association mode is the data stream JOIN mode, the second processing module is further configured to: Determine whether there is target data with the same key value in the target task flow; In the case where it is determined that the target data with the same key value exists, determining whether the target data with the same key value belongs to the same data table; When it is determined that the target data with the same key value corresponds to the same data table, retaining the target data with a later data generation time; When it is determined that the target data with the same key value do not correspond to the same data table, data stream JOIN processing is performed on the data stream corresponding to the data.
15. The multi-data stream parallel processing device according to claim 14, wherein the second processing module is further used for: Determine a JOIN order between the data streams, and associate the data streams according to the JOIN order; When the data stream is associated again according to the JOIN sequence, the association mode is Inner.Join or Right.Join, and if the JoinJOIN association fails, the target data is returned and marked as hidden data; When the data stream is associated again according to the JOIN sequence and the association mode is Left.Join, the target data is returned, and a hidden mark is added to the data in the data stream that does not participate in the association.
16. The device for parallel processing of multiple data streams according to claim 15, wherein when the association mode is a data stream JOIN mode, the third processing module is further configured to: Determine the write merging logic corresponding to the table type information of the data stream, wherein: The write-merge logic includes a write-merge method for a first type of data stream and a second type of data stream, wherein the first type of data stream is a data stream in which the hidden mark exists in the corresponding data, and the second type of data stream is a data stream in which the hidden mark does not exist in the corresponding data; The data stream is persistently written to disk according to the write-merge logic corresponding to the data stream.
17. The apparatus for parallel processing of multiple data streams according to claim 11, wherein when the association mode is a wide table connection mode, the third processing module is further configured to: Determine the target data with the same key value; Determine the generation time of the target data, and determine the target column where the most recently generated target data is located; The data in the target column is used to overwrite the historical data, and after the overwriting, the data associated with the data stream is persisted to the disk.
18. The apparatus for parallel processing of multiple data streams according to claim 10, further comprising a fourth processing module, configured to: Determine the query type of the data query instruction and the table type of the data table corresponding to the data query instruction; The data in the data table is processed according to the query type and the table type.
19. A non-volatile computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, executes the method for parallel processing of multiple data streams as claimed in any one of claims 1 to 9.
20. An electronic device, comprising: A memory and a processor, wherein the memory stores a computer program, and the processor executes the multi-data stream parallel processing method according to any one of claims 1 to 9 when running the computer program.
Citation Information
Patent Citations
Data processing method and device, storage medium and electronic device
CN110727697A
Data processing method and device and computer readable storage medium
CN112765166A
Index updating method and system
CN113468199A
Multi-data-stream parallel processing method and device and nonvolatile storage medium
CN117648342A
Data association query method and apparatus, device, and storage medium
US20230259509A1
Cited By
Task execution method and device, computer equipment, storage medium and program product
CN120295738A