A streaming data batch processing method and device

Through the streaming data batch processing method, the data warehouse performance problem caused by frequent writing in the streaming data storage process is solved, and the unified warehousing of streaming data within the maximum consumption time is realized, which reduces the probability of abnormality in the data warehouse and realizes real-time or quasi-real-time streaming data processing.

CN113722282BActive Publication Date: 2025-05-23CHINA CONSTRUCTION BANK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111015588.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-31
Publication Date
2025-05-23
Estimated Expiration
2041-08-31

AI Technical Summary

Technical Problem

In the process of storing streaming data to a data warehouse, frequent write operations are required, resulting in data warehouse performance problems.

Method used

Through a streaming data batch processing method, the starting message location of the currently unconsumed message in the message partition to be consumed is determined, and the execution time of the next batch job is estimated based on the preset streaming data processing delay, the location of the write message is determined, and the batch job is completed within the maximum consumption time, so as to reduce frequent write operations to the data warehouse.

Benefits of technology

It realizes consumption of at least some messages in the consumption message partition within the maximum consumption time, and uniformly enters the messages consumed within the maximum consumption time, avoids multiple insertion operations in the target data warehouse, effectively reduces the probability of abnormality in the data warehouse, and realizes real-time or quasi-real-time processing of streaming data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113722282B_ABST
    Figure CN113722282B_ABST
Patent Text Reader

Abstract

The present invention discloses a streaming data batch processing method and device, which can consume at least part of the messages in the message partition to be consumed within the maximum consumption time, and uniformly store the messages consumed within the maximum consumption time, that is, it can realize real-time or quasi-real-time processing of at least part of the streaming data in the message partition to be consumed, and it is also unnecessary to perform multiple insertion operations on the target data warehouse during the warehousing operation, thereby effectively reducing the abnormal probability of the target data warehouse.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of streaming data processing, and in particular to a streaming data batch processing method and device. Background Art

[0002] With the development of data processing technology, streaming data processing technology continues to improve.

[0003] Streaming data is a dynamic data set that is infinite in time distribution and quantity. Common streaming data processing scenarios include real-time recommendation and business monitoring. In such scenarios, the value of streaming data decreases over time. It is necessary to process streaming data in real time or quasi-real time and quickly store streaming data in a designated data warehouse to quickly respond to possible business queries.

[0004] Currently, the existing technology may require frequent write operations when storing streaming data in a data warehouse, which may easily lead to performance problems in the data warehouse. Summary of the invention

[0005] In view of the above problems, the present invention provides a streaming data batch processing method and device that overcomes the above problems or at least partially solves the above problems. The technical solution is as follows:

[0006] A streaming data batch processing method, comprising:

[0007] Determine, from the target storage space, the starting message position of the currently unconsumed messages in the to-be-consumed message partition, wherein the messages stored in the to-be-consumed message partition are streaming data;

[0008] Determine the starting message position as the starting consumption position of the current batch job for the currently unconsumed messages in the to-be-consumed message partition;

[0009] Estimate the start time of the next batch job based on the preset streaming data processing delay;

[0010] Determine the message location of the target message written to the to-be-consumed message partition at the start execution time;

[0011] Determine the message position of the target message as the designated final consumption position of the message in the to-be-consumed message partition at the end of the execution of the batch processing job;

[0012] Start executing the batch processing job to consume the currently unconsumed messages in the to-be-consumed message partition;

[0013] During the execution of the batch processing job, when the execution time of the batch processing job reaches the preset maximum consumption time, determining the actual final consumption position of the message in the to-be-consumed message partition by the batch processing job;

[0014] Determine whether the actual final consumption location is consistent with the specified final consumption location. If so, determine that the batch job is executed successfully, and store the consumed messages during the execution of the batch job in the target data warehouse.

[0015] Optionally, after determining that the batch processing job is successfully executed, the method further includes:

[0016] Determine the next position of the actual final consumption position in the to-be-consumed message partition as the starting message position of the currently unconsumed message in the to-be-consumed message partition, and store it in the target storage space;

[0017] Return to the step of determining the starting message position of the currently unconsumed messages in the message partition to be consumed from the target storage space, execute the next batch processing job at the start execution time, and continue to consume the unconsumed messages in the message partition to be consumed.

[0018] Optionally, the method further includes:

[0019] During the execution of the batch processing job, when the consumption position of the messages in the to-be-consumed message partition by the batch processing job reaches the designated final consumption position, the batch processing job is instructed to stop execution.

[0020] Optionally, the starting and executing the batch processing job to consume currently unconsumed messages in the to-be-consumed message partition includes:

[0021] When the number of messages from the starting consumption position to the specified final consumption position in the message partition to be consumed does not exceed a preset threshold, during the execution of this batch job, a consumer is constructed through an RDD partition;

[0022] The messages consumed by the consumer in the message partition to be consumed are saved in the rdd partition.

[0023] Optionally, the starting and executing the batch processing job to consume currently unconsumed messages in the to-be-consumed message partition includes:

[0024] When the number of messages from the starting consumption position to the specified final consumption position in the message partition to be consumed exceeds a preset threshold, all messages from the starting consumption position to the specified final consumption position in the message partition to be consumed are split to obtain multiple message blocks;

[0025] During the execution of the batch job, multiple consumers are constructed through multiple RDD partitions, wherein the RDD partitions correspond to the consumers one by one;

[0026] The messages consumed by each consumer in the message partition to be consumed are saved in the corresponding rdd partition respectively.

[0027] Optionally, storing the consumed messages during the execution of the batch job in a target data warehouse includes:

[0028] Write the consumed messages of the batch job in the to-be-consumed message partition into the cos cache;

[0029] Writing the consumed messages stored in the cos cache into a pre-configured external table, wherein the table structure of the external table is consistent with the table structure of the target data warehouse;

[0030] Insert all consumed messages stored in the external table into the target data warehouse.

[0031] A streaming data batch processing device comprises: a first determining unit, a second determining unit, a first estimating unit, a third determining unit, a fourth determining unit, a first starting unit, a fifth determining unit, a sixth determining unit, a seventh determining unit, and a first storage unit; wherein:

[0032] The first determining unit is configured to execute: determining, from the target storage space, a starting message position of currently unconsumed messages in a to-be-consumed message partition, wherein the messages stored in the to-be-consumed message partition are streaming data;

[0033] The second determining unit is configured to execute: determining the starting message position as the starting consumption position of the current batch job for the currently unconsumed messages in the to-be-consumed message partition;

[0034] The first estimation unit is configured to perform: estimating the start execution time of the next batch processing job based on a preset streaming data processing delay;

[0035] The third determining unit is configured to execute: determining a message position of a target message written into the to-be-consumed message partition at the start execution time;

[0036] The fourth determining unit is configured to execute: determining the message position of the target message as the designated final consumption position of the message in the to-be-consumed message partition at the end of the execution of the batch processing job;

[0037] The first starting unit is configured to execute: starting the execution of the current batch processing job to consume currently unconsumed messages in the to-be-consumed message partition;

[0038] The fifth determining unit is configured to execute: during the execution of the batch processing job, when the execution duration of the batch processing job reaches a preset maximum consumption duration, determining the actual final consumption position of the message in the to-be-consumed message partition by the batch processing job;

[0039] The sixth determining unit is configured to execute: determining whether the actual final consumption location is consistent with the designated final consumption location, and if so, triggering the seventh determining unit;

[0040] The seventh determining unit is configured to execute: determining that the batch processing job is successfully executed;

[0041] The first storage unit is configured to execute: storing the consumed messages during the execution of the batch job in a target data warehouse.

[0042] Optionally, the device further includes: a second storage unit and a first trigger unit; wherein:

[0043] The second storage unit is configured to execute: after determining that the batch job is successfully executed, determine the next position of the actual final consumption position in the to-be-consumed message partition as the starting message position of the currently unconsumed message in the to-be-consumed message partition, and store it in the target storage space;

[0044] The first triggering unit is configured to execute: triggering the first determining unit to execute the next batch processing job at the start execution time, and continuing to consume the unconsumed messages in the to-be-consumed message partition.

[0045] Optionally, the device further includes: a first instruction unit;

[0046] The first instruction unit is configured to execute: during the execution of the batch processing job, when the consumption position of the message in the to-be-consumed message partition by the batch processing job reaches the specified final consumption position, instruct the batch processing job to stop executing.

[0047] Optionally, the first starting unit includes: a first construction unit and a first storage unit;

[0048] The first construction unit is configured to execute: when the number of messages from the starting consumption position to the specified final consumption position in the message partition to be consumed does not exceed a preset threshold, during the execution of the batch job, construct a consumer through an RDD partition;

[0049] The first saving unit is configured to execute: saving the message consumed by the consumer in the to-be-consumed message partition into the rdd partition.

[0050] Optionally, the first starting unit includes: a first splitting unit, a second construction unit and a second storage unit;

[0051] The first splitting unit is configured to execute: when the number of messages from the starting consumption position to the specified final consumption position in the message partition to be consumed exceeds a preset threshold, split all messages from the starting consumption position to the specified final consumption position in the message partition to be consumed to obtain multiple message blocks;

[0052] The second construction unit is configured to execute: during the execution of the batch job, construct multiple consumers through multiple RDD partitions, wherein the RDD partitions correspond to the consumers one by one;

[0053] The second saving unit is configured to execute: saving the messages consumed by each consumer in the to-be-consumed message partition into the corresponding rdd partition respectively.

[0054] Optionally, the first storage unit includes: a first writing unit, a second writing unit and an inserting unit;

[0055] The first writing unit is configured to execute: writing the consumed messages of the batch processing job in the to-be-consumed message partition into the cos cache;

[0056] The second writing unit is configured to execute: writing the consumed message stored in the cos cache into a pre-configured external table, wherein the table structure of the external table is consistent with the table structure of the target data warehouse;

[0057] The inserting unit is configured to execute: inserting all consumed messages stored in the external table into the target data warehouse.

[0058] The streaming data batch processing method and device proposed in this embodiment can determine the starting message position of the currently unconsumed messages in the to-be-consumed message partition from the target storage space, the messages stored in the to-be-consumed message partition are streaming data, the starting message position is determined as the starting consumption position of the currently unconsumed messages in the to-be-consumed message partition for this batch processing job, the start execution time of the next batch processing job is estimated based on the preset streaming data processing delay, the message position of the target message written into the to-be-consumed message partition at the start execution time is determined, the message position of the target message is determined as the designated final consumption position of the message in the to-be-consumed message partition at the end of the execution of this batch processing job, the execution of this batch processing job is started to consume the currently unconsumed messages in the to-be-consumed message partition, during the execution of this batch processing job, when the execution time of this batch processing job reaches the preset maximum consumption time, the actual final consumption position of the message in the to-be-consumed message partition for this batch processing job is determined, and whether the actual final consumption position is consistent with the designated final consumption position is determined, if so, it is determined that the execution of this batch processing job is successful, and the consumed messages during the execution of this batch processing job are stored in the target data warehouse.

[0059] The present invention can consume at least part of the messages in the message partition to be consumed within the maximum consumption time, and uniformly store the messages consumed within the maximum consumption time, that is, it can realize real-time or quasi-real-time processing of at least part of the streaming data in the message partition to be consumed, and there is no need to perform multiple insertion operations on the target data warehouse during the warehousing operation, thereby effectively reducing the abnormal probability of the target data warehouse.

[0060] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0062] Figure 1 A flow chart of a first streaming data batch processing method provided by an embodiment of the present invention is shown;

[0063] Figure 2 A flow chart of a second streaming data batch processing method provided by an embodiment of the present invention is shown;

[0064] Figure 3 A schematic structural diagram of a first streaming data batch processing device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0065] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to enable the scope of the present invention to be fully communicated to those skilled in the art.

[0066] like Figure 1 As shown, this embodiment proposes a first streaming data batch processing method, which may include the following steps:

[0067] S101. Determine the starting message position of currently unconsumed messages in the to-be-consumed message partition from the target storage space, where the messages stored in the to-be-consumed message partition are streaming data;

[0068] The target storage space may be used as a data storage space for storing the location identifier of the starting message location of the currently unconsumed message in the to-be-consumed message partition.

[0069] It should be noted that the present invention does not limit the specific type of the target storage space. For example, the target storage space can be a MySQL database, or a memory or a hard disk.

[0070] The message partition to be consumed may be a partition of a topic in Kafka.

[0071] It should be noted that the message partition to be consumed can store multiple messages arranged in order, that is, streaming data. Specifically, the message partition to be consumed can continuously write streaming data.

[0072] Optionally, the above position identifier may be the sequence number of the message in the message partition to be consumed. It is understandable that the above position identifier may be the offset of the message in Kafka.

[0073] The unconsumed message may be a message that is not currently consumed in the to-be-consumed message partition. The unconsumed message may include one or more messages.

[0074] The starting message position may be the position of a message that is at the front of the sorted list of unconsumed messages.

[0075] S102, determining the starting message position as the starting consumption position of the currently unconsumed messages in the to-be-consumed message partition of this batch job;

[0076] Specifically, the present invention can create a batch processing job, and consume the currently unconsumed messages in the message partition to be consumed through the batch processing job.

[0077] Optionally, the present invention can consume the currently unconsumed messages in the message partition to be consumed through multiple batch processing jobs in an orderly manner. Among them, the present invention can start the execution of the next batch processing job only after the execution of one batch processing job is completed to continue to consume the currently unconsumed messages in the message partition to be consumed. Optionally, when a batch processing job fails to execute, the present invention can first prohibit the start of the next batch processing job, and the technical personnel can check and determine the cause of the processing failure, and perform troubleshooting in time.

[0078] Specifically, the present invention can start from the message located at the starting message position in the message partition to be consumed, and consume the currently unconsumed messages in the message partition to be consumed, thereby avoiding repeated consumption of consumed messages and reducing unnecessary waste of resources.

[0079] It can be understood that the present invention can perform the same operation on the currently unconsumed messages in multiple message partitions to be consumed in the same topic of kafka in parallel by executing batch processing jobs, starting from the messages located at the starting message position in the multiple message partitions to be consumed, and consuming the currently unconsumed messages in each message partition to be consumed. For example, the present invention can first determine the starting message position of the currently unconsumed messages in the first message partition to be consumed, and determine the starting message position of the currently unconsumed messages in the second message partition to be consumed, and then consume the unconsumed messages in the first message partition to be consumed and the second message partition to be consumed in parallel by executing batch processing jobs. Specifically, the batch processing job can start from the message located at the starting message position in the first message partition to be consumed, and consume the currently unconsumed messages in the first message partition to be consumed. The batch processing job can also start from the message located at the starting message position in the second message partition to be consumed, and consume the currently unconsumed messages in the second message partition to be consumed.

[0080] At this time, the target storage space can store the partition identifier of the message partition to be consumed and the position identifier of the starting message position of the currently unconsumed message in the message partition to be consumed. When executing a batch job, the present invention can first obtain the starting message position of the currently unconsumed message in each message partition to be consumed from the target storage space.

[0081] It can be understood that the present invention can consume messages in all message partitions to be consumed in the same topic of Kafka in parallel through batch processing jobs, that is, process the streaming data in all message partitions to be consumed in the same topic of Kafka in parallel, so as to improve the real-time processing rate of streaming data.

[0082] S103, estimating the start execution time of the next batch processing job based on the preset streaming data processing delay;

[0083] The streaming data processing delay may be the delay from the generation of streaming data to the time when the business can query the corresponding data mapped indicators. It should be noted that the streaming data processing delay may be the specified interval between the start execution times of two batch processing jobs.

[0084] It should be noted that the streaming data processing delay can be set by a technician or a user according to actual needs, and the present invention does not limit this. Optionally, the streaming data processing delay can be a delay of minutes or more.

[0085] Specifically, the current time is added with the streaming data processing delay, and the obtained time can be used as the estimated start execution time of the next batch processing job.

[0086] It should be noted that the value obtained by adding the streaming data processing delay to the current time is defined as the estimated start execution time of the next batch job, rather than the actual start execution time. This is because there may be a delay when the batch job is started. There may be a difference between the value obtained by adding the streaming data processing delay to the current time and the actual start execution time of the next batch job.

[0087] S104, determining the message location of the target message written by the to-be-consumed message partition at the time of starting execution;

[0088] The target message may be a message that Kafka writes to the message partition to be consumed when the execution is started.

[0089] Specifically, the present invention can determine the message position of the target message written by Kafka to the message partition to be consumed at the estimated startup execution time by calling the Kafka native API consumer.offsetsForTimes.

[0090] S105, determining the message position of the target message as the designated final consumption position of the message in the consumption message partition at the end of the execution of this batch job;

[0091] The specified final consumption position may be a final consumption position to be reached during the process of consuming the messages in the message partition to be consumed by this batch job.

[0092] It can be understood that, in the process of consuming the messages in the message partition to be consumed in this batch job, the finally consumed message may be the message at the designated final consumption position.

[0093] S106, start executing this batch processing job to consume the currently unconsumed messages in the message partition to be consumed;

[0094] Optionally, the present invention can actually start executing this batch processing job after determining the designated final consumption position, and consume the currently unconsumed messages in the message partition to be consumed.

[0095] Optionally, the present invention may also start executing the batch processing job simultaneously during the process of specifying the final consumption location.

[0096] S107: During the execution of this batch processing job, when the execution time of this batch processing job reaches the preset maximum consumption time, determine the actual final consumption position of the message in the message partition to be consumed by this batch processing job;

[0097] The maximum consumption time may be the maximum consumption time of the messages in the message partition to be consumed by this batch job. It can be understood that the maximum consumption time may be the specified maximum polling time of Kafka.

[0098] It should be noted that the maximum consumption time can be set by technical personnel according to actual working conditions, and the present invention does not limit this.

[0099] The actual final consumption position may be the position of the last message in the messages that have been consumed in the message partition to be consumed when the batch job is finished executing.

[0100] Specifically, the present invention can determine the actual final consumption position of the message in the message partition to be consumed by the batch processing job when the execution time of the batch processing job reaches the maximum consumption time.

[0101] S108. Determine whether the actual final consumption location is consistent with the specified final consumption location. If so, execute step S109; otherwise, prohibit execution of step S109 to avoid unnecessary resource consumption.

[0102] Specifically, during the execution of this batch job, when the execution time of this batch job reaches the maximum consumption time, if the consumption quantity of messages in the message partition to be consumed by this processing job reaches the expectation, that is, the actual final consumption position is consistent with the actual final consumption position, then it can be determined that this batch job is executed successfully.

[0103] Specifically, during the execution of this batch job, when the execution time of this batch job reaches the maximum consumption time, if the consumption quantity of messages in the message partition to be consumed by this processing job does not reach the expected amount, that is, the actual final consumption position does not reach the specified final consumption position, then it can be determined that this batch job has failed. At this time, the present invention can determine that an error has occurred in the execution of this batch job, and can prohibit the execution of subsequent steps, and perform error checking, so as to timely perform troubleshooting to avoid causing greater losses to business operations.

[0104] It should be noted that in the initial job design, when the polling time ends, that is, when the consumption time reaches the maximum consumption time, if no message is consumed, it is considered that there is no message to be consumed. However, the present invention adds a judgment of the actual final consumption position and the specified final consumption position to determine whether a sufficient number of messages have been consumed. If not, the batch job can be determined as an execution failure to prevent users from being unaware when the messages to be consumed are lost.

[0105] Optionally, in order to cope with the busy scenario of Kafka, the present invention can increase the maximum consumption time to prevent data loss or frequent failure of batch jobs; optionally, the present invention can also improve the execution efficiency of batch jobs by reducing the maximum consumption time.

[0106] S109: Determine that the batch processing job is successfully executed, and store the consumed messages during the execution of the batch processing job in the target data warehouse.

[0107] The target data warehouse may be a data warehouse for storing streaming data.

[0108] Specifically, when determining that the execution of this batch processing job is successful, the present invention can store the consumed messages of the message partition to be consumed during the execution of this batch processing job in the target data warehouse, thereby realizing the unified warehousing of the streaming data consumed in this batch without performing multiple insertion operations on the target data warehouse, effectively reducing the probability of abnormalities in the target data warehouse.

[0109] Optionally, when determining that the execution of this batch job has failed, the present invention may first prohibit the storage of consumed messages during the execution of this batch job to the target data warehouse, thereby preventing message loss and avoiding erroneous streaming data from flowing into the target data warehouse, thereby avoiding the impact of erroneous data on business operations.

[0110] It should be noted that the present invention Figure 1In steps S101 to S109 shown, at least part of the messages in the message partition to be consumed can be consumed within the maximum consumption time, and the messages consumed within the maximum consumption time can be uniformly warehoused, that is, real-time or quasi-real-time processing of at least part of the streaming data in the message partition to be consumed can be achieved, and there is no need to perform multiple insertion operations on the target data warehouse during the warehousing operation, thereby effectively reducing the probability of abnormalities in the target data warehouse.

[0111] Optionally, the above method may further include:

[0112] During the execution of this batch processing job, when the consumption position of the message in the message partition to be consumed by this batch processing job reaches the specified final consumption position, the batch processing job is instructed to stop executing.

[0113] Specifically, during the execution of this batch processing job, if the execution time of this batch processing job has not reached the maximum consumption time, that is, the message partition to be consumed has been consumed to the specified final consumption position, then the present invention can end this batch processing job and stop the message consumption of the message partition to be consumed by this batch processing job, thereby avoiding unnecessary resource consumption and improving processing efficiency.

[0114] Specifically, the present invention calls the Kafka native API consumer.seek to implement consumption control of messages in the message partition to be consumed. When the batch processing job consumes the expected number of messages, that is, consumes to the specified final consumption position, the consumption is terminated in advance.

[0115] The streaming data batch processing method proposed in this embodiment can determine the starting message position of the currently unconsumed messages in the to-be-consumed message partition from the target storage space, the messages stored in the to-be-consumed message partition are streaming data, the starting message position is determined as the starting consumption position of the currently unconsumed messages in the to-be-consumed message partition for this batch processing job, the start execution time of the next batch processing job is estimated based on the preset streaming data processing delay, the message position of the target message written into the to-be-consumed message partition at the start execution time is determined, the message position of the target message is determined as the designated final consumption position of the message in the to-be-consumed message partition at the end of the execution of this batch processing job, the execution of this batch processing job is started to consume the currently unconsumed messages in the to-be-consumed message partition, during the execution of this batch processing job, when the execution time of this batch processing job reaches the preset maximum consumption time, the actual final consumption position of the message in the to-be-consumed message partition for this batch processing job is determined, and it is determined whether the actual final consumption position is consistent with the designated final consumption position. If so, it is determined that the execution of this batch processing job is successful, and the consumed messages during the execution of this batch processing job are stored in the target data warehouse. The present invention can consume at least part of the messages in the message partition to be consumed within the maximum consumption time, and uniformly store the messages consumed within the maximum consumption time, that is, it can realize real-time or quasi-real-time processing of at least part of the streaming data in the message partition to be consumed, and there is no need to perform multiple insertion operations on the target data warehouse during the warehousing operation, thereby effectively reducing the abnormal probability of the target data warehouse.

[0116] based on Figure 1 The steps shown are as follows: Figure 2 As shown, this embodiment proposes a second streaming data batch processing method. After the above step S109, the method may further include:

[0117] S201, determining the next position of the actual final consumption position in the to-be-consumed message partition as the starting message position of the currently unconsumed messages in the to-be-consumed message partition, and storing it in the target storage space;

[0118] The next position of the actual final consumption position may be the arrangement position of the message that is arranged first among the currently unconsumed messages in the to-be-consumed message partition after the execution of this batch job is completed.

[0119] Specifically, the present invention can implement the storage of the next position of the actual final consumption position in the target storage space by storing the position identifier of the next position of the actual final consumption position in the target storage space.

[0120] It should be noted that if the batch processing job fails to execute, the present invention can determine that an abnormality occurs in the processing process, prohibit starting the next batch processing job, check the abnormality and perform troubleshooting to avoid causing greater losses to business operations.

[0121] S202, return to the step of determining the starting message position of the currently unconsumed messages in the message partition to be consumed from the target storage space, execute the next batch processing job at the start execution time, and continue to consume the unconsumed messages in the message partition to be consumed.

[0122] Specifically, the present invention can store the arrangement position of the message ranked first among the currently unconsumed messages in the message partition to be consumed in the target storage space after determining that the current batch processing job has been executed successfully, so that when the next batch processing job is started, the next batch processing job can obtain the arrangement position from the target storage space, and start from the arrangement position to consume the currently unconsumed messages in the message partition to be consumed, thereby effectively avoiding repeated consumption of messages in the message partition to be consumed.

[0123] It should be noted that the present invention can start the next batch processing job at the time of starting the execution.

[0124] It can be understood that the present invention can start a batch processing job every time period corresponding to the streaming data processing delay. In this way, by orderly starting multiple batch processing jobs, the consumption of messages in the consumer message partition can be continuously treated, thereby effectively realizing real-time or quasi-real-time processing of streaming data.

[0125] The streaming data batch processing method proposed in this embodiment can continuously consume messages in the consumer message partition by orderly starting multiple batch processing jobs, thereby effectively realizing real-time or quasi-real-time processing of streaming data.

[0126] based on Figure 1 The steps shown in this embodiment propose a third streaming data batch processing method. In this method, step S106 may include:

[0127] When the number of messages from the starting consumption position to the specified final consumption position in the message partition to be consumed does not exceed the preset threshold, a consumer is constructed through one RDD partition during the execution of this batch job;

[0128] Save the messages consumed by the consumer in the message partition to be consumed into the RDD partition.

[0129] Specifically, in the process of creating a batch job, the present invention can submit a computing resource application for executing a Spark task to the resource manager of the shared YARN cluster through a local YARN client. After confirmation by the resource manager, the present invention can submit the application information (the application information can include the user-specified Kafka topic, MPP table information and the corresponding relationship JSON string between the two, the jar package and configuration file on which the Spark job depends, etc.) to the YARN client. At this time, the YARN cluster can allocate the necessary computing resources according to the application information submitted by the YARN client, generate a driver and multiple executor instances, and use them to execute the specified Spark task in a distributed manner and execute the batch job.

[0130] Specifically, after the spark task (i.e., this batch job) is started, it can obtain the actual final consumption position of all partitions (i.e., the partitions of messages to be consumed) of the target topic written to the MySQL database after the last batch job is completed from the MySQL database (i.e., the target storage space) that stores the Kafka offset, and determine it as the start offset of the current batch job for the consumption of the currently unconsumed messages in each partition. After that, the start execution time of the next batch job can be estimated based on the streaming data processing delay, and with this as a parameter, the Kafka native api consumer.offsetsForTimes is called to obtain the offset of the message that Kafka has recently written to each partition when the next batch job is started, and determine it as the specified end offset (i.e., the specified final consumption position) for the consumption of each partition of the target topic by this batch job.

[0131] After obtaining the start offset and end offset of all partitions of the target topic in this batch job, the saprk task can construct a consumer based on the number of partitions of the target topic and the amount of data in each partition. When the number of messages in a single partition is small, one consumer can consume one topic partition.

[0132] Among them, consumers can be constructed by spark's rdd partitions, and one rdd partition can correspond to one consumer.

[0133] Specifically, the present invention can save the messages consumed by each consumer in the corresponding partition into the corresponding rdd partition. For example, the present invention can save the messages consumed by the first consumer in the corresponding first partition into the corresponding first rdd partition, and can save the messages consumed by the second consumer in the corresponding second partition into the corresponding second rdd partition.

[0134] Optionally, in the fourth streaming data batch processing method proposed in this embodiment, step S106 may include:

[0135] When the number of messages from the starting consumption position to the specified final consumption position in the message partition to be consumed exceeds a preset threshold, all messages from the starting consumption position to the specified final consumption position in the message partition to be consumed are split to obtain multiple message blocks;

[0136] During the execution of this batch job, multiple consumers are constructed through multiple RDD partitions, where RDD partitions correspond to consumers one by one;

[0137] The messages consumed by each consumer in the message partition to be consumed are saved in the corresponding rdd partition.

[0138] Among them, when the number of messages in a partition of the target topic exceeds a preset threshold, the unconsumed messages in the partition can be split and consumed by multiple consumers. At this time, Spark can still construct consumers through RDD partitions, and one RDD partition still corresponds to one consumer. The number of RDD partitions constructed by Spark can be more than the number of topic partitions. At this time, the present invention can prevent excessive pressure on the memory when the message to be consumed by a single consumer is too large. Optionally, multiple Spark tasks can consume a partition concurrently to improve consumption efficiency.

[0139] It should be noted that Spark task can call Kafka native API poll to implement the allocation and consumption of consumer partitions.

[0140] Optionally, the Spark task can call the Kafka native API consumer.seek to achieve precise control of the start offset of the consumption partition. When the expected number of messages is consumed in a partition, the consumption of the partition can be ended in advance. The maximum Kafka polling time (i.e., the longest consumption time) can be specified through the input parameter of poll. When the total Kafka polling time reaches the maximum Kafka polling time, and the number of consumed messages in a partition has not reached the expected number, that is, the actual end offset of the partition is inconsistent with the specified end offset, it is determined that the execution of this Spark task has failed to ensure that no messages are lost. The poll parameter can be dynamically adjusted according to the busyness of the Kafka cluster and the amount of data pulled in a single batch, in order to ensure batch data consistency with high efficiency.

[0141] It should be noted that although the upper limit of the number of messages consumed by a single RDD partition has been limited, the spark executor is still at risk of OOM when a single message is large. In this case, you can consider increasing the executor memory during the task configuration phase and setting the single executor behavior to memory priority so that only one RDD partition task is running on it at the same time.

[0142] Optionally, in the fifth streaming data batch processing method proposed in this embodiment, the above step S109 may include:

[0143] Determine that the batch job is executed successfully, and write the consumed messages in the waiting-for-consumption message partition of the batch job into the cos cache;

[0144] Write the consumed messages saved in the cos cache into a pre-configured external table. The table structure of the external table is consistent with that of the target data warehouse.

[0145] Insert all consumed messages saved in the external table into the target data warehouse.

[0146] Specifically, the present invention can store the consumed messages of this batch processing job in the target data warehouse in sequence through the cos cache and the external table during the execution of this batch processing job.

[0147] It can be understood that in the third and fourth streaming data batch processing methods, after the consumption of a partition by this batch processing job is completed, the data of each corresponding RDD partition can be cached in the memory. At this time, the present invention can write the RDD partition data into the COS cache by calling the Spark native API dataFrame.write method. Spark can use the hadoop-cos related dependencies of Apache Hadoop to write to the COS cache.

[0148] Specifically, after all partition messages are written to the cos cache, the spark task can connect to the big data mpp through standard jdbc, use create external table to create an external table with the same table structure as the target data warehouse, and specify the data file directory of the external table as the cos file directory generated in the previous step. After that, you can use the insert in select syntax to insert all external table data into the target data warehouse. After the insertion is complete, you can drop the external table and call hadoop-cos api FileSystem.delete to delete the cos file directory generated in the previous step. If the jdbc sql statement fails to execute, it can be determined that the task has failed. At this time, you also need to delete the cos directory to avoid the accumulation of useless data.

[0149] It should be noted that the present invention consumes messages in the message queue in a batch processing manner, calculates the consumption end position of each batch of jobs according to the loss data processing delay, and can terminate consumption in time when spark consumes to a fixed message position, thereby ensuring the timeliness of the consumption message queue during the batch processing process.

[0150] Among them, the present invention can further improve the efficiency of consuming message queues in batch processing by limiting the amount of data to be consumed by each consumer and using consumers equal to or more than the number of message queue message partitions to concurrently consume message queue messages.

[0151] Among them, in each partition of the target topic, the latest consumption position is independent of the message queue and is stored in a separate target storage space, which can ensure the integrity of consumption data between multiple batches; when a single batch job fails, since the starting displacement is known, the failed job can be repeated without loss of data, which improves the overall reliability of batch processing.

[0152] Among them, the present invention can create an external table and use the object storage transfer data file shared by spark and the data warehouse to realize batch data warehousing, further improving the batch job execution efficiency.

[0153] It should be noted that the existing technology can use stream computing to quickly insert data into a specified data warehouse. Taking stream computing using the flink engine as an example, users can first submit the Flink program to the JobClient, which is then processed, parsed, and optimized by the JobClient and submitted to the JobManager. The JobManager applies for resources from YARN. After YARN allocates resources, the JobManager starts the TaskManager on the corresponding node. The TaskManager starts to start tasks and synchronizes the execution status with the JobManager on a regular basis. The streaming data processed by the TaskManager can be output through the Flink Sink. The Sink is mainly responsible for the output and persistence of real-time calculation results. For example, writing data streams to standard output, files, sockets, and external systems. The Sink capability of Flink mainly implements external storage of data streams by calling the write-related API of the data stream and DataStream.addSink. AddSink supports the standard JDBC interface for writing streaming data into a relational database. If the selected data warehouse provides a JDBC driver, data can be directly streamed into the warehouse through addSink. However, some data warehouses do not have native JDBC interfaces, which increases the difficulty of user adaptation. At the same time, although Flink has provided several implemented Sink Functions, you can also implement a custom Sink by implementing SinkFunction and inheriting RichOutputFormat to call the data warehouse native interface.

[0154] However, the streaming computing engine needs to frequently insert data into the selected data warehouse, which may cause certain performance problems in the data warehouse. For example, if too much data is inserted at a time in the put method of HBase, it may cause the region too busy exception; for example, frequent inserts in Longfu MPP may cause the table to be locked and further insertions may be impossible; for example, Hive submitting data line by line may cause the underlying HDFS to generate too many small files, thus affecting performance. Since it is impossible to eliminate the various impacts of the frequent writing characteristics of streaming data on the data warehouse, the business selection range of data warehouses is limited.

[0155] In addition, the existing technology can also add a transit relational database such as MySQL as the direct data storage of the flink sink, and synchronize the relational database data to the final data warehouse through scheduled batch processing. However, at this time, some detailed queries and multi-dimensional correlation analysis need to be performed in the data warehouse, so the response efficiency of the system depends on the speed of MySQL data entry. Since the streaming data needs to be inserted into MySQL after being processed by flink, and then wait for a batch of data to be accumulated before executing data entry, the lengthening of the data processing link reduces the timeliness of processing streaming data.

[0156] The present invention can consume Kafka based on Spark, and use batch processing to load each batch of streaming data into the selected data warehouse at one time, avoiding the situation where the data warehouse performance is reduced or unavailable due to continuous insertion of data when using the stream computing engine to process streaming data. Moreover, the present invention can ensure the continuity and reliability of batch consumption data by independently saving consumption displacement. And the data file can be transferred through object storage to realize batch data warehousing, shorten the single batch processing time, obtain quasi-real-time streaming data batch processing effect, realize streaming data batch processing of Kafka, and complete the unified warehousing of consumption data at a frequency of minutes, avoiding frequent insert data.

[0157] The streaming data batch processing method proposed in this embodiment can effectively implement batch processing of streaming data.

[0158] and Figure 1 The steps shown correspond to Figure 3 As shown, this embodiment proposes a first streaming data batch processing device, which may include: a first determining unit 101, a second determining unit 102, a first estimating unit 103, a third determining unit 104, a fourth determining unit 105, a first starting unit 106, a fifth determining unit 107, a sixth determining unit 108, a seventh determining unit 109, and a first storage unit 110; wherein:

[0159] The first determining unit 101 is configured to execute: determining, from the target storage space, the starting message position of the currently unconsumed message in the to-be-consumed message partition, where the message stored in the to-be-consumed message partition is streaming data;

[0160] The target storage space may be used as a data storage space for storing the location identifier of the starting message location of the currently unconsumed message in the to-be-consumed message partition.

[0161] It should be noted that the present invention does not limit the specific type of the target storage space. For example, the target storage space can be a MySQL database, or a memory or a hard disk.

[0162] The message partition to be consumed may be a partition of a topic in Kafka.

[0163] It should be noted that the message partition to be consumed can store multiple messages arranged in an orderly manner, that is, streaming data. Specifically, the message partition to be consumed can continuously write streaming data.

[0164] Optionally, the above position identifier may be the sequence number of the message in the message partition to be consumed. It is understandable that the above position identifier may be the offset of the message in Kafka.

[0165] The unconsumed message may be a message that is not currently consumed in the to-be-consumed message partition. The unconsumed message may include one or more messages.

[0166] The starting message position may be the position of a message that is at the front of the sorted list of unconsumed messages.

[0167] The second determining unit 102 is configured to execute: determining the starting message position as the starting consumption position of the currently unconsumed messages in the to-be-consumed message partition of this batch job;

[0168] Specifically, the present invention can create a batch processing job, and consume the currently unconsumed messages in the message partition to be consumed through the batch processing job.

[0169] Optionally, the present invention can consume the currently unconsumed messages in the message partition to be consumed through multiple batch processing jobs in an orderly manner. Among them, the present invention can start the execution of the next batch processing job only after the execution of one batch processing job is completed to continue to consume the currently unconsumed messages in the message partition to be consumed. Optionally, when a batch processing job fails to execute, the present invention can first prohibit the start of the next batch processing job, and the technical personnel can check and determine the cause of the processing failure, and perform troubleshooting in time.

[0170] Specifically, the present invention can start from the message located at the starting message position in the message partition to be consumed, and consume the currently unconsumed messages in the message partition to be consumed, thereby avoiding repeated consumption of consumed messages and reducing unnecessary waste of resources.

[0171] It can be understood that the present invention can perform batch processing jobs to perform the same operation on the currently unconsumed messages in multiple message partitions to be consumed in the same topic of Kafka in parallel, starting from the message located at the starting message position in the multiple message partitions to be consumed, and consuming the currently unconsumed messages in each message partition to be consumed.

[0172] It can be understood that the present invention can consume messages in all message partitions to be consumed in the same topic of Kafka in parallel through batch processing jobs, that is, process the streaming data in all message partitions to be consumed in the same topic of Kafka in parallel, so as to improve the real-time processing rate of streaming data.

[0173] The first estimation unit 103 is configured to perform: estimating the start execution time of the next batch processing job based on the preset streaming data processing delay;

[0174] The streaming data processing delay may be the delay from the generation of streaming data to the time when the business can query the corresponding data mapped indicators. It should be noted that the streaming data processing delay may be the specified interval between the start execution times of two batch processing jobs.

[0175] It should be noted that the streaming data processing delay can be set by a technician or a user according to actual needs, and the present invention does not limit this. Optionally, the streaming data processing delay can be a delay of minutes or more.

[0176] Specifically, the current time is added with the streaming data processing delay, and the obtained time can be used as the estimated start execution time of the next batch processing job.

[0177] It should be noted that the value obtained by adding the streaming data processing delay to the current time is defined as the estimated start execution time of the next batch job, rather than the actual start execution time. This is because there may be a delay when the batch job is started. There may be a difference between the value obtained by adding the streaming data processing delay to the current time and the actual start execution time of the next batch job.

[0178] The third determining unit 104 is configured to perform: determining a message position of a target message written into the to-be-consumed message partition at the start execution time;

[0179] The target message may be a message that Kafka writes to the message partition to be consumed when the execution is started.

[0180] Specifically, the present invention can determine the message position of the target message written by Kafka to the message partition to be consumed at the estimated startup execution time by calling the Kafka native API consumer.offsetsForTimes.

[0181] The fourth determining unit 105 is configured to execute: determining the message position of the target message as the designated final consumption position of the message in the message partition to be consumed when the batch job is finished executing;

[0182] The specified final consumption position may be a final consumption position to be reached during the process of consuming the messages in the message partition to be consumed by this batch job.

[0183] It can be understood that, in the process of consuming the messages in the message partition to be consumed in this batch job, the finally consumed message may be the message at the designated final consumption position.

[0184] The first starting unit 106 is configured to execute: starting the execution of this batch processing job to consume the currently unconsumed messages in the message partition to be consumed;

[0185] Optionally, the present invention can actually start executing this batch processing job after determining the designated final consumption position, and consume the currently unconsumed messages in the message partition to be consumed.

[0186] Optionally, the present invention may also start executing the batch processing job simultaneously during the process of specifying the final consumption location.

[0187] The fifth determining unit 107 is configured to execute: during the execution of the current batch processing job, when the execution time of the current batch processing job reaches the preset maximum consumption time, determine the actual final consumption position of the message in the message partition to be consumed by the current batch processing job;

[0188] The maximum consumption time may be the maximum consumption time of the messages in the message partition to be consumed by this batch job. It can be understood that the maximum consumption time may be the specified maximum polling time of Kafka.

[0189] It should be noted that the maximum consumption time can be set by technical personnel according to actual working conditions, and the present invention does not limit this.

[0190] The actual final consumption position may be the position of the last message in the messages that have been consumed in the message partition to be consumed when the batch job is finished executing.

[0191] Specifically, the present invention can determine the actual final consumption position of the message in the message partition to be consumed by the batch processing job when the execution time of the batch processing job reaches the maximum consumption time.

[0192] The sixth determining unit 108 is configured to execute: determining whether the actual final consumption location is consistent with the specified final consumption location, and if so, triggering the seventh determining unit 109;

[0193] Specifically, during the execution of this batch job, when the execution time of this batch job reaches the maximum consumption time, if the consumption quantity of messages in the message partition to be consumed by this processing job reaches the expectation, that is, the actual final consumption position is consistent with the actual final consumption position, then it can be determined that this batch job is executed successfully.

[0194] Specifically, during the execution of this batch job, when the execution time of this batch job reaches the maximum consumption time, if the consumption quantity of messages in the message partition to be consumed by this processing job does not reach the expected amount, that is, the actual final consumption position does not reach the specified final consumption position, then it can be determined that this batch job has failed. At this time, the present invention can determine that an error has occurred in the execution of this batch job, and can prohibit the execution of subsequent processes, perform error checking, and perform troubleshooting in a timely manner to avoid causing greater losses to business operations.

[0195] It should be noted that in the initial job design, when the polling time ends, that is, when the consumption time reaches the maximum consumption time, if no message is consumed, it is considered that there is no message to be consumed. However, the present invention adds a judgment of the actual final consumption position and the specified final consumption position to determine whether a sufficient number of messages have been consumed. If not, the batch job can be determined as an execution failure to prevent users from being unaware when the messages to be consumed are lost.

[0196] Optionally, in order to cope with the busy scenario of Kafka, the present invention can increase the maximum consumption time to prevent data loss or frequent failure of batch jobs; optionally, the present invention can also improve the execution efficiency of batch jobs by reducing the maximum consumption time.

[0197] The seventh determining unit 109 is configured to execute: determining that the batch processing job is executed successfully;

[0198] The first storage unit 110 is configured to execute: storing the consumed messages during the execution of the batch job in the target data warehouse.

[0199] The target data warehouse may be a data warehouse for storing streaming data.

[0200] Specifically, when determining that the execution of this batch processing job is successful, the present invention can store the consumed messages of the message partition to be consumed during the execution of this batch processing job in the target data warehouse, thereby realizing the unified warehousing of the streaming data consumed in this batch without performing multiple insertion operations on the target data warehouse, effectively reducing the probability of abnormalities in the target data warehouse.

[0201] Optionally, when determining that the execution of this batch job has failed, the present invention may first prohibit the storage of consumed messages during the execution of this batch job to the target data warehouse, thereby preventing message loss and avoiding erroneous streaming data from flowing into the target data warehouse, thereby avoiding the impact of erroneous data on business operations.

[0202] It should be noted that the present invention can consume at least part of the messages in the message partition to be consumed within the maximum consumption time, and uniformly store the messages consumed within the maximum consumption time, that is, it can realize real-time or quasi-real-time processing of at least part of the streaming data in the message partition to be consumed, and there is no need to perform multiple insertion operations on the target data warehouse during the warehousing operation, thereby effectively reducing the probability of abnormalities in the target data warehouse.

[0203] Optionally, the above device may further include: a first instruction unit;

[0204] The first instruction unit is configured to execute: during the execution of this batch processing job, when the consumption position of the message in the message partition to be consumed by this batch processing job reaches the specified final consumption position, instruct this batch processing job to stop executing.

[0205] Specifically, during the execution of this batch processing job, if the execution time of this batch processing job has not reached the maximum consumption time, that is, the message partition to be consumed has been consumed to the specified final consumption position, then the present invention can end this batch processing job and stop the message consumption of the message partition to be consumed by this batch processing job, thereby avoiding unnecessary resource consumption and improving processing efficiency.

[0206] The streaming data batch processing device proposed in this embodiment can consume at least part of the messages in the message partition to be consumed within the maximum consumption time, and uniformly store the messages consumed within the maximum consumption time, that is, it can realize real-time or quasi-real-time processing of at least part of the streaming data in the message partition to be consumed, and there is no need to perform multiple insertion operations on the target data warehouse during the warehousing operation, thereby effectively reducing the abnormal probability of the target data warehouse.

[0207] based on Figure 3 This embodiment proposes a second streaming data batch processing device. The device may also include: a second storage unit and a first trigger unit; wherein:

[0208] The second storage unit is configured to execute: after determining that the batch processing job is successfully executed, determine the next position of the actual final consumption position in the to-be-consumed message partition as the starting message position of the currently unconsumed message in the to-be-consumed message partition, and store it in the target storage space;

[0209] The first triggering unit is configured to execute: triggering the first determining unit 101 to execute the next batch processing job at the start execution time, and continue to consume the unconsumed messages in the to-be-consumed message partition.

[0210] The next position of the actual final consumption position may be the arrangement position of the message that is arranged first among the currently unconsumed messages in the to-be-consumed message partition after the execution of this batch job is completed.

[0211] Specifically, the present invention can implement the storage of the next position of the actual final consumption position in the target storage space by storing the position identifier of the next position of the actual final consumption position in the target storage space.

[0212] It should be noted that if the batch processing job fails to execute, the present invention can determine that an abnormality occurs in the processing process, prohibit starting the next batch processing job, check the abnormality and perform troubleshooting to avoid causing greater losses to business operations.

[0213] Specifically, the present invention can store the arrangement position of the message ranked first among the currently unconsumed messages in the message partition to be consumed in the target storage space after determining that the current batch processing job has been executed successfully, so that when the next batch processing job is started, the next batch processing job can obtain the arrangement position from the target storage space, and start from the arrangement position to consume the currently unconsumed messages in the message partition to be consumed, thereby effectively avoiding repeated consumption of messages in the message partition to be consumed.

[0214] It should be noted that the present invention can start the next batch processing job at the time of starting the execution.

[0215] It can be understood that the present invention can start a batch processing job every time period corresponding to the streaming data processing delay. In this way, by orderly starting multiple batch processing jobs, the consumption of messages in the consumer message partition can be continuously treated, thereby effectively realizing real-time or quasi-real-time processing of streaming data.

[0216] The streaming data batch processing device proposed in this embodiment can continuously consume messages in the consumer message partition by orderly starting multiple batch processing jobs, thereby effectively realizing real-time or quasi-real-time processing of streaming data.

[0217] based on Figure 3 This embodiment proposes a third streaming data batch processing device. In the device, the first starting unit 106 includes: a first construction unit and a first storage unit;

[0218] The first construction unit is configured to execute: when the number of messages from the starting consumption position to the specified final consumption position in the message partition to be consumed does not exceed a preset threshold, during the execution of this batch job, construct a consumer through an RDD partition;

[0219] The first saving unit is configured to execute: saving the message consumed by the consumer in the to-be-consumed message partition into the rdd partition.

[0220] Among them, consumers can be constructed by spark's rdd partitions, and one rdd partition can correspond to one consumer.

[0221] Specifically, the present invention can save the messages consumed by each consumer in the corresponding partition into the corresponding rdd partition.

[0222] Optionally, in the fourth streaming data batch processing device proposed in this embodiment, the first starting unit 106 includes: a first splitting unit, a second construction unit and a second saving unit;

[0223] The first splitting unit is configured to execute: when the number of messages from the starting consumption position to the specified final consumption position in the message partition to be consumed exceeds a preset threshold, split all messages from the starting consumption position to the specified final consumption position in the message partition to be consumed to obtain multiple message blocks;

[0224] The second construction unit is configured to execute: during the execution of this batch job, construct multiple consumers through multiple RDD partitions, wherein the RDD partitions correspond to the consumers one by one;

[0225] The second saving unit is configured to execute: saving the messages consumed by each consumer in the to-be-consumed message partition into corresponding rdd partitions respectively.

[0226] Among them, when the number of messages in a partition of the target topic exceeds a preset threshold, the unconsumed messages in the partition can be split and consumed by multiple consumers. At this time, Spark can still construct consumers through RDD partitions, and one RDD partition still corresponds to one consumer. The number of RDD partitions constructed by Spark can be more than the number of topic partitions. At this time, the present invention can prevent excessive pressure on the memory when the message to be consumed by a single consumer is too large. Optionally, multiple Spark tasks can consume a partition concurrently to improve consumption efficiency.

[0227] It should be noted that Spark task can call Kafka native API poll to implement the allocation and consumption of consumer partitions.

[0228] Optionally, in the fifth streaming data batch processing device proposed in this embodiment, the first storage unit 110 includes: a first writing unit, a second writing unit and an inserting unit;

[0229] The first writing unit is configured to execute: writing the consumed messages in the to-be-consumed message partition of the current batch job into the cos cache;

[0230] The second writing unit is configured to execute: writing the consumed messages stored in the cos cache into a pre-configured external table, where the table structure of the external table is consistent with the table structure of the target data warehouse;

[0231] The insertion unit is configured to execute: inserting all consumed messages stored in the external table into the target data warehouse.

[0232] Specifically, the present invention can store the consumed messages of this batch processing job in the target data warehouse in sequence through the cos cache and the external table during the execution of this batch processing job.

[0233] It can be understood that in the third and fourth streaming data batch processing devices, after the consumption of a partition by this batch processing job is completed, the data of the corresponding rdd partitions can be cached in the memory. At this time, the present invention can write the rdd partition data into the cos cache by calling the spark native api dataFrame.write. Spark can use the hadoop-cos related dependencies of Apache Hadoop to write to the cos cache.

[0234] The streaming data batch processing proposed in this embodiment can effectively implement batch processing of streaming data.

[0235] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0236] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A streaming data batch processing method, It is characterized in that include: Determine, from the target storage space, the starting message position of the currently unconsumed messages in the to-be-consumed message partition, wherein the messages stored in the to-be-consumed message partition are streaming data; Determine the starting message position as the starting consumption position of the current batch job for the currently unconsumed messages in the to-be-consumed message partition; Estimate the start time of the next batch job based on the preset streaming data processing delay; Determine the message location of the target message written to the to-be-consumed message partition at the start execution time; Determine the message position of the target message as the designated final consumption position of the message in the to-be-consumed message partition at the end of the execution of the batch processing job; Start executing the batch processing job to consume the currently unconsumed messages in the to-be-consumed message partition; During the execution of the batch processing job, when the execution time of the batch processing job reaches the preset maximum consumption time, determining the actual final consumption position of the message in the to-be-consumed message partition by the batch processing job; Determine whether the actual final consumption location is consistent with the specified final consumption location. If so, determine that the batch job is executed successfully, and store the consumed messages during the execution of the batch job in the target data warehouse.

2. The method according to claim 1, It is characterized in that After determining that the batch processing job is successfully executed, the method further includes: Determine the next position of the actual final consumption position in the to-be-consumed message partition as the starting message position of the currently unconsumed messages in the to-be-consumed message partition, and store it in the target storage space; Return to the step of determining the starting message position of the currently unconsumed messages in the message partition to be consumed from the target storage space, execute the next batch processing job at the start execution time, and continue to consume the unconsumed messages in the message partition to be consumed.

3. The method according to claim 1, It is characterized in that The method further comprises: During the execution of the batch processing job, when the consumption position of the messages in the to-be-consumed message partition by the batch processing job reaches the designated final consumption position, the batch processing job is instructed to stop execution.

4. The method according to claim 1, It is characterized in that The starting and executing of the batch processing job to consume the currently unconsumed messages in the to-be-consumed message partition includes: When the number of messages from the starting consumption position to the specified final consumption position in the message partition to be consumed does not exceed a preset threshold, during the execution of this batch job, a consumer is constructed through an RDD partition; The messages consumed by the consumer in the message partition to be consumed are saved in the rdd partition.

5. The method according to claim 1, It is characterized in that The starting and executing of the batch processing job to consume the currently unconsumed messages in the to-be-consumed message partition includes: When the number of messages from the starting consumption position to the specified final consumption position in the message partition to be consumed exceeds a preset threshold, all messages from the starting consumption position to the specified final consumption position in the message partition to be consumed are split to obtain multiple message blocks; During the execution of the batch job, multiple consumers are constructed through multiple RDD partitions, wherein the RDD partitions correspond to the consumers one by one; The messages consumed by each consumer in the message partition to be consumed are saved in the corresponding rdd partition respectively.

6. The method according to claim 1, It is characterized in that The storing of the consumed messages during the execution of the batch processing job to the target data warehouse includes: Write the consumed messages of the batch job in the to-be-consumed message partition into the cos cache; Writing the consumed messages stored in the cos cache into a pre-configured external table, wherein the table structure of the external table is consistent with the table structure of the target data warehouse; Insert all consumed messages stored in the external table into the target data warehouse.

7. A streaming data batch processing device, It is characterized in that include: A first determining unit, a second determining unit, a first estimating unit, a third determining unit, a fourth determining unit, a first starting unit, a fifth determining unit, a sixth determining unit, a seventh determining unit, and a first storage unit; wherein: The first determining unit is configured to execute: determining, from the target storage space, a starting message position of currently unconsumed messages in a to-be-consumed message partition, wherein the messages stored in the to-be-consumed message partition are streaming data; The second determining unit is configured to execute: determining the starting message position as the starting consumption position of the current batch job for the currently unconsumed messages in the to-be-consumed message partition; The first estimation unit is configured to perform: estimating the start execution time of the next batch processing job based on a preset streaming data processing delay; The third determining unit is configured to execute: determining a message position of a target message written into the to-be-consumed message partition at the start execution time; The fourth determining unit is configured to execute: determining the message position of the target message as the designated final consumption position of the message in the to-be-consumed message partition at the end of the execution of the batch processing job; The first starting unit is configured to execute: starting the execution of the current batch processing job to consume currently unconsumed messages in the to-be-consumed message partition; The fifth determining unit is configured to execute: during the execution of the batch processing job, when the execution duration of the batch processing job reaches a preset maximum consumption duration, determining the actual final consumption position of the message in the to-be-consumed message partition by the batch processing job; The sixth determining unit is configured to execute: determining whether the actual final consumption location is consistent with the designated final consumption location, and if so, triggering the seventh determining unit; The seventh determining unit is configured to execute: determining that the batch processing job is successfully executed; The first storage unit is configured to execute: storing the consumed messages during the execution of the batch job in a target data warehouse.

8. The device according to claim 7, It is characterized in that The device further comprises: a second storage unit and a first trigger unit; wherein: The second storage unit is configured to execute: after determining that the batch job is successfully executed, determine the next position of the actual final consumption position in the to-be-consumed message partition as the starting message position of the currently unconsumed message in the to-be-consumed message partition, and store it in the target storage space; The first triggering unit is configured to execute: triggering the first determining unit to execute the next batch processing job at the start execution time, and continuing to consume the unconsumed messages in the to-be-consumed message partition.

9. The device according to claim 7, It is characterized in that The device further comprises: a first instruction unit; The first instruction unit is configured to execute: during the execution of the batch processing job, when the consumption position of the message in the to-be-consumed message partition by the batch processing job reaches the specified final consumption position, instruct the batch processing job to stop executing.

10. The device according to claim 7, It is characterized in that The first storage unit includes: a first writing unit, a second writing unit and an inserting unit; The first writing unit is configured to execute: writing the consumed messages of the batch processing job in the to-be-consumed message partition into the cos cache; The second writing unit is configured to execute: writing the consumed message stored in the cos cache into a pre-configured external table, wherein the table structure of the external table is consistent with the table structure of the target data warehouse; The inserting unit is configured to execute: inserting all consumed messages stored in the external table into the target data warehouse.

Citation Information

Patent Citations

  • Message processing method and system based on kafka, storage medium and computer equipment

    CN111314422A

  • Transmission method and system for keeping data consistency

    CN112087501A