Method for starting a data flow processing framework and related apparatus

CN117421183BActive Publication Date: 2026-09-29KINGDEE SMART TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311583750.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2026-09-29
Estimated Expiration
2043-11-24

AI Technical Summary

Technical Problem

[0002]数据流处理框架框架是流式应用程序,处理的数据是无界的,意味着理论上会永远运行下去,但是在一些应用场景中由于网络超时、磁盘坏道、机器故障等,流式应用程序总有挂掉的一天,出于实际需求需要用到数据流处理框架 Checkpoint容错恢复机制来保障任务的正常运行,而相关技术中的数据流处理框架框架在我们修改业务代码后,若要实现任务的恢复,需要重新消费历史数据,导致系统资源的冗余耗费,并且任务恢复所花费的时间较长导致任务恢复效率较低

Benefits of technology

[0062]本申请实施例中,当数据流处理框架中目标算子的并行度发生改变时,为了防止并行度改变而带来的部分子任务没有分区数据的问题,本申请实施例根据并行度更新前目标算子的任务状态,将目标算子对应的业务数据先读取出来,并根据更新后的任务并行度,将原始分区中的业务数据重新划分为新的分区,其中,新的分区数等于更新后的并行度,从而使得并行度更新后的每个算子的子任务对应一个分区,从而保证目标算子在检查点中正常启动数据流处理框架。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117421183B_ABST
    Figure CN117421183B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method for starting a data stream processing framework and a related device, which are used for normally starting the data stream processing framework from a checkpoint. The method comprises: during the process that the data stream processing framework executes a job by an operator, if it is detected that the target operator parallelism degree of the data stream processing framework executing the job is updated from m to n, obtaining the updated parallelism degree n of the target operator, wherein m≠n, m>0 and n>0; determining a preset checkpoint from at least one checkpoint; obtaining the task state corresponding to the target operator from the preset checkpoint, the task state at least including the partition number m and the partition address; obtaining the business data corresponding to the target operator from the m partitions according to the partition address; dividing the business data corresponding to the target operator into n partitions based on the updated parallelism degree n; and distributing the business data in the n partitions to the target operator with the parallelism degree n, so that the target operator in the job starts the data stream processing framework from the preset checkpoint.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and related apparatus for initiating a data stream processing framework. Background Technology

[0002] Data stream processing frameworks are streaming applications that process unbounded data, meaning they can theoretically run indefinitely. However, in some application scenarios, due to network timeouts, disk bad sectors, machine failures, etc., streaming applications will eventually crash. For practical needs, the checkpoint fault tolerance and recovery mechanism of the data stream processing framework is required to ensure the normal operation of tasks. However, in related technologies, if we modify the business code and want to restore the task, we need to re-consume historical data, which leads to redundant consumption of system resources and long task recovery time, resulting in low task recovery efficiency. Summary of the Invention

[0003] This invention provides a method and related apparatus for starting a data stream processing framework, which is used to start the data stream processing framework normally after modifying the business code related to operator parallelism in the data stream processing framework.

[0004] A first aspect of this application provides a method for initiating a data stream processing framework, wherein the data stream processing framework includes at least one checkpoint, the method comprising:

[0005] During the execution of a job by an operator in the data stream processing framework, if it is detected that the parallelism of the target operator executing the job is updated from m to n, the updated parallelism n of the target operator is obtained, wherein m ≠ n, and m > 0, n > 0;

[0006] Determine a preset checkpoint from the at least one checkpoint;

[0007] The task status corresponding to the target operator is obtained from the preset checkpoints, and the task status includes at least the number of partitions m and the partition address;

[0008] Based on the partition address, retrieve the business data corresponding to the target operator from m partitions, where,

[0009] Each degree of parallelism of the target operator corresponds to a subtask, and each subtask corresponds to a partition.

[0010] Based on the updated parallelism n, the business data corresponding to the target operator is divided into n partitions;

[0011] The business data in the n partitions are allocated to the target operator with a parallelism of n, so that the target operator in the job starts the data stream processing framework from the preset checkpoint.

[0012] Preferably, after assigning the task states in the n partitions to the target operator with a parallelism of n, the method further includes:

[0013] When the job starts, based on the parallelism n of the target operator and the task status of the target operator in the state backend, corresponding indexes are assigned to the n partitions of the target operator, so that each subtask of the target operator can use the corresponding index to mark the task status in each partition.

[0014] Preferably, the step of allocating the business data in the n partitions to a target operator with a parallelism of n includes:

[0015] Based on the updated parallelism n, the business data in the m partitions are evenly distributed to the n subtasks of the target operator.

[0016] Preferably, the preset checkpoint includes: a specified checkpoint or the latest checkpoint in the data stream processing framework. Before obtaining the task status corresponding to the target operator from the m partitions of the checkpoint, the method further includes:

[0017] Start the job in the data stream processing framework;

[0018] The preset checkpoint is loaded into the data stream processing framework.

[0019] Preferably, determining the preset checkpoint from the at least one checkpoint includes:

[0020] Determine whether the specified checkpoint has been received;

[0021] If so, then the specified checkpoint is regarded as the preset checkpoint, and the specified checkpoint is determined from the at least one checkpoint;

[0022] If not, the latest checkpoint is considered as the preset checkpoint, and the latest checkpoint is determined from the at least one checkpoint.

[0023] Preferably, before obtaining the task state corresponding to the target operator from the m partitions of the checkpoint, the method further includes:

[0024] Receive the preset save path for the task status in the checkpoint;

[0025] The receiver sets a unique identifier for the at least one operator;

[0026] The identifier of each operator and the corresponding task status of the operator are associated and stored in the m partitions in the form of a hash map.

[0027] Preferably, the checkpoints include full update checkpoints and incremental update checkpoints;

[0028] The methods for obtaining the parallelism n include:

[0029] The parallelism n is received as input, or the parallelism n is calculated by the data stream processing framework.

[0030] Preferably, the application scenarios of the method include:

[0031] This includes updating the data stream processing framework version, updating the data stream processing framework application, A / B testing of the data stream processing framework application, and maintaining and migrating jobs in the data stream processing framework cluster.

[0032] A second aspect of this application provides an apparatus for initiating a data stream processing framework, wherein the data stream processing framework includes at least one checkpoint, the apparatus comprising:

[0033] The acquisition unit is used to acquire the updated parallelism n of the target operator when the parallelism of the target operator is detected to be updated from m to n during the execution of a job by the operator in the data stream processing framework, wherein m ≠ n, and m > 0, n > 0;

[0034] A determining unit is configured to determine a preset checkpoint from the at least one checkpoint;

[0035] The acquisition unit is further configured to acquire the task status corresponding to the target operator from the preset checkpoint, wherein the task status includes at least the number of partitions m and the partition address;

[0036] The acquisition unit is further configured to acquire business data corresponding to the target operator from m partitions according to the partition address, wherein each parallelism of the target operator corresponds to a subtask, and each subtask corresponds to a partition;

[0037] A partitioning unit is used to divide the business data corresponding to the target operator into n partitions based on the updated parallelism n;

[0038] The allocation unit is used to allocate the business data in the n partitions to the target operator with a parallelism of n, so that the target operator in the job starts the data flow processing framework from the preset checkpoint.

[0039] Preferably, the allocation unit is further configured to:

[0040] After assigning the task states in the n partitions to the target operator with a parallelism of n, when the job starts, according to the parallelism of the target operator n and the task state of the target operator in the state backend, corresponding indexes are assigned to the n partitions of the target operator, so that each subtask of the target operator uses the corresponding index to mark the task state in each partition.

[0041] Preferably, the allocation unit is specifically used for:

[0042] Based on the updated parallelism n, the business data in the m partitions are evenly distributed to the n subtasks of the target operator.

[0043] Preferably, the preset checkpoint includes: a specified checkpoint or the latest checkpoint in the data stream processing framework, and the device further includes:

[0044] A startup unit is used to start the job in the data stream processing framework;

[0045] A loading unit is used to load the preset checkpoint in the data stream processing framework.

[0046] Preferably, the device further includes:

[0047] The judgment unit is used to determine whether a specified checkpoint has been received.

[0048] The determining unit is specifically used for:

[0049] Upon receiving a specified checkpoint as input, the specified checkpoint is regarded as the preset checkpoint, and the specified checkpoint is determined from the at least one checkpoint.

[0050] If no specified checkpoint is received, the latest checkpoint is considered as the preset checkpoint, and the latest checkpoint is determined from the at least one checkpoint.

[0051] Preferably, the device further includes:

[0052] The receiving unit is used to receive the storage path of the task status in the preset checkpoint;

[0053] The receiving unit is also configured to receive a unique identifier set for the at least target operator;

[0054] The storage unit is used to perform associated storage of the identifier code of the target operator and the task status of the target operator in the m partitions in the form of a hash map.

[0055] Preferably, the checkpoints include full update checkpoints and incremental update checkpoints;

[0056] The receiving unit is further configured to:

[0057] The device receives the input parallelism n, or receives the parallelism n calculated by the data stream processing framework. Preferably, the device is applied to at least one of the following: updating a data stream processing framework version, updating a data stream processing framework application, A / B testing of a data stream processing framework application, and maintaining and migrating jobs in a data stream processing framework cluster.

[0058] A third aspect of this application provides a computer device including a processor, which, when executing a computer program stored in a memory, implements the method for launching a data stream processing framework provided in the first aspect of this application.

[0059] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is used to implement the method for launching a data stream processing framework provided in the first aspect of this application.

[0060] The fifth aspect of this application provides a computer program product having a computer program stored thereon. When executed by a computer device, the computer program is used to implement the method for launching a data stream processing framework provided in the first aspect of this application.

[0061] As can be seen from the above technical solutions, the embodiments of the present invention have the following advantages:

[0062] In this embodiment, when the parallelism of the target operator in the data stream processing framework changes, in order to prevent the problem of some subtasks not having partitioned data due to the change in parallelism, this embodiment first reads out the business data corresponding to the target operator according to the task status of the target operator before the parallelism update, and then re-divides the business data in the original partition into new partitions according to the updated task parallelism. The number of new partitions is equal to the updated parallelism, so that each subtask of the operator after the parallelism update corresponds to a partition, thereby ensuring that the target operator starts the data stream processing framework normally at the checkpoint. Attached Figure Description

[0063] Figure 1 This is a schematic diagram of an embodiment of the method for initiating a data stream processing framework in this application.

[0064] Figure 2 This is a schematic diagram of another embodiment of the method for initiating a data stream processing framework in this application.

[0065] Figure 3This is a schematic diagram of another embodiment of the method for initiating a data stream processing framework in this application.

[0066] Figure 4 This is a schematic diagram of an embodiment of the apparatus for initiating a data stream processing framework in this application. Detailed Implementation

[0067] This invention provides a method and related apparatus for starting a data stream processing framework, which is used to start the data stream processing framework normally after modifying the business code related to operator parallelism in the data stream processing framework.

[0068] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0069] The terms "first," "second," "third," "fourth," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0070] To facilitate understanding, the technical terms used in the embodiments of this application will be explained below:

[0071] Stream: A data stream processing framework processes and computes data in the form of a data stream. A stream is an ordered sequence of an infinite number of events or records.

[0072] Job: The computational logic of the data stream processing framework is encapsulated as a job. A job consists of one or more operators and is executed on the stream.

[0073] Operator: An operator is a unit of operation in a job, used to transform, process, and compute data streams.

[0074] State: Data stream processing frameworks support stateful computation. State is used to record and maintain the context information of operator data processing. In data stream processing frameworks, state is called "state" and is used to store intermediate results or some cached data. Many operators in data stream processing frameworks rely on certain intermediate results, i.e., state, for computation. Examples include deduplication, CEP (Continuous Encryption Prediction) detection, and Exactly Once.

[0075] Fault Tolerance: Data stream processing frameworks achieve fault tolerance by dividing the data stream into finite, replayable data blocks and automatically recovering in the event of a failure.

[0076] Checkpoint: A checkpoint is a fault-tolerant mechanism that periodically saves the job state to a reliable storage system for recovery in case of failure. In the data stream processing framework, the checkpoint is a fault-tolerant recovery mechanism that ensures that real-time programs can recover themselves even if they encounter an anomaly.

[0077] Parallelism: Data stream processing frameworks can process data streams in parallel by dividing a job into multiple tasks that are executed concurrently, thereby improving throughput and performance.

[0078] In a data flow processing framework, when the parallelism of the target operator of a job increases, the new job will start more operator tasks. However, the state stored in the checkpoint of the data flow framework only covers the number of previous tasks. Therefore, the new task cannot find the corresponding state information, which causes the new job to fail to start normally from the checkpoint.

[0079] For ease of understanding, the method for starting the data stream processing framework in the embodiments of this application will be described below. Please refer to [link to relevant documentation]. Figure 1 One embodiment of the method for initiating a data stream processing framework in this application includes:

[0080] 101. In the process of executing a job through an operator in a data stream processing framework, if it is detected that the parallelism of the target operator executing the job is updated from m to n, the updated parallelism n of the target operator is obtained, wherein m ≠ n, and m > 0, n > 0;

[0081] Generally, data stream processing frameworks (such as Flink) execute computational logic, which is typically encapsulated as a job. A job usually consists of multiple tasks, where each operator executes one task, and each task can be further divided into multiple subtasks. That is, each operator corresponds to multiple subtasks, and the parallelism of each operator corresponds to the number of subtasks. For ease of understanding, the following example illustrates this:

[0082] Suppose that job A is divided into 3 tasks (A1, A2 and A3), and each task is further divided into 4 subtasks. Each task is executed by an operator. Then job A includes 3 operators, and the number of subtasks of each operator is 4. That is, the parallelism of each operator is 4, which means that the parallelism of each operator is equal to the number of subtasks contained in each task.

[0083] In this embodiment, it is assumed that the operator of task A1 in job A is reduce. The parallelism of this operator is changed from 4 to 6. Therefore, in this embodiment, it is necessary to obtain the parallelism 6 of the reduce operator after the update. It should be noted that the parallelism is the number of subtasks. Therefore, the parallelism m>0 and n>0 in this application, and the parallelism m before the update is not equal to the parallelism n after the update.

[0084] Furthermore, the updated parallelism in this application can be manually set by the user or automatically calculated by the data stream processing framework. Therefore, when obtaining the updated parallelism n in this embodiment, it can be the parallelism n received from the input or the parallelism n automatically calculated by the data stream processing framework. There are no specific restrictions on the method of obtaining the updated parallelism n.

[0085] 102. Determine a preset checkpoint from at least one checkpoint;

[0086] In the data stream processing framework, multiple checkpoints are inserted into the data stream to enable data recovery at the point of failure. These checkpoints are used to periodically save the job state (i.e., the task state of multiple operators) in a reliable storage system so that recovery can be performed in case of failure.

[0087] Because each checkpoint stores the state of the operator within a different time period, in order to recover data from a certain checkpoint, a preset checkpoint needs to be determined from multiple checkpoints. The preset checkpoint can be a specified checkpoint or the latest checkpoint determined in chronological order.

[0088] 103. Obtain the task status corresponding to the target operator from the preset checkpoints. The task status includes at least the number of partitions m and the partition address.

[0089] Specifically, the preset checkpoint stores the task status of the target operator within a certain time period. This task status includes the number of partitions m, the storage address of the partitions, and metadata of the business data.

[0090] 104. Based on the partition address, obtain the business data corresponding to the target operator from m partitions, wherein each parallelism degree of the target operator corresponds to a subtask, and each subtask corresponds to a partition;

[0091] Furthermore, in order to achieve load balancing of data across multiple tasks in a job and improve system performance, the data stream processing framework is typically partitioned. This means that the data in the data stream processing framework is divided into multiple partitions, where each partition stores the data volume of one subtask. In other words, the number of partitions equals the number of subtasks, or the number of partitions equals the parallelism of the tasks.

[0092] Once the partition address is obtained, the business data corresponding to the target operator can be retrieved from the m partitions based on the partition address.

[0093] 105. Based on the updated parallelism n, divide the business data corresponding to the target operator into n partitions;

[0094] After obtaining the updated parallelism n and the business data in the m partitions, the business data of the target operator in the m partitions are re-divided into n partitions, where each partition is used to store the business data of a subtask.

[0095] 106. Distribute the business data in the n partitions to the target operator with a parallelism of n, so that the target operator in the job can start the data flow processing framework normally from the preset checkpoint.

[0096] After re-dividing the task states in the original m partitions into n partitions, the business data in the n partitions are distributed to the target operator with a parallelism of n, so that each subtask in the target operator with a parallelism of n corresponds to a partition, which means that the target operator in the job starts the data flow processing framework normally at the preset checkpoint.

[0097] In this embodiment, when the parallelism of the target operator in the data stream processing framework changes, in order to prevent the problem of some subtasks not having partitioned data due to the change in parallelism, this embodiment first reads out the business data corresponding to the target operator according to the task status of the target operator before the parallelism update, and then re-divides the business data in the original partition into new partitions according to the updated task parallelism. The number of new partitions is equal to the updated parallelism, so that each subtask of the operator after the parallelism update corresponds to a partition, thereby ensuring that the target operator starts the data stream processing framework normally at the checkpoint.

[0098] based on Figure 1 In the aforementioned embodiment, after distributing the business data in n partitions to a target operator with a parallelism of n, in order to ensure the normal marking of the task status in each partition at the checkpoint, this embodiment also needs to allocate corresponding indexes to the n partitions of the target operator at job startup based on the parallelism of the target operator n and the task status of the target operator in the status backend, so that each subtask of the target operator can use the corresponding index to mark the task status in each partition.

[0099] In a data stream processing framework, the state backend is a component used to store and manage state data in the application. The state backend stores the task state of each task. Therefore, in this application, based on the parallelism *n* of the target operator and the task state of the target operator in the state backend, corresponding KeyGroup indexes are assigned to the *n* partitions of the target operator. Here, the KeyGroup index is used to mark each partition. For example, when *n*=6, six KeyGroup indexes are needed to mark the six partitions respectively, so that when storing task state in checkpoints, the processing status of the target operator on the data in each partition can be marked. For example, when the target operator is *reduce*, assuming the target operator *reduce* has six parallelisms (i.e., six subtasks, assuming these six subtasks are *reduce-1*, *reduce-2*, *reduce-3*, *reduce-4*, *reduce-5*, and *reduce-6*), then the six partitions need to be assigned corresponding KeyGroup indexes (assuming the indexes of the six partitions are KeyGroup-1, KeyGroup-2, KeyGroup-3, KeyGroup-4, KeyGroup-5, and KeyGroup-6 respectively). Group-6) is used to ensure that marking is performed at checkpoints, and that the task status of each operator can be retrieved from the corresponding partition based on the operator's identifier. Furthermore, after updating the parallelism of the target operator, the business data of each subtask in the target operator needs to be recalculated based on the updated parallelism n. Since the parallelism before the update was m, and the updated parallelism is n, assuming the business data processed by each parallelism before the update was w, the business data of each subtask in the target operator after the update is mw / n, meaning the business data in the m partitions is evenly distributed across the n subtasks.

[0100] Furthermore, it should be noted that when updating data at checkpoints, it can be a full update of all data or an incremental update of the incremental quantity. However, a full update would store a large number of data at each checkpoint when the data volume is large, which would lead to a decrease in system performance. Therefore, in order to improve the real-time performance of the system, incremental updates of data at checkpoints can also be performed. The state data in the data stream processing framework is stored based on state backend components. Among them, state backend components are divided into memory-based state backend (MemoryStateBackend), file system-based state backend (FsStateBackend), and state backend based on RocksDB as the storage medium (RocksDBStateBackend). When performing incremental data updates, the RocksDB component is needed to realize the incremental update of data.

[0101] Furthermore, based on Figure 1 In the described embodiment, before executing step 101, the following steps are also required to load the checkpoints and store the operator task states. Please refer to [link to relevant documentation]. Figure 2 , Figure 2 Another embodiment of the method for initiating the data stream processing framework in this application:

[0102] 201. Start the job in the data stream processing framework;

[0103] It is easy to understand that in execution Figure 1 Prior to the implementation, the job in the data stream processing framework needs to be started first, wherein the job here is... Figure 1 The job in the embodiments is generally an application running in a data stream processing framework, such as an APP.

[0104] 202. Load preset checkpoints in the data stream processing framework. The preset checkpoints include: specified checkpoints or the latest checkpoint in the data stream processing framework.

[0105] After a job is started in the data stream processing framework, preset checkpoints are loaded into the framework. It is easy to understand that in order to ensure the fault tolerance of the data stream processing framework, multiple checkpoints are generally set in the data stream processing framework. In order to realize data recovery, preset checkpoints need to be loaded into the data stream processing framework. Here, the preset checkpoints can be specified checkpoints or the latest checkpoints in the data stream processing framework. The latest checkpoints are the checkpoints that store the latest task status, while the specified checkpoints are the checkpoints set by the user to recover data.

[0106] 203. Receive the save path of the task status in the preset checkpoint;

[0107] Because checkpoints store task states in a permanent storage address rather than in memory to prevent data loss in case of task failure, this application embodiment requires setting a preset storage path for the task states in checkpoints in order to back up the task states in checkpoints. Here, the storage path is a storage address that can be permanently stored.

[0108] 204. Receive the unique identifier set for the target operator;

[0109] In order to mark the task status of the target operator, this embodiment of the application sets a unique identifier for the target operator in the job, so that the target operator and the task status associated with the target operator can be marked according to the identifier of the target operator, so that when the task is executed again, the task status of the target operator can be obtained according to the identifier of the target operator, and the task can be executed again according to the task status of the target operator.

[0110] 205. Associatively store the target operator's identifier and task status using a hash map according to the storage path.

[0111] In order to quickly obtain the task status of the target operator when restoring data later, this application embodiment associates and stores the target operator's identifier code and the target operator's task status in a hash map manner according to the storage path.

[0112] When using a hash map to perform associative storage of the target operator identifier and the target operator's task state, it has at least the following advantages:

[0113] 1. Fast lookup: Hash maps use hash functions to map keys to buckets, enabling fast element lookup.

[0114] 2. Efficient insertion and deletion: Hash maps use hash functions to map elements to buckets, and can also quickly insert and delete elements.

[0115] 3. High space utilization: Hash maps use buckets to store elements, and the size of the buckets can be dynamically adjusted as needed, thus making the space utilization higher.

[0116] 206. During the execution of a job by an operator in the data stream processing framework, if it is detected that the parallelism of the target operator executing the job is updated from m to n, the updated parallelism n of the target operator is obtained, wherein m ≠ n, and m > 0, n > 0;

[0117] Step 201 and Figure 1Step 101 in the embodiment is similar and will not be repeated here. 207. Determine a preset checkpoint from the at least one checkpoint;

[0118] Specifically, corresponding to step 202, the preset checkpoint in this embodiment includes a specified checkpoint or the latest checkpoint in the data stream processing framework. When determining the preset checkpoint from at least one checkpoint in the data stream processing framework, the following steps can be performed:

[0119] When determining the preset checkpoint, it is determined whether the specified checkpoint has been received. The specified checkpoint is generally determined by input. If the specified checkpoint is received, it is regarded as the preset checkpoint and determined from at least one checkpoint. If the specified checkpoint is not received, the latest checkpoint (that is, the checkpoint closest to the current timestamp) is determined as the preset checkpoint and determined from at least one checkpoint.

[0120] 208. Obtain the task status corresponding to the target operator from the preset checkpoints, wherein the task status includes at least the number of partitions m and the partition address;

[0121] 209. Based on the partition address, obtain the business data corresponding to the target operator from m partitions, wherein each parallelism of the target operator corresponds to a subtask, and each subtask corresponds to a partition;

[0122] 210. Based on the updated parallelism n, divide the business data corresponding to the target operator into n partitions;

[0123] 211. Distribute the business data in the n partitions to the target operator with a parallelism of n, so that the target operator in the job starts the data stream processing framework from the preset checkpoint.

[0124] It should be noted that the descriptions of steps 208 to 211 are consistent with... Figure 1 The descriptions of steps 103 to 106 in the embodiments are similar and will not be repeated here.

[0125] In this embodiment, before repartitioning the business data of the target operator, a preset checkpoint is loaded and a save path is set for the preset checkpoint. A unique identifier is set for the target operator, so that the task status of the target operator can be marked according to the unique identifier of the target operator. When loading the preset checkpoint, the checkpoint can be the latest checkpoint or a specified checkpoint, thereby ensuring the flexibility of data recovery and improving the convenience and efficiency of obtaining the task status of the operator.

[0126] based on Figure 1 and Figure 2 The embodiments described herein, in which the method for initiating a data stream processing framework from a checkpoint, can be applied to data stream processing framework version updates, data stream processing framework application updates, A / B testing of data stream processing framework applications, and job maintenance and migration within a data stream processing framework cluster. This application can be used whenever the parallelism of the target operator in the job changes in the aforementioned scenarios. Figure 1 and Figure 2 The method in the embodiment is used to normally start the data stream processing framework from the checkpoint when the parallelism of the target operator changes during the job.

[0127] The following describes in detail the method for launching the data stream processing framework in this application embodiment, using specific scenarios as examples. Please refer to [link / reference]. Figure 3 One embodiment of the method for initiating a data stream processing framework in this application includes:

[0128] The data stream processing framework application starts normally. A preset checkpoint is loaded within the application. This preset checkpoint is either the latest checkpoint when the job's parallelism changes, or a specified checkpoint before the change. If the job's parallelism changes at the preset checkpoint (assuming the parallelism changes from m to n), the job is paused at the preset checkpoint, and the data state is restored from the preset checkpoint. The data state restoration process includes the following steps:

[0129] The task status of the target operator is read from the hash map of the preset checkpoints. Then, based on the partition address in the task status, the business data in the m partitions before the parallelism update is read. Then, according to the new parallelism, the business data in the m partitions is repartitioned, that is, the business data in the m partitions is divided into n partitions. The business data in the n partitions is redistributed to the n subtasks of the target operator, so that the n subtasks reprocess the new processing flow, so as to restore the normal execution process of the data flow.

[0130] The method for initiating a data stream processing framework from a checkpoint in the embodiments of this application has been described in detail above. The apparatus for initiating a data stream processing framework from a checkpoint in the embodiments of this application will be described below. Please refer to [link to relevant documentation]. Figure 4 An embodiment of the apparatus for initiating a data stream processing framework in this application includes:

[0131] The acquisition unit 401 is used to acquire the updated parallelism n of the target operator when the parallelism of the target operator executing the job is detected to be updated from m to n during the execution of the job by the operator in the data stream processing framework, wherein m ≠ n, and m>0, n>0;

[0132] Determining unit 402 is used to determine a preset checkpoint from the at least one checkpoint;

[0133] The acquisition unit 401 is further configured to acquire the task status corresponding to the target operator from the preset checkpoint, wherein the task status includes at least the number of partitions m and the partition address;

[0134] The acquisition unit 401 is further configured to acquire business data corresponding to the target operator from m partitions according to the partition address, wherein each parallelism of the target operator corresponds to a subtask and each subtask corresponds to a partition;

[0135] The partitioning unit 403 is used to divide the business data corresponding to the target operator into n partitions based on the updated parallelism n;

[0136] The allocation unit 404 is used to allocate the business data in the n partitions to the target operator with a parallelism of n, so that the target operator in the job starts the data flow processing framework from the preset checkpoint.

[0137] Preferably, the allocation unit 404 is further configured to:

[0138] After assigning the task states in the n partitions to the target operator with a parallelism of n, when the job starts, according to the parallelism of the target operator n and the task state of the target operator in the state backend, corresponding indexes are assigned to the n partitions of the target operator, so that each subtask of the target operator uses the corresponding index to mark the task state in each partition.

[0139] Preferably, the allocation unit 404 is specifically used for:

[0140] Based on the updated parallelism n, the business data in the m partitions are evenly distributed to the n subtasks of the target operator.

[0141] Preferably, the preset checkpoint includes: a specified checkpoint or the latest checkpoint in the data stream processing framework, and the device further includes:

[0142] The startup unit 405 is used to start the job in the data stream processing framework;

[0143] The loading unit 406 is used to load the preset checkpoint in the data stream processing framework.

[0144] Preferably, the device further includes:

[0145] The judgment unit 407 is used to determine whether the specified checkpoint has been received.

[0146] The determining unit 402 is specifically used for:

[0147] Upon receiving a specified checkpoint as input, the specified checkpoint is regarded as the preset checkpoint, and the specified checkpoint is determined from the at least one checkpoint.

[0148] If no specified checkpoint is received, the latest checkpoint is considered as the preset checkpoint, and the latest checkpoint is determined from the at least one checkpoint.

[0149] Preferably, the device further includes:

[0150] The receiving unit 408 is used to receive the saved path of the task status in the preset checkpoint;

[0151] The receiving unit 408 is also configured to receive a unique identifier set for the at least target operator;

[0152] Storage unit 409 is used to perform associated storage of the identifier code of the target operator and the task status of the target operator in the m partitions in the form of a hash map.

[0153] Preferably, the checkpoints include full update checkpoints and incremental update checkpoints;

[0154] The receiving unit 408 is further configured to:

[0155] The parallelism n is received as input, or the parallelism n is calculated by the data stream processing framework.

[0156] Preferably, the apparatus is used for at least one of the following: updating a data stream processing framework version, updating a data stream processing framework application, A / B testing of a data stream processing framework application, and maintaining and migrating jobs in a data stream processing framework cluster.

[0157] It should be noted that the functions of the above-mentioned units are the same as... Figures 1 to 2 The examples described are similar and will not be repeated here.

[0158] In this embodiment, when the parallelism of the target operator in the data stream processing framework changes, in order to prevent the problem of some subtasks not having partitioned data due to the change in parallelism, this embodiment uses the acquisition unit 401 to read out the business data corresponding to the target operator according to the task status of the target operator before the parallelism update, and according to the updated task parallelism, the allocation unit 404 re-divides the business data in the original partition into new partitions, wherein the number of new partitions is equal to the updated parallelism, so that each subtask of the operator after the parallelism update corresponds to a partition, thereby ensuring that the target operator starts the data stream processing framework normally at the checkpoint.

[0159] This application also provides a computer program product, on which a computer program is stored. When executed by a computer device, the computer program is used to implement this application. Figures 1 to 3 The method for launching the data stream processing framework described in the embodiments.

[0160] The above description of the apparatus for initiating the data stream processing framework in the embodiments of the present invention is from the perspective of modular functional entities. The following description of the computer apparatus in the embodiments of the present invention is from the perspective of hardware processing:

[0161] This computer device is used to implement the function of a device for initiating a data stream processing framework. One embodiment of the computer device in this invention includes:

[0162] Processor and memory;

[0163] When a memory is used to store computer programs, and a processor executes the computer programs stored in the memory, the following steps can be achieved:

[0164] During the execution of a job by an operator in the data stream processing framework, if it is detected that the parallelism of the target operator executing the job is updated from m to n, the updated parallelism n of the target operator is obtained, wherein m ≠ n, and m > 0, n > 0;

[0165] Determine a preset checkpoint from the at least one checkpoint;

[0166] The task status corresponding to the target operator is obtained from the preset checkpoints, and the task status includes at least the number of partitions m and the partition address;

[0167] Based on the partition address, retrieve the business data corresponding to the target operator from m partitions, where,

[0168] Each degree of parallelism of the target operator corresponds to a subtask, and each subtask corresponds to a partition.

[0169] Based on the updated parallelism n, the business data corresponding to the target operator is divided into n partitions;

[0170] The business data in the n partitions are allocated to the target operator with a parallelism of n, so that the target operator in the job starts the data stream processing framework from the preset checkpoint.

[0171] In some embodiments of the present invention, after allocating the business data in the n partitions to the target operator with a parallelism of n, the processor can also be used to implement the following steps:

[0172] When the job starts, based on the parallelism n of the target operator and the task status of the target operator in the state backend, corresponding indexes are assigned to the n partitions of the target operator, so that each subtask of the target operator can use the corresponding index to mark the task status in each partition.

[0173] In some embodiments of the present invention, after allocating the business data in the n partitions to the target operator with a parallelism of n, the processor can also be used to implement the following steps:

[0174] Based on the updated parallelism n, the business data in the m partitions are evenly distributed to the n subtasks of the target operator.

[0175] In some embodiments of the present invention, the preset checkpoint includes: a specified checkpoint or the latest checkpoint in the data stream processing framework. Before obtaining the task status corresponding to the target operator from the m partitions of the checkpoint, the processor can also be used to implement the following steps:

[0176] Start the job in the data stream processing framework;

[0177] Preset checkpoints are loaded into the data stream processing framework.

[0178] In some embodiments of the present invention, the processor may also be used to implement the following steps:

[0179] Determine whether the specified checkpoint has been received;

[0180] If so, then the specified checkpoint is regarded as the preset checkpoint, and the specified checkpoint is determined from the at least one checkpoint;

[0181] If not, the latest checkpoint is considered as the preset checkpoint, and the latest checkpoint is determined from the at least one checkpoint.

[0182] In some embodiments of the present invention, before obtaining the task state corresponding to the target operator from the preset checkpoint, the processor may further perform the following steps:

[0183] Receive the saved path of the task status in the preset checkpoint;

[0184] Receive a unique identifier set for the at least target operator;

[0185] The identifier of the target operator and the task status of the target operator are associated and stored in the m partitions in the form of a hash map.

[0186] In some embodiments of the present invention, the checkpoints include full update checkpoints and incremental update checkpoints;

[0187] In some embodiments of the present invention, the processor may also be used to implement the following steps:

[0188] The parallelism n is received as input, or the parallelism n is calculated by the data stream processing framework.

[0189] In some embodiments of the present invention, the application scenarios of the method include:

[0190] At least one of the following: updating the data stream processing framework version, updating the data stream processing framework application, A / B testing of the data stream processing framework application, and maintaining and migrating jobs in the data stream processing framework cluster.

[0191] It is understood that when the processor in the computer device described above executes the computer program, it can also implement the functions of each unit in the corresponding device embodiments described above, which will not be repeated here. For example, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the device for initiating the data stream processing framework. For example, the computer program can be divided into units in the device for initiating the data stream processing framework described above, and each unit can implement the specific functions described in the corresponding device for initiating the data stream processing framework from a checkpoint.

[0192] The computer device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the processor and memory are merely examples of a computer device and do not constitute a limitation on the computer device. It may include more or fewer components, or a combination of certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.

[0193] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the computer device, connecting various parts of the computer device via various interfaces and lines.

[0194] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital card (SD card), flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0195] The present invention also provides a computer-readable storage medium for implementing the function of a means of initiating a data stream processing framework, wherein a computer program is stored thereon, and when the computer program stored in the computer-readable storage medium is executed by a processor, the processor can perform the following steps:

[0196] During the execution of a job by an operator in the data stream processing framework, if it is detected that the parallelism of the target operator executing the job is updated from m to n, the updated parallelism n of the target operator is obtained, wherein m ≠ n, and m > 0, n > 0;

[0197] Determine a preset checkpoint from the at least one checkpoint;

[0198] The task status corresponding to the target operator is obtained from the preset checkpoints, and the task status includes at least the number of partitions m and the partition address;

[0199] Based on the partition address, retrieve the business data corresponding to the target operator from m partitions, where,

[0200] Each degree of parallelism of the target operator corresponds to a subtask, and each subtask corresponds to a partition.

[0201] Based on the updated parallelism n, the business data corresponding to the target operator is divided into n partitions;

[0202] The business data in the n partitions are allocated to the target operator with a parallelism of n, so that the target operator in the job starts the data stream processing framework from the preset checkpoint.

[0203] In some embodiments of the present invention, after the business data in the n partitions are allocated to the target operator with a parallelism of n, when the computer program stored in the computer-readable storage medium is executed by the processor, the processor may further be used to implement the following steps:

[0204] When the job starts, based on the parallelism n of the target operator and the task status of the target operator in the state backend, corresponding indexes are assigned to the n partitions of the target operator, so that each subtask of the target operator can use the corresponding index to mark the task status in each partition.

[0205] In some embodiments of the present invention, after the business data in the n partitions are allocated to the target operator with a parallelism of n, when the computer program stored in the computer-readable storage medium is executed by the processor, the processor may further be used to implement the following steps:

[0206] Based on the updated parallelism n, the business data in the m partitions are evenly distributed to the n subtasks of the target operator.

[0207] In some embodiments of the present invention, the preset checkpoint includes: a specified checkpoint or the latest checkpoint in the data stream processing framework. Before obtaining the task status corresponding to the target operator from the m partitions of the checkpoint, when the computer program stored in the computer-readable storage medium is executed by the processor, the processor can also be used to implement the following steps:

[0208] Start the job in the data stream processing framework;

[0209] Preset checkpoints are loaded into the data stream processing framework.

[0210] In some embodiments of the present invention, when a computer program stored on a computer-readable storage medium is executed by a processor, the processor may also be used to perform the following steps:

[0211] Determine whether the specified checkpoint has been received;

[0212] If so, then the specified checkpoint is regarded as the preset checkpoint, and the specified checkpoint is determined from the at least one checkpoint;

[0213] If not, the latest checkpoint is considered as the preset checkpoint, and the latest checkpoint is determined from the at least one checkpoint.

[0214] In some embodiments of the present invention, before obtaining the task state corresponding to the target operator from the preset checkpoint, when the computer program stored in the computer-readable storage medium is executed by the processor, the processor may also be used to implement the following steps:

[0215] Receive the saved path of the task status in the preset checkpoint;

[0216] Receive a unique identifier set for the at least target operator;

[0217] The identifier of the target operator and the task status of the target operator are associated and stored in the m partitions in the form of a hash map.

[0218] In some embodiments of the present invention, the checkpoints include full update checkpoints and incremental update checkpoints;

[0219] In some embodiments of the present invention, when a computer program stored on a computer-readable storage medium is executed by a processor, the processor may also be used to perform the following steps:

[0220] The parallelism n is received as input, or the parallelism n is calculated by the data stream processing framework.

[0221] In some embodiments of the present invention, the device can be applied to:

[0222] At least one of the following: updating the data stream processing framework version, updating the data stream processing framework application, A / B testing of the data stream processing framework application, and maintaining and migrating jobs in the data stream processing framework cluster.

[0223] It is understood that if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a corresponding computer-readable storage medium. Based on this understanding, all or part of the processes in the above-described embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above-described method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0224] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0225] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0226] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0227] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for initiating a data stream processing framework, characterized in that, The data stream processing framework includes at least one checkpoint, and the method includes: During the execution of a job by an operator in the data stream processing framework, if it is detected that the parallelism of the target operator executing the job by the data stream processing framework is updated from m to n, the updated parallelism n of the target operator is obtained, wherein m ≠ n, and m > 0, n > 0; Determine a preset checkpoint from the at least one checkpoint; The task status corresponding to the target operator is obtained from the preset checkpoints, and the task status includes at least the number of partitions m and the partition address; Based on the partition address, retrieve the business data corresponding to the target operator from m partitions, where, Each degree of parallelism of the target operator corresponds to a subtask, and each subtask corresponds to a partition. Based on the updated parallelism n, the business data corresponding to the target operator is divided into n partitions; The business data in the n partitions are allocated to the target operator with a parallelism of n, so that the target operator in the job starts the data stream processing framework from the preset checkpoint.

2. The method according to claim 1, characterized in that, After allocating the business data in the n partitions to the target operator with a parallelism of n, the method further includes: When the job starts, based on the parallelism n of the target operator and the task status of the target operator in the state backend, corresponding indexes are assigned to the n partitions of the target operator, so that each subtask of the target operator can use the corresponding index to mark the task status in each partition.

3. The method according to claim 1, characterized in that, The step of allocating business data from the n partitions to a target operator with a parallelism of n includes: Based on the updated parallelism n, the business data in the m partitions are evenly distributed to the n subtasks of the target operator.

4. The method according to claim 1, characterized in that, The preset checkpoints include: a specified checkpoint or the latest checkpoint in the data stream processing framework. Before obtaining the task status corresponding to the target operator from the preset checkpoints, the method further includes: Start the job in the data stream processing framework; The preset checkpoint is loaded into the data stream processing framework.

5. The method according to claim 4, characterized in that, Determining the preset checkpoint from the at least one checkpoint includes: Determine whether the specified checkpoint has been received; If so, then the specified checkpoint is regarded as the preset checkpoint, and the specified checkpoint is determined from the at least one checkpoint; If not, the latest checkpoint is regarded as the preset checkpoint, and the latest checkpoint is determined from the at least one checkpoint.

6. The method according to claim 1, characterized in that, Before obtaining the task state corresponding to the target operator from the preset checkpoint, the method further includes: Receive the saved path of the task status in the preset checkpoint; Receive the unique identifier set for the target operator; The identifier of the target operator and the task status of the target operator are associated and stored in a hash map according to the storage path.

7. The method according to claim 1, characterized in that, The checkpoints include full update checkpoints and incremental update checkpoints; The methods for obtaining the parallelism n include: The parallelism n is received as input, or the parallelism n is calculated by the data stream processing framework.

8. An apparatus for initiating a data stream processing framework, characterized in that, The data stream processing framework includes at least one checkpoint, and the device includes: The acquisition unit is used to acquire the updated parallelism n of the target operator when the parallelism of the target operator is detected to be updated from m to n during the execution of a job by the data stream processing framework through the operator, wherein m ≠ n, and m > 0, n > 0; A determining unit is configured to determine a preset checkpoint from the at least one checkpoint; The acquisition unit is further configured to acquire the task status corresponding to the target operator from the preset checkpoint, wherein the task status includes at least the number of partitions m and the partition address; The acquisition unit is further configured to acquire business data corresponding to the target operator from m partitions according to the partition address, wherein each parallelism of the target operator corresponds to a subtask, and each subtask corresponds to a partition; A partitioning unit is used to divide the business data corresponding to the target operator into n partitions based on the updated parallelism n; The allocation unit is used to allocate the business data in the n partitions to the target operator with a parallelism of n, so that the target operator in the job starts the data flow processing framework from the preset checkpoint.

9. A computer device comprising a processor, characterized in that, When the processor executes a computer program stored in memory, it is used to implement the method of launching a data stream processing framework as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it is used to implement the method of launching a data stream processing framework as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data processing method and related equipment

    CN113900856A

  • Data processing method and device, computer equipment and readable storage medium

    CN115934304A