Flink cluster long-period job operator-level live migration method and device

CN122470291BActive Publication Date: 2026-09-11NEW H3C TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610953656.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-11
Estimated Expiration
2046-06-29

AI Technical Summary

Technical Problem

[0004]然而,在对可用性要求较高的业务场景中,上述方案存在明显的技术局限

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122470291B_ABST
    Figure CN122470291B_ABST
Patent Text Reader

Abstract

The application provides a Flink cluster long-period job operator-level hot migration method and device. The method comprises the following steps: for any operator instance, in the case that an alarm for the operator instance is received, a shadow task instance corresponding to the operator instance is created and started; wherein the alarm is used to indicate that the operator instance meets a hot migration condition, and the operator instance meeting the hot migration condition is determined according to a monitored running state parameter of the operator instance; the running state parameter comprises a performance parameter used to represent the running health degree of the operator instance; the state of the shadow task instance and the operator instance is synchronized through a memory mapping file (MMAP) and a delta residual synchronization mode; and the communication of the shadow task instance and the operator instance is switched. The method can improve the efficiency of hot migration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method and device for operator-level hot migration of long-cycle jobs in Flink clusters. Background Technology

[0002] In core production scenarios such as financial payments and real-time risk control, Apache Flink, as a mainstream distributed streaming computing engine, is widely used to run long-cycle, large-state jobs. As job execution time increases, the runtime characteristics of the Java Virtual Machine (JVM) are prone to "sub-health" issues, such as heap memory leaks and metaspace bloat, which in turn lead to frequent Full Garbage Collections (Full GC), severely impacting job processing latency and throughput.

[0003] The commonly used solution in the industry is a job restart and migration mechanism based on savepoints. This mechanism typically involves operations personnel manually or using automated scripts to trigger Flink's savepoint operation after detecting anomalies through monitoring. This persists the global state of the current job to remote storage (such as HDFS or S3), then stops the old job instance, and finally restarts the job from the savepoint directory. This approach ensures exactly-once consistency of data processing and is a crucial means of guaranteeing the stable operation of stream computing jobs.

[0004] However, in business scenarios with high availability requirements, the above solutions have significant technical limitations. On the one hand, for large-scale jobs at the GB or even TB level, loading the state from remote storage can take tens of seconds to several minutes, resulting in excessively long business interruptions and failing to meet the continuity requirements of financial-grade business. On the other hand, job restarts involve the release and re-application of all computing resources, causing drastic fluctuations in cluster load and potentially affecting other normally operating jobs within the same cluster.

[0005] Therefore, there is an urgent need for a streaming computing job migration method that can both ensure state consistency and achieve low-latency, low-disturbance recovery. Summary of the Invention

[0006] In view of this, this application provides a method and device for operator-level hot migration of long-cycle jobs in Flink clusters.

[0007] According to a first aspect of the embodiments of this application, a method for operator-level hot migration of long-cycle jobs in a Flink cluster is provided, comprising: For any operator instance, upon receiving an alarm for that operator instance, a shadow task instance corresponding to that operator instance is created and started; wherein, the alarm is used to indicate that the operator instance meets the hot migration conditions, and the operator instance meeting the hot migration conditions is determined based on the monitored operating status parameters of the operator instance; the operating status parameters include performance parameters used to characterize the operating health of the operator instance. The shadow task instance and the operator instance are synchronized in state using MMAP and incremental residual synchronization; and... The communication between the shadow task instance and the operator instance is switched.

[0008] According to a second aspect of the embodiments of this application, a Flink cluster long-cycle job operator-level hot migration apparatus is provided, comprising: A creation unit is used to create and start a shadow task instance corresponding to any operator instance when an alarm is received for that operator instance; wherein, the alarm is used to indicate that the operator instance meets the hot migration conditions, and the operator instance meeting the hot migration conditions is determined based on the monitored operating status parameters of the operator instance; the operating status parameters include performance parameters used to characterize the operating health of the operator instance. The synchronization unit is used to synchronize the state of the shadow task instance and the operator instance through MMAP and incremental residual synchronization. The switching unit is used to switch the communication between the shadow task instance and the operator instance.

[0009] According to a third aspect of the embodiments of this application, an electronic device is provided, including a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-readable instructions executable by the processor, the processor being prompted by the machine-readable instructions to perform the method provided in the first aspect.

[0010] According to a fourth aspect of the embodiments of this application, a machine-readable storage medium is provided, storing machine-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method provided in the first aspect.

[0011] Applying the embodiments of this application, for any operator instance, upon receiving an alarm for that operator instance, a shadow task instance corresponding to that operator instance is created and started. The shadow task instance and the operator instance are synchronized in state using MMAP and incremental residual synchronization. Furthermore, communication switching is performed between the shadow task instance and the operator instance. By monitoring the operating status parameters of the operator instance, and determining that the operator instance meets the hot migration conditions based on the monitored operating status parameters, the shadow task instance can be started without stopping the global job. By using MMAP and incremental residual synchronization to synchronize the state of the shadow task instance and the operator instance, the overhead of downloading the entire remote state is reduced, effectively improving state recovery efficiency, thereby improving the efficiency of hot migration. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of a typical Flink cluster architecture; Figure 2 This is a flowchart illustrating a method for operator-level hot migration of long-cycle jobs in a Flink cluster, as provided in an embodiment of this application. Figure 3 This is a schematic diagram illustrating the implementation process of an operator-level hot migration scheme for long-cycle jobs in a Flink cluster, as provided in an embodiment of this application. Figure 4 This is a schematic diagram of the structure of an operator-level hot migration device for long-cycle jobs in a Flink cluster, provided in an embodiment of this application. Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0013] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, some technical terms involved in the embodiments of this application will be explained below.

[0014] 1. Flink (Apache Flink): An open-source distributed streaming computing engine that focuses on real-time processing of data streams.

[0015] 2. TaskManager: The physical execution node of the Flink cluster, responsible for running specific operator tasks of the job, managing computing resources, and interacting with the JobManager.

[0016] 3. JobManager: The core control node of the Flink cluster, responsible for core functions such as job submission, topology scheduling, node monitoring, fault recovery, and cluster command issuance.

[0017] 4. Checkpoint: Flink's native state persistence mechanism, which periodically and asynchronously saves the real-time running state of operators.

[0018] 5. Savepoint: Flink's manually triggered global state snapshot, which is the standard core mechanism for job restart migration in existing technologies.

[0019] 6. Shadow Task (also known as shadow task instance): A backup operator instance defined in this application embodiment, which is dynamically launched by the "sub-healthy" operator, has the same configuration as the original operator instance, and runs in place of the original instance after completing state alignment (or state synchronization).

[0020] 7. LSN (Log Sequence Number): Used to mark the timing of data stream and state synchronization, ensuring strict alignment of data and state between new and old operator instances.

[0021] 8. MMAP (Memory-Mapped File): Maps a disk file to the virtual memory of a process. In this embodiment, it is used for Shadow Task to efficiently preload the baseline state and avoid full data copying.

[0022] 9. VIL (Virtual Input Layer): The lightweight network architecture designed in this application embodiment realizes dynamic binding of network addresses and hot redirection of data streams for downstream operators without restarting.

[0023] 10. JobGraph: A logical description of the topology of job operators in Flink. It defines the data flow relationships and business processing logic between operators and is the foundation for job execution.

[0024] 11. ExecutionGraph: The physical implementation of JobGraph in Flink, which describes the actual execution topology and resource allocation relationship of a job. In this embodiment, it extends the dynamic update interface to enable shadow tasks to be launched without restarting.

[0025] 12. UserFunctionWrapper: A custom operator wrapper class in this application embodiment, used to insert interception logic into the operator call chain to achieve side effect isolation of shadow tasks.

[0026] 13. Copy-on-Write (COW): Copy-on-write is a resource sharing and isolation mechanism. In this embodiment, it is used to asynchronously obtain the real-time state image of the old operator instance to avoid state read-write conflicts.

[0027] 14. Epoch: An incrementing unique identifier used in this application embodiment to preempt the data source connection permission of the Source operator, ensuring the uniqueness and continuity of data consumption.

[0028] 15. RPC (Remote Procedure Call): In this embodiment of the application, it is used for instruction interaction between cluster nodes and for encrypted real-time transmission of residual data between new and old operator instances.

[0029] 16. Exactly-once semantics: The core data consistency requirement of streaming computing, which means that each piece of data is processed only once. In the embodiments of this application, it is used to ensure that this semantics is not violated during the migration process.

[0030] The following section provides an exemplary description of a typical Flink cluster architecture with reference to the accompanying drawings.

[0031] Please see Figure 1 This is a schematic diagram of a typical Flink cluster architecture, such as... Figure 1 As shown, the Flink cluster adopts a standard master-slave architecture, with the core components being the JobManager and TaskManager. The JobManager acts as the scheduling hub, responsible for receiving jobs submitted by clients, transforming the program logic into a logical dependency graph (JobGraph), and further generating a physical execution graph (ExecutionGraph) containing specific parallelism, resource requirements, and data exchange relationships. Each node (ExecutionVertex) in the ExecutionGraph corresponds to an operator instance that can be executed in parallel. The JobManager schedules these instances to appropriate nodes for execution based on the resources (Task Slots) reported by the TaskManager.

[0032] The TaskManager is the worker node that actually performs the computation. Each TaskManager contains multiple Task Slots. Each Slot is the smallest unit of resource allocation. An operator instance will exclusively or share (through a slot sharing mechanism) one Slot to run. Data is transmitted between operator instances in a streaming manner: the upstream instance serializes the data and pushes it to the downstream instance through the network, and the downstream instance deserializes it and continues processing, forming a complete data processing pipeline.

[0033] It should be noted that, Figure 1 For the simplified demonstration of the Flink cluster architecture, operator chain optimization is not shown (i.e., in real-world scenarios, multiple operator instances can run in the same slot as an operator chain).

[0034] To make the above-mentioned objectives, features and advantages of the embodiments of this application more apparent and understandable, the technical solutions of the embodiments of this application will be further described in detail below with reference to the accompanying drawings.

[0035] Please see Figure 2 This is a flowchart illustrating a method for operator-level hot migration of long-cycle jobs in a Flink cluster, provided in an embodiment of this application. This method can be applied to the JobManager in a Flink cluster, such as... Figure 2 As shown, the Flink cluster long-cycle job operator-level hot migration method may include the following steps: Step 201: For any operator instance, upon receiving an alarm for that operator instance, create and start a shadow task instance corresponding to that operator instance; wherein, the alarm is used to indicate that the operator instance meets the hot migration conditions, and the operator instance meeting the hot migration conditions is determined based on the monitored operating status parameters of the operator instance; the operating status parameters include performance parameters used to characterize the operating health of the operator instance.

[0036] In this embodiment of the application, in order to improve the reliability of Flink cluster operation, the health status of operator instances in the cluster can be monitored to determine whether there are operator instances that meet the hot migration conditions.

[0037] For example, a monitoring process can be deployed for each TaskManager node in the Flink cluster to monitor the health of each operator instance in real time and determine whether the operator instance meets the hot migration conditions.

[0038] For example, for any TaskManager node, if it is determined that any operator instance meets the hot migration conditions, an alarm can be reported to the JobManager.

[0039] In one example, determining whether an operator instance meets the hot migration condition based on the monitored runtime status parameters of the operator instance may include: Based on the monitored running status parameters of the operator instance, determine the health status of the operator instance; If the health of the operator instance is less than the preset health threshold, the operator instance is determined to meet the hot migration condition. The operating status parameters include at least one of the following parameters: The time taken for full garbage collection, the utilization rate of metadata space, and the rate of memory leaks.

[0040] For example, the TaskManager node can monitor the running status of operator instances in real time through the monitoring process, and determine the health of operator instances based on the monitored running status parameters.

[0041] For example, the running status parameters include, but are not limited to, at least one of the following parameters: The time taken for full garbage collection, the utilization rate of metadata space, and the rate of memory leaks.

[0042] For example, for any operator instance, if the health of the operator instance is determined to be less than a preset health threshold (which can be set according to the actual scenario), the operator instance can be determined to meet the hot migration condition (also known as the operator instance being in a "sub-healthy" state).

[0043] The "sub-healthy" state refers to a situation where the working surface is still running, but its performance is gradually deteriorating (increased latency, decreased throughput), and may eventually crash or become unavailable, rather than an immediate failure.

[0044] As an example, runtime status parameters may include: full garbage collection time, metaspace utilization, and memory leak rate; The determination of the health of the operator instance based on the monitored running status parameters may include: The health of the operator instance is determined based on its full garbage collection time score, metaspace utilization score, and memory leak rate score. The full garbage collection time score of the operator instance is determined based on the average full garbage collection time of the operator instance within a preset statistical period and the full garbage collection time threshold. The metaspace utilization score of an operator instance is determined based on the current metaspace utilization of the operator instance and the metaspace utilization threshold. The memory leak rate score of an operator instance is determined based on the memory leak rate of the operator instance in a preset statistical period and the memory leak rate threshold.

[0045] For example, runtime parameters include full garbage collection time, metaspace utilization, and memory leak rate.

[0046] For any operator instance, the full garbage collection time, metaspace utilization, and memory leak rate of the operator instance can be monitored to determine the full garbage collection time score, metaspace utilization score, and memory leak rate score of the operator instance, and then determine the health of the operator instance.

[0047] For example, the health of an operator instance can be determined in the following ways: in, , , The weighting coefficients are non-negative and satisfy the following conditions: For example, the possible values ​​are: =0.4, =0.3, =0.3.

[0048] This is the average time taken for Full GC within a preset statistical period. This is the time threshold corresponding to this indicator; This represents the current usage rate of the metaspace. This is the threshold for the utilization rate of the metaspace; To set the operator memory leak rate within a preset statistical period, This is the leakage rate threshold.

[0049] The health level H can be in the range of [0, 1]. When H < 0.8 (taking a health level threshold of 0.8 as an example), the operator instance can be determined to be in a "sub-healthy" state, which satisfies the hot migration condition.

[0050] In this embodiment of the application, for any operator instance, when the JobManager receives an alarm for that operator instance, it can create and start a shadow task instance corresponding to that operator instance (which may be called the target operator instance or the old operator instance).

[0051] For example, the JobManager can call the extended ExecutionGraph dynamic update interface (based on the FlinkSavepoint extension mechanism, without modifying the Flink core source code) to request a slot from the pre-created hot resource pool (reserving idle slots that match the resource specifications of the target operator), launch a Shadow Task with the same configuration and business logic as the original instance, and complete the replacement preparation.

[0052] Step 202: Synchronize the state of the shadow task instance and the operator instance using MMAP and incremental residual synchronization.

[0053] Step 203: Switch the communication between the shadow task instance and the operator instance.

[0054] In this embodiment of the application, when the JobManager creates and starts a Shadow Task corresponding to the old operator instance, it can perform state synchronization (also known as state alignment) and communication switching between the Shadow Task and the old operator instance.

[0055] In this embodiment of the application, in order to improve the efficiency of state synchronization, MMAP and incremental residual synchronization can be used to synchronize the state of Shadow Task and old operator instance, avoiding the overhead of downloading the entire remote state.

[0056] In one example, the above-mentioned method of synchronizing the state of the shadow task instance and the operator instance using MMAP and incremental residual synchronization can include: Control the shadow task instance to preload the operator instance's most recent checkpoint via MMAP and intercept the task instance's external output behavior; The operator instance is controlled to determine the state residual within the migration window and synchronize the state residual to the shadow task instance.

[0057] For example, the Shadow Task can load the most recent checkpoint of an old operator instance in read-only mode based on MMAP.

[0058] If the checkpoint is stored in distributed storage (such as HDFS / S3), the checkpoint file can be cached on the local disk of the node where the Shadow Task is located, and then mapped to the process virtual memory through MMAP. This avoids copying the full checkpoint data to the JVM heap memory and allows direct access through memory address, effectively reducing memory overhead and loading time.

[0059] For example, the external output behavior of the Shadow Task can be intercepted before the Shadow Task successfully replaces the old operator instance, and the interception operation can be canceled after the Shadow Task successfully replaces the old operator instance.

[0060] Considering that the old operator instance may generate some new state data compared to the most recent checkpoint during the window migration period, in order to ensure that the state of the Shadow Task is consistent with that of the old operator instance, it is also necessary to synchronize the state data of the old operator instance during the migration window period (which can be called the state residual) to the Shadow Task.

[0061] As an example, the state residuals within the migration window are determined by this operator instance in the following way: The old operator instance asynchronously creates a first-time read-only snapshot of its own first state through the copy-on-write mechanism of the state backend; where the first moment is the moment when preparation for switching begins. The old operator instance asynchronously creates a second state read-only snapshot of itself at the second moment through the write-on-write copy mechanism of the state backend; where the second moment is the moment when the upstream node is notified to add the LSN alignment sequence number; Based on the first state read-only snapshot, the second state read-only snapshot, and the incremental state data corresponding to the LSN alignment sequence number, the state residual within the migration window period is determined.

[0062] For example, operator instance switching can include three processes: starting preparation for switching, notifying upstream nodes to add LSN alignment sequence numbers, and switching.

[0063] For example, the currently running old operator instance can operate on its own real-time running state: the old operator instance asynchronously creates a read-only snapshot of its own state at the first moment (the moment it begins preparing to switch) through the copy-on-write (COW) mechanism of the state backend (which can be called the first state read-only snapshot), i.e., a real-time state image. This process does not block the normal business processing of the old operator instance, and avoids state read / write conflicts.

[0064] The old operator instance asynchronously creates a second time-space read-only snapshot of its own state (which can be called a second-state read-only snapshot) through the COW mechanism of the state backend, i.e., a real-time state mirror. To avoid read / write conflicts.

[0065] Furthermore, the state residuals within the migration window can be determined based on the first state read-only snapshot, the second state read-only snapshot, and the incremental state data corresponding to the LSN alignment sequence number.

[0066] The incremental data corresponding to the LSN alignment sequence number is used to compensate for the differences in the state added by the old operator after the Shadow Task preloads the baseline state, so as to ensure that the state of the old and new operators is consistent.

[0067] For example, based on LSN time series, the state residuals within the migration window can be calculated in the following way: in, Align the incremental data corresponding to the LSN sequence number.

[0068] For example, the old instance sends residual data to the Shadow Task in real time via an encrypted RPC channel, and verifies the LSN timing during the synchronization process to avoid data out of order.

[0069] In one example, the above-mentioned communication switching between the shadow task instance and the operator instance may include: A dynamic mount command is sent to the downstream operator instance that has data flow interaction with the operator instance, so that the downstream operator instance that receives the dynamic mount command can mount the network link of the shadow task instance through an idle physical port. Based on the operator type of the operator instance, redirect the data input path between the operator instance and the shadow task instance.

[0070] For example, switching communication between the Shadow Task and the old operator instance may include redirecting the network link of the downstream operator, as well as redirecting the data input path.

[0071] For example, JobManager sends a dynamic mount instruction, such as the Dynamic_Mount_Signal preamble, to downstream operator instances that have data flow interactions with the old operator instance, notifying them to prepare for link switching.

[0072] Downstream operator instances can be based on the VIL architecture, reusing the Selector in the existing network thread pool's Reactor model, dynamically detecting and opening new unoccupied physical ports for listening (without blocking existing connections), and registering the Shadow Task's physical address (IP + port) to the local address mapping table to achieve seamless mounting.

[0073] For redirecting data input paths, different methods can be used depending on the type of the old operator instance.

[0074] For example, the type of operator instance can include Source operator instance, non-Source operator instance, such as Transform operator instance or sink operator instance, etc.

[0075] As an example, the above-mentioned data input path redirection between the operator instance and the shadow task instance based on the operator type of the operator instance can include: When the operator instance is a Source operator instance, control the operator instance to stop consuming external data and report the last consumed position; The broadcast increments the epoch so that the shadow task instance can preempt the data source consumption rights based on the epoch and start consuming from the last consumption point.

[0076] For example, when the old operator instance is a Source operator instance, if the old Source receives a migration instruction, it can stop consuming external data sources and record the last consumption point. And report it to JobManager.

[0077] The JobManager broadcasts the incremented Epoch (which can be denoted as Epoch_new), and the Shadow Task preempts the data source consumption rights based on this Epoch.

[0078] Among them, Epoch is unique and increments to avoid permission conflicts.

[0079] If the data source (such as Kafka) natively supports Epoch verification, it can be requested directly; otherwise, requests for old operator instances are rejected through a custom interceptor on the Flink Source side. Shadow Task from Start consuming data to ensure data continuity.

[0080] As an example, the above-mentioned data input path redirection between the operator instance and the shadow task instance based on the operator type of the operator instance can include: When the operator instance is a Transform operator instance or a sink operator instance, an LSN coloring start instruction is sent to the direct upstream operator of the operator instance so that the direct upstream node adds a unique LSN timing mark to each output data stream and sends the colored data stream to the operator instance and the shadow task instance respectively.

[0081] For example, if the old operator instance is a Transform operator instance or a sink operator instance, JobManager can also send an LSN staining start instruction, such as an LSN_Stain_Start trigger instruction, to the direct upstream operator of the old operator instance.

[0082] Among them, the direct upstream operator receives the LSN coloring start instruction through the native RPC communication channel of the Flink cluster, without restarting or blocking services.

[0083] When the upstream operator receives the LSN coloring start command, it performs the LSN coloring operation: it marks each output data stream with a unique LSN timing tag, and sends the colored data streams to the old operator instance and the Shadow Task respectively, ensuring that the data sequences received by the two are consistent.

[0084] For example, the direct upstream operator can generate LSN time series markers based on cluster ID + operator ID + timestamp + auto-incrementing sequence to ensure that the generated time series markers are globally unique.

[0085] In one example, after redirecting the data input path between the operator instance and the shadow task instance based on the operator type of the operator instance, the above may further include: Once it is determined that all downstream operators of the operator instance have completed the data flow switch, update the cluster operator instance mapping table, mark the shadow task instance as a formal instance, destroy the operator instance and reclaim resources; For any downstream operator, the data stream switching of that downstream operator includes: Upon receiving an LSN packet carrying a recovery indication flag from a shadow task instance, the network link with the operator instance is atomically disconnected, and the data flow is routed to the shadow task instance according to the address mapping table; wherein, the LSN packet carrying the recovery indication flag is sent by the shadow task instance after determining that the completion status is aligned.

[0086] For example, when the Shadow Task completes state alignment, it can send an LSN packet carrying a recovery indication flag (such as the Resume_Flag flag) to downstream operator instances to notify them to switch data streams.

[0087] When the downstream operator detects the recovery indication flag, it atomically disconnects the network link with the old operator instance and routes the data stream to the Shadow Task according to the address mapping table.

[0088] For example, the switching delay is ≤50ms (adjustable, without limiting the protection range).

[0089] Once JobManager confirms that all downstream operator instances have been switched over, it can update the cluster operator instance mapping table, mark the Shadow Task as a working instance, destroy the old instance and reclaim resources, and complete the migration.

[0090] It can be seen that, in Figure 1 In the illustrated method flow, for any operator instance, upon receiving an alarm for that operator instance, a shadow task instance corresponding to that operator instance is created and started. The shadow task instance and the operator instance are synchronized in state using MMAP and incremental residual synchronization. Furthermore, communication switching is performed between the shadow task instance and the operator instance. By monitoring the operator instance's running status parameters, and determining that the operator instance meets the hot migration conditions based on the monitored running status parameters, the shadow task instance can be started without stopping the global job. By using MMAP and incremental residual synchronization to synchronize the state of the shadow task instance and the operator instance, the overhead of downloading the entire remote state is reduced, effectively improving state recovery efficiency, thereby enhancing the efficiency of hot migration.

[0091] In one example, the Flink cluster long-cycle job operator-level hot migration method provided in this application embodiment may further include: If a live migration anomaly is confirmed, a live migration rollback is performed.

[0092] For example, the presence of thermal migration anomalies may include, but is not limited to, at least one of the following: The state synchronization between the shadow task instance and the operator instance failed. The state synchronization between the shadow task instance and the operator instance timed out. The communication switch between the shadow task instance and the operator instance failed.

[0093] For example, a failure to synchronize the Shadow Task with the old operator instance state instance may include a failure to load the Shadow Task state, such as a failure to preload the operator instance's most recent checkpoint via MMAP.

[0094] The timeout for synchronizing the Shadow Task with the old operator instance state instance can include a residual merging timeout (e.g., a threshold of 3 seconds).

[0095] The failure to switch communication between the Shadow Task and the old operator instance state instance can include situations where the old operator instance is a Source operator instance, or the Shadow Task fails to preempt data source consumption rights.

[0096] For example, if JobManager determines that a hot migration exception has occurred, it can issue a rollback command, the downstream operator instance cancels the Shadow Task network mount, the old operator instance continues to provide normal service, the migration process is terminated, and it can be triggered again after the exception is resolved.

[0097] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of this application, the technical solutions provided in the embodiments of this application are described below with reference to specific examples.

[0098] Please see Figure 3 This is a schematic diagram illustrating the implementation process of an operator-level hot migration scheme for long-cycle jobs in a Flink cluster, as provided in this embodiment. Figure 3 As shown, the implementation process of this Flink cluster long-cycle job operator-level hot migration scheme can include: 1. Operator instance "Sub-health" state recognition and shadow task hot-launch.

[0099] For example, a monitoring process can be deployed on each TaskManager node to calculate the health of operator instances in real time using a health metric formula, thereby identifying "sub-healthy" states. in, , , The weighting coefficients are non-negative and satisfy the following conditions: For example, the possible values ​​are: =0.4, =0.3, =0.3.

[0100] This is the average time taken for Full GC within a preset statistical period. This is the time threshold corresponding to this indicator; This represents the current usage rate of the metaspace. This is the threshold for the utilization rate of the metaspace; To set the operator memory leak rate within a preset statistical period, This is the leakage rate threshold.

[0101] The health level H can be in the range of [0, 1]. When H < 0.8 (taking a health level threshold of 0.8 as an example), the operator instance can be determined to be in a "sub-healthy" state, and the TaskManager will report an alarm to the JobManager.

[0102] When the JobManager receives an alarm, it calls the extended ExecutionGraph dynamic update interface (based on the Flink Savepoint extension mechanism, without modifying the Flink core source code), requests a slot from the pre-created hot resource pool (reserving idle slots that match the resource specifications of the target operator), and starts a ShadowTask with the same configuration and business logic as the original instance to complete the replacement preparation.

[0103] For example, the above processing can accurately identify and replace abnormal operator instances, avoiding a global restart.

[0104] Among them, the three-dimensional health measurement model (GC, metaspace, memory leak) with adjustable weights can accurately determine sub-health conditions, and the threshold can be dynamically adjusted to adapt to different cluster states.

[0105] Furthermore, based on the Flink Savepoint extended EG dynamic update interface, there is no need to regenerate JG or global scheduling, and the Shadow Task with the same instance as the original operator can be quickly launched without stopping the global job.

[0106] 2. State preloading and side effect shielding.

[0107] Shadow Task is based on MMAP and loads the state of the primitive operator's most recent checkpoint in read-only mode.

[0108] For example, if the checkpoint is stored in distributed storage (such as HDFS / S3), the checkpoint file can be cached on the local disk of the node where the Shadow Task is located, and then mapped to the process virtual memory through MMAP. This avoids copying the full checkpoint data to the JVM heap memory and allows direct access through memory address, effectively reducing memory overhead and loading time.

[0109] Enable side effect isolation mechanism: Insert interception logic into the operator call chain through the UserFunctionWrapper wrapper class; when the Shadow Task context is detected, intercept all external output behaviors (such as database writes and RPC calls) so that it only silently processes data and updates internal state without producing external side effects.

[0110] The above processing can optimize the state loading method and avoid data duplication.

[0111] Specifically, by using the UserFunctionWrapper wrapper class or bytecode enhancement technology, the external output behavior of the shadow task is intercepted, thereby achieving side effect isolation.

[0112] 3. Dynamic mounting of network access paths.

[0113] JobManager can send a Dynamic_Mount_Signal preamble to all downstream operator instances that have data flow interactions with the old operator instance, notifying them to prepare for link switching.

[0114] Downstream operator instances can be based on the VIL architecture, reusing the Selector in the existing network thread pool's Reactor model, dynamically detecting and opening new unoccupied physical ports for listening (without blocking existing connections), and registering the Shadow Task's physical address (IP + port) to the local address mapping table to achieve seamless mounting.

[0115] The above processing can achieve lightweight network adaptation and avoid downstream operator restarts.

[0116] 4. Transfer of ownership at the source end.

[0117] When the old operator instance is a Source operator instance, and the old Source receives a migration instruction, it can stop consuming external data sources and record the last consumption point. And report it to JobManager.

[0118] The JobManager broadcasts the incremented Epoch, and the Shadow Task preempts the data source consumption rights based on that Epoch.

[0119] Among them, Epoch is unique and increments to avoid permission conflicts.

[0120] If the data source (such as Kafka) natively supports Epoch verification, it can be requested directly; otherwise, requests for old operator instances are rejected through a custom interceptor on the Flink Source side. Shadow Task from Start consuming data to ensure data continuity.

[0121] The above processing ensures data consistency for Source operator instances.

[0122] Specifically, data source consumption rights are preempted through incremental Epochs, combined with... Locator records ensure that new Source operator instances start consuming from the breakpoint, guaranteeing data continuity and exactly-once semantics.

[0123] 5. Incremental residuals are closed.

[0124] For example, after confirming that the shadow task has completed preloading, side effect isolation, and network path mounting, JobManager can not only send the Dynamic_Mount_Signal instruction to the downstream operator instance in the manner described above, but also send the LSN_Stain_Start trigger instruction to all direct upstream operators of the old operator instance.

[0125] The direct upstream operator receives the instruction through the Flink cluster's native RPC communication channel, without requiring a restart or causing any service disruption.

[0126] When the upstream operator receives the LSN_Stain_Start trigger instruction, it performs the LSN staining operation: it marks each output data stream with a unique LSN timing mark, and sends the stained data stream to the old operator instance and the Shadow Task respectively, ensuring that the data sequence received by the two is consistent.

[0127] Based on this, the currently running old operator instance can manipulate its own real-time running state: the old operator instance asynchronously creates a read-only snapshot of its own state at the first moment through the copy-on-write (COW) mechanism of the state backend, that is, a real-time state mirror. This process does not block the normal business processing of the old operator instance, and avoids state read / write conflicts.

[0128] The old operator instance asynchronously creates a second-time read-only snapshot of its own state, i.e., a real-time state mirror, through the COW mechanism of the state backend. To avoid read / write conflicts.

[0129] Based on LSN time series, the state residuals within the migration window are calculated using the following method: in, Align the incremental data corresponding to the LSN sequence number.

[0130] For example, the old instance sends residual data to the Shadow Task in real time via an encrypted RPC channel, and verifies the LSN timing during the synchronization process to avoid data out of order.

[0131] Through the above processing, incremental state synchronization is achieved, ensuring the consistency of state between the old and new operator instances.

[0132] 6. Atomic switching and resource recycling.

[0133] For example, when the Shadow Task completes state alignment, it can send an LSN packet carrying the Resume_Flag flag to downstream operator instances to notify them to switch data streams.

[0134] If the downstream operator recognizes the Resume_Flag flag, it atomically disconnects the network link with the old operator instance and routes the data stream to the Shadow Task according to the address mapping table.

[0135] For example, the switching delay is ≤50ms (example value, adjustable, does not limit the protection range).

[0136] Once JobManager confirms that all downstream operator instances have been switched over, it can update the cluster operator instance mapping table, mark the Shadow Task as a working instance, destroy the old instance and reclaim resources, and complete the migration.

[0137] Through the above processing, uninterrupted switching (extremely low interruption time, on the order of milliseconds) can be achieved, and resources can be automatically reclaimed.

[0138] 7. Abnormal rollback mechanism.

[0139] For example, if the Shadow Task fails to load, the residual merge times out (e.g., more than 3 seconds), or the new Source fails to preempt permissions during the migration process, the JobManager issues a rollback command, the downstream operator instance cancels the Shadow Task network mount, the old operator instance continues to provide normal service, the migration process terminates, and can be triggered again after the exception is resolved.

[0140] The above processing enables rollback in case of migration anomalies, improving the reliability of Flink cluster operation.

[0141] The above process achieves operator-level hot migration with no business interruption (extremely low service interruption time, and business is basically unaware of it), high consistency, and high efficiency through core steps such as "sub-health identification - shadow revive - state synchronization - network mounting - permission preemption - atomic switching", and includes a multi-dimensional anomaly rollback mechanism.

[0142] As can be seen, this embodiment achieves operator-level hot-plugging, with extremely low latency (≤50ms) only during the atomic switching phase. Compared with the existing Savepoint restart scheme, it reduces service interruption time from minutes to milliseconds, with virtually no impact on the business.

[0143] Secondly, by using MMAP memory mapping and incremental residual synchronization, the overhead of downloading the entire remote state is avoided, and the state recovery efficiency is improved by more than 90%.

[0144] Finally, through the built-in multi-dimensional rollback mechanism, in the event of Shadow Task state loading failure, residual merging timeout, or new Source failing to preempt permissions, JobManager can issue a rollback command, and the original operator instance continues to provide services, ensuring uninterrupted business operations.

[0145] The method provided in this application has been described above. The apparatus provided in this application is described below: Please see Figure 4 This is a schematic diagram of a Flink cluster long-cycle job operator-level hot migration device provided in an embodiment of this application. The Flink cluster long-cycle job operator-level hot migration device can be implemented in the JobManager of the Flink cluster, such as... Figure 4 As shown, the Flink cluster long-cycle job operator-level hot migration device may include: A creation unit is used to create and start a shadow task instance corresponding to any operator instance when an alarm is received for that operator instance; wherein, the alarm is used to indicate that the operator instance meets the hot migration conditions, and the operator instance meeting the hot migration conditions is determined based on the monitored operating status parameters of the operator instance; the operating status parameters include performance parameters used to characterize the operating health of the operator instance. The synchronization unit is used to synchronize the state of the shadow task instance and the operator instance through memory-mapped file MMAP and incremental residual synchronization. The switching unit is used to switch the communication between the shadow task instance and the operator instance.

[0146] For example, the specific implementation process of operator-level hot migration of long-cycle jobs in the Flink cluster can be found in the relevant descriptions in the above method embodiments.

[0147] Please see Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. The electronic device may include a processor 501 and a machine-readable storage medium 502 storing machine-executable instructions. The processor 501 and the machine-readable storage medium 502 can communicate via a system bus 503. Furthermore, by reading and executing the machine-executable instructions in the machine-readable storage medium 502 corresponding to the operator-level hot migration logic for long-cycle jobs in a Flink cluster, the processor 501 can execute the Flink cluster long-cycle job operator-level hot migration method described above.

[0148] The machine-readable storage medium 502 mentioned herein can be any electronic, magnetic, optical, or other physical storage device that can contain or store information such as executable instructions, data, etc. For example, a machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid-state drive, any type of storage disk (such as optical disc, DVD, etc.), or similar storage media, or combinations thereof.

[0149] This application also provides a machine-readable storage medium including machine-executable instructions, such as... Figure 5 The machine-readable storage medium 502 in the message transmission device contains machine-executable instructions that can be executed by the processor 501 in the message transmission device to implement the Flink cluster long-cycle job operator-level hot migration method described above.

Claims

1. A method for operator-level hot migration of long-cycle jobs in a Flink cluster, characterized in that, include: For any operator instance, upon receiving an alarm for that operator instance, a shadow task instance corresponding to that operator instance is created and started; wherein, the alarm is used to indicate that the operator instance meets the hot migration conditions, and the operator instance meeting the hot migration conditions is determined based on the monitored operating status parameters of the operator instance; the operating status parameters include performance parameters used to characterize the operating health of the operator instance. The state of the shadow task instance and the operator instance is synchronized using memory-mapped files (MMAP) and incremental residual synchronization; and... The communication between the shadow task instance and the operator instance is switched. The step of synchronizing the state of the shadow task instance and the operator instance through memory-mapped file MMAP and incremental residual synchronization includes: Control the shadow task instance to preload the operator instance's most recent checkpoint via MMAP, and intercept the external output behavior of the shadow task instance; The operator instance is controlled to determine the state residual within the migration window and synchronize the state residual to the shadow task instance.

2. The method according to claim 1, characterized in that, Based on the monitored runtime status parameters of the operator instance, it is determined that the operator instance meets the hot migration conditions, including: Based on the monitored running status parameters of the operator instance, determine the health status of the operator instance; If the health of the operator instance is less than the preset health threshold, the operator instance is determined to meet the hot migration condition. The operating status parameters include at least one of the following parameters: The time taken for full garbage collection, the utilization rate of metadata space, and the rate of memory leaks.

3. The method according to claim 2, characterized in that, The operational status parameters include: full garbage collection time, metaspace utilization, and memory leak rate; The process of determining the health of an operator instance based on its monitored operational status parameters includes: The health of the operator instance is determined based on its full garbage collection time score, metaspace utilization score, and memory leak rate score. The full garbage collection time score of the operator instance is determined based on the average full garbage collection time of the operator instance within a preset statistical period and the full garbage collection time threshold. The metaspace utilization score of an operator instance is determined based on the current metaspace utilization of the operator instance and the metaspace utilization threshold. The memory leak rate score of an operator instance is determined based on the memory leak rate of the operator instance in a preset statistical period and the memory leak rate threshold.

4. The method according to claim 1, characterized in that, The state residuals within the migration window are determined by the operator instance in the following way: The old operator instance asynchronously creates a first-time read-only snapshot of its own first state through the copy-on-write mechanism of the state backend; where the first moment is the moment when preparation for switching begins. The old operator instance asynchronously creates a second state read-only snapshot of itself at the second moment through the write-on-write copy mechanism of the state backend; where the second moment is the moment when the upstream node is notified to add the LSN alignment sequence number; Based on the first state read-only snapshot, the second state read-only snapshot, and the incremental state data corresponding to the log sequence number LSN alignment sequence number, determine the state residual within the migration window period; And / or, The operator instance synchronizes the state residual to the shadow task instance, including: The operator instance synchronizes the state residual to the shadow task instance via an encrypted remote procedure call (RPC) channel.

5. The method according to claim 1, characterized in that, The communication switching between the shadow task instance and the operator instance includes: A dynamic mount command is sent to the downstream operator instance that has data flow interaction with the operator instance, so that the downstream operator instance that receives the dynamic mount command can mount the network link of the shadow task instance through an idle physical port. Based on the operator type of the operator instance, redirect the data input path between the operator instance and the shadow task instance.

6. The method according to claim 5, characterized in that, The step of redirecting the data input path between the operator instance and the shadow task instance based on the operator type of the operator instance includes: When the operator instance is a data source (Source) operator instance, control the operator instance to stop consuming external data and report the last consumed position; The epoch is broadcast incremented so that the shadow task instance can preempt the data source consumption right based on the epoch and start consuming from the last consumption point.

7. The method according to claim 5, characterized in that, The step of redirecting the data input path between the operator instance and the shadow task instance based on the operator type of the operator instance includes: When the operator instance is a Transform operator instance or a sink operator instance, an LSN coloring start instruction is sent to the direct upstream operator of the operator instance so that the direct upstream operator adds a unique LSN timing mark to each output data stream and sends the colored data stream to the operator instance and the shadow task instance respectively.

8. The method according to claim 5, characterized in that, After redirecting the data input path between the operator instance and the shadow task instance based on the operator type of the operator instance, the process further includes: Once it is determined that all downstream operators of the operator instance have completed the data flow switch, update the cluster operator instance mapping table, mark the shadow task instance as a formal instance, destroy the operator instance and reclaim resources; For any downstream operator, the data stream switching of that downstream operator includes: Upon receiving an LSN packet carrying a recovery indication flag from a shadow task instance, the network link with the operator instance is atomically disconnected, and the data flow is routed to the shadow task instance according to the address mapping table; wherein, the LSN packet carrying the recovery indication flag is sent by the shadow task instance after determining that the completion status is aligned.

9. The method according to any one of claims 1-8, characterized in that, The method further includes: If a live migration anomaly is confirmed, a live migration rollback is performed. The presence of thermal migration anomalies includes at least one of the following: The state synchronization between the shadow task instance and the operator instance failed. The state synchronization between the shadow task instance and the operator instance timed out. The communication switch between the shadow task instance and the operator instance failed.

10. A Flink cluster long-cycle job operator-level hot migration device, characterized in that, include: A creation unit is used to create and start a shadow task instance corresponding to any operator instance when an alarm is received for that operator instance; wherein, the alarm is used to indicate that the operator instance meets the hot migration conditions, and the operator instance meeting the hot migration conditions is determined based on the monitored operating status parameters of the operator instance; the operating status parameters include performance parameters used to characterize the operating health of the operator instance. The synchronization unit is used to synchronize the state of the shadow task instance and the operator instance through memory-mapped file MMAP and incremental residual synchronization. The switching unit is used to switch the communication between the shadow task instance and the operator instance; The synchronization unit synchronizes the state of the shadow task instance and the operator instance using a memory-mapped file (MMAP) and incremental residual synchronization method, including: Control the shadow task instance to preload the operator instance's most recent checkpoint via MMAP, and intercept the external output behavior of the shadow task instance; The operator instance is controlled to determine the state residual within the migration window and synchronize the state residual to the shadow task instance.

11. An electronic device, characterized in that, The method includes a processor and a machine-readable storage medium storing machine-readable instructions executable by the processor, which in turn cause the processor to perform the method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Instance migration method and device

    CN110069470A

  • Method, device and system for virtual machines migration

    WO2017075989A1