Node migration method and device
By receiving resource release instructions in a distributed cluster, determining redundant resources and migrating tasks to other work nodes, the problems of business outage and task delay caused by resource release are solved, and rapid task migration and resource utilization are achieved.
Patent Information
- Application Number
- CN202510156330.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-02
AI Technical Summary
In the pooled deployment mode of distributed clusters, when resources running in node devices with low resource utilization are released, it may lead to long-term outages and task delays during business failure repair.
By receiving resource release instructions for the work node to be evicted, redundant resources to be applied, and the target application work node is determined in the work node to be applied in the work node to be applied in the distributed cluster, the target task component is created, the task to be run for the work node to be evicted, and these tasks are run in the work node to be evicted.
This method can quickly migrate tasks without waiting for troubleshooting of the work node to be evict, greatly shortening the interruption time of the job or task, and reducing task delay.
Smart Images

Figure CN119922083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a node migration method. The present invention also relates to a node migration device, a computing device, a computer-readable storage medium, and a computer program product. Background Art
[0002] At present, in the pooled deployment mode of distributed clusters, in order to ensure resource utilization, some resources running in node devices with low resource utilization will be released. However, if the released resources happen to be the resources used by the business running on the working node, the resources corresponding to the business need to be repaired when they are evicted. During the business fault repair period, the business will no longer process data, resulting in a long period of interruption time for the business, that is, no data output. When the business is restored, it is also necessary to start processing from the data in the recent past, resulting in task delays. Therefore, there is an urgent need for a method to solve the technical problems of long business interruption time and task delays. Summary of the invention
[0003] In view of this, an embodiment of this specification provides a node migration method. One or more embodiments of this specification also relate to a node migration device, a computing device, a computer-readable storage medium and a computer program product to solve the technical defects existing in the prior art.
[0004] According to a first aspect of an embodiment of this specification, a node migration method is provided, which is applied to a management node in a distributed cluster, wherein the distributed cluster includes the management node and at least two working nodes; The method comprises: Receiving a resource release instruction for a working node to be evicted, wherein the resource release instruction carries a resource parameter to be evicted of the working node to be evicted; Determine the redundant resources to be applied for based on the to-be-evicted resource parameters, and determine the target application work node among the to-be-applied work nodes of the distributed cluster, wherein the to-be-applied work nodes do not include the to-be-evicted work node; Creating a target task component corresponding to the redundant resource to be applied for in the target application work node; Obtaining the tasks to be run of the working node to be evicted; Based on the target task component, the task to be run is run in the target application work node.
[0005] According to a second aspect of an embodiment of this specification, there is provided a node migration device, which is applied to a management node in a distributed cluster, wherein the distributed cluster includes the management node and at least two working nodes; The device comprises: A receiving module is configured to receive a resource release instruction for a working node to be evicted, wherein the resource release instruction carries a resource parameter to be evicted of the working node to be evicted; A determination module is configured to determine the redundant resources to be applied for based on the to-be-evicted resource parameters, and determine a target application work node among the to-be-applied work nodes of the distributed cluster, wherein the to-be-applied work nodes do not include the to-be-evicted work node; A creation module, configured to create a target task component corresponding to the redundant resource to be applied for in the target application work node; An acquisition module is configured to acquire the to-be-run tasks of the to-be-evicted working node; The running module is configured to run the to-be-run task in the target application work node based on the target task component.
[0006] According to a third aspect of an embodiment of this specification, a computing device is provided, including: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned node migration method are implemented.
[0007] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps of the above-mentioned node migration method are implemented.
[0008] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instruction, which implements the steps of the above-mentioned node migration method when executed by a processor.
[0009] An embodiment of the present specification realizes that during the operation of a job or task in a working node, a resource release instruction for a working node to be evicted is received, redundant resources to be applied for are determined through the resource parameters to be evicted carried in the resource release instruction, and a target application working node for creating a target task component corresponding to the redundant resources to be applied for is determined in the working node to be applied for in the distributed cluster, so that after the target task component is successfully created in the target application working node, the task to be run of the working node to be evicted can be re-run in the target application working node based on the target task component, without waiting for the time to repair the fault of the working node to be evicted, thereby greatly shortening the interruption time of the job or task and reducing the task delay. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1is a schematic diagram of a node migration method provided by an embodiment of this specification; Figure 2 is a flowchart of a node migration method provided by an embodiment of this specification; Figure 3 It is a process flow chart of querying resource application progress provided by an embodiment of this specification; Figure 4 is an exception handling flow chart provided by an embodiment of this specification; Figure 5 is a process flow chart of a node migration method provided by an embodiment of this specification; Figure 6 It is a flow chart of a node migration method provided by an embodiment of this specification; Figure 7 It is a structural schematic diagram of a node migration device provided by an embodiment of this specification; Figure 8 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION
[0011] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.
[0012] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0013] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0014] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0015] First, the terms involved in one or more embodiments of this specification are explained.
[0016] Apache Flink: abbreviated as Flink, Flink is a distributed stream processing framework that supports efficient, reliable, and scalable stream data processing and batch data processing.
[0017] TaskManager: abbreviated as TM, usually refers to the component responsible for managing and scheduling tasks in a distributed computing framework. It is responsible for receiving tasks from the JobManager and executing them.
[0018] JobManager: abbreviated as JM, is one of the core components in the distributed computing framework, mainly responsible for the management and coordination of the entire job.
[0019] Operator: It is a data processing operator in the Flink application, which is used to implement specific data processing logic.
[0020] Task: A Task is the smallest execution unit of a Flink application. It is a data processing logic composed of one or more operators.
[0021] Task Slot: TaskManager has one or more Task Slots, and Task Slot can execute one or more Tasks.
[0022] Job: A program responsible for performing a series of business logic processing on data.
[0023] Checkpoint: Also known as snapshot, Flink's checkpoint is a mechanism for ensuring data consistency in stream processing applications, periodically saving application states to persistent storage. Through checkpoint, Flink can recover from the last successful savepoint in the event of a failure, thereby achieving fault tolerance and high availability.
[0024] Failover: When a Flink Job encounters some exceptions during operation, the task will be restarted. During the restart, the task will be automatically restored from the last successful checkpoint. Failover will cause the task to be interrupted, and the impact time is generally about 3 minutes.
[0025] Flink Rest Api: Flink Rest Api refers to a set of RESTful APIs provided by the Flink framework. These APIs can be used to manage and monitor Flink clusters. Flink's REST API can be used on any client with HTTP / HTTPS access rights, such as web browsers, command line tools, scripts, and applications.
[0026] Kubernetes: K8s for short, is an open source container orchestration platform used to automate the deployment, expansion, and management of containerized applications. It helps developers run and maintain distributed systems more efficiently through cluster management, load balancing, and self-healing capabilities.
[0027] Node: virtual machine, physical machine. A Node is a single working machine in a K8s cluster. It can be a virtual machine or a physical machine. It is responsible for running Pods. Each Node contains Kubelet and container runtime components to ensure that Pods run as expected and maintain the health of the cluster.
[0028] Pod: A container group. A Pod is the smallest deployable computing unit in K8s. It usually contains one or more closely related containers and shares network and storage resources. A Pod runs as a single entity in a cluster, providing more efficient resource management and rationality. When Flink is deployed on a K8s cluster, one worker node corresponds to one Pod.
[0029] Out Of Memory, or OOM for short, refers to an error or exception when the system or application is unable to continue running due to the exhaustion of available memory when allocating memory. Usually, when the operating system or virtual machine detects that there is not enough memory to allocate, an OOM error is thrown.
[0030] The method provided in the embodiment of the present application is applied to the pooled deployment scenario of the Flink framework. The resource pooled deployment scenario is relative to the exclusive deployment mode. The following introduces the difference between pooled deployment and exclusive deployment.
[0031] Exclusive deployment means that in the current Flink resource management mode, resource pool business isolation is achieved by exclusively using physical machines in the cluster, and several machines are divided in the same K8s cluster for use by specific business parties. The machine resources of each business party are relatively independent, and only Pods of the Flink computing engine can be run on the machine. Exclusive deployment will bring the following problems: 1. Too many unallocated resources: Because it is an exclusive deployment, after the business party applies for a certain amount of machine resources, usually there is no way to use up all the resources in the resource pool. For example, the business party applies for 1000 cores of resources, but actually only 800 cores are used, and the remaining 200 cores of resources are idle.
[0032] 2. Resource fragmentation: The current machine models are relatively small, generally 32c128g models, which easily cause resource fragmentation. For example, a machine can provide 32 cores of resources. When 30 cores have been allocated, a 3-core pod is requested. The unallocated 2 cores in the machine cannot meet the demand, resulting in resource fragmentation.
[0033] 3. Too many small resource pools: There are many small resource pools for small businesses. There are only one or two machines in a small resource pool, but in fact only 5-10 core resources are used for operations, resulting in a waste of resources.
[0034] 4. The decommissioning process is complicated: When business needs change and a machine needs to be decommissioned or scaled down, the machine to be decommissioned needs to delete all pod-related jobs and restart them. This usually involves restarting all tasks in the resource pool, which is costly and affects business availability.
[0035] 5. Low startup efficiency: It involves manual operations by multiple parties, including business work orders, communication on whether there is a corresponding machine model, startup, labeling, and delivery to users. This process may involve multiple round-trip communications, which is time-consuming and labor-intensive.
[0036] Under the premise that exclusive deployment has the above problems, a pooled deployment solution is proposed. In the pooled deployment mode, a pod of a resource pool can be started on any Node in the K8s cluster. At the same time, a Node can run both the Pod of the Flink computing engine and the non-Flink Pod.
[0037] In order to ensure the overall resource utilization, the pooled cluster will monitor the remaining resources of the cluster in real time. When there are fewer allocatable resources, some machines (Nodes) will be automatically added to the cluster to reduce the overall cluster load to a reasonable level. When there are more allocatable resources, it means that the cluster is currently relatively idle. For the Node with low resource utilization in the cluster, the Pod running on it will be forcibly expelled to achieve the purpose of emptying the machine, that is, no Pod will run on the Node, so that the low-utilization machine can be safely recycled. In the resource pooling deployment mode, each business party shares the same K8s cluster and continuously submits new tasks (creates new Pods) or stops tasks (releases Pods) to the cluster. Therefore, at certain times, the number of Pods running on some machines (Nodes) in the cluster is small, resulting in a low overall resource utilization rate for the Node. In order to improve the overall resource utilization rate of the pooled cluster, it is necessary to evict the Pods running on these machines with low utilization rates, thereby draining the machine and then removing it from the cluster.
[0038] When draining a machine, there is a problem. If the Pod to be evicted happens to be the Pod corresponding to a working node (TaskManager) of a Flink task, then when the Pod is forcibly evicted, a working node of the Flink task will be lost. At this time, the Flink task will have a failover, reapply for a Pod of a new working node, and restart the tasks running in the Pod corresponding to the deleted working node to restore the tasks. Generally, this process will last about 3 minutes. During this period, the Flink task will no longer process data, that is, the business will be affected and suspended for about 3 minutes. For some core tasks, a 3-minute interruption time is unacceptable, but Pod expulsion is also inevitable in the pooled deployment mode.
[0039] Based on this, a node migration method is provided in this specification. This specification also involves a node migration device, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.
[0040] See also Figure 1 , Figure 1 A schematic diagram of a node migration method provided according to an embodiment of the present specification is shown. In actual applications, a user can submit a job to be executed to a management node in a distributed cluster. After receiving the job submitted by the user, the management node will create a central component for the job on the management node of the distributed cluster ( Figure 1(not shown in the figure), the central component will create corresponding task components in one or more working nodes in the distributed cluster according to the resource configuration of the job. Figure 1 As shown in the figure, for the job submitted by the user, the management node (specifically, the central component in the management node) creates task component 1 in worker node 1 and task component 2 in worker node 2. After task component 1 and task component 2 are created, the job submitted by the user will run in the distributed cluster. In this process, the eviction manager will inspect the resource utilization rate of each worker node in the distributed cluster. After detecting and determining that the resource utilization rate of worker node 2 is low, the eviction manager will send a resource release instruction for worker node 2 to the central component in the management node, so that the central component releases the Pod running in worker node 2 according to the received resource release instruction, thereby draining worker node 2.
[0041] The eviction manager is part of the kubelet component in the K8s cluster. It monitors the resource usage of the node and evicts (terminates) the Pod on the node when resources are insufficient to release resources and ensure the stable operation of the node. The eviction manager is deployed on each node of the k8s cluster together with kubelet. It is a functional module inside kubelet and is responsible for monitoring the resource usage of the node.
[0042] After receiving the resource release instruction for worker node 2, the central component will generate a query eviction processing progress request for worker node 2, and send the generated query eviction processing progress request to the eviction manager, so that the eviction manager can query the resource release processing progress and processing results for worker node 2 according to the query eviction processing progress request. The central component can determine the number of task components that need to be evicted and released according to the resource release instruction, and apply for an equal number of redundant task components (i.e., target task components). After the target task component is successfully applied for, another worker node for creating the target task component is determined in the distributed cluster. Figure 1 In this example, the central component creates a target task component in the working node 3 (i.e. Figure 1 After task component 3 is successfully created, the central component will shield worker node 2 so that no other jobs or tasks will be scheduled to worker node 2 in the subsequent process. Then, the central component obtains the tasks to be run on task component 2 from worker node 2.
[0043] After obtaining the task to be run, the task running information of the task to be run is obtained, that is, the task snapshot information obtained after the task to be run is snapshotted. If the configuration time of the task running information of the task to be run is far away from the current time, the historical task running information of the task to be run may be read during the process of restarting and recovering the task to be run, resulting in more data to be processed when restarting and recovering the task to be run. Based on this, in actual applications, it can be determined whether it is necessary to perform snapshot processing on the task to be run based on the completion time of the historical task running information of the task to be run and the completion time of the reference task running information (that is, the task running information most recently from the current time), so that the task to be run can be restarted based on the task running information of the task to be run. Since the working node 2 has been shielded, the restarted task to be run will only be re-run in the task component 3 in the working node 3. At this time, the task to be run running in the task component 2 can be transferred to the task component 3 for execution.
[0044] For the eviction manager, the eviction status of worker node 2 can be queried from the central component according to the query eviction processing progress request, that is, the resource release processing progress and processing result of worker node 2 can be queried. If the resource release processing result of worker node 2 is successful, the Pod in worker node 2 is deleted. If the resource release processing result of worker node 2 is failed or timed out, the eviction manager abandons the resource release processing and records the reason for the failure of resource release.
[0045] An embodiment provided in the present specification realizes that during the running of a job or task in a working node, the resources corresponding to the job or task are released. By applying for resources equivalent to the resources to be released and creating the applied resources on other working nodes, the job or task can be re-run in other working nodes without waiting for the time to repair the fault of the original working node, which greatly shortens the interruption time of the job or task and reduces task delays.
[0046] See also Figure 2 , Figure 2 A flowchart of a node migration method provided according to an embodiment of the present specification is shown. The node migration method is applied to a management node in a distributed cluster. The distributed cluster includes a management node and at least two working nodes. Specifically, the following steps are included: Step 202: Receive a resource release instruction for a working node to be evicted, wherein the resource release instruction carries resource parameters of the working node to be evicted.
[0047] In the pooled cluster mode of actual application, in order to ensure resource utilization, the resources of the worker nodes with low resource utilization or high load will be released. If the resource utilization of a worker node is low, the Pod on the worker node can be released and recreated on other worker nodes. In addition, if the overall load on a worker node is high, the Pod processing throughput on the worker node is significantly lower than that on other worker nodes, resulting in service delays. In this case, the Pod on the worker node also needs to be released.
[0048] Among them, the work node to be evicted refers to the work node that needs to be evicted and released of resources, which can be specifically a work node with low resource utilization or high load. The work node to be evicted is often determined by the eviction manager through inspection. The resource release instruction refers to the instruction generated by the eviction manager for eviction and release of resources for the work node to be evicted. The resource release instruction carries the resource parameters to be evicted of the work node to be evicted. The resource parameters to be evicted refer to the resource parameters that need to be evicted and released by the work node to be evicted. Furthermore, the resource parameters to be evicted can specifically be the Pod list in the work node to be evicted.
[0049] Specifically, the management node can perform subsequent resource expulsion and release processing on the working node to be evicted by receiving the resource release instruction for the working node to be evicted. It should be noted that before receiving the resource release instruction for the working node to be evicted, the central component has been created in the management node. Based on this, the node migration method can be specifically applied to the central component of the management node. In order to ensure that the eviction manager can know the processing progress and processing results of the working node to be evicted, the management node can feedback the processing progress of the working node to be evicted to the eviction manager after receiving the resource release instruction. The specific implementation is as follows: In a specific implementation manner provided by the present application, after receiving a resource release instruction for a working node to be evicted, the method further includes: A request for querying the eviction processing progress is generated, and the request for querying the eviction processing progress is sent to the initiator of the resource release instruction, so that the initiator queries the execution progress of the resource release instruction according to the request for querying the eviction processing progress.
[0050] The request to query the eviction processing progress refers to a request made by the initiator of the resource release instruction to obtain the resource eviction release processing progress and processing results of the working node to be evicted. The initiator of the resource release instruction may be the eviction manager or other terminals that generate resource release instructions for the working node to be evicted.
[0051] Specifically, after receiving the resource release instruction of the working node to be evicted, the management node may generate a query eviction processing progress request for the working node to be evicted, and send the generated query eviction processing progress request to the initiator of the resource release instruction. The initiator of the resource release instruction may query the execution progress of the resource release instruction, that is, the processing progress of the working node to be evicted, through the obtained query eviction processing progress request.
[0052] Furthermore, the initiator can query the management node for the processing progress of the working node to be evicted by polling. If the result is in progress, wait for a fixed interval (for example, 10 seconds) and try the next poll. If the result is completed, it means that the management node has completed the application of the redundant resources to be applied and the restart switching of the task. At this time, the resources to be evicted can be directly deleted. If the result is failure or timeout, it is considered that the resource release instruction processing has failed and the eviction is directly abandoned.
[0053] The management node feeds back the progress of the eviction processing of the working node to be evicted to the initiator, so that the initiator can timely perceive the eviction processing status of the working node to be evicted. If an exception occurs during the eviction processing, the initiator can respond and handle it in time, thereby improving the perception efficiency of abnormal eviction in the case of abnormal eviction processing.
[0054] Step 204: determining redundant resources to be applied for based on the parameters of the resources to be evicted, and determining a target application work node among the work nodes to be applied for in the distributed cluster, wherein the work nodes to be applied for do not include the work node to be evicted.
[0055] After receiving the resource release instruction for the work node to be evicted, the resource parameters of the work node to be evicted can be obtained from the resource release instruction, and the redundant resources to be applied for can be determined based on the resource parameters to be evicted. To ensure the rationality of resource utilization, redundant resources to be applied that are equivalent to the resource parameters to be evicted can be applied. The redundant resources to be applied can be understood as the redundant backup resources corresponding to the Pods that need to release resources in the work node to be evicted.
[0056] In actual applications, when a task is started, the user will define the resource configuration in advance. In the Flink scenario, Flink's current resource management mode is declarative resource management, which emphasizes defining and managing resource requirements through declarative configuration rather than dynamically adjusting resources in an imperative manner. Through declarative management, users only need to define the desired state without worrying about how to achieve this state. For example, taking Flink as an example, 5 task components (i.e. TaskManager Pods) are required, and each TaskManager Pod can run 2 task slots (i.e. Slots). Just send instructions to Flink's resource management module to apply for 10 Slots, and Flink will automatically calculate that these 10 Slots need to apply for 5 TaskManager Pods from the K8s cluster. The K8s cluster will create 5 TaskManager Pods.
[0057] If the management node receives a resource release instruction at this time, it determines the resource parameters to be evicted according to the resource release instruction, which is to evict TaskManager Pod1 and TaskManager Pod2 of the five TaskManager Pods. After receiving the resource release instruction, it will inform Flink's resource management module that in addition to the five TaskManager Pods, two additional TaskManager Pods need to be applied for, that is, four task slots. At this time, the redundant resources to be applied are four task slots, that is, two TaskManager Pods.
[0058] In practical applications, the application process of the redundant resources to be applied is asynchronous. To ensure the accuracy of the redundant resources to be applied, it is necessary to further determine whether the application of the redundant resources to be applied is completed. In the method provided in the embodiment of the present application, after determining to apply for the redundant resources to be applied, the application progress of the redundant resources to be applied can be queried to determine whether the application of the redundant resources to be applied is completed.
[0059] Based on this, in a specific implementation manner provided in this specification, after determining the redundant resources to be applied for based on the resource parameters to be evicted, the method further includes: Query the resource application progress of the redundant resources to be applied for; If the resource application progress is not completed, determining whether the resource application progress has timed out; If the resource application progress has not timed out, continue to query the resource application progress according to a preset time interval until the resource application progress is completed or the resource application progress times out; When the resource application progress times out, it is determined that the resource release instruction fails to execute.
[0060] The resource application progress specifically refers to the application progress of the redundant resources to be applied for. The preset time interval refers to a preset time interval for querying the resource application progress of the redundant resources to be applied for. For example, the preset time interval can be set to 10 seconds, 15 seconds, etc.
[0061] Specifically, after applying for the redundant resources to be applied, query the resource application progress of the redundant resources to be applied, and determine whether the application of the redundant resources to be applied has been completed according to the resource application progress. If the redundant resources to be applied have not been applied for, that is, the resource application progress has not been completed, determine whether the resource application progress has timed out. If the resource application progress has not timed out, the resource application progress of the redundant resources to be applied can continue to be queried according to the preset time interval until the resource application progress of the redundant resources to be applied is completed or the resource application progress times out. If the resource application progress times out, it can be determined that the resource release instruction for the working node to be evicted has failed to execute, and the reason for the failure is that the resource application progress of the redundant resources to be applied has timed out.
[0062] It should be noted that to determine whether the resource application progress has timed out, you can set the time range within which the resource application progress can be applied in advance, and make a judgment based on the time range. For example, it is pre-set that the application of redundant resources to be applied for must be completed within 30 seconds, and the timing starts when the application for redundant resources to be applied for is initiated. The resource application progress is queried at the preset time interval. If the resource application progress has not been completed after more than 30 seconds, it is determined that the resource application progress has timed out. Determining whether the resource application progress has timed out allows the cluster to quickly understand the resource application status and avoid cluster anomalies caused by resource application timeouts.
[0063] The following describes the specific execution process of querying the resource application progress of the redundant resources to be applied for.
[0064] In a specific implementation manner provided in this specification, querying the resource application progress of the redundant resource to be applied for includes: Determine the initial registered resource amount of the currently running task and the redundant resource amount of the redundant resources to be applied for; Query the target registered resource amount of the currently running task; The resource application progress of the redundant resources to be applied for is queried according to the initial registered resource amount, the redundant resource amount and the target registered resource amount.
[0065] Among them, the currently running task refers to the task that is running when the management node receives the resource release instruction. The initial registered resource amount refers to the amount of resources registered in the central component resource manager of the management node when the currently running task is started. Specifically, it can be the number of task components or the number of containers corresponding to the task components. The redundant resource amount is the amount of resources to be applied for redundant resources. The target registered resource amount refers to the amount of resources registered in the central component resource manager of the management node after applying for the redundant resources to be applied for.
[0066] Specifically, the initial registered resource amount of the currently running task and the redundant resource amount of the redundant resources to be applied for are determined respectively, and after applying for the redundant resources to be applied for, the target registered resource amount of the currently running task in the central component resource manager is queried. Furthermore, the resource application progress of the redundant resources to be applied for can be queried based on the initial registered resource amount, the redundant resource amount and the target registered resource amount.
[0067] Further, in a specific implementation manner provided in this specification, querying the resource application progress of the redundant resources to be applied for according to the initial registered resource amount, the redundant resource amount and the target registered resource amount includes: Calculate the resource sum of the initial registered resource amount and the redundant resource amount; When the resource sum is equal to the target registered resource amount, determining that the resource application progress is completed; When the resource sum is not equal to the target registered resource amount, it is determined that the resource application progress is incomplete.
[0068] Specifically, after determining the initial registered resource amount of the currently running task and the redundant resource amount of the redundant resource to be applied, the resource sum result between the initial registered resource amount and the redundant resource amount is calculated, and it is determined whether the resource sum result is equal to the target registered resource amount. If the resource sum result is equal to the target registered resource amount, it means that the redundant resource to be applied has been applied for, that is, the resource application progress is completed; if the resource sum result is not equal to the target registered resource amount, it means that the redundant resource to be applied has not been applied for, that is, the resource application progress is not completed.
[0069] Furthermore, combined with Figure 3 Describe the progress of query resource application. Figure 3 , Figure 3 FIG. 1 shows a process flow chart of querying resource application progress according to an embodiment of this specification. Figure 3As shown in the figure, 10 slots are required to execute the current running task, and each TaskManager Pod corresponds to 2 slots. The management node starts the current running task and creates 5 TaskManager Pods. At this time, the management node receives the resource release instruction and needs to evict and release two of the 5 TaskManager Pods. At this time, it can be determined that an additional 4 slots (corresponding to 2 Pods) are required, that is, the redundant resources to be applied are 4 slots.
[0070] In order to improve the accuracy of querying the resource application progress of the redundant resources to be applied, a resource in place checker can be created to query the resource application progress of the redundant resources to be applied. When starting the current running task, 5 TaskManager Pods are created, that is, the initial registered resource amount is 5 TaskManager Pods, and two Pods need to be expelled and released, that is, the redundant resource amount of the redundant resources to be applied is 2 TaskManager Pods. Therefore, if the application of the redundant resources to be applied is completed, the target registered resource amount registered to the central component resource manager should be 5+2=7 TaskManager Pods. Based on this, it is determined whether the number of TaskManager Pods registered to the central component (i.e. JobManager) resource manager is 7. If so, it means that the application of the redundant resources to be applied is successful, and the subsequent process can be continued; if not, it can be further determined whether the resource application progress of the redundant resources to be applied has timed out. If not, the resource application progress of the redundant resources to be applied can be queried according to the preset time interval. If it has timed out, it can be determined that the application of the redundant resources to be applied has failed, and the reason for the failure is that the resource application progress of the redundant resources to be applied has timed out. At this time, the resource release instruction timeout can be set.
[0071] Furthermore, after the application for the redundant resources is successful, it is necessary to determine the target application work node for using the redundant resources in each application work node in the distributed cluster, so as to create a task component corresponding to the redundant resources in the target application work node.
[0072] The work nodes to be applied for refer to the work nodes in the distributed cluster except the work nodes to be evicted. In practical applications, to evict the Pods in the work nodes to be evicted, it is necessary to create corresponding Pods in other work nodes except the work nodes to be evicted. Therefore, the work nodes except the work nodes to be evicted are regarded as the work nodes to be applied for.
[0073] The target application work node is used to create the target task component based on the redundant resources to be applied. The target task component is the task component created based on the redundant resources to be applied. Using the above example, the redundant resources to be applied are 2 Pods, that is, 4 Slots. Based on this, the number of TaskManager Pods can be determined to be 2, and the 2 TaskManager Pods are the target task components. In actual applications, the management node can determine the target application work node among the work nodes to be applied based on the resource utilization rate of each work node in the distributed cluster.
[0074] An embodiment provided in the present specification can determine which task components need to be evicted based on the parameters of the resources to be evicted carried in the resource release instruction, and then determine the redundant resources to be applied for that need to be additionally applied for based on the resources corresponding to the task components, and after determining the redundant resources to be applied for, further determine the target application work node among the work nodes to be applied for in the distributed cluster, so that in the subsequent process, the target task component corresponding to the redundant resources to be applied can be created in another work node, which facilitates the subsequent migration of the task components in the work node to be evicted to the target work node, thereby reducing the delay of the task running in the target task component.
[0075] Step 206: Create a target task component corresponding to the redundant resource to be applied for in the target application work node.
[0076] After determining the redundant resources to be applied for and the target application work node, the target task component corresponding to the redundant resources to be applied for can be created in the target application work node. The target task component is used to process the running tasks assigned to the task component in the work node to be evicted.
[0077] Since the task components of the work node to be evicted are still registered in the central component manager, if the currently running task is restarted or other tasks are scheduled, the task may be scheduled to the work node to be evicted again. To avoid this problem, the work node to be evicted needs to be shielded. By shielding the work node to be evicted, the currently running task is prevented from being scheduled to the task component of the work node to be evicted again after restarting.
[0078] Based on this, in a specific implementation manner provided in this specification, after creating the target task component corresponding to the redundant resource to be applied for in the target application work node, the method further includes: Shield the working node to be evicted.
[0079] Furthermore, in actual applications, there are various situations in which tasks are scheduled. Therefore, the worker nodes to be evicted can be shielded according to different situations, as follows: In a specific implementation manner provided in this specification, shielding the working node to be evicted includes: rejecting the task slot registration request submitted by the to-be-evicted working node; and / or Release the idle task slots in the working node to be evicted; and / or Release the running task slot of the running task in the working node to be evicted.
[0080] The task slot registration request is a request to register a task slot with the central component resource manager of the management node. An idle task slot is a task slot that is not running any task. A run task is a task that has been completed. A run task slot is a task slot that has run a task.
[0081] Specifically, if the resource manager of the central component receives a task slot registration request sent from a work node to be evicted, the task slot registration request can be rejected to ensure that the task slot from any task component in the work node to be evicted will not be re-registered in the resource manager of the central component. If there are still idle task slots in the task component of the work node to be evicted, the idle task slots can be released. If it is detected that the task has ended, and the task slot where the task is running is located on the task component of the work node to be evicted, the running task slot is released.
[0082] In one embodiment provided in this specification, after creating the target task component corresponding to the redundant resources to be applied for in the target application work node, the work node to be evicted will be shielded to prevent the task from being rescheduled to the work node to be evicted. It should be noted that shielding the work node to be evicted means shielding the tasks scheduled to the work node to be evicted, preventing the task from being scheduled to enter the work node to be evicted, and will not prevent the task from being transferred out of the work node to be evicted.
[0083] Step 208: Obtain the tasks to be run of the working node to be evicted.
[0084] After creating the target task component corresponding to the redundant resources to be applied for in the target application work node and shielding the work node to be evicted, the tasks to be run of the work node to be evicted can be obtained, so that in the subsequent process, based on the created target task component, the tasks to be run can be rerun in the target application work node. It should be noted that the number of tasks to be run can be one or more, and this manual does not impose any limitation on the number of tasks to be run. Tasks to be run can be understood as tasks performed by task components that need to be evicted in the work node to be evicted.
[0085] In actual applications, after restarting a task, Flink will automatically read the status from the last successful checkpoint and restore the task. If the checkpoint time interval configured for the task is long (for example, 10 minutes), when the task to be run is restarted at the 9th minute, only the data of the last checkpoint (9 minutes ago) can be loaded during recovery. This means that the task to be run needs to rerun the tasks within 9 minutes, resulting in a large amount of duplicate data, which increases the workload of restarting the task. In order to solve this problem, in the method provided in the embodiment of the present application, a checkpoint will be actively triggered before restarting the task to be run, that is, a snapshot of the task to be run is taken, and the task is restarted based on the snapshot data, thereby reducing the amount of data processing and shortening the time to restart the task.
[0086] Based on this, in a specific implementation manner provided in this specification, after obtaining the tasks to be run of the working node to be evicted, the method further includes: Performing snapshot processing on the task to be run to obtain task running information of the task to be run; The task running information is saved.
[0087] The task running information refers to the snapshot information obtained after taking a snapshot of the task to be run, which may include the status of each operator of the task to be run, the current data processing progress of the task to be run, and other information.
[0088] Specifically, after obtaining the pending tasks of the working node to be evicted, a snapshot is taken of the pending tasks to obtain task running information of the pending tasks, and the task running information is saved so as to quickly restart the pending tasks based on the task running information of the pending tasks in subsequent processes.
[0089] Snapshot is a data backup technology used to capture the state of a storage volume at a certain moment so that data can be restored when needed. In K8s, snapshots can be created for the storage volumes used by Pods, which is especially important for data persistence and disaster recovery. Snapshots allow data to be quickly restored when a Pod fails or needs to be upgraded, ensuring business continuity and stability. In the method provided in the embodiment of the present application, when data is subsequently recovered, the running task can be quickly restarted based on the snapshot, reducing duplicate data and improving restart efficiency.
[0090] An embodiment provided in the present specification realizes that before triggering the restart of the pending tasks, a snapshot operation is actively performed on the pending tasks of the evicted working node to obtain the task running information of the pending tasks, so that when the tasks are subsequently restored, the tasks can be restored based on the relatively new snapshot information, thereby reducing the amount of data that is repeatedly processed, shortening the delay time of the tasks, and improving the efficiency and accuracy of restarting the pending tasks in the subsequent process.
[0091] Step 210: Based on the target task component, run the task to be run in the target application work node.
[0092] After creating the target task component in the target application work node, you can run the task to be run in the target task component of the target application work node. As mentioned above, in order to improve the efficiency and accuracy of restarting the task to be run, you can restart and run the task to be run based on the task running information of the task to be run. The specific implementation method is as follows: In a specific implementation manner provided in this specification, based on the target task component, running the task to be run in the target application work node includes: Based on the target task component, restart the task to be run in the target application work node; Execute the task to be executed according to the task execution information.
[0093] Specifically, the task to be run is restarted according to the task running information. In actual applications, the work node to be evicted has been blocked, and the task to be run cannot be scheduled to the work node to be evicted after restart. At the same time, a target task component has been created in the target work node. The function and parameters of the target task component are the same as those of the task component in the work node to be evicted. Therefore, after restarting the task to be run, the task to be run can be scheduled to the target task component. Restart and execute the task to be run in the target task component. In this way, the task to be run in the work node to be evicted can be converted to the target application work node for execution.
[0094] In the method provided in the embodiment of the present application, the snapshot information of the task to be run is obtained, and the storage volume of the task to be run is restored based on the snapshot information. According to the created target task component, the storage volume is loaded into the target task component, thereby realizing the operation of scheduling the restarted task to be run into the target task component.
[0095] In actual applications, the migration of pending tasks will last for 3 to 4 minutes. This duration refers to the time from when Flink's JobManager receives the eviction request to when the pending tasks on the working node to be evicted are restarted. During this period of time, there may be some abnormalities in the entire node migration process, resulting in the failure of node migration to execute normally. At this time, in order to improve the feedback speed, the requester is informed of the abnormality in a timely manner, so that the requester can perceive the abnormality in the node migration process in a timely manner and determine the abnormality, without having to wait until the final step of the node migration process to determine the abnormality.
[0096] Abnormalities during node migration mainly include two aspects: One is that an exception occurs during the execution of any link in the node migration process. The task migration process itself requires a series of complex processes, including applying for redundant resources to be applied, obtaining tasks to be run on the working nodes to be evicted, blocking the working nodes to be evicted, triggering checkpoints and waiting for checkpoints to complete, etc. During the execution of any link in these steps, an exception may occur, resulting in failure to complete the execution normally.
[0097] Second, during the node migration process, the task itself may need to be repaired due to other reasons. For example, dirty data that cannot be processed may occur during the node migration, or some TaskManager Pods may exit due to OOM (Out of Memory) exceptions. In this case, you can monitor the task status and determine the exception based on the task status.
[0098] See also Figure 4 , Figure 4 FIG. 2 shows an exception handling flow chart provided according to an embodiment of the present specification. Figure 4 As shown, after receiving the resource release instruction, the management node applies for the redundant resources to be applied for. At this time, it can be determined whether the process of applying for the redundant resources to be applied for is abnormal. If so, it is determined that the result of the expulsion processing of the working node to be evicted is a failure. If not, continue to obtain the tasks to be run of the working node to be evicted. Determine whether the process of obtaining the tasks to be run of the working node to be evicted is abnormal. If so, it is determined that the result of the expulsion processing of the working node to be evicted is a failure. If not, continue to perform snapshot processing on the acquired tasks to be run. Determine whether the process of performing snapshot processing on the tasks to be run is abnormal. If so, it is determined that the result of the expulsion processing of the working node to be evicted is a failure. If not, continue to execute the operation of restarting the tasks to be run. Determine whether the process of restarting the tasks to be run is abnormal. If so, it is determined that the result of the expulsion processing of the working node to be evicted is a failure. If not, it can be determined that the result of the expulsion processing of the working node to be evicted is successful.
[0099] In addition, the running status of the task can also be monitored. If the running status of the task changes to a failed / restarted / cancelled state, the eviction processing result of the work node to be evicted is determined to be a failure; if the running status of the task does not change to a failed / restarted / cancelled state, the running status of the task continues to be monitored.
[0100] By detecting anomalies in any link of the node migration process and monitoring the running status of the task, anomalies that may occur during the node migration process can be detected in a timely manner. The anomalies are fed back to the requester, who can immediately perceive the request failure and give up the eviction without having to wait until the timeout.
[0101] The node migration method provided in the present specification is applied to a management node in a distributed cluster, wherein the distributed cluster includes the management node and at least two working nodes; the method includes: receiving a resource release instruction for a working node to be evicted, wherein the resource release instruction carries resource parameters to be evicted of the working node to be evicted; determining redundant resources to be applied for based on the resource parameters to be evicted, and determining a target application working node among each working node to be applied for in the distributed cluster, wherein the working node to be applied for does not include the working node to be evicted; creating a target task component corresponding to the redundant resources to be applied for in the target application working node; obtaining a task to be run of the working node to be evicted; and running the task to be run in the target application working node based on the target task component.
[0102] An embodiment of the present specification realizes that during the operation of a job or task in a working node, a resource release instruction for a working node to be evicted is received, redundant resources to be applied for are determined through the resource parameters to be evicted carried in the resource release instruction, and a target application working node for creating a target task component corresponding to the redundant resources to be applied for is determined in the working node to be applied for in the distributed cluster, so that after the target task component is successfully created in the target application working node, the task to be run of the working node to be evicted can be re-run in the target application working node based on the target task component, without waiting for the time to repair the fault of the working node to be evicted, thereby greatly shortening the interruption time of the job or task and reducing the task delay.
[0103] The following combination Figure 5 , the node migration method is further described. Figure 5 A flowchart of a node migration method according to an embodiment of the present specification is shown, which specifically includes the following steps: Step 502: Receive a resource release instruction for a working node to be evicted, wherein the resource release instruction carries resource parameters of the working node to be evicted.
[0104] Step 504: Generate a query eviction processing progress request, and send the query eviction processing progress request to the initiator of the resource release instruction.
[0105] Step 506: Determine the redundant resources to be applied for based on the parameters of the resources to be evicted.
[0106] Step 508: Query the resource application progress of the redundant resources to be applied for.
[0107] Step 510: When the resource application progress is not completed, determine whether the resource application progress has timed out. When the resource application progress has not timed out, continue to query the resource application progress according to a preset time interval until the resource application progress is completed.
[0108] Step 512: Determine a target application work node among all the work nodes to be applied for in the distributed cluster, wherein the work nodes to be applied for do not include the work node to be evicted.
[0109] Step 514: Create a target task component corresponding to the redundant resource to be applied for in the target application work node.
[0110] Step 516: Shield the working node to be evicted.
[0111] Step 518: Obtain the tasks to be run of the working node to be evicted.
[0112] Step 520: Snapshot the task to be executed to obtain task execution information of the task to be executed.
[0113] Step 522: Restart and execute the task to be run in the target task component according to the task running information.
[0114] An embodiment of the present specification realizes that during the operation of a job or task in a working node, a resource release instruction for a working node to be evicted is received, redundant resources to be applied for are determined through the resource parameters to be evicted carried in the resource release instruction, and a target application working node for creating a target task component corresponding to the redundant resources to be applied for is determined in the working node to be applied for in the distributed cluster, so that after the target task component is successfully created in the target application working node, the task to be run of the working node to be evicted can be re-run in the target application working node based on the target task component, without waiting for the time to repair the fault of the working node to be evicted, thereby greatly shortening the interruption time of the job or task and reducing the task delay.
[0115] Combine the following Figure 6 , taking the application of this specification in the Flink scenario as an example, the node migration method provided by this application is further explained. Figure 6 This is a flow chart of a node migration method provided by an embodiment of this specification. In this embodiment, the distributed stream processing framework Flink is deployed in a K8s cluster, and the K8s cluster deploys the Flink framework in a pooled deployment manner, including a management node, a working node 1, a working node 2, and a working node 3 in the K8s cluster.
[0116] Step 602: The user submits a Flink task to the management node.
[0117] In this implementation, the user can submit a Flink task to the Flink streaming framework deployed in the K8s cluster through the control terminal.
[0118] Step 604: The management node creates corresponding TaskManagers in worker node 1 and worker node 2 according to the Flink task.
[0119] After receiving a Flink task, the management node will first create the Flink JobManager component (the central manager of the task) in the management node. The JobManager component will create the corresponding TaskManager based on the resource configuration of the Flink task (the number and specifications of TaskManagers required). Figure 6 As shown, the management node sends instructions to the worker node 1 and the worker node 2, and creates TaskManager in the worker node 1 and the worker node 2 respectively.
[0120] Step 606: Worker node 2 sends an eviction request to the management node.
[0121] K8s's kubelet component is deployed in Worker Node 1 and Worker Node 2. The eviction manager in the kubelet component monitors the usage of each worker node. During the daily inspection of Worker Node 2, it is found that the resource utilization rate in Worker Node 2 is low, and Worker Node 2 needs to be drained. At this time, the FlinkTaskManager Pod deployed on Worker Node 2 needs to be evicted. Migrate the Flink TaskManager Pod on Worker Node 2 to other work nodes for continued execution. At this time, the eviction manager in Worker Node 2 generates a list of Pods to be evicted based on the Flink TaskManagerPod corresponding to the Flink task, and generates an eviction request based on the list of Pods to be evicted, and sends the eviction request to the management node. Inform the management node that the relevant Pods in Worker Node 2 need to be evicted.
[0122] Step 608: The management node returns a receiving identifier corresponding to the eviction request.
[0123] After receiving the eviction request, the JobManager in the management node will synchronously return a receiving identifier (request-id) for the eviction request to the eviction manager on the working node 2. The receiving identifier is used to subsequently track the execution result of the eviction request.
[0124] Step 610: The management node sends resource application information of redundant resources to be applied to the working node 3.
[0125] The management node will also asynchronously execute the eviction request in worker node 2. Specifically, the resource information of the redundant resources to be applied for will be determined based on the information in the list of Pods to be evicted, so as to facilitate the migration of the tasks processed by the Flink TaskManager Pod in worker node 2 in the subsequent processing. The management node will also select worker node 3 as the target worker node from the worker nodes other than worker node 2. The selection of worker node 3 as the target worker node can be based on the processing principle of load balancing in the K8s cluster, and this process is not further limited here.
[0126] After determining that the working node 3 is the target working node, the management node sends resource application information for the redundant resources to be applied to the working node 3.
[0127] Step 612: Working node 3 allocates redundant resources and creates corresponding TaskManagers, and sends a creation success message to the management node.
[0128] After receiving the resource application information, worker node 3 allocates the corresponding redundant resources in worker node 3 and creates the corresponding redundant TaskManager based on the redundant resources. After the creation is successful, a message of successful creation is sent to the management node to inform the management node that the redundant TaskManager corresponding to the Flink TaskManager to be evicted in worker node 2 has been created in worker node 3.
[0129] Step 614: The management node blocks the TaskManager in the working node 2 and takes a snapshot of the Flink task on the working node 2.
[0130] After the JobManager in the management node determines that the redundant TaskManager in worker node 3 has been created successfully, it will blacklist the TaskManager in worker node 2 to ensure that no Flink tasks will be scheduled to the TaskManager in worker node 2 in the future. At the same time, it will perform snapshot processing on the TaskManager in worker node 2 and back up the list of running tasks in the TaskManager, which includes the running information of the tasks being run by the TaskManager in worker node 2.
[0131] Step 616: The management node restarts the task in the task list, and schedules the restarted task to the working node 3.
[0132] The management node restarts the tasks in the TaskManager of worker node 2 according to the information in the task list. Since the TaskManager in worker node 2 has been blocked, the restarted tasks cannot be scheduled to worker node 2. After the tasks in the task list are restarted, they will be automatically scheduled to the redundant TaskManager created in worker node 3, thus completing the migration of Flink tasks from worker node 2 to worker node 3.
[0133] Step 618: Worker node 2 polls the eviction status of the Flink TaskManager Pod, and after determining that the eviction is complete, ends the eviction task for the Flink TaskManager Pod.
[0134] After receiving the receiving identifier returned by the management node, the eviction manager in the worker node 2 will periodically ask the JobManager in the management node for the processing result of the eviction request based on the receiving identifier. Once the eviction success information sent by the management node is received, the task of eviction of the Flink TaskManager Pod will be terminated. If the eviction fails, the eviction manager will abandon the eviction and record the reason for the eviction failure for subsequent analysis.
[0135] An embodiment of the present specification realizes that during the operation of a job or task in a working node, a resource release instruction for a working node to be evicted is received, redundant resources to be applied for are determined through the resource parameters to be evicted carried in the resource release instruction, and a target application working node for creating a target task component corresponding to the redundant resources to be applied for is determined in the working node to be applied for in the distributed cluster, so that after the target task component is successfully created in the target application working node, the task to be run of the working node to be evicted can be re-run in the target application working node based on the target task component, without waiting for the time to repair the fault of the working node to be evicted, thereby greatly shortening the interruption time of the job or task and reducing the task delay.
[0136] Corresponding to the above method embodiment, this specification also provides a node migration device embodiment, Figure 7 The schematic diagram of the structure of a node migration device provided by an embodiment of the present specification is shown. The node migration device is applied to a management node in a distributed cluster, and the distributed cluster includes the management node and at least two working nodes, such as Figure 7 As shown, the device comprises: The receiving module 702 is configured to receive a resource release instruction for a working node to be evicted, wherein the resource release instruction carries a resource parameter to be evicted of the working node to be evicted; The determination module 704 is configured to determine the redundant resources to be applied for based on the to-be-evicted resource parameters, and determine the target application work node among the to-be-applied work nodes of the distributed cluster, wherein the to-be-applied work nodes do not include the to-be-evicted work node; A creation module 706 is configured to create a target task component corresponding to the redundant resource to be applied for in the target application work node; An acquisition module 708 is configured to acquire the to-be-run tasks of the to-be-evicted working node; The running module 710 is configured to run the to-be-run task in the target application work node based on the target task component.
[0137] Optionally, the device further includes a query module configured to: Query the resource application progress of the redundant resources to be applied for; If the resource application progress is not completed, determining whether the resource application progress has timed out; If the resource application progress has not timed out, continue to query the resource application progress according to a preset time interval until the resource application progress is completed or the resource application progress times out; When the resource application progress times out, it is determined that the resource release instruction fails to execute.
[0138] Optionally, the query module is further configured to: Determine the initial registered resource amount of the currently running task and the redundant resource amount of the redundant resources to be applied for; Query the target registered resource amount of the currently running task; The resource application progress of the redundant resources to be applied for is queried according to the initial registered resource amount, the redundant resource amount and the target registered resource amount.
[0139] Optionally, the query module is further configured to: Calculate the resource sum of the initial registered resource amount and the redundant resource amount; When the resource sum is equal to the target registered resource amount, determining that the resource application progress is completed; When the resource sum is not equal to the target registered resource amount, it is determined that the resource application progress is incomplete.
[0140] Optionally, the device further includes a shielding module configured to: Shield the working node to be evicted.
[0141] Optionally, the shielding module is further configured as follows: rejecting the task slot registration request submitted by the to-be-evicted working node; and / or Release the idle task slots in the working node to be evicted; and / or Release the running task slot of the running task in the working node to be evicted.
[0142] Optionally, the device further includes a snapshot module configured to: Performing snapshot processing on the task to be run to obtain task running information of the task to be run; The task running information is saved.
[0143] Optionally, the operation module 710 is further configured to: Based on the target task component, restart the task to be run in the target application work node; Execute the task to be executed according to the task execution information.
[0144] Optionally, the device further includes a sending module configured to: A request for querying the eviction processing progress is generated, and the request for querying the eviction processing progress is sent to the initiator of the resource release instruction, so that the initiator queries the execution progress of the resource release instruction according to the request for querying the eviction processing progress.
[0145] The node migration device provided in the present specification is applied to a management node in a distributed cluster, wherein the distributed cluster includes the management node and at least two working nodes; the device includes: a receiving module, configured to receive a resource release instruction for a working node to be evicted, wherein the resource release instruction carries a resource parameter to be evicted of the working node to be evicted; a determining module, configured to determine a redundant resource to be applied for based on the resource parameter to be evicted, and determine a target application working node among each working node to be applied for in the distributed cluster, wherein the working node to be applied for does not include the working node to be evicted; a creating module, configured to create a target task component corresponding to the redundant resource to be applied for in the target application working node; an acquiring module, configured to acquire a task to be run of the working node to be evicted; and a running module, configured to run the task to be run in the target application working node based on the target task component.
[0146] An embodiment of the present specification realizes that during the operation of a job or task in a working node, a resource release instruction for a working node to be evicted is received, redundant resources to be applied for are determined through the resource parameters to be evicted carried in the resource release instruction, and a target application working node for creating a target task component corresponding to the redundant resources to be applied for is determined in the working node to be applied for in the distributed cluster, so that after the target task component is successfully created in the target application working node, the task to be run of the working node to be evicted can be re-run in the target application working node based on the target task component, without waiting for the time to repair the fault of the working node to be evicted, thereby greatly shortening the interruption time of the job or task and reducing the task delay.
[0147] The above is a schematic scheme of a node migration device of this embodiment. It should be noted that the technical scheme of the node migration device and the technical scheme of the node migration method described above are of the same concept, and the details not described in detail in the technical scheme of the node migration device can be found in the description of the technical scheme of the node migration method described above.
[0148] Figure 8 The structure block diagram of a computing device 800 provided according to an embodiment of the present specification is shown. The components of the computing device 800 include but are not limited to a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and a database 850 is used to store data.
[0149] The computing device 800 also includes an access device 840 that enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 840 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.
[0150] In one embodiment of the present specification, the above components of the computing device 800 and Figure 8 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 8 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0151] The computing device 800 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 800 may also be a mobile or stationary server.
[0152] The processor 820 implements the steps of the node migration method when executing the computer program / instructions.
[0153] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the node migration method described above are of the same concept, and the details not described in detail in the technical scheme of the computing device can be found in the description of the technical scheme of the node migration method described above.
[0154] An embodiment of the present specification further provides a computer-readable storage medium storing a computer program / instruction, which implements the steps of the node migration method as described above when the computer program / instruction is executed by a processor.
[0155] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the node migration method described above are of the same concept, and the details not described in detail in the technical scheme of the storage medium can be found in the description of the technical scheme of the node migration method described above.
[0156] An embodiment of the present specification also provides a computer program product, including a computer program / instruction, which implements the steps of the above-mentioned node migration method when executed by a processor.
[0157] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of the computer program product and the technical solution of the node migration method described above are of the same concept, and the details not described in detail in the technical solution of the computer program product can be found in the description of the technical solution of the node migration method described above.
[0158] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0159] The computer program / instruction includes a computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. The computer readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.
[0160] It should be noted that, for the convenience of description, the aforementioned method embodiments are all described as a series of action combinations, but those skilled in the art should be aware that this specification is not limited by the order of the actions described, because according to this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this specification.
[0161] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0162] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A node migration method, characterized in that: Applied to a management node in a distributed cluster, the distributed cluster comprising the management node and at least two working nodes; The method comprises: Receiving a resource release instruction for a working node to be evicted, wherein the resource release instruction carries a resource parameter to be evicted of the working node to be evicted; Determine the redundant resources to be applied for based on the to-be-evicted resource parameters, and determine the target application work node among the to-be-applied work nodes of the distributed cluster, wherein the to-be-applied work nodes do not include the to-be-evicted work node; Creating a target task component corresponding to the redundant resource to be applied for in the target application work node; Obtaining the tasks to be run of the working node to be evicted; Based on the target task component, the task to be run is run in the target application work node.
2. The method according to claim 1, characterized in that After determining the redundant resources to be applied for based on the resource parameters to be evicted, the method further includes: Query the resource application progress of the redundant resources to be applied for; If the resource application progress is not completed, determining whether the resource application progress has timed out; If the resource application progress has not timed out, continue to query the resource application progress according to a preset time interval until the resource application progress is completed or the resource application progress times out; When the resource application progress times out, it is determined that the resource release instruction fails to execute.
3. The method according to claim 2, characterized in that Query the resource application progress of the redundant resource to be applied, including: Determine the initial registered resource amount of the currently running task and the redundant resource amount of the redundant resources to be applied for; Query the target registered resource amount of the currently running task; The resource application progress of the redundant resources to be applied for is queried according to the initial registered resource amount, the redundant resource amount and the target registered resource amount.
4. The method according to claim 3, characterized in that According to the initial registered resource amount, the redundant resource amount and the target registered resource amount, querying the resource application progress of the redundant resource to be applied for includes: Calculate the resource sum of the initial registered resource amount and the redundant resource amount; When the resource sum is equal to the target registered resource amount, determining that the resource application progress is completed; When the resource sum is not equal to the target registered resource amount, it is determined that the resource application progress is incomplete.
5. The method according to claim 1, characterized in that After creating the target task component corresponding to the redundant resource to be applied for in the target application work node, the method further includes: Shield the working node to be evicted.
6. The method according to claim 5, characterized in that Shielding the working node to be evicted includes: rejecting the task slot registration request submitted by the to-be-evicted working node; and / or Release the idle task slots in the working node to be evicted; and / or Release the running task slot of the running task in the working node to be evicted.
7. The method according to claim 1, characterized in that After obtaining the tasks to be run of the working node to be evicted, the method further includes: Performing snapshot processing on the task to be run to obtain task running information of the task to be run; The task running information is saved.
8. The method according to claim 7, characterized in that Based on the target task component, running the task to be run in the target application work node includes: Based on the target task component, restart the task to be run in the target application work node; Execute the task to be executed according to the task execution information.
9. The method according to claim 1, characterized in that After receiving a resource release instruction for the working node to be evicted, the method further includes: A request for querying the eviction processing progress is generated, and the request for querying the eviction processing progress is sent to the initiator of the resource release instruction, so that the initiator queries the execution progress of the resource release instruction according to the request for querying the eviction processing progress.
10. A node migration device, characterized in that: Applied to a management node in a distributed cluster, the distributed cluster comprising the management node and at least two working nodes; The device comprises: A receiving module is configured to receive a resource release instruction for a working node to be evicted, wherein the resource release instruction carries a resource parameter to be evicted of the working node to be evicted; A determination module is configured to determine the redundant resources to be applied for based on the to-be-evicted resource parameters, and determine a target application work node among the to-be-applied work nodes of the distributed cluster, wherein the to-be-applied work nodes do not include the to-be-evicted work node; A creation module, configured to create a target task component corresponding to the redundant resource to be applied for in the target application work node; An acquisition module is configured to acquire the to-be-run tasks of the to-be-evicted working node; The running module is configured to run the to-be-run task in the target application work node based on the target task component.
11. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. It is characterized in that when the computer program / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 9 are implemented.
12. A computer-readable storage medium storing a computer program / instruction, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
13. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.