Node offline control method and device, equipment and medium

CN122845651APending Publication Date: 2026-09-29GUANGZHOU HUANJUMARK NETWORK INFORMATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611069454.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-17
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

在该方案中,节点被标记为下线后即于下次心跳时被通知停止,未设置等待任务完成的步骤,因此无法适用于需要确保异步任务完整执行的场景

Benefits of technology

[0014]相较于传统技术,本申请通过节点配置表与任务状态表的协同配合,结合延迟等待机制与任务忽略名单的动态管理,在执行节点响应下线指令时实现了安全、精准且高效的优雅下线。执行节点接收下线指令后主动将自身在节点配置表中的节点状态更新为失活状态,使调度器停止向该节点分配新异步任务,并通过预设时长的延迟等待消除因状态同步滞后导致的新任务误分配风险。延迟等待结束后,执行节点周期性查询任务状态表中本执行节点对应的执行中任务记录,并对存在前置依赖阻塞的任务进行识别后将其加入任务忽略名单以跳过无效等待,仅等待前置依赖已满足且确实能够在本节点上完成的任务。当确认无执行中任务记录时,执行节点自行执行下线操作,全程无需人工干预。本申请在保证所有可完成的任务确实完成的前提下,跳过了因前置依赖阻塞而无法在本节点上完成的任务,实现了任务完整性与下线效率之间的最佳平衡,显著提升了分布式异步任务系统中节点下线的安全性、效率和可维护性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845651A_ABST
    Figure CN122845651A_ABST
Patent Text Reader

Abstract

This application relates to a node offline control method, apparatus, device, and medium. The method includes: updating the node status of the execution node in the node configuration table maintained by the control node to an inactive state to prevent the scheduler in the control node from assigning new asynchronous tasks to the execution node; performing a delay wait according to a preset delay duration; after the wait is completed, periodically querying the task status table maintained by the control node for task records whose execution status is "in execution" and are not in the task ignore list; when an in-execution task record exists, determining whether the corresponding asynchronous task has any unfinished prerequisite dependent tasks; if so, skipping the wait for the asynchronous task and adding it to the task ignore list; otherwise, continuing to wait for completion; when no in-execution task record exists, performing a offline operation. This application achieves a safe, autonomous, efficient, and graceful offline of asynchronous task nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of microservice architecture technology, and in particular to a node offline control method, apparatus, device, and medium thereof. Background Technology

[0002] In distributed microservice architectures, asynchronous task processing systems are widely used in e-commerce, finance, and the Internet of Things (IoT). These systems typically consist of a control node and multiple execution nodes. The control node runs a scheduler responsible for distributing user-submitted asynchronous tasks to the execution nodes for processing. The execution nodes are responsible for executing the specific tasks and feeding the results back to the control node. To support high concurrency and large-scale task processing, execution nodes are usually deployed in a cluster, with the number of nodes dynamically adjusted through elastic scaling mechanisms to adapt to changes in business load. When an execution node needs to undergo version upgrades, configuration updates, or resource reclamation, it must be safely taken offline from the cluster to prevent interruptions or loss of tasks being processed due to forced process termination.

[0003] In traditional technologies, taking an execution node offline is typically achieved by directly stopping the process or container. While this approach is simple to operate, it has significant technical drawbacks. When there are still asynchronous tasks running on the execution node, directly stopping the process will forcibly interrupt these tasks, preventing their states from being updated correctly. Processed data may become inconsistent, or even result in data loss. For long-chain business processes involving multiple stages and dependencies, the interruption of a task at one stage can halt the entire business chain, leading to extremely high recovery costs.

[0004] To address these issues, the industry has proposed several graceful shutdown solutions. A common approach involves traffic removal through a service registry. This means deregistering the node from the registry before shutting it down, and then stopping the node process once the gateway or load balancer detects the deregistration. While this solution ensures no new requests are routed to the node at the network request level, it only considers the completion status of requests at the transport layer and cannot detect the actual completion status of asynchronous tasks at the business layer. In asynchronous task scenarios, task states are often persistently stored in a database and may execute across threads or coroutines, detaching from the original process's memory space. Therefore, the completion status of requests at the transport layer does not necessarily indicate the completion of asynchronous tasks at the business layer.

[0005] Another approach is to achieve graceful shutdown through consumer-side confirmation. This involves the service provider proactively notifying all consumers to update their local cache lists after pre-deregistration, and then stopping the node only after all consumers confirm that they have stopped sending requests to it. This approach can effectively guarantee the integrity of the call chain in remote procedure call scenarios, but it relies on consumer-side status feedback, and its task concept corresponds to a single remote procedure call, rather than an asynchronous task with complex dependencies at the business layer. For asynchronous tasks with long execution times, distributed across nodes, and with dependencies, the consumer side cannot perceive the actual execution status of these tasks.

[0006] Another approach involves setting up an offline application programming interface (API) within the resource scheduling framework. Administrators call this interface to mark target nodes as offline, and the scheduling framework notifies the node to stop upon its next heartbeat. This approach focuses on node lifecycle management at the resource scheduling level, improving operational convenience rather than waiting for ongoing tasks on the node to complete. In this approach, the node is notified to stop immediately upon its next heartbeat after being marked offline, without any step to wait for task completion. Therefore, it is unsuitable for scenarios requiring guaranteed complete execution of asynchronous tasks.

[0007] In summary, traditional technologies, when execution nodes go offline, either force a direct stop causing task interruption, or perform traffic removal at the network transport or resource scheduling layer without being aware of the actual completion status of asynchronous tasks at the business layer, or lack the ability to identify dependencies between tasks, resulting in either waiting for or skipping all tasks. When asynchronous tasks are persistently stored in a database, executed across threads and nodes, and have complex dependencies, traditional technologies cannot achieve graceful, self-driven shutdown of execution nodes while ensuring the safe completion of all existing tasks. Summary of the Invention

[0008] The purpose of this application is to solve at least one of the above-mentioned problems by providing a node offline control method and corresponding apparatus, devices, non-volatile readable storage media, and computer program products.

[0009] According to one aspect of this application, a node offline control method is provided, comprising: In response to the offline command, the node status of this execution node in the node configuration table maintained by the control node is updated to inactive state to prevent the scheduler in the control node from assigning new asynchronous tasks to this execution node; A delay waiting period is performed according to a preset delay time to compensate for the propagation lag of the inactive state from the node configuration table to the scheduler; After the delay period ends, the task records in the task status table maintained by the control node that are in execution and are not in the task ignore list are periodically queried as query results; When the query results contain a record of an executing task, it is determined whether the corresponding asynchronous task has any unfinished pre-dependent tasks. If so, the waiting for the asynchronous task is skipped and the asynchronous task is added to the task ignore list; otherwise, the asynchronous task is waited for to be completed. If the query results do not contain any records of tasks in progress, the current execution node will be taken offline.

[0010] According to another aspect of this application, a node offline control device is provided, comprising: The instruction response module is configured to respond to offline instructions by updating the node status of this execution node in the node configuration table maintained by the control node to an inactive state, so as to prevent the scheduler in the control node from assigning new asynchronous tasks to this execution node. The delayed execution module is configured to perform a delayed wait according to a preset delay duration to compensate for the propagation lag of the inactive state from the node configuration table to the scheduler; The task polling module is configured to periodically query the task records in the task status table maintained by the control node after the delay wait ends, where the execution status of the current execution node is "in execution" and it is not in the task ignore list, as the query result. The busy time processing module is configured to, when the query result contains a record of an executing task, determine whether the corresponding asynchronous task has any unfinished pre-dependent tasks. If so, skip waiting for the asynchronous task and add it to the task ignore list; otherwise, continue to wait for the asynchronous task to complete. The idle-time offline module is configured to perform an offline operation on the current execution node when the query results do not contain any records of tasks in progress.

[0011] According to another aspect of this application, an electronic device is provided, including a central processing unit and a memory, wherein the central processing unit is configured to invoke and run a computer program stored in the memory to perform the steps of the method described in this application.

[0012] According to another aspect of this application, a non-volatile readable storage medium is provided, which stores a computer program implemented according to the node offline control method in the form of computer-readable instructions, wherein the computer program, when invoked by a computer, executes the steps included in the method.

[0013] According to another aspect of this application, a computer program product is provided, comprising a computer program / instructions that, when executed by a processor, implement the steps of the method.

[0014] Compared to traditional technologies, this application achieves a safe, accurate, and efficient graceful shutdown when execution nodes respond to shutdown commands through the coordinated operation of node configuration tables and task status tables, combined with a delayed waiting mechanism and dynamic management of task ignore lists. Upon receiving a shutdown command, the execution node proactively updates its node status in the node configuration table to an inactive state, causing the scheduler to stop allocating new asynchronous tasks to that node. A preset delay period eliminates the risk of misallocation of new tasks due to state synchronization lag. After the delay period, the execution node periodically queries the task status table for the corresponding executing task record, identifies tasks with pre-defined dependencies that are blocking it, adds them to the task ignore list to skip invalid waiting, and only waits for tasks whose pre-defined dependencies are satisfied and can indeed be completed on the node. When no executing task record is confirmed, the execution node automatically performs the shutdown operation without manual intervention. This application, while ensuring that all completeable tasks are indeed completed, skips tasks that cannot be completed on the node due to pre-defined dependencies, achieving an optimal balance between task integrity and shutdown efficiency, significantly improving the security, efficiency, and maintainability of node shutdown in distributed asynchronous task systems. Attached Figure Description

[0015] Figure 1 This is an exemplary network architecture for this application; Figure 2 This is a flowchart illustrating one embodiment of the node offline control method of this application; Figure 3 This is a schematic block diagram of the node offline control device of this application; Figure 4 This is a schematic diagram of the structure of an electronic device used in this application. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0017] This application's embodiments can be applied to various online business platforms employing a distributed microservice architecture, particularly platforms that rely on asynchronous task systems to handle complex business processes, such as e-commerce platforms, social networking platforms, financial service platforms, and content distribution platforms. In these platforms, the asynchronous task system typically serves as a core component of the backend processing, undertaking key functions such as order processing, data synchronization, message push, file transcoding, and report generation. With the continuous expansion of business scale, the number of tasks that asynchronous task systems need to handle can reach millions, the types of tasks are becoming increasingly diverse, and the dependencies between tasks are becoming increasingly complex. To support high concurrency and large-scale task processing, asynchronous task systems typically adopt an architecture combining a control node and multiple execution nodes. The control node is responsible for task scheduling and allocation, while the execution nodes are responsible for executing specific tasks. When an execution node needs to undergo version upgrades, configuration updates, or resource reclamation, how to safely remove the node from the cluster without interrupting the currently executing asynchronous tasks becomes a key technical issue for ensuring business continuity and data integrity. This application's embodiments provide a technical solution for such application scenarios, enabling the graceful self-driving shutdown of execution nodes while ensuring the safe completion of all existing asynchronous tasks.

[0018] Please see Figure 1 , Figure 1 This is a schematic diagram of the architecture of a node offline control system provided in an embodiment of this application. The system may include a client 80, a first server 81 as a control node, and a second server 82 as an execution node. The client 80 may be a terminal device used by maintenance personnel to initiate offline commands to the offline control system and to receive and display the execution status of the offline process. The first server 81 may be a server carrying a scheduler and a database instance, running one or more scheduling processes and a database management system. The database management system may also be deployed on other servers. The scheduling process is responsible for allocating asynchronous tasks to each execution node according to the node status of each execution node in the node configuration table. The database management system maintains a node configuration table and a task status table. The second server 82 is a server carrying asynchronous task execution processes, running one or more execution processes. These execution processes are responsible for receiving and executing the asynchronous tasks allocated by the scheduler and feeding back the task execution results to the task status table. The client 80, the first server 81, and the second server 82 communicate via a network, which may be a wired or wireless network, such as a local area network (LAN), a wide area network (WAN), or the Internet.

[0019] In this embodiment, the node configuration table records the node status of each execution node, including active and inactive states. An active state indicates that the execution node is running normally and can receive new asynchronous tasks, while an inactive state indicates that the execution node is about to go offline and should no longer receive new asynchronous tasks. The task status table records the execution status of each asynchronous task and its associated execution node. Execution status includes states such as "in execution" and "completed." Records in the task status table are associated with records in the node configuration table through node identifiers. Asynchronous tasks may have prerequisite dependencies; subsequent asynchronous tasks are in a waiting-to-be-triggered state until the prerequisite asynchronous task completes. The task status table records the dependencies between tasks through a prerequisite task identifier field.

[0020] The node offline control method of this application can be implemented as a computer program, which is installed and runs on the second server 82 as an offline control unit independent of the execution process. The processor of the offline control unit executes the computer program to implement the various steps in the method. During operation, the offline control unit needs to communicate with the scheduling process and database management system on the first server 81 to update the node status in the node configuration table, query the task records in the task status table, and synchronize the update records of the task ignore list. To support the reliable execution of the offline process, the offline control unit may also include a local storage module specifically for storing the offline progress and the task ignore list. This local storage module can be an embedded database or file system, used to persistently record information such as the current offline progress status, the list of tasks added to the task ignore list, and the number of polling cycles executed. In actual deployment, the offline control unit can run on the same physical server or virtual server as the execution process, or it can run on a separate server; this application does not impose any restrictions on this.

[0021] Taking an e-commerce platform as an example, the order processing flow of this platform can be decomposed into multiple asynchronous tasks with dependencies. From submission to final completion, an order needs to go through multiple processing stages, such as payment verification, inventory deduction, logistics order generation, and shipping notification. Each stage corresponds to an asynchronous task, and the task of the next stage can only be triggered and executed after the task of the previous stage is completed. These asynchronous tasks are allocated by the scheduler according to the load and status of each execution node. Asynchronous tasks at different stages may be assigned to different execution nodes for execution. When an execution node needs to go offline, there may be multiple asynchronous tasks at different processing stages on that node at the same time, and the prerequisite tasks of some of these tasks may not have been completed on other execution nodes. The technical solution of this application is aimed at this scenario of long-chain asynchronous tasks with complex dependencies. Through the coordinated cooperation of the node configuration table and the task status table, combined with the delay waiting mechanism and the dynamic management of the task ignore list, the execution node can skip tasks that cannot be completed on its own node due to prerequisite dependencies, and safely, autonomously, and efficiently perform the offline operation, provided that all completeable tasks are indeed completed.

[0022] It should be noted that the order processing flow of the e-commerce platform described above is merely an illustrative example. In actual applications, different platforms can flexibly set the types and dependencies of asynchronous tasks according to their own business characteristics, and this application does not impose any restrictions on this. As long as there are prerequisite dependencies between asynchronous tasks and the task states are persistently stored in the database, the technical solution of this application is applicable.

[0023] The technical solution of this application will be further described in detail below with reference to specific embodiments.

[0024] Please see Figure 2 According to the node offline control method provided in this application, it can be implemented as a computer program product and run on an electronic device such as a server that acts as an execution node. This execution node is mainly responsible for executing the steps of the method, specifically including the following steps: Step S5100: In response to the offline command, update the node status of this execution node in the node configuration table maintained by the control node to the inactive state, so as to prevent the scheduler in the control node from assigning new asynchronous tasks to this execution node. The shutdown command can be triggered in several ways. In one embodiment, the shutdown command originates from an active shutdown operation initiated by operations personnel through a client. The operations personnel select the target execution node in the client and trigger the shutdown command, which is then sent to the corresponding execution node. In another embodiment, the shutdown command originates from the automatic scheduling of the cluster management system and can be triggered by the control node or other management nodes. For example, when the cluster management system detects that the resource utilization of an execution node is lower than a preset threshold, it automatically triggers the shutdown command for that execution node to reclaim resources. Alternatively, when an execution node needs to be upgraded, the cluster management system automatically sends a shutdown command to the execution node to be updated in the rolling update strategy. The shutdown command must at least include the node identifier of the target execution node so that the execution node receiving the command can confirm that it is the target of the shutdown operation.

[0025] Upon receiving a shutdown command, the execution node first performs a node status update operation. The execution node calls the status update interface provided by the control node to update its own node status in the corresponding node record in the node configuration table to an inactive state. The node configuration table is stored on the control node side, specifically in a database maintained by the control node, and is used to record the node status of all execution nodes in the cluster. Each node record in the node configuration table contains at least a node identifier field and a node status field. The node identifier field uniquely identifies an execution node, and the node status field indicates the current state of the execution node. Node status includes at least an active state and an inactive state. An active state indicates that the execution node is running normally and can receive new asynchronous tasks, while an inactive state indicates that the execution node is about to go offline and should no longer receive new asynchronous tasks.

[0026] When allocating new asynchronous tasks, the scheduler in the control node queries the node configuration table to check the node status of each execution node, assigning the new asynchronous task only to execution nodes with an active status. Once an execution node updates its status to inactive, the scheduler will no longer consider that execution node as a candidate in subsequent task allocation decisions, thus preventing the scheduler from assigning new asynchronous tasks to that execution node. This mechanism cuts off the inflow of new tasks at the source, ensuring that execution nodes do not have their offline time extended or new tasks interrupted due to receiving new tasks during the offline process.

[0027] There are several ways to implement the node state update operation. In one embodiment, the executing node calls the application programming interface provided by the control node via the Hypertext Transfer Protocol (HTTP), sending a request containing its own node identifier to the control node. Upon receiving the request, the control node looks up the corresponding node record in the node configuration table and updates its node state field to an inactive state. In another embodiment, the executing node calls the state update service provided by the control node via a remote procedure call, and the state update service performs the update operation on the node configuration table. After completing the node state update, the control node returns a confirmation response confirming the update success to the executing node.

[0028] Step S5200: Perform a delay wait according to a preset delay duration to compensate for the propagation lag of the inactive state from the node configuration table to the scheduler; After an execution node updates its node status to inactive in the node configuration table, it immediately enters a delayed waiting phase. The execution node waits for a preset delay duration to compensate for the propagation lag in synchronizing the inactive status from the node configuration table to the scheduler. In a distributed system, the node configuration table is stored on the control node side, and the scheduler needs to read the node status information in the node configuration table when allocating new asynchronous tasks. However, the scheduler does not perceive every status change in the node configuration table in real time; instead, it obtains the latest node status through polling or event listening mechanisms. When an execution node updates its node status to inactive, the scheduler may need a polling cycle or event propagation cycle to detect this change. Within this propagation lag window, the scheduler may still allocate new asynchronous tasks to the execution node based on the expired node status, resulting in the new task being incorrectly received and failing to complete properly because the execution node is about to go offline. Therefore, this application reserves a sufficient time window for the propagation of the inactive state by performing a preset delay wait after the state update, ensuring that the scheduler has fully perceived the change in the node state before the delay wait ends, thereby eliminating the risk of misassignment of new tasks from the root.

[0029] The preset delay duration can be set in several ways. In one embodiment, the preset delay duration is set to a fixed value, such as 20 seconds. This fixed value can be set according to the polling interval of the scheduler's node configuration table, for example, by setting it to a value greater than the polling interval, to ensure that the scheduler completes at least one full poll and obtains the updated node status during the delay waiting period. In another embodiment, the preset delay duration is dynamically obtained by the execution node after receiving the offline command, based on the scheduler's polling interval, and a value greater than this polling interval is used as the preset delay duration. For example, the execution node obtains the scheduler's polling interval by calling the configuration query interface provided by the control node. If the polling interval is 10 seconds, the preset delay duration is set to 12 seconds or 15 seconds to ensure that the delay waiting ends only after the scheduler completes at least one poll. The specific value of the preset delay duration can be configured and adjusted according to the actual deployment of the system, and this application does not impose any restrictions on it.

[0030] During the delay period, the execution node does not initiate any query requests or perform any other operations related to the offline process. This prevents the execution node from prematurely querying the task status table before the scheduler detects its inactive state, thus preventing the offline process from being unnecessarily extended due to newly assigned tasks in the query results. After the delay period ends, the execution node can then enter the subsequent periodic query phase. There are several ways to implement the delay. In one embodiment, the execution node implements the delay by calling a sleep function provided by the operating system, such as calling the `sleep` function in Linux to pause the current thread for a specified duration. In another embodiment, the execution node implements the delay through a timer mechanism, such as setting a timer to trigger a callback function after a preset delay, in which the subsequent periodic query process is initiated. During the delay period, the execution process on the execution node continues to run normally, and the asynchronous tasks being executed are unaffected by the delay, thus ensuring the parallelism between the offline process and task execution.

[0031] Step S5300: After the delay wait ends, periodically query the task records in the task status table maintained by the control node that are in execution and do not belong to the task ignore list, and use them as query results; After the delay period ends, the execution node enters a periodic query phase. Using its own node identifier as the query condition, the execution node periodically queries the task status table maintained by the control node for task records whose execution status is "in progress" and which are not on the task ignore list. The query results are used as the basis for subsequent offline decisions. The task status table is stored on the control node side, specifically in a database maintained by the control node, and records the execution status of all asynchronous tasks in the cluster and their respective execution nodes. Each task record in the task status table contains at least a task identifier field, a node identifier field, an execution status field, and a preceding task identifier field. The task identifier field uniquely identifies an asynchronous task; the node identifier field identifies the execution node currently responsible for executing the task; the execution status field indicates the current execution status of the asynchronous task; and the preceding task identifier field records the identifiers of the preceding tasks that the asynchronous task depends on. The execution status includes at least two states: "in progress" and "completed." "In progress" indicates that the asynchronous task is being processed by the corresponding execution node, and "completed" indicates that the asynchronous task has been completed.

[0032] Periodic queries are initiated when the delay period ends and the execution node initiates its first query request. Subsequent queries are initiated at preset polling intervals until the offline condition is met or a timeout condition is triggered. The specific value of the polling interval can be configured according to the actual system deployment. In one embodiment, the polling interval is set to a fixed value, such as 10 seconds, causing the execution node to initiate a query request to the control node every 10 seconds. In another embodiment, the polling interval is dynamically adjusted based on the system load. For example, when there are many running task records corresponding to this execution node in the task status table, the polling interval is appropriately shortened to improve response sensitivity; when there are few running task records, the polling interval is appropriately lengthened to reduce the access pressure on the database.

[0033] The construction of query conditions can include filtering in two dimensions. The first dimension uses the node identifier of the current execution node as a condition to filter task records belonging to that execution node. The second dimension uses the execution status as "in progress" as a condition to filter out task records that have not yet been completed. After obtaining the query results, the execution node further filters the results based on its local task ignore list, removing task records belonging to the task ignore list from the query results. The final query results are the task records corresponding to the current execution node whose execution status is "in progress" and which do not belong to the task ignore list. There are several ways to implement periodic queries. In one embodiment, the execution node calls the task query application programming interface provided by the control node through the Hypertext Transfer Protocol, carrying the node identifier of the current execution node in the request. After receiving the request, the control node constructs a database query statement based on the node identifier and execution status, retrieves task records that meet the conditions from the task status table, and returns the original query results to the execution node. After receiving the original query results, the execution node filters the results based on its local task ignore list, removing task records belonging to the task ignore list to obtain the final query results. In another embodiment, the execution node invokes the task query service provided by the control node via remote procedure call. The task query service performs a database query operation and returns the results. When returning the query results, the control node may also include dependency information for all asynchronous tasks on this execution node, so that the execution node can determine the prerequisite dependent tasks in subsequent steps.

[0034] To avoid excessive pressure on the database during periodic queries, a caching mechanism can be introduced on the control node side. In one embodiment, after receiving a query request, the control node first checks if a valid query result exists in its local cache. If it does, the cached result is returned directly without accessing the database; otherwise, a database query is performed, and the result is stored in the cache before being returned. The cache expiration time can be set to a value that matches the polling interval. For example, when the polling interval is 10 seconds, the cache expiration time can be set to 8 seconds or 10 seconds, ensuring that there is at least one actual database query in each polling cycle, while avoiding multiple repeated database queries within the polling interval. In another embodiment, the control node introduces a distributed cache in addition to the local cache, forming a two-level caching architecture. Two-level queries are performed when necessary to further improve query performance.

[0035] There are two termination conditions for periodic queries. The first is an empty query result, meaning there is no task record in the task status table corresponding to this execution node that is in execution and not on the task ignore list. In this case, the execution node can safely perform the shutdown operation. The second is a query timeout, meaning the execution node fails to obtain an empty query result within the preset total timeout period. In this case, the execution node triggers the timeout handling process. The specific value of the total timeout period can be configured according to the actual deployment of the system, for example, set to 2 hours to ensure that the shutdown process does not continue indefinitely in extreme cases.

[0036] Step S5400: When the query result contains a record of an executing task, determine whether the corresponding asynchronous task has any unfinished pre-dependent tasks. If so, skip waiting for the asynchronous task and add it to the task ignore list; otherwise, continue to wait for the asynchronous task to complete. After the delay period ends, the execution node enters a periodic query phase, using the query results as the basis for offline decisions. When a query result contains records of tasks in progress, the execution node needs to perform a prerequisite dependency check on each record to determine whether the asynchronous task truly needs to wait for its completion on this execution node, or whether it can skip the wait because its prerequisite tasks have not yet been completed.

[0037] Precursor tasks are other asynchronous tasks that must be completed before the current asynchronous task can begin execution. In business processes with multiple processing stages, the asynchronous task of the later stage can only be triggered and executed after the asynchronous task of the previous stage is completed. The asynchronous task of the previous stage is the prerequisite task of the asynchronous task of the later stage. For example, in the order processing flow of an e-commerce platform, an order needs to go through multiple processing stages from submission to final completion, such as payment verification, inventory deduction, logistics order generation, and shipping notification. Each stage corresponds to an asynchronous task. The prerequisite task for the logistics order generation task is the inventory deduction task, and the prerequisite task for the inventory deduction task is the payment verification task. If the inventory deduction task has not been completed, even if the logistics order generation task has been assigned to an execution node by the scheduler and is in the execution state, it cannot actually proceed because its required input data is not yet ready.

[0038] When an execution node determines whether a corresponding asynchronous task has any incomplete pre-dependent tasks, it first needs to obtain the pre-dependent task information. In one embodiment, each task record in the task status table includes a pre-dependent task identifier field, which records the identifier of the pre-dependent task that the current asynchronous task depends on. The execution node extracts the pre-dependent task identifier for each executing task record from the query results, and then looks up the execution status of the corresponding pre-dependent task in the task status table based on the pre-dependent task identifier. In another embodiment, before each round of periodic queries begins, the execution node loads the task records of all asynchronous tasks on this execution node and the pre-dependent task identifiers in the task records from the task status table, generates a dependency relationship snapshot, and then determines the completion status of the pre-dependent tasks of each executing asynchronous task based on this dependency relationship snapshot. The dependency relationship snapshot records the dependency relationship topology of all asynchronous tasks on this execution node at the current moment. The execution node can quickly find the pre-dependent tasks and their completion status of each asynchronous task in this snapshot without initiating a separate query request to the control node for each determination.

[0039] The completion status of a prerequisite task depends on its current execution status and the execution node to which it belongs. In one embodiment, if a prerequisite task is executed on an execution node other than the current execution node and its status is "in execution," then the prerequisite task is determined to be incomplete. This is because the completion of a prerequisite task executed on another execution node is not controlled by the current execution node, and the current execution node cannot accelerate its completion by waiting. Therefore, even if the current asynchronous task is in an "in execution" state, it cannot truly move forward. In another embodiment, if a prerequisite task is executed on the current execution node and its status is "in execution," then the prerequisite task is determined to be incomplete. However, the current execution node can further wait for the prerequisite task to complete before determining whether the current asynchronous task can be skipped. This is because the prerequisite task and the current asynchronous task on the current execution node are executed on the same execution node, and the current execution node can promote the execution of the current asynchronous task by waiting for the prerequisite task to complete.

[0040] When an execution node determines that a corresponding asynchronous task has incomplete prerequisite tasks, it skips waiting for that asynchronous task and adds it to the task ignore list. Skipping the wait means that the execution node no longer waits for the asynchronous task to complete. In subsequent periodic queries, the asynchronous task is excluded from the query results because it has been added to the task ignore list, and the execution node will not check its prerequisites again. After adding the asynchronous task to the task ignore list, the waiting state of the asynchronous task on this execution node is lifted, and the execution node can focus more on waiting for tasks whose prerequisites have been satisfied and can indeed be completed on this node. Asynchronous tasks added to the task ignore list are not discarded or interrupted; they remain in the running state in the task status table, and the scheduler decides whether to reassign them to other execution nodes for continued execution based on subsequent allocation strategies.

[0041] When an execution node determines that a corresponding asynchronous task has no incomplete pre-dependent tasks, it continues to wait for the asynchronous task to complete. Continuing to wait means that the execution node will still include this asynchronous task in the query results of the next round of periodic queries and continuously monitor its execution status until its execution status changes from "in progress" to "completed," or a timeout condition is triggered. During this waiting period, the execution node repeatedly initiates periodic queries at preset polling intervals. After each query, it re-determines the pre-dependent tasks in the "in progress" records in the query results and dynamically adjusts the waiting strategy based on the determination results. If an asynchronous task's pre-dependent tasks are still incomplete in subsequent polls, the execution node can add it to the task ignore list and skip the waiting; if an asynchronous task's pre-dependent tasks are completed in subsequent polls, the execution node can remove it from the task ignore list and resume waiting.

[0042] Through the aforementioned pre-dependency judgment mechanism, execution nodes can distinguish between truly independent tasks that require waiting and passively waiting tasks that are temporarily unable to proceed due to pre-dependency blocking. For the former, execution nodes patiently wait to ensure business integrity; for the latter, execution nodes skip them to avoid unnecessary blocking, thereby minimizing offline waiting time while ensuring the safe completion of all existing tasks.

[0043] Step S5500: When the query result does not contain any records of tasks in execution, perform the offline operation of this execution node.

[0044] The query result shows no records of tasks in execution, indicating that all asynchronous tasks on this execution node that are in execution and not on the task ignore list have been completed, and this execution node is ready to be safely shut down. At this point, the execution node performs a shutdown operation, safely exiting the cluster.

[0045] There are several ways to implement the offline operation. In one embodiment, the offline operation terminates the service process running on the current execution node. After the service process terminates, the execution node no longer receives new asynchronous tasks or executes any task processing logic. In another embodiment, the offline operation sends an offline completion notification to the control node, which then removes the execution node from the cluster's management list. Upon receiving the offline completion notification, the control node updates the node status of the execution node in its node configuration table or deletes its node record from the node configuration table. In yet another embodiment, the offline operation includes first sending an offline completion notification to the control node, and then terminating the service process after receiving confirmation from the control node, to ensure that the control node has synchronized the offline status of the execution node.

[0046] Before performing the shutdown operation, the execution node can also perform relevant preparatory work to ensure a smooth and safe shutdown process. In one embodiment, after confirming that the query results do not contain any ongoing task records, the execution node first generates a shutdown confirmation report. The shutdown confirmation report includes an overview of the processing of all asynchronous tasks on this execution node, such as the number of completed tasks, the number of tasks added to the task ignore list and the reasons for being skipped, the total time taken for this shutdown process, etc. Then, the shutdown confirmation report is sent to the control node for archiving for subsequent auditing and traceability. In another embodiment, after generating the shutdown confirmation report, the execution node sends the shutdown confirmation report to the control node, which verifies whether there are any potential risks in the shutdown confirmation report, such as whether the shutdown of this node will cause downstream dependent tasks on other execution nodes to be permanently unable to complete. If the verification passes, the execution node continues to perform the shutdown operation; if the verification fails, the execution node stops the shutdown operation and restores its node status to an active state.

[0047] After an execution node goes offline, it will no longer be included in the cluster management system to which it belongs, and the scheduler will no longer use it as a target node for task allocation. Asynchronous tasks added to the task ignore list on the execution node are still in execution and have not been interrupted. Therefore, the scheduler can reassign these tasks to other execution nodes according to a preset strategy during subsequent task allocation, ensuring that these tasks are not discarded. The results of completed tasks on the execution node are recorded in the task status table and are unaffected by the execution node going offline.

[0048] In one embodiment, after an execution node performs an offline operation, its container orchestration platform detects that the execution node has stopped and automatically starts a new execution node instance to replace the offline execution node, thereby maintaining a stable total number of execution nodes in the cluster. After the new execution node instance starts and registers with the control node, the scheduler begins to allocate new asynchronous tasks to it, and at the same time, reallocates asynchronous tasks that were previously added to the task ignore list to the new execution node or other existing execution nodes for continued execution as needed.

[0049] In some embodiments, during periodic queries, it is also necessary to handle situations where query timeouts or query anomalies occur. In embodiments where a timeout handling process is triggered to handle query timeouts, a query timeout refers to the execution node failing to obtain an empty query result within a preset total timeout period, meaning that there are always task records in the execution node that are in execution and not on the task ignore list. The specific value of the total timeout period can be configured according to the actual deployment of the system, for example, set to one hour or two hours, to ensure that the offline process does not continue indefinitely in extreme cases. When a query times out, the execution node triggers the timeout handling process, sending a timeout alarm message to the control node. The timeout alarm message includes at least the node identifier of the execution node, a detailed list of tasks that are still in execution, and the waiting time, so that maintenance personnel can understand the reason for the blockage of the offline process and perform manual intervention.

[0050] In this embodiment where query anomalies trigger the exception handling process, a query anomaly refers to the inability of the execution node to obtain query results due to network failure, unavailability of the control node service, or other reasons when initiating a periodic query. When a query anomaly occurs, the execution node re-initiates the query request according to a preset retry strategy, for example, a maximum of five retries, with the interval between each retry increasing progressively. If the query succeeds within the preset number of retries, the execution node continues the normal periodic query process; if the query still fails after exceeding the preset number of retries, the execution node sends an anomaly alarm message to the control node. The anomaly alarm message includes at least the node identifier of the execution node, the anomaly type, and the time of the last failed attempt, so that maintenance personnel can promptly investigate and handle the issue. Timeout alarm messages and anomaly alarm messages can be sent to the maintenance group via an instant messaging interface, enabling visualized maintenance under unattended conditions.

[0051] As can be understood from the above embodiments, this application, through the coordinated operation of the node configuration table and the task status table, combined with the delay waiting mechanism and the dynamic management of the task ignore list, achieves a safe, accurate, and efficient graceful shutdown when executing the node response shutdown command, and achieves the following beneficial effects, including but not limited to: First, this application implements a secure, self-driven offline mechanism for execution nodes in asynchronous task scenarios. Upon receiving a offline command, the execution node proactively updates its node status in the node configuration table to an inactive state, causing the scheduler to stop assigning new asynchronous tasks to that node. After the status update, a preset delay is performed, providing ample time for the inactive state to synchronize from the node configuration table to the scheduler, thus eliminating the risk of misassignment of new tasks due to delayed status synchronization. After the delay, the execution node periodically queries the task status table to determine whether it can safely stop. The entire offline process is autonomously driven by the execution node, independent of real-time notifications from external components. Even if the control node experiences a brief failure, the execution node that has initiated the offline process can still independently complete subsequent steps, significantly improving the system's robustness.

[0052] Secondly, this application implements selective waiting based on task dependency awareness, effectively shortening offline waiting time while ensuring task integrity. In long-chain business scenarios, multiple asynchronous tasks in the same business process have sequential dependencies. Among the tasks in execution on an execution node, some may be in a state of being unable to proceed because their predecessor dependencies have not yet been completed on other nodes. This application identifies tasks that cannot proceed due to blockage caused by their predecessor dependencies by determining whether there are any unfinished predecessor dependencies in the executing task, and adds them to a task ignore list, thereby skipping the invalid waiting for these tasks and only waiting for those tasks whose predecessor dependencies have been satisfied and can indeed be completed on the current node. This mechanism prevents the execution node from being blocked indefinitely by waiting for a task that can never be completed, and also prevents the loss of tasks that should have been completed on the current node by skipping the wait, ensuring the consistency of business data while guaranteeing the efficiency of cluster operation and maintenance.

[0053] Furthermore, this application automates and enhances the operational friendliness of the offline process. Upon receiving the offline command, the execution node automatically performs steps such as status update, delayed waiting, periodic querying, dependency judgment, ignore list update, and offline operation, all without manual intervention. Simultaneously, through dynamic maintenance of the task ignore list, the execution node can adaptively adjust its waiting strategy, eliminating the need for operations personnel to configure which tasks to skip beforehand. This reduces operational complexity, minimizes the possibility of human error, and makes the management of node offline / offline operations in large-scale clusters more efficient and reliable.

[0054] Based on any embodiment of the method in this application, a delay waiting is performed according to a preset delay duration to compensate for the propagation lag of the inactive state from the node configuration table to the scheduler, including: Step S5210: Obtain the polling interval of the scheduler for the node configuration table; The polling interval of the scheduler's node configuration table is the time interval between two consecutive reads of the node configuration table. The scheduler does not perceive every change in the node status in the node configuration table in real time; instead, it reads the node status of each execution node from the node configuration table according to a fixed polling period to obtain the latest node status information and make task allocation decisions accordingly. The specific value of the polling interval is determined by the system configuration, and can be set to, for example, 5 seconds, 10 seconds, or 15 seconds. After receiving a shutdown command and updating its own node status to inactive, an execution node first needs to obtain the scheduler's polling interval to determine an appropriate delay time.

[0055] There are several ways to obtain the polling interval. In one embodiment, the execution node obtains the scheduler's polling interval by calling the configuration query interface provided by the control node. This interface returns the value of the polling interval currently used by the scheduler. In another embodiment, the polling interval is pre-stored in the execution node's local configuration file, and the execution node reads this value directly from the local configuration file. In yet another embodiment, the polling interval is sent as a parameter by the operations and maintenance personnel when initiating a shutdown command, and the execution node parses and obtains this value from the shutdown command.

[0056] Step S5220: Use a value greater than the polling interval as the preset delay time, and within the preset delay time, this execution node will not initiate any query request.

[0057] After the execution node obtains the scheduler's polling interval, it sets a preset delay time greater than this interval. The principle for setting the preset delay time is to ensure that the scheduler completes at least one full poll and obtains the updated node status within the delay waiting period. For example, if the scheduler's polling interval is 10 seconds, the preset delay time can be set to 12, 15, or 20 seconds to ensure that the delay wait ends only after the scheduler completes at least one poll. If the scheduler's polling interval is 5 seconds, the preset delay time can be set to 8 or 10 seconds. The difference between the preset delay time and the polling interval can be adjusted appropriately based on the system's network latency and status propagation time. For example, a predetermined margin of 2 to 5 seconds can be added to the polling interval to cope with propagation delays caused by network fluctuations. Within the preset delay time, the execution node does not initiate any query requests or perform any other operations related to the offline process. The execution node implements delayed waiting by calling the sleep function provided by the operating system or setting a timer. For example, in Linux, calling the sleep function pauses the current thread for a specified duration, or setting a timer to trigger a callback function after a preset delay. During the delayed waiting period, the execution process on the execution node continues to run normally, and the asynchronous tasks being executed are not affected by the delayed waiting, thus ensuring the parallelism between the offline process and task execution.

[0058] Through the above embodiments, this application further defines the specific method for determining the preset delay duration, thus quantitatively linking the delay waiting duration with the scheduler's polling interval. The execution node obtains the scheduler's polling interval and uses a duration greater than this value as the preset delay duration, ensuring that the scheduler completes at least one full poll and obtains the updated node status during the delay waiting period. This minimizes the waiting time while ensuring that the inactive state is fully perceived by the scheduler, balancing the safety and efficiency of the offline process.

[0059] Based on any embodiment of the method in this application, determining whether the corresponding asynchronous task has any unfinished pre-dependent tasks includes: Step S5411: Before each round of periodic query begins, generate a dependency snapshot based on the task records of all asynchronous tasks on the current execution node in the task status table and the predecessor task identifiers in the task records. Before each round of periodic queries begins, the execution node first retrieves task records for all asynchronous tasks on its side from the task status table maintained by the control node. As mentioned earlier, each task record in the task status table contains at least a task identifier field, a node identifier field, an execution status field, and a predecessor task identifier field. The predecessor task identifier field records the identifiers of the predecessor tasks that the current asynchronous task depends on. If the current asynchronous task has no predecessor tasks, the value of this field is empty or a preset null value. The execution node extracts the predecessor task identifiers from all task records and constructs a directed acyclic graph with task identifiers as nodes and predecessor dependencies as edges, serving as a snapshot of the dependency relationships.

[0060] The dependency snapshot records the dependency topology between all asynchronous tasks on the current execution node at the current moment. For example, if there are four asynchronous tasks on the current execution node, with task identifiers T1001, T1002, T1003, and T1004, where the predecessor task of T1002 is T1001, the predecessor task of T1003 is T1002, and T1004 has no predecessor task identifier, then the dependency snapshot records the dependency links from T1001 to T1002, from T1002 to T1003, and the topology of T1004 as an independent node.

[0061] Dependency snapshots are generated before each round of periodic queries to ensure that the snapshot reflects the latest dependency state. After generating the dependency snapshot, the execution node uses this snapshot to determine the completion status of subsequent prerequisite tasks without needing to send a separate query request to the control node for each determination, thus reducing the number of communications with the control node.

[0062] Step S5412: Based on the dependency snapshot, determine the completion status of the preceding dependent tasks of each executing asynchronous task in the dependency snapshot; For each asynchronous task in execution, the execution node first looks up the identifier of the preceding task in the dependency snapshot, and then queries the task status table to find the current execution status of the corresponding preceding dependent task based on the preceding task identifier. The execution status of the preceding dependent task includes an in-process status and a completed status. If the execution status of the preceding dependent task in the task status table is completed, it is determined that the preceding dependent task has been completed, the preceding dependency of the current asynchronous task has been satisfied, and it can continue to wait for the current asynchronous task to complete. If the execution status of the preceding dependent task in the task status table is in process, it is necessary to further determine its completion status based on the execution node to which the preceding dependent task belongs. For example, for task T1003, its preceding task identifier is T1002. The execution node queries the task status table for the execution status of T1002. If the execution status of T1002 is completed, it is determined that the preceding dependent task of T1003 has been completed; if the execution status of T1002 is in process, it is necessary to further determine the execution node to which T1002 belongs.

[0063] Step S5413: If the prerequisite task is executed on another execution node outside this execution node and its status is "in execution", then it is determined that the prerequisite task has not been completed. When an execution node determines that a prerequisite task for an ongoing asynchronous task is in progress, and that the execution node to which the prerequisite task belongs is not its own, the execution node considers the prerequisite task incomplete. This is because the prerequisite task is executed on other execution nodes, and its completion is not controlled by the current execution node; the current execution node cannot accelerate its completion by waiting. Even if the current asynchronous task is in progress, it cannot truly move forward because its prerequisite task is incomplete. For example, in the order processing flow of an e-commerce platform, the prerequisite task for the logistics order generation task is the inventory deduction task. If the inventory deduction task is executed on another execution node and has not yet been completed, the logistics order generation task on this execution node, even if it has been assigned and is in progress, cannot actually generate a logistics order because the required inventory deduction result is not yet ready. At this time, it is meaningless for the current execution node to continue waiting for the logistics order generation task, because its completion depends on the completion of the prerequisite task on other nodes. Therefore, the execution node determines that the prerequisite task is incomplete and adds the current asynchronous task to the task ignore list, skipping the wait for it.

[0064] Step S5414: If the preceding dependent task is being executed on this execution node and its status is "in execution", then wait for the preceding dependent task to complete before determining whether the current asynchronous task can be skipped.

[0065] When an execution node determines that a prerequisite task for an ongoing asynchronous task is also in progress, and that the execution node to which the prerequisite task belongs is the current execution node, the execution node decides that although the prerequisite task is not yet complete, it can push forward the execution of the current asynchronous task by waiting for the prerequisite task to complete. This is because the prerequisite task and the current asynchronous task on the same execution node are executed on the same execution node. The completion of the prerequisite task will release the input data or resources required by the current asynchronous task, allowing the current asynchronous task to continue. For example, if both the inventory deduction task and the logistics order generation task are executed on the current execution node, once the inventory deduction task is completed, the logistics order generation task can immediately obtain the inventory deduction result and begin execution. Therefore, the execution node continues to wait for the prerequisite task to complete. After the prerequisite task is completed, it re-evaluates whether the current asynchronous task can be skipped. If the prerequisite task is completed and the prerequisites of the current asynchronous task are satisfied, the execution node continues to wait for the current asynchronous task to complete. If the prerequisite task is completed but there are still other unfinished prerequisite tasks for the current asynchronous task, the execution node re-evaluates based on the new dependency relationship status.

[0066] Through the above embodiments, this application defines a specific implementation method for determining prerequisite dependency tasks. The execution node generates a dependency snapshot before each round of periodic queries, topologicalizing the dependency relationships between tasks. This avoids initiating a separate query request to the control node for each determination, reducing communication overhead. Simultaneously, by distinguishing between prerequisite dependency tasks in this execution node and other execution nodes, the execution node can accurately identify which tasks truly need to wait and which tasks can be skipped due to prerequisite dependency blocking, thereby minimizing offline waiting time while ensuring task integrity.

[0067] Based on any embodiment of the method in this application, after adding the asynchronous task to the task ignore list, the method includes: Step S5421: Synchronize the update record of the task ignore list to the control node, so that the control node can broadcast the update record to other execution nodes currently in the offline process; After adding an asynchronous task to the task ignore list, this execution node generates an update record for the task ignore list. This update record includes at least the task identifier of the asynchronous task added to the ignore list, the addition time, and the node identifier of this execution node. This execution node sends this update record to the control node. Upon receiving it, the control node identifies other execution nodes in the cluster that are currently in the offline process. Execution nodes in the offline process are those that have received the offline command and are executing periodic queries and pre-dependency checks. The control node can identify the target execution node by querying the node configuration table for execution nodes with an inactive status, or by maintaining a list of execution nodes in the offline process. The control node only broadcasts the update record to these other execution nodes in the offline process, not to all execution nodes in the cluster, to avoid unnecessary communication overhead.

[0068] There are several ways to broadcast. In one embodiment, the control node pushes update records to other execution nodes in the offline process through a publish-subscribe message queue, and each execution node in the offline process subscribes to the relevant message topic. In another embodiment, other execution nodes in the offline process periodically poll the interface provided by the control node to obtain the latest task ignore list update records, for example, by initiating a synchronization request to the control node every 5 seconds to obtain the update records added since the last synchronization. Step S5422: After receiving the update records, the other execution nodes add the corresponding asynchronous tasks in the update records to the task ignore list of the other execution nodes, so that when the other execution nodes are in the offline process and query the task status table, the asynchronous tasks in the task ignore list will not be included in the query results.

[0069] After receiving the update record broadcast by the control node, other execution nodes parse the task identifier of the asynchronous task added to the task ignore list from the update record, and then add the task identifier to their own task ignore list. Each execution node maintains a unified task ignore list, which includes both asynchronous tasks that are skipped and waiting identified by the execution node itself through pre-dependency checks during its own offline process, and asynchronous tasks that are skipped and waiting received from other execution nodes through the control node's broadcast.

[0070] For this execution node, the same mechanism described above can be used to receive update records from other execution nodes. When another execution node adds an asynchronous task to its task ignore list and synchronizes the update record to the control node, the control node will also broadcast the update record to this execution node, which is currently in the offline process. After receiving the update record, this execution node adds the corresponding asynchronous task to its own task ignore list, so that in subsequent periodic queries by this execution node, the asynchronous task will also be filtered out from the query results. Through this two-way synchronization mechanism, all execution nodes in the cluster that are in the offline process can share each other's task ignore list information.

[0071] When another execution node is in the offline process and periodically queries the task status table, its query conditions include its own node identifier and execution status as "in execution". After obtaining the original query results, the execution node filters the results according to its own task ignore list, removing task records belonging to the task ignore list from the query results to ensure that these tasks are not included in the final query results.

[0072] For example, after execution node A adds task T2001 to its task ignore list, the control node broadcasts this update to execution nodes B and C, which are currently in the offline process. Execution nodes B and C then add T2001 to their respective task ignore lists. When execution node B performs subsequent periodic queries, it will exclude T2001 from its query conditions, even if T2001's node identifier in the task status table is execution node B and its execution status is "in progress," it will not be included in the query results. Through this mechanism, asynchronous tasks skipped on an execution node will not be repeatedly waited for by other execution nodes in the offline process, avoiding redundant waiting across the cluster.

[0073] Through the above embodiments, this application further defines a cross-node synchronization mechanism for the task ignore list. Any execution node synchronizes the updated record of the task ignore list to the control node, which then broadcasts it to other execution nodes. This allows each execution node to be aware of which asynchronous tasks have been skipped during its own offline process, thereby preventing the same asynchronous task from being repeatedly waited for by multiple execution nodes. This mechanism extends the scope of the ignore list from a single execution node to the entire cluster, eliminating redundant waiting across the cluster and further improving the overall efficiency when multiple execution nodes go offline simultaneously in a large-scale cluster.

[0074] Based on any embodiment of the method in this application, continuing to wait for the asynchronous task to complete includes: Step S5431: While waiting for the asynchronous task to complete, monitor the waiting time of the asynchronous task. When the waiting time exceeds the preset single task timeout threshold, send a dependency timeout report to the control node. While an execution node continues to wait for an asynchronous task, it starts a timer associated with that asynchronous task to monitor the elapsed waiting time. The elapsed waiting time is calculated from the moment the execution node first identifies the asynchronous task as a waiting object; that is, from the moment the execution node determines that the asynchronous task has no outstanding pre-dependent tasks and decides to continue waiting for its completion, the accumulated time is recorded. The preset single-task timeout threshold is the maximum allowed waiting time for a single asynchronous task, and its specific value can be configured based on the asynchronous task type, historical execution time, and business requirements for response time. For example, in the order processing flow of an e-commerce platform, the normal execution time of a payment verification task is 200 to 500 milliseconds, and its single task timeout threshold can be set to 5 seconds; the normal execution time of an inventory deduction task is 1 to 3 seconds, and its single task timeout threshold can be set to 15 seconds; the normal execution time of a logistics order generation task is 2 to 10 seconds, and its single task timeout threshold can be set to 30 seconds; the normal execution time of a shipping notification task is 500 to 2 seconds, and its single task timeout threshold can be set to 10 seconds.

[0075] The principle for setting the single-task timeout threshold is to add sufficient buffer margin to the normal execution time. This avoids accidental timeouts due to normal execution time fluctuations while ensuring that the transfer mechanism can be triggered promptly when a task truly malfunctions. When the waiting time recorded by the timer exceeds the preset single-task timeout threshold, the execution node determines that the asynchronous task has timed out, and then generates a dependency timeout report and sends it to the control node. The dependency timeout report includes at least the task identifier of the timed-out asynchronous task, the identifiers of the task's preceding dependent tasks, the waiting time, and the preset single-task timeout threshold, so that the control node can perform subsequent transfer processing accordingly.

[0076] Step S5432: The control node marks the corresponding asynchronous task and all its downstream dependent tasks as pending transfer status based on the dependency timeout report; After receiving a dependency timeout report from the execution node, the control node parses the task identifier of the timed-out asynchronous task from the report. The control node first searches the task status table for the task record of the asynchronous task to confirm that it is still in execution. Then, based on the preceding task identifier field in the task status table, the control node recursively searches for all downstream dependent tasks of the asynchronous task. Downstream dependent tasks refer to other asynchronous tasks that have the timed-out task as a preceding task; that is, tasks that directly or indirectly depend on the timed-out task to execute. For example, if the preceding task identifier for task T1002 is T1001, and the preceding task identifier for task T1003 is T1002, then the downstream dependent task of T1002 includes T1003. The control node marks the execution status or additional status field of the timed-out asynchronous task and all its downstream dependent tasks in the task status table as pending transfer. The pending transfer status indicates that the execution node currently hosting these tasks can no longer effectively execute them, and the control node needs to reschedule other execution nodes to take over execution. After being marked as pending transfer, these tasks retain their original task records in the task status table, but the control node and scheduler can recognize that they need to be reassigned.

[0077] Step S5433: The control node reassigns the asynchronous tasks marked as pending transfer to other execution nodes and updates the execution node to which the corresponding task record belongs in the task status table. The control node reassigns asynchronous tasks marked as pending transfer to other execution nodes one by one. The target execution nodes for reassignment are other currently active execution nodes in the cluster that are capable of executing such tasks. When selecting a target execution node, the control node can consider factors such as the current load of each execution node, the task queue length, and historical execution efficiency to assign tasks to the most suitable execution node. For each reassigned asynchronous task, the control node updates its own node identifier field in the task status table, replacing the original own execution node identifier with the node identifier of the newly assigned target execution node. Simultaneously, its execution status is reset to "in execution," and the pending transfer mark is removed. By updating the own node identifier, ownership of the asynchronous task is transferred from the original execution node to the new execution node, which will be responsible for executing the task in subsequent task processing. For downstream dependent tasks of timed-out tasks, the control node also reassigns them in the same way, ensuring that tasks throughout the entire dependency chain can be transferred to new execution nodes for continued execution.

[0078] Step S5434: After the execution node receives the transfer completion confirmation message from the control node, it adds the asynchronous task to the task ignore list and continues subsequent queries.

[0079] After the control node completes the reallocation of all asynchronous tasks marked as pending transfer, it sends a transfer completion confirmation message to the original execution node. This message includes at least a list of task identifiers for successfully transferred asynchronous tasks and the node identifier of the target execution node to which each task was assigned. Upon receiving the confirmation message, the original execution node adds the successfully transferred asynchronous tasks to its task ignore list. With these tasks added to the ignore list, they will be filtered out of the query results during subsequent periodic queries, and the original execution node will no longer wait for their completion. The original execution node then continues with subsequent periodic queries, performing pre-dependency checks and waiting on the remaining in-process task records in the query results. Through this mechanism, asynchronous tasks that were initially blocked due to timeouts are safely transferred to other execution nodes for continued execution, allowing the original execution node to release its waiting block and continue its offline process.

[0080] Through the above embodiments, this application introduces a single-task-level timeout monitoring and forced transfer mechanism. While waiting for asynchronous tasks to complete, the execution node independently times each awaited asynchronous task. When the waiting time of an asynchronous task exceeds the single-task timeout threshold matching its type, a dependency timeout report is triggered, and the control node reassigns the task and all its downstream dependent tasks to other execution nodes. This mechanism solves the problem of execution nodes waiting indefinitely due to a single asynchronous task abnormally freezing, upgrading passive waiting to proactive disaster recovery. While ensuring the integrity of the dependency chain, it effectively avoids single-point failures blocking the entire shutdown process, further improving the robustness and reliability of the shutdown process.

[0081] Based on any embodiment of the method in this application, after determining whether the corresponding asynchronous task has any unfinished pre-dependent tasks, the method includes: Step S5441: Subscribe to dependency change events published by the control node, including the completion of a prerequisite dependency task, the reassignment of a prerequisite dependency task, or the cancellation of a prerequisite dependency task; After initiating periodic queries, each execution node sends a subscription request to the control node to subscribe to dependency change events related to itself. Dependency change events refer to state changes in the task status table related to the pre-dependencies of asynchronous tasks on this execution node, and include at least three types. The first type is the pre-dependency task completion event, generated by the control node when the execution status of a pre-dependency task of an asynchronous task changes from "in execution" to "completed." The second type is the pre-dependency task reassignment event, generated by the control node when a pre-dependency task of an asynchronous task is reassigned to another execution node due to timeout or other reasons. The third type is the pre-dependency task cancellation event, generated by the control node when a pre-dependency task of an asynchronous task is cancelled due to business reasons.

[0082] There are several ways to subscribe. In one embodiment, the execution node subscribes to dependency change event topics published by the control node through a message queue, such as using message middleware like Apache Kafka or RabbitMQ. The control node publishes event messages to the corresponding topic, and the execution node consumes event messages from that topic. In another embodiment, the execution node requests the latest dependency change events from the control node using long polling, for example, by sending an event query request to the control node every 2 seconds. The control node returns immediately when there are new events, and maintains the connection until timeout when there are no new events.

[0083] Step S5442: When the dependency change event is received, pause the current periodic query and re-determine the waiting status of all asynchronous tasks in execution on this execution node according to the changed dependency. Upon receiving a dependency change event from the control node, the execution node immediately pauses its current periodic queries. Pausing periodic queries aims to avoid making incorrect waiting decisions based on outdated dependencies during dependency changes. The execution node parses the specific details of the change from the event message, including the change type, the identifiers of the involved asynchronous tasks, and the changed dependency status. Then, based on the changed dependencies, the execution node re-performs the pre-dependency check for all asynchronous tasks currently in execution on its node. The re-checking process is the same as the pre-dependency check logic in step S5400; that is, for each executing asynchronous task, the current completion status of its pre-dependent tasks is determined based on its pre-dependent task identifier in the task status table. Since the dependencies have changed, previous waiting or skipping decisions based on old dependencies may no longer be applicable, thus requiring a comprehensive re-evaluation. For example, if a prerequisite task completion event is received, asynchronous tasks that were previously added to the task ignore list because the prerequisite task was not completed should resume waiting if their prerequisites are now satisfied. If a prerequisite task cancellation event is received, asynchronous tasks that were previously waiting because the prerequisite task was not completed should be added to the task ignore list if their prerequisites are no longer possible to be completed.

[0084] Step S5443: Based on the result of the reassessment, update the task ignore list. Asynchronous tasks that were originally skipped and waiting are removed from the task ignore list and resume waiting after their prerequisites are completed. Asynchronous tasks that were originally waiting are added to the task ignore list after their prerequisites are canceled. The execution node updates the task ignore list based on the reassessment results. The update operation includes two scenarios. The first scenario is that an asynchronous task that was previously skipped from the task ignore list is removed and resumed waiting after its prerequisites are completed. When an asynchronous task was previously added to the task ignore list and skipped from waiting because its prerequisites were not completed, if a prerequisite completion event is received, it indicates that the prerequisites for the asynchronous task are now satisfied, and the asynchronous task has the conditions to continue execution. The execution node removes the asynchronous task from the task ignore list and re-includes it in the query results in subsequent periodic queries, resuming its waiting. The second scenario is that an asynchronous task that was previously continued to wait is added to the task ignore list after its prerequisites are cancelled. When an asynchronous task was previously continued to wait because its prerequisites were satisfied, if a prerequisite cancellation event is received, it indicates that the prerequisites for the asynchronous task are impossible to complete, and the asynchronous task has lost its basis for continued execution. The execution node adds the asynchronous task to the task ignore list, skipping its waiting, thus avoiding blocking the offline process due to waiting for a task that will never complete. Through the above update operations, the task ignore list always reflects the latest dependency status, ensuring that the waiting strategy is consistent with the actual dependencies.

[0085] Step S5444: Resume the periodic query and use the updated task ignore list for subsequent queries.

[0086] After updating the task ignore list, the execution node resumes the paused periodic queries. The resumed periodic queries use the updated task ignore list to filter the query results. Asynchronous tasks removed from the task ignore list reappear in subsequent query results, and the execution node continues to wait for their completion. Asynchronous tasks added to the task ignore list are filtered out in subsequent query results, and the execution node no longer waits for them. After resuming periodic queries, the execution node continues to execute subsequent queries and judgments according to the original polling interval until the offline condition is met or a timeout condition is triggered. Through an event-driven dynamic re-judgment mechanism, the execution node can respond to changes in dependencies in real time, avoiding the failure of the waiting strategy due to changes in dependencies.

[0087] Through the above embodiments, this application introduces an event-driven re-judgment mechanism for dependency changes. Execution nodes subscribe to dependency change events published by the control node. When dependencies change, periodic queries are paused, the waiting status of all executing asynchronous tasks is reassessed, and the task ignore list is updated, achieving real-time adaptation of offline decisions to the dynamic environment. This mechanism solves the problem of outdated waiting strategies caused by dynamic changes in dependencies during execution, enabling execution nodes to flexibly adjust waiting strategies based on the latest dependency status, further improving the intelligence and adaptability of the offline process.

[0088] Based on any embodiment of the method in this application, before performing the offline operation of this execution node, the following steps are included: Step S5501: Generate a dependency integrity audit report. The audit report includes the completion status of all completed asynchronous tasks and their downstream dependent tasks on this execution node, as well as the asynchronous tasks added to the task ignore list and the reasons for being skipped. After confirming that no tasks are running in the query results and before performing the shutdown operation, the execution node can generate a dependency integrity audit report. This report can be a structured data file that records the processing status of all asynchronous tasks involved in the entire shutdown process by this execution node, allowing the control node to verify the security of the shutdown operation.

[0089] The audit report contains at least two parts. The first part lists the completion status of all completed asynchronous tasks and their downstream dependent tasks on the current execution node. For each completed asynchronous task, the audit report records its task identifier, task type, completion time, and the identifiers and completion status of all its downstream dependent tasks. The completion status of downstream dependent tasks includes completed, executing, or not triggered. The second part lists asynchronous tasks added to the task ignore list and the reasons for their skipping. For each asynchronous task added to the task ignore list, the audit report records its task identifier, task type, addition time, and the reason for skipping. Reasons for skipping include the incompleteness of preceding dependent tasks on other execution nodes, the cancellation of preceding dependent tasks, and single-task timeout triggering a transfer. In addition, the audit report may also include auxiliary information such as the node identifier of the current execution node, the total time taken for the offline process, and the number of polling iterations. The audit report can be generated by the execution node by compiling historical data from its locally maintained offline progress records and task status tables.

[0090] Step S5502: Send the audit report to the control node, and the control node verifies whether there is a risk in the audit report that downstream dependent tasks may be permanently unable to complete due to the offline status of this node, and then returns the verification result; The execution node sends the generated dependency integrity audit report to the control node. Upon receiving the audit report, the control node verifies its contents, focusing on whether there is a risk that certain downstream dependent tasks may permanently fail to complete due to the execution node's shutdown. The specific verification logic is as follows: The control node iterates through the list of completed asynchronous tasks in the audit report. For each completed asynchronous task, it checks the completion status of its downstream dependent tasks. If a completed asynchronous task has downstream dependent tasks, and the completion status of these downstream dependent tasks is "not triggered" or "in execution," it further checks the execution node to which the downstream dependent task currently belongs. If the execution node to which the downstream dependent task currently belongs is this execution node, and the downstream dependent task has not been added to the task ignore list, there is a risk that the downstream dependent task may permanently fail to complete due to the execution node going offline. This is because the downstream dependent task depends on asynchronous tasks already completed on this execution node, and it is itself on this execution node; after this execution node goes offline, the downstream dependent task will lose its execution environment.

[0091] The control node also iterates through the list of asynchronous tasks that have been added to the task ignore list in the audit report. For each skipped asynchronous task, it checks whether the reason for skipping is reasonable and whether there is a risk that its downstream dependent tasks will be disconnected due to the offline status of this execution node.

[0092] After verification is complete, the control node generates a verification result, which includes either a pass or fail status, along with a detailed verification description. The control node then returns the verification result to the execution node.

[0093] Step S5503: If the verification result is successful, then perform the offline operation; otherwise, stop the offline operation and restore the node status of this execution node to the active state so as to receive new asynchronous tasks again.

[0094] After receiving the verification result from the control node, the execution node performs the corresponding processing based on the verification result. If the verification result is successful, it indicates that there is no risk that downstream dependent tasks will be permanently unable to complete due to the execution node's shutdown, and the execution node continues to perform the shutdown operation, safely exiting the cluster. If the verification result is unsuccessful, it indicates that there is a risk that some downstream dependent tasks will be permanently unable to complete due to the execution node's shutdown, and the execution node stops the shutdown operation.

[0095] When a shutdown operation is terminated, the execution node first calls the status update interface provided by the control node to restore its node status from inactive to active in the node configuration table. After the node status is restored to active, the scheduler will re-select the execution node as a candidate target in subsequent task allocation decisions and begin allocating new asynchronous tasks to it. Simultaneously, the execution node resumes normal operation, continuing to execute existing asynchronous tasks and receive new asynchronous tasks. Through this rollback mechanism, this application avoids the risk of data loss due to incorrect shutdown decisions, ensuring the security of the shutdown operation.

[0096] In the above embodiments, after confirming that there are no tasks in execution, the execution node first generates a dependency integrity audit report, and the control node verifies whether there is a risk that downstream dependent tasks will be permanently unable to complete due to the node's shutdown. The shutdown operation is only performed if the verification passes; if the verification fails, the shutdown is aborted and the node status is restored to active. This mechanism upgrades graceful shutdown from a best-effort approach to a verifiable and rollback-capable highly reliable process, fundamentally eliminating the risk of data loss due to shutdown and further improving the security and reliability of the shutdown process.

[0097] Please see Figure 3 According to one aspect of this application, a node offline control device includes an instruction response module 5100, a delayed execution module 5200, a task polling module 5300, a busy-hour processing module 5400, and an idle-hour offline module 5500. The instruction response module 5100 is configured to respond to an offline instruction by updating the node status of the current execution node in the node configuration table maintained by the control node to an inactive state, thereby preventing the scheduler in the control node from assigning new asynchronous tasks to the current execution node. The delayed execution module 5200 is configured to perform a delayed wait according to a preset delay duration to compensate for the propagation lag of the inactive state from the node configuration table to the scheduler. The task polling module 5300 is configured to periodically query the task status table maintained by the control node after the delay wait ends, and the task record corresponding to the current execution node with the execution status of being in progress and not belonging to the task ignore list as the query result; the busy time processing module 5400 is configured to, when the query result contains a task record in progress, determine whether the corresponding asynchronous task has any unfinished pre-dependent tasks. If so, skip the wait for the asynchronous task and add the asynchronous task to the task ignore list; otherwise, continue to wait for the asynchronous task to complete; the idle time offline module 5500 is configured to, when the query result does not contain a task record in progress, perform the offline operation of the current execution node.

[0098] Based on any embodiment of the device in this application, the delayed execution module 5200 includes: an interval acquisition module, configured to acquire the polling interval duration of the scheduler for the node configuration table; and a duration setting module, configured to use a value greater than the polling interval duration as the preset delay duration, during which the execution node does not initiate any query request.

[0099] Based on any embodiment of the apparatus in this application, the busy-hour processing module 5400 includes: a dependency determination module, configured to generate a dependency snapshot before the start of each round of periodic query, based on the task records of all asynchronous tasks on the current execution node in the task status table and the identifiers of the preceding tasks in the task records; a completion identification module, configured to determine the completion status of the preceding dependent tasks of each executing asynchronous task in the dependency snapshot based on the dependency snapshot; an external processing module, configured to determine that if the preceding dependent task is executed on another execution node outside the current execution node and its status is "in execution", then the preceding dependent task is not completed; and an internal processing module, configured to further wait for the preceding dependent task to complete before determining whether the current asynchronous task can be skipped if the preceding dependent task is executed on the current execution node and its status is "in execution".

[0100] Based on any embodiment of the device in this application, the busy time processing module 5400 includes: a synchronous broadcast module, configured to synchronize the update record of the task ignore list to the control node, so that the control node broadcasts the update record to other execution nodes currently in the offline process; and an external call module, configured to add the corresponding asynchronous task in the update record to the task ignore list of the other execution node after the other execution node receives the update record, so that when the other execution node queries the task status table in the offline process, the asynchronous task in the task ignore list will not be included in the query result.

[0101] Based on any embodiment of the device in this application, the busy time processing module 5400 includes: a waiting monitoring module, configured to monitor the waiting time of the asynchronous task while waiting for it to complete, and send a dependency timeout report to the control node when the waiting time exceeds a preset single task timeout threshold; a pending transfer marking module, configured to mark the corresponding asynchronous task and all its downstream dependent tasks as pending transfer status by the control node according to the dependency timeout report; a pending transfer reassignment module, configured to reassign the asynchronous tasks marked as pending transfer status to other execution nodes by the control node, and update the execution node to which the corresponding task record belongs in the task status table; and a transfer completion processing module, configured to add the asynchronous task to the task ignore list and continue subsequent queries after the execution node receives the transfer completion confirmation message from the control node.

[0102] Based on any embodiment of the device in this application, the busy-hour processing module 5400 includes: an event subscription module, configured to subscribe to dependency change events published by the control node, the dependency change events including the completion of a prerequisite dependency task, the reassignment of a prerequisite dependency task, or the cancellation of a prerequisite dependency task; a query pause module, configured to pause the current periodic query when the dependency change event is received, and re-determine the waiting status of all asynchronous tasks in execution on the current execution node according to the changed dependency; a list update module, configured to update the task ignore list according to the result of the re-determination, wherein asynchronous tasks that were originally skipped from waiting are removed from the task ignore list and resume waiting after their prerequisite dependencies are completed, and asynchronous tasks that were originally continued to wait are added to the task ignore list after their prerequisite dependencies are cancelled; and a query recovery module, configured to recover the periodic query and use the updated task ignore list for subsequent queries.

[0103] Based on any embodiment of the device in this application, the idle-time offline module 5500 includes: a report generation module, configured to generate a dependency integrity audit report, the audit report including the completion status of all completed asynchronous tasks and their downstream dependent tasks on the execution node, as well as asynchronous tasks added to the task ignore list and the reasons for being skipped; a control verification module, configured to send the audit report to the control node, the control node verifying whether there is a risk in the audit report that downstream dependent tasks may be permanently unable to complete due to the offline status of the node and then returning the verification result; and a result processing module, configured to execute the offline operation if the verification result is passed; otherwise, the offline operation is stopped, the node status of the execution node is restored to the active state, so as to receive new asynchronous tasks again.

[0104] Another embodiment of this application also provides an electronic device. For example... Figure 4 The diagram shows the internal structure of an electronic device. This electronic device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable, non-volatile storage medium stores an operating system, a database, and computer-readable instructions. The database can store information sequences, and when executed by the processor, the computer-readable instructions enable the processor to implement a node offline control method.

[0105] The processor of this electronic device provides computing and control capabilities to support the operation of the entire device. The memory of this electronic device can store computer-readable instructions, which, when executed by the processor, cause the processor to perform the node offline control method of this application. The network interface of this electronic device is used for communication with a terminal.

[0106] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0107] In this embodiment, the processor is used to execute... Figure 3 The specific functions of each module are described, and the memory stores the program code and various data required to execute the above modules or sub-modules. The network interface is used to realize data transmission between user terminals or servers. In this embodiment, the non-volatile readable storage medium stores the program code and data required to execute all modules in the node offline control device of this application, and the server can call the server's program code and data to execute the functions of all modules.

[0108] This application also provides a non-volatile readable storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the node offline control method of any embodiment of this application.

[0109] This application also provides a computer program product, including a computer program / instructions that, when executed by one or more processors, implement the steps of the method described in any embodiment of this application.

[0110] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM).

Claims

1. A node offline control method, characterized in that, include: In response to the offline command, the node status of this execution node in the node configuration table maintained by the control node is updated to inactive state to prevent the scheduler in the control node from assigning new asynchronous tasks to this execution node; A delay waiting period is performed according to a preset delay time to compensate for the propagation lag of the inactive state from the node configuration table to the scheduler; After the delay period ends, the task records in the task status table maintained by the control node that are in execution and are not in the task ignore list are periodically queried as query results; When the query results contain records of tasks in execution, it is determined whether the corresponding asynchronous task has any unfinished prerequisite dependent tasks. If so, the waiting for the asynchronous task is skipped and the asynchronous task is added to the task ignore list; otherwise, the asynchronous task is waited for to be completed. If the query results do not contain any records of tasks in progress, the current execution node will be taken offline.

2. The node offline control method according to claim 1, characterized in that, Perform a delayed wait according to a preset delay duration to compensate for the propagation lag of the inactive state from the node configuration table to the scheduler, including: Obtain the polling interval of the scheduler for the node configuration table; The preset delay time is a value greater than the polling interval. During the preset delay time, this execution node does not initiate any query requests.

3. The node offline control method according to claim 1, characterized in that, Determine if the corresponding asynchronous task has any unfinished prerequisite tasks, including: Before each round of periodic queries begins, a dependency snapshot is generated based on the task records of all asynchronous tasks on the current execution node in the task status table and the predecessor task identifiers in the task records. Based on the dependency snapshot, determine the completion status of the preceding dependent tasks of each executing asynchronous task in the dependency snapshot; If a prerequisite task is executed on an execution node other than this execution node and its status is "in execution", then the prerequisite task is determined to be incomplete. If a prerequisite task is being executed on this execution node and its status is "in execution", then we will wait for the prerequisite task to complete before determining whether the current asynchronous task can be skipped.

4. The node offline control method according to claim 1, characterized in that, After adding the asynchronous task to the task ignore list, the following is included: The update record of the task ignore list is synchronized to the control node, so that the control node can broadcast the update record to other execution nodes that are currently in the offline process; When other execution nodes receive the update record, they add the corresponding asynchronous task in the update record to the task ignore list of the other execution node, so that when the other execution node queries the task status table in the offline process, the asynchronous task in the task ignore list will not be included in the query result.

5. The node offline control method according to claim 1, characterized in that, Continue waiting for the asynchronous task to complete, including: While waiting for the asynchronous task to complete, the waiting time of the asynchronous task is monitored. When the waiting time exceeds the preset single task timeout threshold, a dependency timeout report is sent to the control node. Based on the dependency timeout report, the control node marks the corresponding asynchronous task and all its downstream dependent tasks as pending transfer. The control node reassigns asynchronous tasks marked as pending transfer to other execution nodes and updates the execution node to which the corresponding task record belongs in the task status table. Once this execution node receives the transfer completion confirmation message from the control node, it adds the asynchronous task to the task ignore list and continues subsequent queries.

6. The node offline control method according to claim 1, characterized in that, After determining whether the corresponding asynchronous task has any unfinished prerequisite tasks, the process includes: Subscribe to dependency change events published by the control node, including the completion of a prerequisite dependency task, the reassignment of a prerequisite dependency task, or the cancellation of a prerequisite dependency task; When the dependency change event is received, the current periodic query is paused, and the waiting status of all asynchronous tasks in execution on this execution node is reassessed based on the changed dependency. Based on the results of the reassessment, the task ignore list is updated. Asynchronous tasks that were originally skipped and were waiting are removed from the task ignore list and resume waiting after their prerequisites are completed. Asynchronous tasks that were originally waiting are added to the task ignore list after their prerequisites are canceled. Resume the periodic query and use the updated task ignore list for subsequent queries.

7. The node offline control method according to any one of claims 1 to 6, characterized in that, Before performing the offline operation on this execution node, the following steps are required: Generate a dependency integrity audit report, which includes the completion status of all completed asynchronous tasks and their downstream dependent tasks on this execution node, as well as the asynchronous tasks added to the task ignore list and the reasons for being skipped; The audit report is sent to the control node, which verifies whether there is a risk in the audit report that downstream dependent tasks may be permanently unable to complete due to the offline status of this node, and then returns the verification result. If the verification result is successful, the offline operation is executed; otherwise, the offline operation is aborted, and the node status of this execution node is restored to the active state so as to receive new asynchronous tasks again.

8. A node offline control device, characterized in that, include: The instruction response module is configured to respond to offline instructions by updating the node status of this execution node in the node configuration table maintained by the control node to an inactive state, so as to prevent the scheduler in the control node from assigning new asynchronous tasks to this execution node. The delayed execution module is configured to perform a delayed wait according to a preset delay duration to compensate for the propagation lag of the inactive state from the node configuration table to the scheduler; The task polling module is configured to periodically query the task status table maintained by the control node after the delay wait ends, and the task records whose execution status is "in execution" and which are not in the task ignore list corresponding to the current execution node are used as query results. The busy time processing module is configured to, when the query result contains a record of an executing task, determine whether the corresponding asynchronous task has any unfinished pre-dependent tasks. If so, skip waiting for the asynchronous task and add it to the task ignore list; otherwise, continue to wait for the asynchronous task to complete. The idle-time offline module is configured to perform an offline operation on the current execution node when the query results do not contain any records of tasks in progress.

9. An electronic device comprising a central processing unit and a memory, characterized in that, The central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the method as described in any one of claims 1 to 7.

10. A non-volatile readable storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to any one of claims 1 to 7, which, when invoked by a computer, executes the steps included in the corresponding method.