Streaming task scheduling method, system and equipment and storage medium

By analyzing the logs of the exception subtask nodes, positioning the interrupt location and restoring execution according to the preset strategy, the problem that the streaming task scheduling system is difficult to recover after the exception occurs, achieving more efficient and stable task processing.

CN120196410APending Publication Date: 2025-06-24SHANGHAI CTRIP DIGITAL INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510269297.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When the existing distributed task scheduling system handles streaming tasks, it is difficult to recover from the interrupted node after an exception occurs, resulting in waste of resources, repeated data processing or omission, and low task processing efficiency.

Method used

Position the interrupt location by parsing the subtask execution log of the exception subtask node, resume execution from the interrupt location of the exception subtask node according to the preset retry strategy, and skip the subtask node that has been successfully executed.

Benefits of technology

Avoid global restarts, save system resources, improve task continuity and fault tolerance, and ultimately improve task processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196410A_ABST
    Figure CN120196410A_ABST
Patent Text Reader

Abstract

The invention provides a streaming task scheduling method, system and device and a storage medium, and the method comprises the steps: sequentially scheduling sub-task nodes for execution according to a scheduling sequence, recording a sub-task execution log of each sub-task node in real time, collecting node operation data of each sub-task node, and storing the node operation data of each sub-task node; when at least one sub-task node is detected to be abnormal based on the node operation data, the interruption position is positioned by analyzing the sub-task execution log of the abnormal sub-task node, execution is recovered from the interruption position of the abnormal sub-task node according to the preset retry strategy, and the successfully executed sub-task node is skipped, so that global restart is avoided, and the efficiency of the system is improved. And system resources are saved, the task continuity is improved, and finally the task processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Currently, most distributed task scheduling systems regard the overall task as a continuous process and often adopt global restart or simple preset retry strategies in case of anomalies. Taking Figure 1 the financial daily cut-off business shown as an example, there are strict dependencies between the links of transaction data sorting, balance update, bill generation, and data archiving. However, if the entire task is restarted due to an anomaly in one of the subtasks, problems such as resource waste, duplicate data processing, or omission will occur.

[0003] It should be noted that the information disclosed in the above background art section is only used to strengthen the understanding of the background of the present invention, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0004] Aiming at the problems in the prior art, the purpose of the present invention is to provide a streaming task scheduling method, system, device, and storage medium, which overcome the difficulties of the prior art and can solve the technical problem of low streaming task processing efficiency.

[0005] The first aspect of the present disclosure provides a streaming task scheduling method, which includes:

[0006] In response to a task execution request, obtain the target task registration information, split the target task into multiple subtask nodes according to the target task registration information, and establish the scheduling order of each subtask node;

[0007] Schedule the execution of each subtask node in sequence according to the scheduling order, and record the subtask execution log of each subtask node in real time;

[0008] Collect the node operation data of each subtask node. When at least one subtask node is detected to be abnormal based on the node operation data, locate the interruption position by parsing the subtask execution log of the abnormal subtask node;

[0009] Resume execution from the interruption position of the abnormal subtask node according to the preset retry strategy, and skip the successfully executed subtask nodes.

[0010] Optionally, the target task registration information includes at least one of the task name, task type, business logic, dependencies between the subtask nodes, input parameters, execution environment, preset retry strategy, and warning rules of the target task, where the dependencies between the subtask nodes are used to establish the scheduling order of each subtask node.

[0011] Optionally, before responding to the task execution request and obtaining the target task registration information, the streaming task scheduling method further includes:

[0012] Receive and respond to a first user input to obtain the target task registration information;

[0013] Receive and respond to a second user input to obtain the task execution request.

[0014] Optionally, the resuming execution from the interruption position of the abnormal subtask node according to a preset retry policy includes:

[0015] Wait for a preset time interval before retrying, and record the retry times and recovery status during the retry;

[0016] If the retry times of the abnormal subtask node exceed a preset upper limit and the recovery status shows failure, automatically construct and send a warning message.

[0017] Optionally, the locating the task interruption position by parsing the subtask execution log of the abnormal subtask node includes:

[0018] Analyze the timestamps, error codes, and exception descriptions in the subtask execution log to determine the operation position of the last successful execution as the interruption position.

[0019] Optionally, a directed acyclic graph structure is used to establish the scheduling order of each subtask node.

[0020] Optionally, the collecting the node operation data of each subtask node includes:

[0021] Collect at least one operation metric such as the resource utilization rate, memory occupancy, network latency, I / O metric, and program error code of each subtask node as the node operation data for real-time anomaly detection.

[0022] Optionally, the streaming task scheduling method includes the following abnormal subtask node detection methods:

[0023] Compare the at least one metric with corresponding preset operation thresholds, and determine whether the corresponding subtask node is abnormal according to the comparison result.

[0024] Optionally, the resuming execution from the interruption position of the abnormal subtask node according to a preset retry policy includes:

[0025] According to the preset retry policy, restart the abnormal subtask node from the interruption position; and update the retry times and recovery status during the retry.

[0026] A second aspect of the present disclosure provides a streaming task scheduling system, which includes:

[0027] The task splitting module, in response to a task execution request, obtains target task registration information, splits the target task into multiple subtask nodes according to the target task registration information, and establishes a scheduling order for each subtask node;

[0028] The task execution module schedules the execution of each subtask node in sequence according to the scheduling order, and records the subtask execution logs of each subtask node in real time;

[0029] The exception detection module collects the node operation data of each subtask node. When at least one subtask node is detected to be abnormal based on the node operation data, it locates the interruption position by parsing the subtask execution log of the abnormal subtask node;

[0030] The task retry module resumes execution from the interruption position of the abnormal subtask node according to a preset retry policy, and skips the subtask nodes that have been successfully executed.

[0031] The third aspect of the present disclosure provides an electronic device, which includes:

[0032] A processor;

[0033] A memory, in which executable instructions of the processor are stored;

[0034] Wherein, the processor is configured to execute the steps of the streaming task scheduling method according to any of the above embodiments by executing the executable instructions.

[0035] The fourth aspect of the present disclosure provides a computer-readable storage medium for storing a program, and when the program is executed, it implements the steps of the streaming task scheduling method according to any of the embodiments.

[0036] The streaming task scheduling method, system, device and storage medium provided by the embodiments of the present disclosure have the following advantages:

[0037] During the execution of any target task, when an abnormal subtask node is detected, the interruption position is located by parsing the subtask execution log of the abnormal subtask node, and the execution is resumed only from the interruption position of the abnormal subtask node according to a preset retry policy, and the successfully executed subtask nodes are automatically skipped, thereby avoiding global restart, saving system resources, and improving the continuity of the task, and finally improving the task processing efficiency.

[0038] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects and advantages of the present invention will become more obvious.

[0040] Figure 1 A schematic diagram showing the dependency relationships among the steps in a financial daily cut operation of the prior art.

[0041] Figure 2 A flowchart showing a streaming task scheduling method provided by an embodiment of the present disclosure;

[0042] Figure 3 An architecture diagram of a streaming task scheduling system provided by an embodiment of the present disclosure;

[0043] Figure 4 A schematic diagram showing the module structure of a streaming task scheduling system provided by an embodiment of the present disclosure;

[0044] Figure 5 A schematic diagram showing the structure of an electronic device provided by an embodiment of the present disclosure;

[0045] Figure 6 It is a schematic diagram of the structure of a computer program product according to an embodiment of the present disclosure. Detailed implementation manners

[0046] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments.

[0047] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0048] As described in the background art, when the existing task scheduling engines currently available handle streaming tasks, there is often a problem that it is difficult to resume from the interruption node after the task is interrupted. Specifically, once an exception occurs at a certain step node of the task, the recovery of the entire task flow must start from the very beginning node, resulting in low task processing efficiency and making it difficult to ensure the reliability and accuracy of the process.

[0049] Through retrieval, it is found that although the technology of resume from breakpoint is widely used in the field of file transfer, there is still a lack of a technical solution for local recovery of each subtask in a streaming task scheduling system in the existing technology. Therefore, the embodiments of the present disclosure propose a breakpoint retry method based on fine-grained task splitting, real-time monitoring, and log parsing, aiming to only resume and execute the abnormal subtask nodes to achieve local recovery, and significantly improve the execution efficiency and stability of the overall streaming task.

[0050] The embodiments of the present disclosure propose a streaming task scheduling method, which not only includes task registration, splitting, execution, and logging, but also combines real-time collection of node operation data, anomaly detection, breakpoint positioning, and local recovery based on a preset retry policy, so that when an anomaly occurs during task execution, only the abnormal subtask nodes are resumed, skipping the successfully executed parts, and improving the system continuity and fault tolerance. This technology is applicable to occasions with high requirements for task continuity, such as financial daily cut-off, batch processing, and big data processing.

[0051] Specifically, as Figure 2 shown, the embodiments of the present disclosure provide a streaming task scheduling method, and this streaming task scheduling method is a streaming task scheduling system. The streaming task scheduling method of this embodiment includes but is not limited to the following steps:

[0052] Step 210: In response to a task execution request, obtain the target task registration information, split the target task into multiple subtask nodes according to the target task registration information, and establish the scheduling order of each subtask node;

[0053] Step 220: Schedule the execution of each subtask node in sequence according to the scheduling order, and record the subtask execution log of each subtask node in real time;

[0054] Step 230: Collect the node operation data of each subtask node, and when at least one subtask node is detected to be abnormal based on the node operation data, locate the interruption position by parsing the subtask execution log of the abnormal subtask node;

[0055] Step 240: Resume the execution from the interruption position of the abnormal subtask node according to the preset retry policy, and skip the successfully executed subtask nodes.

[0056] Using the above streaming task scheduling method, during the execution of any target task, when a certain subtask node is detected to be abnormal, the interruption position is located by parsing the subtask execution log of the abnormal subtask, and the execution is resumed only from the interruption position of the abnormal subtask node according to the preset retry policy, and the successfully executed subtask nodes are automatically skipped, so as to avoid global restart, save system resources, improve the task continuity, and finally improve the task processing efficiency.

[0057] This includes the following technical effects:

[0058] 1) Ensure the continuity of streaming task scheduling: The function of breakpoint retry significantly reduces the waste of resources and time, which is particularly important for tasks that need to run for a long time or resource-intensive tasks.

[0059] 2) Enhance system stability: The retry policy improves the fault tolerance mechanism of the system and ensures the integrity of task execution in case of exceptions.

[0060] 3) Improve user experience: The breakpoint recovery function reduces the manual intervention of developers in tasks and lowers the cost of manual operation and maintenance.

[0061] In the embodiment of the present disclosure, when performing step 210, the task execution request may come from user clicks, scheduled task triggers, or external system calls.

[0062] In one embodiment, the target task registration information is stored in advance. In response to the task execution request, the system first obtains the target task registration information stored in the database or configuration file, and this information includes at least one of the task name, task type, business logic, the dependency relationship between each sub-task node, input parameters, execution environment, preset retry policy, and warning rules.

[0063] Subsequently, the system divides the entire target task into multiple sub-task nodes according to these target task registration information. For example, in the financial day cut scenario, the target task can be split into sub-task nodes such as "transaction data classification", "balance update", "bill generation", and "data backup and archiving"; and according to the actual business process, based on the above dependency relationship, a scheduling order is constructed.

[0064] In this way, only after the previous task is completed can the subsequent task be executed, thus ensuring the correctness of the business process.

[0065] In another embodiment, the target task registration information can be dynamically registered or registered during the execution of the target task. For example, before obtaining the target task registration information in response to the task execution request, the streaming task scheduling method further includes:

[0066] Receiving and responding to the first user input to obtain the target task registration information;

[0067] Receiving and responding to the second user input to obtain the task execution request.

[0068] In this embodiment, the first user input is used to represent the user-configured target task registration information. For example, the user is allowed to dynamically register the target task through a graphical interface or an API interface and configure the target task registration information. At this time, the system presents a graphical interface to the user, and the graphical interface displays an input window or menu options to receive the first user input. Alternatively, the system provides an API interface that allows the user to upload a configuration file with the target task registration information.

[0069] The second user input is used to generate a task execution request, such as the user clicking or an interface icon being touched.

[0070] In an alternative approach, a Directed Acyclic Graph (DAG) structure is used to establish the scheduling order of each subtask node. A DAG structure is a special graph structure that consists of a set of vertices (subtask nodes) and directed edges, and there are no cycles in the graph. In other words, a DAG is a directed graph where the edges have a direction and it is not possible to start from a vertex and return to that vertex along the direction of the edges.

[0071] In task scheduling, there may be dependencies between subtask nodes, and a DAG can be used to represent the scheduling (execution) order of these subtask nodes. For example, if subtask node A must be completed before subtask node B, it can be represented by a directed edge A → B. Topological sorting can be used to determine the scheduling order of subtask nodes to ensure that all dependencies are satisfied.

[0072] In the embodiment of the present disclosure, after the scheduling order of each subtask node is established, step 220 is executed. After the target task is split, each subtask node is started one by one according to the established scheduling order (such as the topological sorting of subtask nodes in a DAG).

[0073] During the operation of each subtask node, a predefined business logic function is called, and at the same time, the execution status (including information such as start time, end time, execution result, error code, etc.) is recorded in real time to form a detailed subtask execution log. These subtask execution logs are not only used for retrospective recording of the completion status of the target task but also provide key data basis for subsequent fault detection and breakpoint recovery.

[0074] During the execution of the target task, step 230 is executed. Specifically, the node operation data of each subtask node during execution is collected in real time, such as at least one operation metric among resource utilization rate (such as CPU utilization rate), memory occupancy, network latency, disk I / O, and program error code, for real-time anomaly detection.

[0075] In an alternative embodiment, the operation metrics included in the above node operation data are compared with corresponding preset operation thresholds, and it is determined whether the corresponding sub-task node is abnormal according to the comparison result.

[0076] When it is detected that the operation metrics of a certain sub-task node are abnormal (for example, the CPU usage rate suddenly rises or the response time is abnormally extended), the system determines that the sub-task node is abnormal and regards it as an abnormal sub-task node. Subsequently, the system automatically retrieves the sub-task execution log of the abnormal sub-task node and parses information such as the timestamp, error code, and exception description in the log, so as to accurately locate the specific interruption position in the corresponding sub-task. Specifically, the position of the last successfully executed operation is determined as the interruption position.

[0077] Using this embodiment, this process can ensure that the root cause of the exception is quickly captured and provide an accurate basis for local recovery.

[0078] In the embodiment of the present disclosure, after determining the interruption position of the abnormal sub-task node, the system restores the abnormal sub-task node according to a preset retry policy.

[0079] Optionally, the preset retry policy can be one or more of a fixed interval retry, an exponential backoff retry, or an adaptive retry mode. Among them, the fixed interval retry (Fixed Interval Retry) is the simplest retry policy. In this mode, the time interval between each retry operation is fixed.

[0080] The exponential backoff retry is a more intelligent retry policy, and the time interval of each retry will gradually increase. An exponential growth method can be adopted, for example, the time interval of each retry is twice that of the previous one.

[0081] The adaptive retry is a more flexible retry policy, which dynamically adjusts the time interval and number of retries according to the type of exception of the current abnormal sub-task node.

[0082] In this embodiment, the system retries the sub-task node detected as abnormal, and does not need to re-execute those sub-task nodes that have been successfully completed. In this way, the recovery operation only starts from the interruption position, which not only shortens the task recovery time but also avoids repeated processing of the completed sub-tasks. If the recovery is still not successful after several consecutive retries, the system can also trigger a warning notification to require the operation and maintenance personnel to intervene and handle it.

[0083] In this way, resume the execution from the interruption position of the abnormal subtask node according to the preset retry policy, which specifically includes: restart the abnormal subtask node from the interruption position according to the preset retry policy; and update the retry count and status during the retry process.

[0084] Optionally, wait for a preset time interval before the retry, and record the retry count and recovery status during the retry;

[0085] When the retry count of the abnormal subtask node exceeds the preset upper limit and the recovery status shows failure, automatically construct and send a warning message. This warning message is used to notify the operation and maintenance personnel. For example, send a warning notice to the preset operation and maintenance personnel via email, text message or instant messaging tool.

[0086] The following takes the financial end-of-day processing (End-of-Day Processing or Daily Cut-Off) as an application scenario of the target task to elaborate on the specific implementation method of the streaming task scheduling method.

[0087] In the financial end-of-day processing task, four subtask nodes of transaction data classification, balance update, bill generation, and data backup and archiving need to be executed daily. When a database query exception occurs at the "balance update" node in the traditional scheduling system, the entire end-of-day task will be restarted from the beginning, affecting the overall processing efficiency and data consistency.

[0088] The end-of-day task scheduling method of this embodiment includes the following steps:

[0089] 1. Registration and splitting of the end-of-day task.

[0090] Respond to the end-of-day task execution request, and obtain the task registration information pre-entered by the bank, including:

[0091] Task name: "End-of-day data processing task";

[0092] Subtask nodes: A. Transaction data classification; B. Balance update; C. Bill generation; D. Data backup and archiving;

[0093] Dependency relationship: A→B→C→D;

[0094] Preset retry policy: For example, for the "balance update" node, the fixed retry interval is set to 5 seconds, and the maximum retry count is 3 times.

[0095] The above task registration information is stored in the system in a standardized JSON format, and the task is automatically split to establish a DAG structure.

[0096] 2. Task execution and log recording.

[0097] The scheduler starts the sub - task nodes one by one in the order of dependencies.

[0098] The sub - task node A executes normally and generates detailed sub - task execution logs;

[0099] During the execution of sub - task node B, an error occurs due to database response latency. "DB_ERR_1001" is recorded in the sub - task execution log, and the execution failure time is marked.

[0100] All sub - task execution logs are stored in the log database in real - time for subsequent queries.

[0101] 3. Node monitoring and anomaly detection.

[0102] Meanwhile, the monitoring module collects the node operation data of each sub - task node every second.

[0103] For sub - task node B, the node operation data shows that its response time exceeds the preset operation threshold and the memory occupancy is abnormal;

[0104] Combining these metrics, the system automatically determines that an anomaly has occurred in sub - task node B and transmits the anomaly information along with the sub - task execution date to the breakpoint recovery module.

[0105] 4. Breakpoint location and selective retry.

[0106] After receiving the anomaly information, the breakpoint recovery module parses the sub - task execution log of sub - task node B and determines that the anomaly occurred in the middle of the balance update operation (for example, an error occurred during the second database query);

[0107] According to the preset retry policy, the system initiates the first retry after waiting for 5 seconds (or other time intervals), and resumes the execution of sub - task node B only from the breakpoint position, without repeating the execution of sub - task node A.

[0108] If the first retry is still not successful, the system continues to wait and perform the second and even the third retry. If all three retries fail, the system generates a warning message.

[0109] 5. Warning and manual intervention.

[0110] When sub - task node B still cannot be recovered after 3 consecutive retries, the system automatically triggers a warning notification, sending text messages and emails to the on - duty operation and maintenance personnel, indicating "Abnormal balance update, please check the database connection".

[0111] After the operation and maintenance personnel intervene and process according to the warning information, they adjust the database parameters and restart sub - task node B until the task resumes normal execution.

[0112] 6. Event collection and feedback.

[0113] During the entire task execution process, all operations, retry records, and warning events of the task are collected and archived, and the system generates a detailed task report to provide a basis for subsequent performance optimization.

[0114] Through this embodiment, only the "balance update" node is retried at the breakpoint, which greatly shortens the recovery time, avoids restarting the global task, and ensures that data is not processed repeatedly. The warning notification mechanism further ensures that manual intervention is timely when the retry is ineffective, overall improving the execution efficiency and data consistency of the daily cut task.

[0115] To implement the above-mentioned streaming task scheduling method, this embodiment provides a system architecture diagram of a streaming task scheduling engine, as Figure 3 shown. The streaming task scheduling system based on this streaming task scheduling engine includes five parts: a task registration module 31, a task execution module 32, a task retry module 33, a warning notification module 34, and an event center 35.

[0116] The main technical function points of each module are described as follows:

[0117] (1) Task registration module 31:

[0118] A) Task definition: Allows users or developers to create tasks, including task names, task types, execution logics, dependencies, etc.

[0119] B) Parameter configuration: Supports detailed parameter configuration for tasks, such as input parameters, retry policies, warning methods, etc., to meet the requirements of different scenarios.

[0120] C) Dynamic registration and cancellation: Supports dynamic registration and cancellation of tasks, enabling the system to flexibly adapt to changing requirements.

[0121] (2) Task execution module 32:

[0122] A) Task scheduling: Automatically schedules the execution of tasks according to task configurations (corresponding to the above-mentioned task execution requests, such as triggering by timing, API calls, etc.) to ensure that tasks run on time.

[0123] B) Records the execution process, results, and any exception information of the task and its corresponding sub-task nodes, and generates sub-task execution logs for easy problem tracking and performance analysis.

[0124] C) Sub-task node monitoring: Real-time monitors the running status of sub-task nodes, including node running data such as the execution progress of sub-task nodes and the exception information of sub-task node execution.

[0125] (3) Task retry module 33:

[0126] A) Retry Policy Maintenance: Provide multiple different retry policies for users and developers to choose from;

[0127] B) Node Failure Detection: Monitor the health status of subtask nodes and promptly detect subtask nodes that need to be retried.

[0128] C) Breakpoint Recovery and Retry: Trigger the retry logic of the task based on the configured preset retry policy and resume from the subtask node where the task was interrupted.

[0129] (4) Warning Notification Module 34:

[0130] A) Abnormality Warning: Send a warning notification promptly when the task execution fails.

[0131] B) Notification Channels: Support multiple notification channels, such as emails, text messages, instant messaging tools, phones, etc., to ensure that information can be quickly conveyed to the operation and maintenance personnel.

[0132] C) Customizable Notification Rules: Allow users to customize notification rules according to their needs, such as the advance notice of node failures, notification frequencies, etc.

[0133] (5) Event Center 35:

[0134] A) Event Collection: Asynchronously collect important events that occur within the event collection client, such as the status of subtask execution, execution information, etc.

[0135] B) Event Processing: Support an event-driven programming model and allow users to define event processing logic, such as logging of subtask node logs, status updates, etc.

[0136] Through the collaborative work of these five main modules, the system can not only schedule and execute tasks efficiently and reliably, but also respond quickly in the face of failures or abnormalities, resume running from the task interruption node after the failure is recovered, and ensure business continuity and reliability.

[0137] Figure 4 It is a module schematic diagram of an embodiment of the streaming task scheduling system provided by the present disclosure. As Figure 4 shown, the streaming task scheduling system 400 of the present disclosure includes but is not limited to:

[0138] Task Splitting Module 410, in response to a task execution request, obtains target task registration information, splits the target task into multiple subtask nodes according to the target task registration information, and establishes the scheduling order of each subtask node;

[0139] Task Execution Module 420, schedules the execution of each subtask node in sequence according to the scheduling order, and records the subtask execution logs of each subtask node in real time;

[0140] Anomaly detection module 430 collects the node operation data of each sub-task node. When at least one sub-task node is detected to be abnormal based on the node operation data, it locates the task interruption position by parsing the sub-task execution log of the abnormal sub-task node.

[0141] Task retry module 440 resumes execution from the interruption position of the abnormal sub-task node according to a preset retry policy and skips the sub-task nodes that have been successfully executed.

[0142] In an alternative embodiment, the target task registration information includes at least one of the task name, task type, business logic, the dependency relationship between each sub-task node, input parameters, execution environment, preset retry policy, and warning rule of the target task, where the dependency relationship between each sub-task node is used to establish the scheduling order of each sub-task node.

[0143] In an alternative embodiment, before obtaining the target task registration information in response to a task execution request, the task splitting module 410 is further configured to:

[0144] Receive and respond to a first user input to obtain the target task registration information;

[0145] Receive and respond to a second user input to obtain the task execution request.

[0146] In an alternative embodiment, the task retry module 440:

[0147] Waits for a preset time interval before retrying and records the retry count and recovery status during the retry;

[0148] If the retry count of the abnormal sub-task node exceeds a preset upper limit and the recovery status shows failure, an early warning message is automatically constructed and sent.

[0149] In an alternative embodiment, the anomaly detection module 430 is specifically configured to:

[0150] Analyze the timestamps, error codes, and anomaly descriptions in the sub-task execution log to determine the position of the last successfully executed operation as the interruption position.

[0151] In an alternative embodiment, a directed acyclic graph structure is used to establish the scheduling order of each sub-task node.

[0152] In an alternative embodiment, the anomaly detection module 430 is specifically configured to:

[0153] Collect at least one of the resource utilization rate, memory occupancy, network latency, I / O metrics, and program error codes of each sub-task node as the node operation data for real-time anomaly detection.

[0154] In an alternative embodiment, the anomaly detection module 430 specifically performs the following anomaly sub-task node detection method:

[0155] Compare the at least one metric with the corresponding preset running threshold, and determine whether the corresponding sub-task node is abnormal according to the comparison result.

[0156] In an alternative embodiment, the task retry module 440 is specifically configured to:

[0157] Restart the abnormal sub-task node from the interruption position according to the preset retry policy; and update the retry count and recovery status during the retry process.

[0158] During the execution of any target task in this system, when an abnormal sub-task node is detected, the interruption position is located by parsing the sub-task execution log of the abnormal sub-task, and the execution is resumed only from the interruption position of the abnormal sub-task node according to the preset retry policy, and the successfully executed sub-task nodes are automatically skipped, thereby avoiding global restart, saving system resources, and improving the continuity of the task, and finally improving the task processing efficiency.

[0159] This system consists of a task splitting module 410, a task execution module 420, an anomaly detection module 430, and a task retry module 440. Each module can be interconnected through a data interface to jointly implement a streaming task scheduling solution. This system can be embedded into a front-end development tool as an independent software module, or can provide an interface as a cloud service for developers to remotely call.

[0160] An embodiment of the present invention further provides a streaming task scheduling device, including a processor and a memory in which executable instructions of the processor are stored. Wherein, the processor is configured to complete the steps of the streaming task scheduling method by executing the executable instructions.

[0161] As shown above, when using the streaming task scheduling device of the present invention, during the execution of any target task, when an abnormal sub-task node is detected, the interruption position is located by parsing the sub-task execution log of the abnormal sub-task, and the execution is resumed only from the interruption position of the abnormal sub-task node according to the preset retry policy, and the successfully executed sub-task nodes are automatically skipped, thereby avoiding global restart, saving system resources, and improving the continuity of the task, and finally improving the task processing efficiency. At the same time, the breakpoint retry function significantly reduces the waste of resources and time, which is particularly important for tasks that need to run for a long time or resource-intensive tasks, and improves the fault tolerance mechanism of the system, ensuring the integrity of task execution in case of anomalies. In addition, the breakpoint recovery function reduces the manual intervention of developers in tasks and reduces the cost of manual operation and maintenance.

[0162] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, method, or program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to herein as "circuitry", "module", or "platform".

[0163] Figure 5 is a schematic structural diagram of the streaming task scheduling device of the present invention. The following refers to Figure 5 to describe the electronic device 500 according to this embodiment of the present invention. Figure 5 The electronic device 500 shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0164] As Figure 5 shown, the electronic device 500 is presented in the form of a general-purpose computing device. The components of the electronic device 500 may include, but are not limited to: at least one processing unit 510, at least one storage unit 520, a bus 530 connecting different platform components (including the storage unit 520 and the processing unit 510), a display unit 540, etc.

[0165] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 510, so that the processing unit 510 executes the steps according to various exemplary embodiments of the present invention described in the above-mentioned streaming task scheduling method part of this specification. For example, the processing unit 510 can execute the steps as Figure 2 shown in.

[0166] The storage unit 520 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 521 and / or a cache storage unit 522, and may further include a read-only storage unit (ROM) 523.

[0167] The storage unit 520 may further include a program / utility 524 having a set (at least one) of program modules 525. Such program modules 525 include, but are not limited to: a processing system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples.

[0168] The bus 530 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.

[0169] The electronic device 500 can also communicate with one or more external devices 501 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 500, and / or communicate with any device that enables the electronic device 500 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 550. Moreover, the electronic device 500 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 560. The network adapter 560 can communicate with other modules of the electronic device 500 through the bus 530. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.

[0170] Embodiments of the present disclosure also provide a computer-readable storage medium for storing a program, and the steps of a streaming task scheduling method are implemented when the program is executed. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above-mentioned streaming task scheduling method part of this specification.

[0171] As Figure 6 shown, the computer program product 600 for implementing the above method according to an embodiment of the present invention can adopt a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the computer program product of the present invention is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0172] A computer program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0173] The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted with any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0174] The program code for performing the processing of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).

[0175] In summary, the object of the present invention is to provide a streaming task scheduling method, system, device, and storage medium. During the execution of any target task, when an abnormal sub-task node is detected, the interruption position is located by parsing the sub-task execution log of the abnormal sub-task, and the execution is resumed only from the interruption position of the abnormal sub-task node according to a preset retry policy, and the successfully executed sub-task nodes are automatically skipped, thereby avoiding global restart, saving system resources, and improving the continuity of the task, and finally improving the task processing efficiency.

[0176] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as falling within the protection scope of the present invention.

Claims

1. A streaming task scheduling method, characterized in that: include: In response to a task execution request, obtaining target task registration information, splitting the target task into multiple subtask nodes according to the target task registration information, and establishing a scheduling order for each subtask node; Scheduling the execution of each subtask node in turn according to the scheduling order, and recording the subtask execution log of each subtask node in real time; Collecting node operation data of each subtask node, and when at least one subtask node is detected to be abnormal based on the node operation data, locating the interruption position by parsing the subtask execution log of the abnormal subtask node; The execution is resumed from the interruption position of the abnormal subtask node according to a preset retry strategy, and the subtask nodes that have been successfully executed are skipped.

2. The streaming task scheduling method according to claim 1, characterized in that: The target task registration information includes the task name of the target task, task type, business logic, dependencies between the subtask nodes, input parameters, execution environment, preset retry strategy and at least one of early warning rules, wherein the dependencies between the subtask nodes are used to establish the scheduling order of the subtask nodes.

3. The streaming task scheduling method according to claim 1, characterized in that: Before obtaining the target task registration information in response to the task execution request, the streaming task scheduling method further includes: Receiving and responding to a first user input, obtaining the target task registration information; The task execution request is received and obtained in response to a second user input.

4. The streaming task scheduling method according to claim 1, characterized in that: The resuming execution from the interruption position of the abnormal subtask node according to a preset retry strategy includes: Wait for a preset time interval before retrying, and record the number of retries and recovery status during the retry period; If the number of retries of the abnormal subtask node exceeds a preset upper limit and the recovery status shows failure, an early warning message is automatically constructed and sent.

5. The method according to claim 1, characterized in that The locating the task interruption position by parsing the subtask execution log of the abnormal subtask node includes: The timestamp, error code and exception description in the subtask execution log are analyzed to determine the last successfully executed operation position as the interruption position.

6. The streaming task scheduling method according to claim 1, characterized in that: A directed acyclic graph structure is used to establish the scheduling order of each subtask node.

7. The streaming task scheduling method according to claim 1, characterized in that: The collecting of node operation data of each subtask node includes: At least one operating indicator of the resource utilization rate, memory occupancy, network delay, I / O index and program error code of each subtask node is collected as the node operation data to perform real-time anomaly detection.

8. The streaming task scheduling method according to claim 7, characterized in that: The streaming task scheduling method includes the following abnormal subtask node detection method: The at least one indicator is compared with a corresponding preset operating threshold, and whether the corresponding subtask node is abnormal is determined according to the comparison result.

9. The streaming task scheduling method according to claim 1, characterized in that: The resuming execution from the interruption position of the abnormal subtask node according to a preset retry strategy includes: According to the preset retry strategy, the abnormal subtask node is restarted from the interruption position; and the number of retries and the recovery status are updated during the retry process.

10. A streaming task scheduling system, characterized in that: include: The task splitting module, in response to the task execution request, obtains the target task registration information, splits the target task into multiple subtask nodes according to the target task registration information, and establishes a scheduling order for each subtask node; The task execution module sequentially schedules the execution of each subtask node according to the scheduling order, and records the subtask execution log of each subtask node in real time; an abnormality detection module, which collects node operation data of each subtask node, and locates the interruption position by parsing the subtask execution log of the abnormal subtask node when at least one subtask node is detected to be abnormal based on the node operation data; The task retry module resumes execution from the interruption position of the abnormal subtask node according to a preset retry strategy, and skips the subtask nodes that have been successfully executed.

11. An electronic device, characterized in that: include: processor; a memory storing executable instructions of the processor; The processor is configured to execute the steps of the streaming task scheduling method according to any one of claims 1 to 9 by executing the executable instructions.

12. A computer-readable storage medium for storing a program, characterized in that: When the program is executed, the steps of the streaming task scheduling method described in any one of claims 1 to 9 are implemented.

Citation Information

Cited By

  • Task flow branch node failure accurate retry method, system and equipment

    CN121807500A

  • Task decomposition arrangement and exception retry method and system for workflow engine

    CN121961495A