Workflow engine system and task scheduling method for super computing cluster

Through the workflow engine system for supercomputing clusters, visual orchestration and automated scheduling are achieved, which solves the shortcomings of task scheduling methods in existing technologies, improves the efficiency and flexibility of scientific research tasks, and is suitable for various scientific research scenarios.

CN120723397APending Publication Date: 2025-09-30COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510772181.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing supercomputing cluster task scheduling methods lack general and visual task orchestration capabilities, are difficult to adapt to multiple task types and heterogeneous computing resources, lack workflow templates and node reuse mechanisms, and cannot meet the customized and efficient execution needs of multi-role and cross-domain scientific research users.

Method used

A workflow engine system for supercomputing clusters is provided, including a drag-and-drop orchestration design module, a virtualized task encapsulation and adaptation module, and a workflow automation execution module, which realizes visual orchestration, unified encapsulation and automated scheduling, and supports customized task management for multi-role and cross-domain scientific research users.

Benefits of technology

It improves the efficiency of scientific research task process construction, scheduling flexibility and execution stability in supercomputing clusters, has good scalability and cross-platform adaptability, and is suitable for various scientific research scenarios such as artificial intelligence, big data, material simulation, and meteorological analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723397A_ABST
    Figure CN120723397A_ABST
Patent Text Reader

Abstract

The invention provides a hypercomputing cluster-oriented workflow engine system and task scheduling method, and the system comprises a drag-and-drop arrangement design module which is used for providing a visual workflow arrangement environment so as to respond to a drag-and-drop arrangement operation of a user to obtain a target workflow; the target workflow comprises a plurality of tasks and a dependency relationship among the plurality of tasks; the virtualization task packaging and adaptation module is used for packaging a plurality of tasks in the target workflow by utilizing a unified packaging mechanism to obtain a plurality of target containers; and the workflow automatic execution module is used for managing and distributing hardware resources in the super computing cluster so as to schedule and execute tasks in the target containers according to the dependency relationship. Through the system provided by the invention, the complex workflow can be efficiently executed and managed in a super computing cluster environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more particularly to a workflow engine system and task scheduling method for supercomputing clusters. Background Art

[0002] With the widespread application of big data and artificial intelligence technologies in scientific research, engineering, energy and other fields, the complexity and computing density of scientific research computing tasks are constantly increasing. Traditional script-based, manually submitted task scheduling methods have gradually exposed many problems in dealing with high concurrency, multi-task dependencies and multi-resource collaboration scenarios.

[0003] Supercomputers (High Performance Computing, HPC) are high-performance computing systems typically used to handle large-scale, complex computing tasks. In supercomputing cluster environments, researchers often need to frequently switch between different models, parameters, data sources, and scheduling strategies. This is especially true for tasks such as deep learning training, high-throughput simulations, and large-scale data preprocessing. Existing task scheduling methods have the following major drawbacks:

[0004] Lack of universal and visual task orchestration capabilities, requiring users to manually maintain complex scripts and dependency graphs;

[0005] It is difficult to adapt to multiple task types and heterogeneous computing resources, and the task scheduling efficiency is low;

[0006] The lack of workflow templates and node reuse mechanisms leads to repeated process construction, resulting in low scientific research efficiency;

[0007] There is a lack of system support for functions such as containerized deployment, multi-user isolation, and retry of task failures.

[0008] Although some scheduling systems (such as Slurm) currently provide job dependency management functions, they tend to be implemented at the bottom level and lack a user-friendly visual configuration interface and integrated node logic expression. They cannot meet the needs of multi-role, cross-domain scientific research users for customized, intelligent, and efficient execution processes. Summary of the Invention

[0009] The embodiments of the present application provide a workflow engine system and task scheduling method for a supercomputing cluster, which can efficiently execute and manage complex workflows in a supercomputing cluster environment.

[0010] In a first aspect, an embodiment of the present application provides a workflow engine system for a supercomputing cluster, the system comprising:

[0011] The drag-and-drop orchestration design module provides a visual workflow orchestration environment to respond to the user's drag-and-drop orchestration operations to obtain the target workflow; the target workflow includes multiple tasks and the dependencies between the multiple tasks;

[0012] The virtualization task encapsulation and adaptation module is used to encapsulate multiple tasks in the target workflow using a unified encapsulation mechanism to obtain multiple target containers;

[0013] A workflow automation execution module is used to manage and allocate hardware resources in a supercomputing cluster to schedule and execute tasks in multiple target containers based on dependencies.

[0014] Therefore, the present invention proposes a customized workflow engine system for supercomputing clusters. This system provides users with a modular, graphical, and reusable workflow construction and automatic scheduling platform, thereby enabling the efficient organization and operation of various computing tasks such as AI training, reasoning, and data processing in a supercomputing environment.

[0015] In a second aspect, an embodiment of the present application provides a task scheduling method for a supercomputing cluster. The method is executed by a system including a drag-and-drop orchestration design module, a virtualized task encapsulation and adaptation module, and a process template support module. The method includes:

[0016] In response to a user's drag-and-drop arrangement operation in the drag-and-drop arrangement design module, a target workflow is generated; the target workflow includes multiple tasks and dependencies between the multiple tasks;

[0017] The virtualized task encapsulation and adaptation module is used to encapsulate multiple tasks separately to obtain multiple target containers; the encapsulation adopts a unified encapsulation strategy;

[0018] Leveraging process template support modules, hardware resources in a supercomputing cluster are managed and allocated to schedule and execute tasks in multiple target containers based on dependencies.

[0019] It can be understood that the beneficial effects of the second aspect mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1The following is a diagram showing the architecture of a workflow engine system for supercomputing clusters provided by an embodiment of the present application;

[0022] Figure 2 A schematic diagram of a workflow engine system for supercomputing clusters provided by an embodiment of the present application is shown;

[0023] Figure 3 A schematic diagram of a workflow provided by an embodiment of the present application is shown;

[0024] Figure 4 A schematic diagram of the task node structure provided in an embodiment of the present application is shown;

[0025] Figure 5 A schematic diagram of node attribute settings provided by an embodiment of the present application is shown;

[0026] Figure 6 A schematic diagram of a task scheduling method for a supercomputing cluster provided in an embodiment of the present application is shown;

[0027] Figure 7 A flowchart of a workflow management system for a supercomputing cluster provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0029] In the description of the embodiments of this application, any embodiment or design scheme using "exemplary," "for example," or "for example" should not be understood as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.

[0030] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly identifying the technical features being referred to. Thus, a feature specified as "first" or "second" may explicitly or implicitly include one or more of such features. The terms "include," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized. "Multiple" can mean one or more, where "multiple" means two or more.

[0031] To address the problems in existing technologies, the present invention proposes a customized workflow engine system for supercomputing clusters. This system provides users with a modular, graphical, and reusable workflow construction and automatic scheduling platform, enabling the efficient organization and operation of various computing tasks such as AI training, inference, and data processing in a supercomputing environment.

[0032] Figure 1 FIG1 shows a system architecture diagram of a workflow engine for supercomputing clusters provided by an embodiment of the present application. Figure 1 As shown in FIG, the workflow engine system can be deployed as a master node on any supercomputer in a supercomputing cluster. The master node includes a web server (Web Server), a web user interface (Web UI) and a database (MySQL).

[0033] The Web UI provides a visual interactive interface through which users can submit computation flows consisting of multiple computing tasks. The Web Server is responsible for processing user requests, providing an interface for submitting computing tasks, and scheduling and allocating tasks to achieve load balancing across computing nodes. MySQL is used to store and manage computing task information, such as task status and results.

[0034] Each supercomputer in a supercomputing cluster has pre-deployed compute nodes. Each compute node has an executor (server), which is responsible for allocating and invoking the supercomputer's various hardware resources (resources), such as the central processing unit (CPU), graphics processing unit (GPU), and memory, to execute the computing tasks assigned to that compute node. Compute nodes have data dependencies.

[0035] Based on this architecture, users can customize task node parameters and command logic through a graphical interface and connect task nodes by dragging and dropping to form a directed acyclic graph (DAG). The system automatically generates executable cluster scheduling plans based on dependencies and integrates with scheduling systems (such as Slurm) to automate job management.

[0036] In the process of allocating supercomputer hardware resources based on computing task requirements, the master node receives computing tasks submitted by users. The computing nodes use a load balancing mechanism to ensure even distribution of tasks, ultimately allocating the required hardware resources to the task. This architecture effectively utilizes the computing power of the supercomputer to handle large-scale, complex computing tasks.

[0037] Based on the above content, a workflow engine system for supercomputing clusters proposed in this application is introduced in detail.

[0038] Figure 2 FIG. 1 shows a schematic diagram of a workflow engine system for a supercomputing cluster provided by an embodiment of the present application. Figure 2 As shown, the workflow engine system 200 includes the following components:

[0039] The drag-and-drop arrangement design module 210 is used to provide a visual workflow arrangement environment to obtain a target workflow in response to the user's drag-and-drop arrangement operation; the target workflow includes multiple tasks and dependencies between the multiple tasks.

[0040] Exemplarily, the drag-and-drop layout design module 210 provides a workflow editing function based on a graphical user interface, which can be further divided into:

[0041] (1) Node connection authority verification module: Controls the establishment of connections between nodes corresponding to tasks through the node connection mechanism, ensuring the logical compliance and executability of the workflow and avoiding circular dependencies or illegal task chains.

[0042] Specifically, for the node connection authority verification module, when the connection establishment between the nodes corresponding to the tasks is controlled by the node connection mechanism, in response to the user trying to connect the nodes corresponding to the two tasks by dragging in the graphical editing engine, the module uses the node connection mechanism to check whether the attempted connection meets the node type compatibility.

[0043] This module also defines the node types supported by the system. Node types include data processing nodes, analysis nodes, and decision nodes. Each node type can have different functions and outputs. Data processing nodes are responsible for performing tasks such as data cleansing, formatting, and conversion to prepare data for subsequent analysis. Analysis nodes execute decision logic based on analysis results or other conditions, such as selecting a different process path or triggering a specific action. Decision nodes enable workflows to flexibly adapt to various business needs while maintaining a logical and efficient process.

[0044] This module is also used to analyze compatibility requirements between different node types.

[0045] For example, when using the node connection mechanism to check whether the attempted connection meets the node type compatibility, the following factors may be considered:

[0046] Data dependencies, which ensure that nodes in a workflow receive the correct types of input data to perform their tasks. For example, a node might only accept data of a specific format or type (e.g., text, numbers, images, etc.);

[0047] Data flow requirements define the direction and order in which data flows from one node to another in the workflow. For example, some nodes may need to be executed after other nodes to ensure a logical order of data processing.

[0048] Resource dependencies ensure that the resources required for node execution are met. These resources include, but are not limited to, computing resources, storage resources, and network resources. Resource dependencies ensure that nodes can execute in the appropriate environment, preventing execution failures due to insufficient resources. Some nodes may require specific computing resources, and only nodes that can provide these resources can connect to them.

[0049] Thus, data flow requirements, data dependencies, and resource dependencies jointly ensure the integrity and correctness of data in the workflow, preventing data from being mishandled or lost, and ensuring the efficient and correct execution of the workflow. Through these mechanisms, workflow management systems can automatically execute complex data processing tasks while maintaining high efficiency and accuracy.

[0050] After the module 210 is executed, the node connection mechanism is checked, and if the check result is yes, the connection attempt is allowed. Otherwise, the connection attempt is rejected and a corresponding error message is provided.

[0051] (2) Graphical editing engine: Users can drag and drop task nodes to place, connect, move, and delete them, build a job graph in real time, and achieve clear expression of task flows.

[0052] (3) Custom attribute configuration module: allows users to configure task node parameters, such as input variables, command templates, dependent file path rules, computing resource requirements, etc., to prepare for subsequent scheduling.

[0053] (4) Process template support module: used to encapsulate the completed process into a template for easy reuse, improving construction efficiency and standardization.

[0054] The virtualized task encapsulation and adaptation module 220 is used to encapsulate multiple tasks in the target workflow respectively using a unified encapsulation mechanism to obtain multiple target containers.

[0055] For example, in order to adapt to the environmental dependencies and operating modes of different scientific research tasks, a unified task containerization encapsulation mechanism is provided in the virtualization task encapsulation and adaptation module 220, specifically including the following modules:

[0056] (1) Multi-task support module: supports the encapsulation and deployment of multiple tasks such as AI training, data cleaning, scientific computing, and model reasoning;

[0057] (2) Image loading strategy module: allows users to select or upload a specific model environment image and automatically mount the running environment;

[0058] (3) Reusable execution template module: Task execution parameters can be exported as reusable templates for easy sharing within the organization;

[0059] (4) Environmental isolation and resource limitation control module: Each task is encapsulated as an independent container and runs in isolation to ensure that resources do not conflict.

[0060] The workflow automation execution module 230 is used to manage and allocate hardware resources in the supercomputing cluster to schedule and execute tasks in multiple target containers according to dependency relationships.

[0061] Exemplarily, the workflow automation execution module 230 is the core of system scheduling and execution, responsible for task execution sequence derivation, resource allocation scheduling, and monitoring feedback. The module 230 includes:

[0062] (1) Scheduler policy adaptation module: connects to HPC scheduling systems such as Slurm, and has built-in policy logic such as job submission, resource allocation, and priority management;

[0063] (2) Distributed execution and DAG scheduling engine: Builds DAG based on node dependencies, automatically derives parallel and sequential execution logic, and implements concurrent scheduling of nodes across partitions;

[0064] Specifically, by constructing a directed acyclic graph, the engine can clearly represent the dependencies between tasks. Each node represents a task, while the edges represent the dependencies between tasks. The engine also automatically analyzes the dependencies between tasks and derives a reasonable execution order. This includes determining which tasks can be executed in parallel and which tasks need to be executed sequentially. Based on the derived execution logic, the scheduling engine can determine how tasks are executed, optimize resource utilization, and improve execution efficiency. In distributed systems, the engine can implement concurrent scheduling across different computing partitions or nodes, allowing tasks to be executed simultaneously on multiple computing resources, thereby speeding up overall processing speed.

[0065] (3) Task scheduling module: supports periodic task scheduling configuration, suitable for scenarios such as daily training and periodic simulation;

[0066] (4) Dependency node reasoning module: automatically builds a dependency tree to ensure the correct and orderly execution of tasks;

[0067] (5) Task failure retry module: configure the number of failed retries and the interval time to enhance system stability;

[0068] (6) Conditional branching and decision logic module: The node execution results can trigger process branches, supporting dynamic path selection based on output results in complex processes.

[0069] In addition, the workflow engine system 200 also includes the following modules:

[0070] Log and monitoring module 240: used to equip the system with complete log and visual monitoring capabilities.

[0071] In this module 240, the following functions are mainly implemented:

[0072] Log collection and query: The running logs of each node task are centrally managed and can be filtered by task number, user or time;

[0073] Graphical presentation of running status: During the task running process, users can view the status, running time, exception details, etc. of each node in the graph;

[0074] Abnormal alarm mechanism: Task failure, resource abnormality, etc. support triggering real-time alarm notifications.

[0075] Template and authority management module 250 (including Figure 2 The "Template and Reuse" and "Permission and Role Management" modules in the project are used to focus on the sedimentation and reuse of task processes and ensure the standardization of organizational-level permission management.

[0076] Specifically, the templating and permission management module 250 is used to save or reuse target workflow templates in response to user templating operations and to establish a permission authentication mechanism for permission management. The permission authentication mechanism includes the following three aspects: control over the connection relationship between nodes corresponding to tasks, control over task execution, and control over access to task-related data.

[0077] Control over the connections between nodes corresponding to tasks ensures that only authorized users can establish or modify connections between nodes. Control over task execution ensures that only users with the appropriate permissions can initiate or manage task execution. Control over access to task-related data ensures that users can only access data within their permissions.

[0078] In this module 250, the following functions are mainly implemented:

[0079] Template management: Users can save built processes as templates, and support version control, permission allocation, and editing and updating;

[0080] Create a permission authentication mechanism: configure roles based on the organizational structure to control permissions such as node connection, task execution, and data access;

[0081] Execution permission audit: Administrators can view execution history and permission change records to improve system compliance and controllability.

[0082] Exemplarily, in addition to creating a target workflow, the user's drag-and-drop arrangement operation is also used to:

[0083] Design templates to quickly create similar workflow templates;

[0084] Version management to track and roll back to older versions;

[0085] Reuse templates to quickly create new workflows.

[0086] Together, these features enhance the flexibility and efficiency of workflow management systems, enabling users to more flexibly respond to diverse work requirements while maintaining workflow consistency and manageability. Through drag-and-drop orchestration, users can intuitively and efficiently design, manage, and reuse workflows, thereby improving overall work efficiency.

[0087] Therefore, the embodiment of the present application effectively improves the efficiency of scientific research task process construction, scheduling flexibility and execution stability in supercomputing clusters through highly modular and visual design, has good scalability and cross-platform adaptability, and is suitable for various scientific research scenarios such as artificial intelligence, big data, material simulation, and meteorological analysis.

[0088] The following will be Figures 3-5 As an example, the specific application of the drag-and-drop layout design module 210 is described in detail.

[0089] Figure 3 The following is a schematic diagram of the workflow provided by the embodiment of the present application. Figure 3 As shown, the software interface provided by the drag-and-drop orchestration design module 210 is shown, which is used to design and execute various computing tasks. On the left side of the interface is a graphical editing engine that displays different functional modules available for user operation, such as "Model Selection-Region", "Data Assimilation-Region", "Execution Script", "Model Forward", etc. These modules represent different tasks, and users can build workflows by clicking or dragging them to the visual design area on the right. A simple workflow example is shown on the right side of the interface, which contains nodes corresponding to several tasks, which are connected by lines to form a target workflow represented by a directed acyclic graph. The directed acyclic graph includes task node types, their execution order, and dependencies.

[0090] When a user attempts to connect the nodes corresponding to two tasks by dragging them in the graphical editing engine, the node connection permission verification module uses the node connection mechanism to check the compatibility of the node types being attempted. If the check results in compatibility, the attempted connection is allowed. Otherwise, the connection is rejected and an appropriate error message is displayed.

[0091] like Figure 3 The target workflow shown includes five nodes: "Model Selection - Region," "Data Assimilation - Region," "Model Forward Modeling," "Execute Command," and "Import Results." The "Model Selection - Region" node selects different models or model parameters based on specific conditions or parameters and is a decision node. The "Data Assimilation - Region" node combines observational data with model predictions to improve the model's state estimate and is a data processing node. The "Model Forward Modeling" node predicts output results based on input data and is an analysis node. The "Execute Command" node involves running scripts or programs to process data and is a data processing node. The "Import Results" node involves importing external data or computational results into the system and is a data processing node. The execution order and dependencies between these nodes meet the node type compatibility requirements.

[0092] Figure 4 The following is a schematic diagram showing the task node structure provided by the embodiment of the present application. Figure 4 As shown, in the software interface provided by the drag-and-drop layout design module 210, an "option setting" window pops up, which displays the function of the custom attribute configuration module, where the user can construct a specific task node by configuring parameters.

[0093] The following options can be configured in this window:

[0094] Node Name: Currently set to "Model Selection - Region", which indicates the name of the specific task node that the user is configuring.

[0095] Select node icon: Users can select an icon to represent this node. The currently selected icon is "Icon 6".

[0096] Retry on failure: Users can choose whether to retry when a task fails. The options are "Yes" or "No", and the current selection is "No". Users can configure the parameters of this node to cooperate with the task execution of the task failure retry module.

[0097] Is it a result node: Users can specify whether this node is the final result node. The options are also "Yes" or "No". The current selection is "No".

[0098] Select Connectable Nodes: Users can select which nodes can connect to the current node. Options include "Select All" and a range of specific node types, such as "Model Selection - Region," "Data Assimilation - Region," and "Execute Script." Currently, all options are selected. Users can set the node connection mechanism by configuring two node parameters, which determine the data flow requirements for each node type.

[0099] Finally, there are two buttons at the bottom of the window: "Cancel" and "Confirm". Users can save the settings by clicking "Confirm" or abandon the changes by clicking "Cancel".

[0100] Figure 5 Schematic diagram of node attribute setting provided by the embodiment of the present application is shown. Figure 5 As shown, a pop-up window provided by the drag-and-drop arrangement design module 210 is displayed for configuring detailed information of the task node.

[0101] The following is a detailed description of each configuration item in the pop-up window:

[0102] Select File Type: This is a drop-down menu where you select the file type. Options include parameters (e.g., Java, Python, Spark, custom script, etc.). You can configure the node connection mechanism by configuring the node parameters, which determine the data dependencies of the node type.

[0103] Driver Data Upload: This option is used to upload files for driver data processing. Observation Data Upload: This option is used to upload observation data files. Carbon Pool Initial Value Upload: This option is used to upload files for initial carbon pool values. Users can set the node connection mechanism by configuring these three node parameters. These parameters will determine the resource dependencies of the node type.

[0104] Regional Resolution (km): Enter a value to specify the regional resolution of the data processing, in kilometers. Function: Select the function type from this drop-down menu. Model: Select the model type from this drop-down menu. At the bottom of the window, there are two buttons: Return: Click this button to cancel the current operation and close the window. Submit: Click this button to save the configuration and submit it.

[0105] Figure 6 A schematic diagram of a task scheduling method for a supercomputing cluster provided by an embodiment of the present application is shown. Figure 2 The system shown in the figure is executed. The system includes a drag-and-drop orchestration design module, a virtualization task encapsulation and adaptation module, and a process template support module. The method mainly includes the following steps:

[0106] Step S601 : generating a target workflow in response to a user's drag-and-drop arrangement operation in a drag-and-drop arrangement design module; the target workflow includes multiple tasks and dependencies between the multiple tasks.

[0107] Exemplarily, the orchestration operation also includes: designing templates, which are used to quickly create similar workflows; version management, which is used to track and roll back to old versions; and template reuse, which is used to quickly create new workflows.

[0108] Step S602 : Encapsulate multiple tasks separately using a virtualized task encapsulation and adaptation module to obtain multiple target containers; a unified encapsulation strategy is used for the encapsulation.

[0109] Step S603 : Using the process template support module, manage and allocate hardware resources in the supercomputing cluster to schedule and execute tasks in multiple target containers according to dependency relationships.

[0110] Exemplarily, scheduling includes scheduling based on a scheduling policy set by a user.

[0111] Execution includes making conditional branch decisions based on the results of task execution to determine which task to execute next; automatically executing at the scheduled time; and checking whether the task is successfully executed; if the task fails, it will be retried according to the configured retry mechanism. If the task is successfully executed, the task script will be visualized and stored in the database for subsequent reuse.

[0112] Furthermore, the method further includes: collecting and monitoring task execution logs so that users can view task status.

[0113] Figure 7 FIG. 1 shows a flow chart of a workflow management system for a supercomputing cluster provided by an embodiment of the present application. Figure 7 The figure shows a flowchart of a workflow management system, describing the entire process from user registration and login to task execution and monitoring. The following are the steps in the flowchart:

[0114] Step S701, register / log in.

[0115] For example, the user first needs to register or log in to the system, which is the first step in using the system.

[0116] Step S702: process design layer.

[0117] Exemplarily, the user enters the process design layer and starts designing the workflow.

[0118] Step S703: Customize the node.

[0119] For example, users can create custom nodes, which can be specific tasks or operations.

[0120] Step S704: Design DAG.

[0121] For example, a user designs a logical structure of a workflow, namely a DAG, and defines dependencies between tasks.

[0122] Step S705: template design.

[0123] For example, users can design templates that can be used to quickly create similar workflows.

[0124] Step S706: version management.

[0125] Exemplarily, the system supports version management of workflows, and users can track and roll back to older versions.

[0126] Step S707: template reuse.

[0127] For example, users can reuse existing templates to quickly create new workflows.

[0128] Step S708: task packaging and control.

[0129] For example, the user encapsulates the designed workflow into tasks and controls them.

[0130] Step S709: supporting multiple types of tasks.

[0131] Exemplarily, the system supports multiple types of tasks, such as data processing, model training, etc.

[0132] Step S710: policy configuration.

[0133] For example, users can configure task execution policies, such as resource allocation, priority, etc.

[0134] Step S711, automated scheduling layer.

[0135] Exemplarily, the system enters the automated scheduling layer and prepares to execute tasks.

[0136] Step S712: DAG scheduling execution.

[0137] Exemplarily, the system automatically schedules and executes tasks according to the designed DAG.

[0138] Step S713: conditional branch decision.

[0139] For example, the system makes a conditional branch decision based on the result of task execution to determine which task to execute next.

[0140] Step S714, scheduled task.

[0141] For example, the system supports scheduled tasks, which can be automatically executed at a predetermined time.

[0142] Step S715: log collection / monitoring.

[0143] Exemplarily, the system collects and monitors task execution logs so that users can view task status.

[0144] Step S716: Determine whether the task is successful.

[0145] Illustratively, the system checks whether the task was successfully executed.

[0146] Step S717: Failure retry mechanism.

[0147] For example, if a task fails, the system will retry according to the configured retry mechanism.

[0148] Step S718: The visual script is stored in the database.

[0149] For example, successfully executed task scripts can be visualized and stored in a library for subsequent reuse.

[0150] After the task is completed, the process ends. Figure 7 It demonstrates the entire process of a complete workflow management system from design to execution to monitoring, covering multiple links such as user interaction, task scheduling, execution, monitoring and result processing.

[0151] It is understandable that the size of the sequence number of each step in the above embodiment does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. In addition, in some possible implementations, the steps in the above embodiment can be selectively executed according to actual conditions, and can be partially executed or fully executed, which is not limited here. In addition, all or part of any features in the above embodiment can be freely and arbitrarily combined without contradiction. The combined technical solution is also within the scope of this application.

[0152] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state drive (SSD)).

[0153] It is understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not intended to limit the scope of the embodiments of the present application. It should be understood that in the embodiments of the present application, the order of the sequence numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0154] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this application should be included in the scope of protection of this application.

Claims

1. A workflow engine system for supercomputing clusters, characterized in that: The system comprises: A drag-and-drop orchestration design module is used to provide a visual workflow orchestration environment to obtain a target workflow in response to a user's drag-and-drop orchestration operation; the target workflow includes multiple tasks and dependencies between the multiple tasks; The virtualization task encapsulation and adaptation module is used to encapsulate multiple tasks in the target workflow using a unified encapsulation mechanism to obtain multiple target containers; The workflow automation execution module is used to manage and allocate hardware resources in the supercomputing cluster to schedule and execute tasks in the multiple target containers according to the dependency relationship.

2. The system according to claim 1, wherein: The system further comprises: The log monitoring module provides log and visualization monitoring capabilities, including log collection and query, graphical presentation of operating status, and abnormal alarm mechanism; The modularization and permission management module is used to save or reuse the template of the target workflow in response to the user's template operation, and to formulate a permission authentication mechanism for permission management; the permission authentication mechanism includes control over the connection relationship between the nodes corresponding to the task, control over the execution of the task, and control over access to data related to the task.

3. The system according to claim 1, wherein: The drag-and-drop layout design module includes: A node connection authority verification module is used to control the establishment of connections between nodes corresponding to the task through a node connection mechanism; A graphical editing engine, configured to perform at least one of the following operations on the node in response to a user's dragging action: placing, connecting, moving, and deleting, so as to construct the target workflow; A custom attribute configuration module, configured to configure the parameters of the node according to user setting parameters, wherein the parameters include at least one of the following: input variables, command templates, dependent file path rules, and computing resource requirements; The process template support module is used to encapsulate the constructed target workflow into a template.

4. The system according to claim 3, characterized in that The controlling the establishment of connections between the nodes corresponding to the tasks through a node connection mechanism includes: In response to a user attempting to connect the nodes corresponding to two tasks respectively by dragging in the graphical editing engine, the node connection mechanism is used to check whether the attempted connection complies with node type compatibility.

5. The system according to claim 4, characterized in that The controlling the establishment of connections between the nodes corresponding to the tasks through the node connection mechanism also includes: If the result of the check is yes, allowing the attempted connection to be performed; Otherwise, the connection attempt is rejected and a corresponding error message is provided.

6. The system according to claim 4, characterized in that The node types include data processing nodes, analysis nodes, and decision nodes; the node type compatibility includes the node's data dependency, data flow requirements, and resource dependency conditions.

7. The system according to claim 1, wherein: The virtualization task encapsulation and adaptation module includes: A multi-task type support module, used to support the encapsulation and deployment of multiple types of tasks; the multiple types of tasks include at least one of the following: AI training, data cleaning, scientific computing, and model reasoning; An image loading strategy module is used to start a specific model environment image based on user selection or upload, and automatically configure and mount the operating environment required by the model environment image; A reusable execution template module, used to export the execution parameters of the task into a reusable template; The environment isolation and resource limitation control module is used to encapsulate the task into an independent container.

8. The system according to claim 1, wherein: The workflow automation execution module includes: A scheduler policy adaptation module, configured to cooperate with the scheduling system to determine a scheduling policy, wherein the scheduling policy includes at least one of the following: job submission, resource allocation, and priority management; A distributed execution and directed acyclic graph scheduling engine, configured to construct a directed acyclic graph based on the dependency relationships of the nodes corresponding to the tasks, and to implement distributed scheduling of the tasks according to the directed acyclic graph; A task timing module is used to configure a periodic plan for the task; the periodic task plan includes daily training and regular simulation; A dependency node reasoning module, configured to construct a dependency tree to sequentially execute the plurality of tasks; A task failure retry module is used to configure the number of failed retries and the interval time of the task; The conditional branch and decision logic module is used to perform dynamic path selection based on the execution results of the nodes corresponding to the tasks.

9. The system according to claim 1, wherein: The drag-and-drop arrangement operation is also used to: Design templates to design new workflow templates; Version management to track and roll back to older versions of workflows; Reuse templates to quickly create new workflows.

10. A task scheduling method for a supercomputing cluster, characterized in that: The method is executed by a system including a drag-and-drop orchestration design module, a virtualization task encapsulation and adaptation module, and a process template support module. The method includes: In response to the user's drag-and-drop arrangement operation in the drag-and-drop arrangement design module, a target workflow is generated; the target workflow includes a plurality of tasks and dependency relationships between the plurality of tasks; The plurality of tasks are respectively encapsulated by using the virtualization task encapsulation and adaptation module to obtain a plurality of target containers; the encapsulation adopts a unified encapsulation strategy; The process template support module is utilized to manage and allocate hardware resources in a supercomputing cluster so as to schedule and execute tasks in the plurality of target containers according to the dependency relationships.

Citation Information

Cited By

  • Tasking arrangement method for realizing system deployment through declarative template

    CN120929223A

  • Distributed workflow task scheduling method and system

    CN121597375A

  • A distributed workflow task scheduling method and system

    CN121597375B