Cross-domain workflow scheduling system for dynamic resource arrangement

By designing a cross-domain workflow scheduling system for dynamic resource orchestration, the problem of low resource management and scheduling efficiency in cross-domain scientific computing is solved, and unified resource management and stable task execution are achieved.

CN120104345AActive Publication Date: 2025-06-06国家超级计算天津中心

Patent Information

Application Number
CN202510562912.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-06
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively manage and schedule heterogeneous computing resources in cross-domain scientific computing, resulting in low resource utilization efficiency, high overhead of data transmission and task scheduling, and a single point of failure risk.

Method used

A cross-domain workflow scheduling system for dynamic resource orchestration is designed, including resource abstraction subsystem, task granular molecular system, workflow scheduling subsystem and data preheating management subsystem. By uniformly managing resources of different platform, the overhead of data transmission and task scheduling is reduced and the stability of cross-domain task completion is improved.

Benefits of technology

It realizes unified management of resources on different platforms, reduces the overhead of data transmission and task scheduling, and improves the stability of cross-domain task completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104345A_ABST
    Figure CN120104345A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-domain workflow scheduling system for dynamic resource arrangement, and the system comprises a resource abstraction subsystem which is used for collecting all resource platform information and network topology information, and constructing a resource directory and all adaptation layers; the task granularity division subsystem is used for receiving a target task, determining task characteristics of the target task, the cost of each task and an initial dependency graph, determining a target granularity according to the cost of each task, and optimizing the initial dependency graph to obtain a target dependency graph; the workflow arrangement subsystem is used for determining an initial execution plan according to the target dependency graph and the resource directory, determining a corresponding adaptation layer after each target execution platform receives target data, and submitting a corresponding subtask to each target execution platform; and the data preheating management subsystem is used for determining the transmission priority of the cross-platform sub-tasks and sending the target data to each target execution platform in combination with the network topology information, so that the effect of uniformly managing different platform resources is realized, and the stability of the cross-domain tasks is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automated computing, and in particular to a cross-domain workflow scheduling system for dynamic resource orchestration. Background Art

[0002] As the scale of scientific computing continues to expand, a single computing cluster or supercomputing center often cannot meet the resource requirements of large-scale scientific computing workflows. In fields such as material science, weather simulation, and fusion energy research, it is often necessary to use heterogeneous computing resources distributed in different geographical locations to complete complex computing tasks.

[0003] At present, cross-domain scientific computing mainly adopts manual coordination, point-to-point integration or centralized scheduling. Among them, the manual coordination method is inefficient, prone to errors, and difficult to deal with complex workflows; the point-to-point integration method lacks versatility and requires redevelopment of interfaces when adding new computing platforms; the centralized scheduling method is inefficient when network conditions are poor or when firewalls are isolated between computing platforms, and there is a risk of single point failure.

[0004] In view of this, the present invention is proposed. Summary of the invention

[0005] In order to solve the above technical problems, the present invention provides a cross-domain workflow scheduling system with dynamic resource orchestration, which realizes unified management of resources on different platforms, reduces the overhead of data transmission and task scheduling, and improves the stability of cross-domain task completion.

[0006] An embodiment of the present invention provides a cross-domain workflow scheduling system for dynamic resource orchestration, the system comprising:

[0007] The resource abstraction subsystem is used to collect information about each resource platform and network topology, build a resource directory based on the information about each resource platform, and build an adaptation layer between each task submission system type and each resource platform type;

[0008] The task granularity division subsystem is used to receive the target task, determine the task characteristics corresponding to the target task, the cost of each task and the initial dependency graph, determine the target granularity according to the cost of each task, and optimize the initial dependency graph according to the target granularity to obtain the target dependency graph;

[0009] The workflow orchestration subsystem is used to obtain task characteristics, resource catalogs, and target dependency graphs, determine the initial execution plan based on the target dependency graph and resource catalogs, and determine the adaptation layer corresponding to each target execution platform after each target execution platform in the initial execution plan receives the target data, and submit the corresponding subtasks to each target execution platform based on each adaptation layer, so that each target execution platform can jointly execute the target task;

[0010] The data preheating management subsystem is used to obtain network topology information, target dependency graph and initial execution plan, determine the transmission priority of cross-platform subtasks in the initial execution plan, and send corresponding target data to each target execution platform according to each transmission priority and network topology information.

[0011] The embodiments of the present invention have the following technical effects:

[0012] Through the resource abstraction subsystem, the information of each resource platform and the network topology information are collected, the resource directory is constructed, and the adaptation layers are constructed. Through the task granularity division subsystem, the target granularity is calculated, and the initial dependency graph is converted into a target dependency graph. Through the workflow orchestration subsystem, the initial execution plan is determined in combination with the task characteristics, the resource directory and the target dependency graph. Through the data preheating management subsystem, the target data is pre-transmitted across platforms. After the target execution platforms in the initial execution plan receive the target data through the workflow orchestration subsystem, the adaptation layer corresponding to each target execution platform is determined, and the corresponding subtasks are submitted to each target execution platform to schedule the execution of the target task, thereby realizing unified management of resources on different platforms, reducing the overhead of data transmission and task scheduling, and improving the stability of cross-domain task completion. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0014] Figure 1 It is a structural diagram of a cross-domain workflow scheduling system for dynamic resource orchestration provided by an embodiment of the present invention;

[0015] Figure 2 It is a structural diagram of another cross-domain workflow scheduling system for dynamic resource orchestration provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present invention.

[0017] The cross-domain workflow scheduling system with dynamic resource orchestration provided by the embodiment of the present invention is mainly applicable to the situation of unifying the resources of various computing platforms and performing reasonable resource orchestration and workflow scheduling for target tasks.

[0018] Figure 1 Schematic diagram of a cross-domain workflow scheduling system for dynamic resource orchestration provided by an embodiment of the present invention. Figure 1 The cross-domain workflow scheduling system for dynamic resource orchestration specifically includes: a resource abstraction subsystem 110, a task granularity division subsystem 120, a workflow orchestration subsystem 130 and a data preheating management subsystem 140.

[0019] The resource abstraction subsystem 110 is used to collect information about each resource platform and network topology information, build a resource directory based on the information about each resource platform, and build an adaptation layer between each task submission system type and each resource platform type.

[0020] Among them, resource platform information is the local resource information of the platform where heterogeneous computing resources distributed in different supercomputing centers and cloud platforms are located. It can be the local resource information of resource platforms in different clusters and domains, such as CPU (Central Processing Unit), GPU (Graphics Processing Unit), memory, storage, network bandwidth, etc. Network topology information is the network topology structure and network status information formed between various resource platforms. The resource directory is a directory used to integrate and count the information of various resource platforms, which can be understood as a global resource view. The task submission system type is the system type used when submitting the target task to be executed. The resource platform type is the type of different resource platforms. The adaptation layer can include task submission, monitoring and management interfaces.

[0021] Specifically, the resource abstraction subsystem 110 can collect information about each resource platform and network topology at regular intervals, and integrate the information about each resource platform into a resource directory, so as to facilitate subsequent resource allocation in combination with target tasks. In addition, an adaptation layer between different task submission system types and different resource platform types can be pre-built to facilitate direct calls during task submission, monitoring, and management.

[0022] The task granularity division subsystem 120 is used to receive the target task, determine the task characteristics corresponding to the target task, the cost of each task and the initial dependency graph, determine the target granularity according to the cost of each task, and optimize the initial dependency graph according to the target granularity to obtain the target dependency graph.

[0023] Among them, the target task is a task that requires resource allocation and execution. It is understandable that the target task is a workflow description-related file uploaded by the user, which may include a workflow description file, a task code file, an environment configuration file, a data description file, a resource preference configuration file, etc. Task characteristics include computationally intensive, communication intensive, and I / O (Input / Output) intensive. Task cost is the processing cost of the target task in terms of computation, communication, etc. The initial dependency graph is a task dependency graph between subtasks determined by analyzing the submitted workflow description file. The target granularity is used to describe the size of the target task, that is, the workload contained in each subtask. The target dependency graph is a task dependency graph between subtasks obtained by merging or splitting subtasks based on the target granularity.

[0024] Specifically, the task granularity division subsystem 120 receives the target task submitted by the user, analyzes the task code file in the target task, determines the task characteristics and the cost of various types of tasks, analyzes the workflow description file in the target task, and constructs the initial dependency graph. By comprehensively analyzing the cost of each task, the target granularity of the target task can be calculated, and it is determined whether the target granularity is within a reasonable range. If not, it is necessary to optimize the initial dependency graph, such as merging subtasks or splitting subtasks, and the optimized initial dependency graph is used as the target dependency graph.

[0025] It is understandable that the workflow description file is the core file for the system to understand the entire computing process. It is usually in JSON (JavaScript Object Notation) or YAML (YAML Ain't a Markup Language, data serialization) format, and usually contains task identification, task name, task description, subtask list (the subtask list contains subtask identification, subtask name, subtask description, subtask code file path, input data, output data, required computing resources such as CPU and memory, estimated execution time, task priority, other subtasks that the subtask depends on), data entity relationship, operation restriction conditions, etc. The task code file is the execution code corresponding to each task that the user needs to provide. It can be in various programming languages. It needs to contain clear task description comments and metadata, good error handling and logging, support configuration of input / output paths through environment variables or command line parameters, and return a clear exit code to indicate task success or failure. The environment configuration file is the environment dependency configuration that the user needs to provide for the task. The data description file usually contains data identification, data name, data type, data size, and data file metadata information. Resource preference profiles are user-provided resource preferences to guide the system's resource allocation, which can include preferred resources, avoided resources, and performance preferences.

[0026] The workflow orchestration subsystem 130 is used to obtain task characteristics, resource directories, and target dependency graphs, determine an initial execution plan based on the target dependency graph and resource directory, and after each target execution platform in the initial execution plan receives the target data, determine the adaptation layer corresponding to each target execution platform, and submit corresponding subtasks to each target execution platform based on each adaptation layer, so that each target execution platform can jointly execute the target task.

[0027] The initial execution plan is an execution plan that determines the workflow according to the target dependency graph and allocates resources based on the task characteristics and resource directory. The target execution platform is the resource platform that needs to execute subtasks after resource allocation. The target data is the data required for task execution that is transmitted across platforms.

[0028] Specifically, the task characteristics and target dependency graph determined by the task granularity division subsystem 120 can be obtained through the workflow orchestration subsystem 130, and the new workflow can be constructed and resources allocated in combination with the resource directory constructed by the resource abstraction subsystem 110 to obtain an initial execution plan. In addition, after each target execution platform in the initial execution plan receives the target data, that is, when the cross-platform data preparation is completed, according to the task submission system type of the target task and each target execution platform, the adaptation layer pre-constructed in the resource abstraction subsystem 110 is searched to obtain the adaptation layer corresponding to each target execution platform, and then, for the target execution platform, the corresponding subtask is submitted to the target execution platform based on the corresponding adaptation layer, so that each target execution platform can jointly execute the target task.

[0029] The data preheating management subsystem 140 is used to obtain network topology information, target dependency graph and initial execution plan, determine the transmission priority of cross-platform subtasks in the initial execution plan, and send corresponding target data to each target execution platform according to each transmission priority and network topology information.

[0030] Among them, the cross-platform subtask is a subtask that needs to switch resource platforms for execution, that is, the previous subtask is executed on one resource platform, and the current subtask is executed on another resource platform. The current subtask is a cross-platform subtask. The transmission priority is the priority of the target data required by the cross-platform subtask.

[0031] Specifically, the data preheating management subsystem 140 can obtain network topology information from the resource abstraction subsystem 110, obtain the target dependency graph from the task granularity division subsystem 120, and obtain the initial execution plan from the workflow orchestration subsystem 130. By analyzing the target dependency graph and the initial execution plan, each cross-platform subtask in the initial execution plan can be obtained, and the transmission priority can be calculated and determined for these cross-platform subtasks. Then, combined with each transmission priority and network topology information, it is determined whether parallel transmission can be performed, and the corresponding target data is sent to each target execution platform in a parallel transmission mode or a non-parallel transmission mode, so that the required target data can be transmitted to the target execution platform in advance before each subtask is executed, thereby reducing the waiting time for data transmission.

[0032] It is understandable that the data preheating management subsystem 140 runs after the initial execution plan is generated and determined in the workflow orchestration subsystem 130, and before the adaptation layer corresponding to each target execution platform is determined and the corresponding subtask is submitted to each target execution platform based on each adaptation layer. The data preheating management subsystem 140 analyzes the task dependency, the location of the dependent target data, and each target execution platform in the initial execution plan, and determines which data needs to be migrated or copied (target data) to perform the preheating operation.

[0033] Figure 2 It is a structural diagram of another cross-domain workflow scheduling system for dynamic resource orchestration provided by an embodiment of the present invention.

[0034] like Figure 2 As shown, the resource abstraction subsystem 110 includes: a resource proxy module 111 , a resource directory module 112 , a resource adaptation module 113 and a network topology module 114 .

[0035] The resource agent module 111 is used to collect information of each resource platform based on a first preset period.

[0036] Among them, the first preset period is a preset period for collecting resource platform information, which can be set according to needs.

[0037] Specifically, the resource agent module 111 deploys a lightweight resource agent on each resource platform, which is responsible for collecting resource platform information based on a first preset period and regularly reporting it to the resource directory module 112 .

[0038] Exemplarily, a multi-threading technology can be used to implement the resource platform information collection function, and a timed task mechanism can be used to call a system command to collect resource platform information every first preset period (such as 5 minutes). The collected resource platform information is formatted and encapsulated into a data packet such as JSON format, which can be reported to the resource directory module 112 via HTTP (Hypertext Transfer Protocol) / HTTPS (Hypertext Transfer Protocol Secure) protocol. To ensure data transmission security, SSL / TLS (Secure Sockets Layer / Transport Layer Security) encrypted communication can be used.

[0039] The resource directory module 112 is used to receive information of each resource platform and construct a resource directory according to the information of each resource platform.

[0040] The resource directory module 112 includes a resource query interface and a resource allocation interface.

[0041] Specifically, the global resource view is maintained through the resource directory module 112, including information of each resource platform, that is, resource type, specification, availability status and other information, and can provide a resource query interface and a resource allocation interface.

[0042] Exemplarily, a relational database (such as MySQL) is used to store the information of each resource platform and build a global resource view. The database table design includes fields such as platform identification, resource type, resource specification, available status, and update time. The resource query interface can support queries based on platform identification, resource type, available status, and other conditions. The resource allocation interface uses a transaction mechanism to ensure the atomicity of resource allocation and avoid resource conflicts. At the same time, a caching mechanism can be introduced, such as Redis (RemoteDictionary Server, remote dictionary service), to cache commonly used resource platform information, reduce database query pressure, and improve query efficiency.

[0043] The resource adaptation module 113 is used to construct an adaptation group according to each task submission system type and each resource platform type, and pre-construct an adaptation layer corresponding to each adaptation group.

[0044] Among them, the adaptation group is a combination of a task submission system type and a resource platform type.

[0045] Specifically, each task submission system type is combined with each resource platform type to obtain multiple adaptation groups, and then, for each adaptation group, the task submission, monitoring and management interface corresponding to the adaptation group is pre-built as the corresponding adaptation layer. That is, for different types of task submission systems, such as Slurm (Simple Linux Utility for Resource Management), PBS (Portable Batch System), Kubernetes (open source container orchestration platform), etc., a unified task submission, monitoring and management interface is provided.

[0046] For example, develop corresponding adaptation layers for task submission systems on different platforms. Take Slurm as an example, use the command line tools provided by Slurm to submit, monitor, and manage tasks. Encapsulate the unified task submission, monitoring, and management interfaces into Python modules, and call the corresponding implementation modules of different resource platforms through the importlib (dynamic import module and check module) dynamic import mechanism. When submitting a task, automatically call the corresponding task submission interface according to the resource platform type queried in the resource directory.

[0047] The network topology module 114 is used to obtain the network topology structure and collect the network status information corresponding to each network node in the network topology structure, and determine the network topology information according to the network topology structure and the network status information corresponding to each network node.

[0048] Specifically, the network topology structure is obtained, each node in the network topology structure is used as a network node, the network status information of each network node, such as network bandwidth, availability, etc., is collected, and the network status information of each network node is written into each network node of the network topology structure to obtain the network topology information.

[0049] like Figure 2 As shown, the task granularity division subsystem 120 includes: a task characteristic analysis module 121, a dependency graph construction module 122, a granularity optimization module 123, a dependency graph update module 124 and a critical path identification module 125.

[0050] The task characteristic analysis module 121 is used to receive the target task, determine the number of each type of operation based on the task code file in the target task, determine the task characteristic based on the number of each type of operation, determine the task calculation cost in the task cost based on the task code file, determine the task communication cost in the task cost based on the data description file in the target task, and determine the task switching cost in the task cost based on the historical execution time data of similar tasks corresponding to the target task.

[0051] The number of type operations includes the number of computing operations, the number of transmission operations, and the number of I / O operations. The task cost includes the task computing cost, the task communication cost, and the task switching cost. The task computing cost is used to describe the computational complexity of the target task, the task communication cost is used to describe the estimated cost of data dependency in the target task, and the task switching cost is used to describe the time cost from the start of task execution to the start of scheduling. The historical execution time data is the time from the start of task execution to the start of scheduling.

[0052] Specifically, based on the target task uploaded by the user, the task code file in the target task is analyzed to determine the number of operations of each type in the target task, and the type with the largest number of operations is used as the task feature to determine the basic resource requirements of the target task. In addition, the task calculation cost in the task cost is obtained based on the analysis of the task code file, and the task communication cost in the task cost is estimated based on the data dependency in the data description file. Similar tasks of the target task are selected from the historical tasks, and the task switching overhead in the task cost is obtained based on the historical data statistics of similar tasks. Specifically, the difference between the task start execution time and the task start scheduling time of similar tasks is used as the historical execution time data, and the average value of each historical execution time data is used as the task switching overhead.

[0053] Exemplarily, a method combining static code analysis and dynamic runtime monitoring is used to analyze task characteristics. Static code analysis is to parse the syntax tree of the task code file and count the number of different types of operations to preliminarily determine whether the task characteristics of the target task are compute-intensive, I / O-intensive, or communication-intensive. Dynamic runtime monitoring is to use system performance monitoring tools to collect real-time resource usage during task runtime during the task trial run phase, and further accurately determine the basic resource requirements of the task.

[0054] The dependency graph construction module 122 is used to receive the target task, determine each initial node, the node attributes corresponding to each initial node, each initial directed edge and the edge attributes corresponding to each initial directed edge according to the initial workflow in the target task, and construct the initial dependency graph according to each initial node, the node attributes corresponding to each initial node, each initial directed edge and the edge attributes corresponding to each initial directed edge.

[0055] The initial workflow is the workflow description file in the target task. The initial nodes are the subtask nodes in the initial workflow. Node attributes include subtask ID, task priority, required computing resources, estimated execution time, etc. The initial directed edge is used to describe the edge structure of the dependency of each subtask node, and the edge attributes include data dependency information, etc.

[0056] Specifically, the target task is received, and the initial workflow in the target task is analyzed to obtain each initial node, the node attributes corresponding to each initial node, each initial directed edge, and the edge attributes corresponding to each initial directed edge. Each initial node, the node attributes corresponding to each initial node, each initial directed edge, and the edge attributes corresponding to each initial directed edge are used to construct a directed acyclic graph (DAG) of dependencies between subtasks to obtain an initial dependency graph, which serves as the basis for subsequent task division and workflow arrangement.

[0057] Exemplarily, based on the workflow description file, a graph data structure (such as an adjacency list) is used to construct a directed acyclic graph (DAG) of dependencies between tasks, that is, an initial dependency graph. During the construction process, the task dependency information in the workflow description file is parsed, an initial node is created for each subtask, and directed edges between the initial nodes are established based on the task dependency information, that is, initial directed edges. At the same time, corresponding node attributes and edge attributes are added to each initial node and initial directed edge, such as subtask identification, task priority, required computing resources, estimated execution time, data dependency information, etc., for subsequent analysis and processing.

[0058] The granularity optimization module 123 is used to receive the task calculation cost, the task communication cost and the task switching cost, take the sum of the task communication cost and the task switching cost as the process value, and take the quotient of the task calculation cost and the process value as the target granularity.

[0059] The process value is the sum of the task communication cost and the task switching cost.

[0060] Specifically, the task calculation cost, task communication cost and task switching cost determined by the task characteristic analysis module 121 are received, the task communication cost and the task switching cost are summed to obtain the process value, and the task calculation cost is divided by the process value to obtain the target granularity, and the real-time resource status is not considered here. The target granularity can be calculated according to the following formula:

[0061] G opt = C comp / (C comm + C overhead )

[0062] Among them, G opt is the target particle size, C comp Calculate the cost for the task, C comm is the task communication cost, C overhead is the task switching cost.

[0063] The dependency graph updating module 124 is used to receive the initial dependency graph and the target granularity, determine the adjustment method of the initial dependency graph according to the preset granularity range and the target granularity, and adjust the initial dependency graph based on the adjustment method to obtain the target dependency graph.

[0064] The preset granularity range is a pre-set granularity range used to determine whether to adjust the initial dependency graph. The adjustment method includes a task merging method and a task splitting method. The target dependency graph is the adjusted initial dependency graph or the initial dependency graph that does not need to be adjusted.

[0065] Specifically, the initial dependency graph constructed by the dependency graph construction module 122 and the target granularity determined by the granularity optimization module 123 are received, and it is determined whether the target granularity is within the preset granularity range. If not, when the target granularity is greater than the preset granularity range, the adjustment method of the initial dependency graph is determined to be the task merging method, and when the target granularity is less than the preset granularity range, the adjustment method of the initial dependency graph is determined to be the task splitting method. Then, the initial dependency graph can be adjusted and optimized based on the determined adjustment method to obtain the target dependency graph. If the target granularity is within the preset granularity range, the initial dependency graph can be directly used as the target dependency graph.

[0066] For example, for a computationally intensive target task, if the target granularity is large and the dependencies between subtasks allow, multiple adjacent small subtasks can be merged into a large subtask to reduce the task scheduling overhead. For an I / O intensive target task, if the target granularity is small, splitting a large subtask into multiple small subtasks can improve resource utilization. During the subtask merging and splitting process, the initial dependency graph is updated to obtain the target dependency graph to ensure that the task execution order is correct.

[0067] Based on the above example, the dependency graph updating module 124 is further used to:

[0068] Initialize the number of adjustments to determine the relationship between the target particle size and the preset particle size range;

[0069] In response to the target granularity being greater than the upper limit of the preset granularity range, determining that the adjustment mode of the initial dependency graph is the task merging mode; determining whether the number of adjustments is equal to the preset number; in response to the number of adjustments being less than the preset number, taking two adjacent initial nodes with the smallest data dependency as merging nodes according to the initial dependency graph, merging the merging nodes to obtain a new initial node, updating the initial dependency graph and the number of adjustments, determining the target granularity of the updated initial dependency graph, and returning to execute the step of determining the relationship between the target granularity and the preset granularity range; in response to the number of adjustments being equal to the preset number, taking the updated initial dependency graph as the target dependency graph;

[0070] In response to the target granularity being within a preset granularity range, determining the initial dependency graph as the target dependency graph;

[0071] In response to the target granularity being less than the lower limit of the preset granularity range, determining that the adjustment method of the initial dependency graph is the task splitting method; judging whether the number of adjustments is equal to the preset number; in response to the number of adjustments being less than the preset number, taking the initial node containing the parallel task in the initial dependency graph as the split node, splitting the split node into at least two new initial nodes, updating the initial dependency graph and the number of adjustments, determining the target granularity of the updated initial dependency graph, and returning to execute the step of judging the relationship between the target granularity and the preset granularity range; in response to the number of adjustments being equal to the preset number, taking the updated initial dependency graph as the target dependency graph.

[0072] The preset granularity range is an interval including an upper limit and a lower limit. The number of adjustments is the count of the current iteration merge or split. The preset number is the upper limit of the preset number of adjustments. The merge node is the two adjacent initial nodes with the minimum data dependency. The split node is the initial node with parallel tasks.

[0073] Specifically, the number of adjustments is initialized, that is, the number of adjustments is set to zero. Determine the relationship between the target granularity and the preset granularity range. If the target granularity is greater than the upper limit of the preset granularity range, it means that the target granularity is too large, and the adjustment method of the initial dependency graph can be determined as the task merging method, and then, determine whether the number of adjustments is equal to the preset number of times. If the number of adjustments is less than the preset number of times, the initial dependency graph is analyzed, and two adjacent initial nodes with the smallest data dependency are determined, and the two initial nodes are used as merged nodes. The two merged nodes are merged to obtain a new initial node, and the initial dependency graph is updated. The task cost corresponding to the updated initial dependency graph is analyzed and calculated, and then the target granularity is updated, and the number of adjustments is increased by one. At this time, the step of determining the relationship between the target granularity and the preset granularity range can be returned to enter the next iteration. If the number of adjustments is equal to the preset number of times, it means that the upper limit of the adjustment is reached, and the updated initial dependency graph is directly used as the target dependency graph. If the target granularity is within the preset granularity range, it means that the target granularity has met the requirements, and the current initial dependency graph is directly used as the target dependency graph. If the target granularity is less than the lower limit of the preset granularity range, it means that the target granularity is too small, and the adjustment method of the initial dependency graph is determined to be the task splitting method, and then, it is determined whether the number of adjustments is equal to the preset number of times. If the number of adjustments is less than the preset number of times, it is analyzed whether there is an initial node containing parallel tasks in the initial dependency graph, and the initial node is used as a split node. The split node is split into at least two new initial nodes according to the number of parallel tasks, and the initial dependency graph is updated. The task cost corresponding to the updated initial dependency graph is analyzed and calculated, and then the target granularity is updated, and the number of adjustments is increased by one. At this time, it can return to the step of determining the relationship between the target granularity and the preset granularity range and enter the next iteration. If the number of adjustments is equal to the preset number of times, it means that the upper limit of the adjustment is reached, and the updated initial dependency graph is directly used as the target dependency graph. If the target granularity is within the preset granularity range, it means that the target granularity has met the requirements, and the current initial dependency graph is directly used as the target dependency graph.

[0074] Exemplarily, for a target task whose target granularity is greater than the upper limit of the preset granularity range (such as 1.5, etc.), the adjustment method is considered to be task merging: identify adjacent subtasks with small data dependencies in the initial dependency graph, evaluate the target granularity after merging the subtask nodes, and if it is within the preset granularity range, end the adjustment; if it is not within the preset granularity range, continue to adjust, and the number of adjustments is not greater than the preset number. According to the adjustment, update the initial dependency graph, update the nodes of the new subtask and their dependencies to the initial dependency graph, and obtain the target dependency graph. For a target task whose target granularity is less than the lower limit of the preset granularity range (such as 0.5, etc.), consider the adjustment method of task splitting: analyze the internal structure of the task in combination with the initial dependency graph, identify the parallelizable part, generate a splitting strategy, including splitting nodes and data flow methods, evaluate the target granularity after splitting, and if it is within the preset granularity range, end the adjustment; if it is not within the preset granularity range, continue to adjust, and the number of adjustments is not greater than the preset number. According to the adjustment, update the initial dependency graph, replace the split nodes with the initial nodes of multiple subtasks, and update their dependencies.

[0075] The critical path identification module 125 is used to receive the target dependency graph, determine the earliest start time and the latest start time corresponding to each target node in the target dependency graph according to the target dependency graph, determine the time float corresponding to each target node according to the earliest start time and the latest start time corresponding to each target node, and determine the critical path and the path weight corresponding to each target node according to the time float corresponding to each target node, and add the critical path and the path weight corresponding to each target node to the target dependency graph.

[0076] Among them, the target node is each subtask node in the target dependency graph. The earliest start time is the latest completion time of all predecessor subtasks. The latest start time is the latest subtask start time that does not affect the workflow completion time. The time float is the difference between the latest start time and the earliest start time. The critical path is the path composed of target nodes with a time float of 0. The path weight is used to describe the importance of the path.

[0077] Specifically, the target dependency graph in the dependency graph update module 124 is received, the target dependency graph is analyzed, and the earliest start time and the latest start time corresponding to each target node in the target dependency graph are calculated. For each target node, the difference between the latest start time and the earliest start time corresponding to the target node is used as the time float corresponding to the target node. The path composed of each target node with a time float of 0 is used as the critical path, and each time float is analyzed to calculate and determine the path weight corresponding to each target node. In addition, the critical path and the path weight corresponding to each target node are added to the target dependency graph to provide information for subsequent workflow arrangement.

[0078] Based on the above example, the critical path identification module 125 is further used to:

[0079] For each target node, the difference between the latest start time and the earliest start time corresponding to the target node is used as the time float corresponding to the target node;

[0080] Construct the critical path based on each target node with a time float of 0;

[0081] According to the time floating amount corresponding to each target node, determine the maximum value of the time floating amount;

[0082] For each target node, the path weight corresponding to the target node is determined according to the preset maximum weight, the maximum value of the time float and the time float corresponding to the target node.

[0083] The maximum time float is the maximum value of the time floats corresponding to each target node. The preset maximum weight is the maximum value that the preset path weight can obtain, such as 10. For example, if the path weight ranges from 0 to 10, then the preset maximum weight is 10. The larger the path weight, the more important the subtask corresponding to the target node.

[0084] Specifically, for each target node, the difference between the latest start time and the earliest start time corresponding to the target node is used as the time float corresponding to the target node. Each target node with a time float of 0 is constructed into a critical path. The maximum value of the time float corresponding to each target node is used as the maximum value of the time float. For each target node, a corresponding relationship between the preset maximum weight and the maximum value of the time float is established, and then the time float corresponding to the target node is converted into the corresponding weight, which is the path weight corresponding to the target node.

[0085] Exemplarily, the following properties are calculated for each target node: EST (earliest start time): the latest completion time of all predecessor subtasks. EFT (earliest finish time): EST plus the task execution time. LST (latest start time): the start time of the latest subtask that does not affect the completion time of the workflow. LFT (latest finish time): LST plus the execution time of the subtask. Specifically, it can be calculated in the following way: First, calculate EST and EFT (forward traversal), for the entry target node (target node without predecessor subtask): EST=0; for other target nodes: EST=max{EFT(target node of all predecessor subtasks)}; for all target nodes: EFT=EST+estimated execution time of the target node. Secondly, calculate LFT and LST (reverse traversal). For the target node of the exit (the target node without subsequent subtasks): LFT = the earliest completion time of the entire workflow (the maximum value of EFT among all the target nodes of the exit); for other target nodes: LFT = min{LST (the target node of all subsequent subtasks)}; for all target nodes: LST = LFT – the estimated execution time of the target node. Calculate the time float slack = LST – EST ​​for each target node. If the time float is 0, that is, the target node with slack = 0, it is the node on the critical path. The path weight of the target node can be calculated by the following formula:

[0086]

[0087] in, is the maximum value of time floating amount, is the time float of the target node, is the preset maximum weight, is the path weight of the target node. If slack=0, it is a node on the critical path, and the path weight can be considered to be ; if slack = slack max , indicating that this task is the least critical and the path weight is 0.

[0088] Finally, the critical path and the weight of each path are marked in the target dependency graph, and the subtask sequence of the critical path can be output.

[0089] like Figure 2 As shown, the workflow scheduling subsystem 130 includes: a task priority determination module 131, a workflow construction module 132, a resource matching and allocation module 133 and an execution plan generation module 134.

[0090] The task priority determination module 131 is used to obtain the target dependency graph and the initial workflow in the target task, and determine the designated priority corresponding to each target node in the target dependency graph based on the initial workflow; for each target node, the task priority of the target node is determined based on the estimated execution time in the node attributes of the target node, the designated priority corresponding to the target node, and the path weight corresponding to the target node.

[0091] The specified priority is the priority assigned by the user in the initial workflow. The estimated execution time is the estimated time required to execute the subtask of the target node. The task priority is used to describe the order of the subtasks corresponding to each target node.

[0092] Specifically, the target dependency graph in the dependency graph update module 124 and the initial workflow in the target task are received. Since the initial workflow contains the specified priority of each subtask, the specified priority corresponding to each target node in the target dependency graph can be determined. For each target node, the estimated execution time in the node attributes of the target node, the specified priority corresponding to the target node, and the path weight corresponding to the target node are calculated to determine the task priority of the target node, so as to comprehensively determine the execution priority of the subtask corresponding to each target node based on multiple factors, that is, the task priority, so as to provide a decision basis for subsequent resource allocation.

[0093] For example, the task priority can be calculated by the following formula:

[0094] Task priority = α × specified priority + β × path weight + γ × estimated execution time score

[0095] Among them, the specified priority comes from the workflow description file (initial workload) and can be normalized to between 0 and 10. The estimated execution time score is the estimated execution time (unit: seconds) mapped to between 0 and 10, which can be a linear normalized value based on the maximum / minimum duration. α, β, γ are preset coefficients that can be configured according to system policies. For example, the preset coefficient β corresponding to the path weight is higher, which can be α=0.2, β=0.5, and γ=0.3.

[0096] Estimated execution time score = 10×(1-(tt min ) / (t max -t min ))

[0097] Among them, t is the estimated execution time of the subtask corresponding to the target node, t min is the shortest estimated execution time of all subtasks to be executed (e.g. 10 seconds), maxIt is the longest estimated execution time of all tasks to be executed (such as 360 seconds). The range of the estimated execution time score: the shorter the subtask, the higher the score (closer to the preset maximum value of 10), and the longer the subtask, the lower the score (closer to the preset minimum value of 0), that is, the shorter the subtask, the higher the priority.

[0098] Optionally, dynamic priority adjustment can be performed to calculate new task priorities by regularly recalculating the estimated execution time and the estimated execution time score according to the execution progress of the target task and changes in resource platform information.

[0099] The workflow construction module 132 is used to construct a target workflow according to the target dependency graph and the task priority of each target node.

[0100] The target workflow is a workflow that is standardized according to the target dependency graph and the task priority of each target node in accordance with the standard workflow description language.

[0101] Specifically, a standard workflow description language (such as JSON, YAML, etc.) is used to describe the target dependency graph and the workflow corresponding to the task priority of each target node, supporting the expression of task dependencies, resource requirements, data dependencies, etc.

[0102] For example, a standard workflow description language is defined to ensure the accuracy and consistency of workflow description. When describing scientific computing workflows, nested structures are supported, which can express complex task dependencies, resource requirements, data dependencies, and task priorities in detail. For example, for a complex task containing multiple subtasks, the execution order, dependencies, and respective resource requirements of the subtasks can be defined in the workflow description.

[0103] The resource matching and allocation module 133 is used to obtain the task characteristics and the task priority of each target node; obtain the resource directory, and determine the real-time resource information of each candidate resource platform based on the resource directory; determine the target execution platform corresponding to each target node based on the task characteristics, the real-time resource information of each candidate resource platform and the task priority of each target node.

[0104] The candidate resource platforms are the resource platforms recorded in the resource directory. The real-time resource information is the real-time information of the candidate resource platforms. The target execution platform is the candidate resource platform allocated to the target node.

[0105] Specifically, the task characteristics analyzed by the task characteristic analysis module 121 and the task priority of each target node determined in the task priority determination module 131 are obtained, and the resource directory constructed in the resource directory module 112 is obtained. The real-time resource information of each candidate resource platform can be obtained from the resource directory. By comprehensively analyzing the task characteristics, the real-time resource information of each candidate resource platform and the task priority of each target node, a corresponding target execution platform can be allocated to each target node, so as to allocate the most suitable resource platform to each subtask in combination with the task resource requirements and the current available resource status.

[0106] Exemplarily, the real-time resource information in the resource directory is obtained from the resource abstraction subsystem 110, including the resource type, specification, and available status of each resource platform. According to the resource requirements of each task in the target dependency graph provided by the task granularity partitioning subsystem 120, qualified resources are searched. The matching degree between task characteristics and resource characteristics can be considered: high-performance CPU / GPU resources are allocated preferentially for computing-intensive tasks; high-bandwidth storage resources are allocated preferentially for I / O-intensive tasks. In addition, resource reservation and allocation can be performed: resources are reserved for subtasks with high task priority, resource allocation requests are submitted to the target execution platform through the resource abstraction subsystem 110, and task queues are maintained for different resource pools (high task priority is given priority).

[0107] The execution plan generation module 134 is used to obtain the target execution platform and target workflow corresponding to each target node and determine the initial execution plan; after each target execution platform in the initial execution plan receives the target data, the adaptation layer corresponding to each target execution platform is determined, and based on each adaptation layer, the subtask corresponding to the corresponding target node is submitted to each target execution platform.

[0108] Among them, the initial execution plan includes the task allocation plan and the execution schedule.

[0109] Specifically, the target dependency graph corresponding to the target workflow and the priority of each task are identified, and the resource allocation result, that is, the target execution platform corresponding to each target node, can generate an initial execution plan. After each target execution platform in the initial execution plan receives the target data, it can be determined that the corresponding target execution platform can execute the corresponding task. Therefore, the adaptation layer corresponding to each target execution platform is determined, and the subtask corresponding to the corresponding target node is submitted to each target execution platform based on each adaptation layer, so that each target execution platform can execute the corresponding task.

[0110] Exemplarily, the target dependency graph is received, and the corresponding target execution platform and estimated execution time are determined for each task in combination with the resource matching and allocation results. Then, a multi-objective optimization algorithm can be used to generate an initial execution plan: taking NSGA-II (Nondominated Sorting Genetic Algorithm II) as an example for multi-objective optimization, the specific steps are as follows: first, the objective function is constructed with the goal of minimizing the completion time, maximizing the resource utilization, and minimizing the data transmission cost. Then, the initial population is constructed, that is, multiple feasible scheduling schemes are randomly generated, and then the fitness is evaluated, the performance of each scheme on each target is calculated respectively, non-dominated sorting is performed, the schemes are divided into different Pareto frontier levels, the congestion is calculated, the diversity of the solution set is maintained, a new generation of population is generated by selection, crossover, and mutation, and iterative optimization is performed, that is, the above steps are repeated until convergence or the maximum number of iterations is reached. From the non-dominated solution set finally obtained, the most suitable scheduling scheme is selected according to the current system state and user preferences as the initial execution plan. This multi-objective optimization method can take into account multiple performance indicators at the same time and generate a balanced and efficient scheduling scheme.

[0111] Based on the above example, the workflow orchestration subsystem 130 further includes: a dynamic scheduling module 135 and a workflow variation module 136 .

[0112] The dynamic scheduling module 135 is used to determine whether each target execution platform meets at least one rescheduling condition. If so, the target execution platform and the rescheduling condition triggered by the target execution platform are fed back to the execution plan generation module so that the execution plan generation module updates the initial execution plan.

[0113] Among them, the rescheduling conditions are the conditions that trigger the update of the initial execution plan, such as: the task execution time deviates significantly from expectations, resource availability changes, network conditions change significantly, users adjust task priorities, etc.

[0114] Specifically, each target execution platform is monitored to determine in real time whether the target execution platform meets at least one rescheduling condition. If so, rescheduling is required. Therefore, the target execution platform and the rescheduling condition triggered by the target execution platform are fed back to the execution plan generation module, so that the execution plan generation module updates the initial execution plan to facilitate subsequent execution according to the updated initial execution plan.

[0115] Exemplarily, the task execution status and resource status are monitored in real time. By deploying a monitoring agent on the task execution node to collect information such as task execution time and resource usage, logs and historical data can be recorded. Rescheduling is triggered when the target execution platform meets at least one rescheduling condition, such as when the task execution time deviates significantly from expectations (such as exceeding 20% ​​of the expected time), resource availability changes (such as CPU usage exceeds 80% or memory remaining is less than 20%, etc.), network conditions change significantly (such as network bandwidth is lower than the set threshold) or the user adjusts the task priority. When rescheduling, an event-driven scheduling algorithm is used. According to the type of triggering event and the current system status (the target execution platform and the rescheduling condition triggered by the target execution platform), the problematic tasks are recalculated based on the multi-objective optimization algorithm. The dynamic adjustment includes the target execution platform used for task execution and the expected execution time.

[0116] The workflow variation module 136 is used to, during the execution process of the target task, in response to a change in the algorithm of at least one target node in the target dependency graph, update the target execution platform of the corresponding target node according to the changed algorithm and resource directory, and feed back the updated algorithm and target execution platform of the target node to the execution plan generation module so that the execution plan generation module updates the initial execution plan; in response to a change in the configuration parameters of at least one target node in the target dependency graph, issue new configuration parameters to the target execution platform corresponding to the target node whose configuration parameters have changed.

[0117] Specifically, during the execution of the target task according to the initial execution plan, if the algorithm of at least one target node in the target dependency graph changes, it is necessary to reallocate the corresponding target execution platform for the corresponding target node in combination with the changed algorithm and resource directory, and it is necessary to feed back the updated algorithm and target execution platform of the target node to the execution plan generation module so that the execution plan generation module updates the initial execution plan. During the execution of the target task according to the initial execution plan, if the configuration parameters of at least one target node in the target dependency graph change, new configuration parameters are issued to the target execution platform corresponding to the target node whose configuration parameters change, so that the target execution platform is reconfigured to facilitate the execution of the subtask corresponding to the target node according to the new configuration parameters. The workflow variation module 136 supports dynamic adjustment of the workflow structure during task execution, such as replacing the algorithm implementation for a task, adjusting parameter configuration, etc.

[0118] Exemplarily, during the execution process of the initial execution plan corresponding to the target workflow, the workflow structure can be dynamically adjusted by modifying the constructed target dependency graph. For example, when replacing the algorithm implementation for a task, according to the resource requirements and interface specifications of the new algorithm, it is necessary to reallocate the corresponding target execution platform for the task and modify the execution code of the task; when adjusting the parameter configuration, the corresponding parameter value can be modified directly on the target node of the constructed target dependency graph, and the relevant task nodes and their corresponding target execution platforms are notified to reload the configuration parameters.

[0119] like Figure 2 As shown, the data preheating management subsystem 140 includes: a target data determination module 141, a transmission priority determination module 142, a multi-path transmission module 143 and a data recording module 144.

[0120] The target data determination module 141 is used to obtain the target dependency graph and the initial execution plan, determine the cross-platform subtask in the initial execution plan, and use the data required by the cross-platform subtask as the target data.

[0121] Specifically, according to the target dependency graph provided by the task granularity division subsystem 120 and the initial execution plan in the workflow orchestration subsystem 130, the data dependency relationship between the tasks in the target workflow is analyzed, wherein the task relationship and data size are included in the target dependency graph, and the target execution platform corresponding to the task is included in the initial execution plan. Based on this, cross-platform subtasks can be identified, and the data that needs to be transmitted across platforms can be determined as the target data.

[0122] The transmission priority determination module 142 is used to determine the transmission priority of the target data for each group of target data based on the preset weight, the data size of the target data, the task priority of the cross-platform subtask corresponding to the target data, and the network bandwidth between the two target execution platforms corresponding to the target data.

[0123] Among them, the preset weights are the corresponding weights obtained by training based on the data size, task priority, network bandwidth and corresponding real priority of each target data recorded in the historical records.

[0124] Specifically, the network bandwidth between the two target execution platforms can be obtained through the resource abstraction subsystem 110, and the data size and task priority of each target data can be determined based on the initial execution plan, etc. The transmission priority of the target data can be obtained by weighted summation of the data size, network status, task execution plan and their corresponding preset weights.

[0125] Exemplarily, a weighted score based on access frequency may be used to determine the transmission priority of target data, and the calculation method is as follows:

[0126] P = α × data size + β × network bandwidth + γ × task priority

[0127] Among them, P is the transmission priority of the target data, α, β, γ are the preset weights corresponding to the data size, network bandwidth and task priority, which are obtained by training the data size, task priority, network bandwidth and corresponding real priority of each target data in the historical records. For example, the result range of the transmission priority is 0.0~1.0.

[0128] The multi-path transmission module 143 is used to obtain network topology information, and send corresponding target data to each target execution platform based on incremental synchronization according to the transmission priority of each target data and the network topology information.

[0129] Specifically, the resource abstraction subsystem 110 can obtain network topology information, analyze at least one transmission path and transmission order of each target data according to the transmission priority of each target data and the network topology information, and use the incremental synchronization method to transmit the target data to the corresponding target execution platform according to the transmission path and transmission order corresponding to each target data. Multiple transmission paths can be used to transmit target data in parallel to improve transmission efficiency, and when transmitting data, the incremental synchronization method is used to transmit only the changed part of the data using the data difference algorithm, which can reduce the amount of transmitted data.

[0130] Exemplarily, based on the network topology information, a network routing algorithm (such as the Dijkstra algorithm) is used to calculate multiple transmission paths. During the data transmission process, multi-threading technology is used to transmit data in parallel, and each thread is responsible for the data transmission of one path. In order to avoid network congestion, a flow control mechanism is introduced to dynamically adjust the data transmission rate of each thread according to the network bandwidth utilization. At the same time, a data check and retransmission mechanism is used to ensure the integrity and accuracy of data transmission. The version information of the data can be recorded, and a data change log is generated each time the data is updated to record the modified content of the data. During data transmission, only the changed part of the data is transmitted according to the version information and the data change log. For example, a data differential algorithm is used to calculate data changes to improve the efficiency of incremental synchronization.

[0131] The data recording module 144 is used to record the data size, task priority, network bandwidth and actual priority of each target data transmitted during the transmission of each target data, and update the preset weight according to the data size, task priority, network bandwidth and actual priority of each target data transmitted.

[0132] The real priority is the priority actually used when each target data is transmitted.

[0133] Specifically, for each target data, when the target data is transmitted to the corresponding target execution platform, the data size, task priority, network bandwidth and real priority of the target data are recorded. Then, the preset weights are periodically trained and adjusted using the stored data size, task priority, network bandwidth and real priority of each target data, and the trained and adjusted preset weights are applied to the transmission priority determination module 142.

[0134] Exemplarily, the data size (such as MB), task priority (0-10), network bandwidth (such as Mbps) and real priority (the transmission priority order actually used in scheduling can be expressed as a decimal, with a value range of 0.0-1.0, and the larger the value, the higher the priority) of the recorded target data are used as historical data. The historical data is used as sample data, and regression analysis is used to fit each preset weight (α, β, γ). For example, the least squares method can be used to fit each preset weight. Before regression, each input variable is normalized to a uniform scale (0-1), so that the trained α+β+γ≈1. Before the initial run, a batch of sample data can be prepared (at least dozens to hundreds of sample data), and each preset weight can be calculated first, and then each preset weight can be retrained periodically after the run. It should be noted that the sample data is normalized before the training of the preset weights.

[0135] Based on the above example, the cross-domain workflow scheduling system for dynamic resource orchestration further includes: a cross-domain fault-tolerant subsystem 150, such as Figure 2 As shown, the cross-domain fault-tolerant subsystem 150 includes: a fault detection and replacement module 151 and a key data storage module 152.

[0136] The fault detection and replacement module 151 is used to perform fault detection on each resource platform during the execution of the target task. When at least one of the target execution platforms corresponding to the target task fails, the replacement execution platform corresponding to each failed target execution platform is determined based on the information of each resource platform, and each failed target execution platform and the corresponding replacement execution platform are sent to the workflow orchestration subsystem to update the initial execution plan.

[0137] The alternative execution platform is a platform that can replace the target execution platform to execute the corresponding subtask.

[0138] Specifically, during the execution of the target task, a fault detection mechanism can be configured for each resource platform to perform fault detection in real time, so as to facilitate the timely identification of platform or network failures. When at least one of the target execution platforms corresponding to the target task fails, it is necessary to replace the resource platform. Combined with the information of each resource platform in the resource abstraction subsystem 110, search for candidate resource platforms similar to each failed target execution platform as the corresponding alternative execution platform. In addition, each failed target execution platform and the corresponding alternative execution platform are sent to the workflow orchestration subsystem 130, so that each failed target execution platform is replaced with the corresponding alternative execution platform through the workflow orchestration subsystem 130, and the initial execution plan is updated.

[0139] Exemplarily, a distributed heartbeat detection mechanism can be deployed, and a heartbeat detection agent can be deployed on the nodes of each resource platform and the nodes of the network link. The heartbeat detection agent can send heartbeat messages to other nodes at regular intervals (such as 10 seconds), and at the same time, receive heartbeat messages from other nodes. If the heartbeat message of a node is not received within a certain period of time (such as 30 seconds), it is determined that the node has a fault. Furthermore, a distributed hash table technology can be used to manage the information of the heartbeat detection agent to ensure the reliability and scalability of the heartbeat detection, so as to facilitate timely discovery of the failure of the nodes of the resource platform and the nodes of the network link. When the target execution platform fails and is unavailable, the available alternative execution platform list is searched and obtained from the resource abstraction subsystem 110, and then, the resource matching algorithm (such as the Hungarian algorithm) can be used to select the most suitable alternative execution platform according to the resource requirements of the corresponding subtasks and the resource platform information of each candidate resource platform in the alternative execution platform list. Then, according to the requirements of the task submission system of the alternative execution platform, the task submission script is regenerated, and the workflow orchestration subsystem 130 updates the initial execution plan. If the corresponding target execution platform is marked in the target dependency graph, the task granularity division subsystem 120 can update the target dependency graph.

[0140] The key data saving module 152 is used to incrementally save the data of each target node on the key path in the target dependency graph according to a second preset period.

[0141] The second preset period is a preset period for incrementally saving the data of each target node on the critical path, and can be set according to demand.

[0142] Specifically, according to the second preset period, the intermediate states during the execution of the subtasks corresponding to the target nodes on the critical path are periodically saved to support task recovery.

[0143] Exemplarily, the intermediate states of the subtasks corresponding to each target node on the critical path during execution are periodically saved, and the information is stored in a file system or database. During the task execution process, the intermediate state of the task (such as memory data, file pointer, task execution progress, etc.) is saved to a checkpoint file or database table every second preset period (such as 10 minutes) or after the subtasks corresponding to each target node on the critical path are completed. In order to improve the saving efficiency, an incremental saving method is adopted to save only the data that has changed compared with the last time.

[0144] Based on the above example, the cross-domain fault-tolerant subsystem 150 also includes: a downgraded execution module 153, which is used to determine the task level corresponding to each target node based on the critical path in the target dependency graph and the task priority corresponding to each target node; for each target node, in response to insufficient resources in the target execution platform corresponding to the target node, a downgrade strategy is determined based on the task level of the target node and the resource status in the target execution platform, and based on the downgrade strategy, the subtask corresponding to the target node is downgraded in the target execution platform.

[0145] The task level is used to describe the importance of the subtasks corresponding to each target node, which may include core, important, secondary, etc. The resource status includes memory status, CPU usage, system load, etc.

[0146] Specifically, the task level corresponding to each target node in the critical path of the target dependency graph is determined as the core, and the remaining target nodes are graded according to the corresponding task priority to obtain the task level corresponding to each target node. For each target node, when the resources in the target execution platform corresponding to the target node are insufficient, it is determined whether the downgrade process can be performed in combination with the task level of the target node. If it can, the resource status in the corresponding target execution platform is considered to determine the downgrade strategy to control the use of the functional downgrade strategy in the target execution platform to execute the subtask corresponding to the target node. When resources are seriously insufficient, a downgrade strategy is formulated according to the task level of the target node and the resource status in the target execution platform to support downgrade execution and ensure core functions. The calculation accuracy can be reduced or the calculation steps can be reduced through downgrade to reduce resource consumption. During the downgraded execution process, the user is notified that the workflow is being downgraded and the downgrade operation log is recorded for subsequent analysis and recovery.

[0147] For example, if the task priority ranges from 0 to 10, the larger the value, the higher the priority. Then, a core task is a task with a task priority of >= 7 or a target node on the critical path, indicating that it must be completed and cannot be downgraded; an important task is a task with a task priority of >3 and <7 and a target node that is not on the critical path, indicating that it can be moderately downgraded to ensure basic functions; a secondary task is a task with a task priority of <= 3, indicating that it can be significantly downgraded or suspended. Downgrade strategies may include: reducing algorithm accuracy (reducing the number of iterations, convergence accuracy or sampling rate), reducing features (temporarily shutting down non-core functional modules), reducing data size (reducing the amount of data processed, such as reducing resolution and sampling processing), reducing resource requirements (adjusting algorithm parameters, reducing memory and CPU requirements), etc. For example, a decision mechanism can be formulated: Level 1 (mild): CPU usage > 80% or memory usage > 70%, a downgrade strategy of data sampling and reduced visualization frequency will be adopted; Level 2 (moderate): available memory < 2GB or I / O congestion is obvious, a downgrade strategy of data sampling, reduced visualization frequency, reduced calculation accuracy and module restriction will be adopted; Level 3 (severe): system load > 95% or scheduling failure times > N (preset failure times), only core tasks are retained and all non-critical processes are closed. This downgrade mechanism can ensure that the most core scientific computing tasks can be completed under resource constraints and the availability of basic functions is guaranteed.

[0148] The present invention has the following technical effects: through the resource abstraction subsystem, the information of each resource platform and the network topology information are collected, the resource directory is constructed, and each adaptation layer is constructed; through the task granularity division subsystem, the target granularity is calculated, and the initial dependency graph is converted into a target dependency graph; through the workflow orchestration subsystem, the task characteristics, the resource directory and the target dependency graph are combined to determine the initial execution plan; through the data preheating management subsystem, each target data is pre-transmitted across platforms; after each target execution platform in the initial execution plan receives the target data through the workflow orchestration subsystem, the adaptation layer corresponding to each target execution platform is determined, and the corresponding subtask is submitted to each target execution platform to schedule the execution of the target task, thereby realizing unified management of resources on different platforms, reducing the overhead of data transmission and task scheduling, and improving the stability of cross-domain task completion.

[0149] It should be noted that the terms used in the present invention are only for describing specific embodiments, rather than limiting the scope of the present application. Unless the context clearly indicates an exception, the words "one", "a", "a kind of" and / or "the" do not specifically refer to the singular, but may also include the plural. The terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of further restrictions, the elements defined by the sentence "include one..." do not exclude the presence of other identical elements in the process, method or device including the elements.

[0150] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments may still be modified, or some or all of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A cross-domain workflow scheduling system with dynamic resource orchestration, characterized in that: include: The resource abstraction subsystem is used to collect information about each resource platform and network topology, build a resource directory based on the information about each resource platform, and build an adaptation layer between each task submission system type and each resource platform type; The task granularity division subsystem is used to receive a target task, determine the task characteristics corresponding to the target task, the cost of each task and the initial dependency graph, determine the target granularity according to the cost of each task, and optimize the initial dependency graph according to the target granularity to obtain a target dependency graph; The workflow orchestration subsystem is used to obtain the task characteristics, the resource directory and the target dependency graph, determine the initial execution plan according to the target dependency graph and the resource directory, and after each target execution platform in the initial execution plan receives the target data, determine the adaptation layer corresponding to each target execution platform, and submit the corresponding subtask to each target execution platform based on each adaptation layer, so that each target execution platform jointly executes the target task; The data preheating management subsystem is used to obtain the network topology information, the target dependency graph and the initial execution plan, determine the transmission priority of the cross-platform subtasks in the initial execution plan, and send the corresponding target data to each target execution platform according to each transmission priority and the network topology information.

2. The system according to claim 1, characterized in that The resource abstraction subsystem includes: A resource agent module, used for collecting information of each resource platform based on a first preset period; The resource directory module is used to receive the information of each resource platform and construct a resource directory according to the information of each resource platform; the resource directory module includes a resource query interface and a resource allocation interface; The resource adaptation module is used to construct an adaptation group according to each task submission system type and each resource platform type, and pre-construct an adaptation layer corresponding to each adaptation group; The network topology module is used to obtain the network topology structure and collect the network status information corresponding to each network node in the network topology structure, and determine the network topology information according to the network topology structure and the network status information corresponding to each network node.

3. The system according to claim 1, characterized in that The task granularity division subsystem includes: A task characteristic analysis module is used to receive a target task, determine the number of operations of each type according to a task code file in the target task, determine the task characteristic according to the number of operations of each type, determine the task calculation cost in the task cost according to the task code file, determine the task communication cost in the task cost according to a data description file in the target task, and determine the task switching cost in the task cost according to historical execution time data of similar tasks corresponding to the target task; A dependency graph construction module is used to receive the target task, determine each initial node, the node attribute corresponding to each initial node, each initial directed edge and the edge attribute corresponding to each initial directed edge according to the initial workflow in the target task, and construct an initial dependency graph according to each initial node, the node attribute corresponding to each initial node, each initial directed edge and the edge attribute corresponding to each initial directed edge; a granularity optimization module, configured to receive the task computing cost, the task communication cost, and the task switching cost, take the sum of the task communication cost and the task switching cost as a process value, and take the quotient of the task computing cost and the process value as a target granularity; A dependency graph updating module is used to receive the initial dependency graph and the target granularity, determine an adjustment method of the initial dependency graph according to a preset granularity range and the target granularity, and adjust the initial dependency graph based on the adjustment method to obtain a target dependency graph; wherein the adjustment method includes a task merging method and a task splitting method; A critical path identification module is used to receive the target dependency graph, determine the earliest start time and the latest start time corresponding to each target node in the target dependency graph according to the target dependency graph, determine the time float corresponding to each target node according to the earliest start time and the latest start time corresponding to each target node, and determine the critical path and the path weight corresponding to each target node according to the time float corresponding to each target node, and add the critical path and the path weight corresponding to each target node to the target dependency graph.

4. The system according to claim 3, characterized in that The dependency graph updating module is further used for: Initialize the adjustment times and determine the relationship between the target particle size and the preset particle size range; In response to the target granularity being greater than an upper limit of the preset granularity range, determining that the adjustment mode of the initial dependency graph is a task merging mode; and determining whether the number of adjustments is equal to a preset number; In response to the adjustment number being less than the preset number, according to the initial dependency graph, two adjacent initial nodes with the smallest data dependency are used as merge nodes, the merge nodes are merged to obtain a new initial node, the initial dependency graph and the adjustment number are updated, the target granularity of the updated initial dependency graph is determined, and the step of determining the relationship between the target granularity and the preset granularity range is returned to be executed; In response to the adjustment number being equal to the preset number, taking the updated initial dependency graph as the target dependency graph; In response to the target granularity being within the preset granularity range, determining the initial dependency graph as a target dependency graph; In response to the target granularity being smaller than a lower limit of the preset granularity range, determining that the adjustment method of the initial dependency graph is a task splitting method; Determining whether the adjustment number is equal to a preset number; In response to the number of adjustments being less than the preset number, taking the initial node containing the parallel task in the initial dependency graph as a split node, splitting the split node into at least two new initial nodes, updating the initial dependency graph and the number of adjustments, determining the target granularity of the updated initial dependency graph, and returning to execute the step of determining the relationship between the target granularity and the preset granularity range; In response to the adjustment number being equal to the preset number, the updated initial dependency graph is used as the target dependency graph.

5. The system according to claim 3, characterized in that The critical path identification module is further used for: For each target node, the difference between the latest start time and the earliest start time corresponding to the target node is used as the time float corresponding to the target node; Constructing a critical path according to each target node whose time float is 0; According to the time floating amount corresponding to each target node, determine the maximum value of the time floating amount; For each target node, the path weight corresponding to the target node is determined according to the preset maximum weight, the maximum value of the time floating amount and the time floating amount corresponding to the target node.

6. The system according to claim 1, characterized in that The workflow orchestration subsystem includes: A task priority determination module is used to obtain a target dependency graph and an initial workflow in a target task, and determine the designated priority corresponding to each target node in the target dependency graph according to the initial workflow; for each target node, determine the task priority of the target node according to the estimated execution time in the node attributes of the target node, the designated priority corresponding to the target node, and the path weight corresponding to the target node; A workflow construction module, used to construct a target workflow according to the target dependency graph and the task priority of each target node; The resource matching and allocation module is used to obtain the task characteristics and the task priority of each target node; obtain the resource directory, and determine the real-time resource information of each candidate resource platform according to the resource directory; determine the target execution platform corresponding to each target node according to the task characteristics, the real-time resource information of each candidate resource platform and the task priority of each target node; The execution plan generation module is used to obtain the target execution platform corresponding to each target node and the target workflow to determine the initial execution plan; after each target execution platform in the initial execution plan receives the target data, the adaptation layer corresponding to each target execution platform is determined, and based on each adaptation layer, the subtask corresponding to the corresponding target node is submitted to each target execution platform.

7. The system according to claim 6, characterized in that The workflow orchestration subsystem further includes: A dynamic scheduling module is used to determine, for each target execution platform, whether the target execution platform meets at least one rescheduling condition, and if so, to feed back the target execution platform and the rescheduling condition triggered by the target execution platform to the execution plan generation module, so that the execution plan generation module updates the initial execution plan; The workflow variation module is used to update the target execution platform of the corresponding target node according to the changed algorithm and the resource directory in response to a change in the algorithm of at least one target node in the target dependency graph during the execution of the target task, and feed back the updated algorithm and target execution platform of the target node to the execution plan generation module so that the execution plan generation module updates the initial execution plan; in response to a change in the configuration parameters of at least one target node in the target dependency graph, issue new configuration parameters to the target execution platform corresponding to the target node whose configuration parameters have changed.

8. The system according to claim 1, characterized in that The data preheating management subsystem includes: A target data determination module, used to obtain the target dependency graph and the initial execution plan, determine the cross-platform subtask in the initial execution plan, and use the data required by the cross-platform subtask as target data; A transmission priority determination module, for determining, for each group of target data, the transmission priority of the target data according to a preset weight, the data size of the target data, the task priority of the cross-platform subtask corresponding to the target data, and the network bandwidth between the two target execution platforms corresponding to the target data; A multi-path transmission module, used to obtain the network topology information, and send corresponding target data to each target execution platform based on incremental synchronization according to the transmission priority of each target data and the network topology information; The data recording module is used to record the data size, task priority, network bandwidth and real priority of each target data that has been transmitted during the transmission of each target data, and update the preset weight according to the data size, task priority, network bandwidth and real priority of each target data that has been transmitted.

9. The system according to claim 1, characterized in that It also includes a cross-domain fault-tolerant subsystem, the cross-domain fault-tolerant subsystem includes: A fault detection and replacement module is used to perform fault detection on each resource platform during the execution of the target task. When at least one of the target execution platforms corresponding to the target task fails, the replacement execution platform corresponding to each failed target execution platform is determined based on the information of each resource platform, and each failed target execution platform and the corresponding replacement execution platform are sent to the workflow orchestration subsystem to update the initial execution plan; The key data saving module is used to incrementally save the data of each target node on the key path in the target dependency graph according to a second preset period.

10. The system according to claim 9, characterized in that The cross-domain fault-tolerant subsystem further includes: A downgraded execution module is used to determine the task level corresponding to each target node based on the critical path in the target dependency graph and the task priority corresponding to each target node; for each target node, in response to insufficient resources in the target execution platform corresponding to the target node, a downgrade strategy is determined based on the task level of the target node and the resource status of the target execution platform, and based on the downgrade strategy, the subtask corresponding to the target node is downgraded and executed in the target execution platform.

Citation Information

Patent Citations

  • Heterogeneous computing resource virtualization processing method, electronic equipment and storage medium

    CN118626275A

  • Data analysis method and system based on visual modeling and job arrangement scheduling

    CN118819903A

  • Intelligent operation and maintenance method and system for managing different types of computing resources

    CN119645614A

  • Cloud computing parallel task optimization scheduling method based on priority dependency graph

    CN119806776A

  • Intelligent computing-oriented method, system and apparatus for scheduling distributed training tasks

    WO2024060789A1

Cited By

  • Task processing method and device and storage medium

    CN120386640A

  • Intelligent parallel task scheduling and monitoring system

    CN121144050A

  • Intelligent parallel task scheduling and monitoring system

    CN121144050B

  • Intelligent underwater sound task scheduling method based on multi-objective optimization

    CN121998383A

  • A multi-objective optimization-based underwater acoustic task intelligent scheduling method

    CN121998383B