Cross-domain Workflow Scheduling System for Dynamic Resource Orchestration
Through the cross-domain workflow scheduling system with dynamic resource orchestration, the problem of inefficient resource management in cross-domain scientific computing is solved, unified management and stable scheduling of heterogeneous computing resources is realized, data transmission and task scheduling overhead is reduced, and task completion is improved.
Patent Information
- Application Number
- CN202510562912.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The prior art is inefficient in cross-domain scientific computing, making it difficult to effectively manage heterogeneous computing resources distributed in different geographical locations, and there are problems with single point failure risk and resource scheduling instability.
It provides a cross-domain workflow scheduling system for dynamic resource orchestration, including resource abstraction subsystem, task granular molecular system, workflow scheduling subsystem and data preheating management subsystem. By building resource catalogs and adaptation layers, it optimizes the dependency diagram, determines the execution plan, and performs cross-platform data transmission and task scheduling.
It realizes unified management of resources on different platforms, reduces data transmission and task scheduling overhead, and improves the stability and efficiency of cross-domain tasks.
Smart Images

Figure CN120104345B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automated computing, and particularly to a cross-domain workflow scheduling system for dynamic resource orchestration. Background Art
[0002] With the continuous expansion of the scale of scientific computing, a single computing cluster or supercomputer center often struggles to meet the resource requirements of large-scale scientific computing workflows. In fields such as materials science, weather simulation, and fusion energy research, it is often necessary to utilize heterogeneous computing resources distributed in different geographical locations to complete complex computing tasks.
[0003] Currently, cross-domain scientific computing mainly adopts manual coordination, point-to-point integration, or centralized scheduling. Among them, the manual coordination method is inefficient, error-prone, and difficult to handle complex workflows; the point-to-point integration method lacks generality, and new interfaces need to be developed when adding new computing platforms; the centralized scheduling method is inefficient when network conditions are poor or there is a firewall isolation between computing platforms, and there is a risk of single-point failure.
[0004] In view of this, the present invention is specifically proposed. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a cross-domain workflow scheduling system for dynamic resource orchestration, which realizes unified management of resources on different platforms, reduces the overhead of data transmission and task scheduling, and improves the stability of cross-domain task completion.
[0006] An embodiment of the present invention provides a cross-domain workflow scheduling system for dynamic resource orchestration, which includes:
[0007] A resource abstraction subsystem, configured to collect information of each resource platform and network topology information, construct a resource directory according to the information of each resource platform, and construct an adaptation layer between each task submission system type and each resource platform type;
[0008] A task granularity division subsystem, configured to receive a target task, determine the task characteristics, each task cost, and an initial dependency graph corresponding to the target task, determine a target granularity according to each task cost, and optimize the initial dependency graph according to the target granularity to obtain a target dependency graph;
[0009] A workflow orchestration subsystem, configured to obtain task characteristics, a resource directory, and a target dependency graph, determine an initial execution plan according to the target dependency graph and the resource directory, and determine an adaptation layer corresponding to each target execution platform after each target execution platform in the initial execution plan receives target data, and submit corresponding subtasks to each target execution platform based on each adaptation layer, so that each target execution platform jointly executes the target task;
[0010] The data preheating management subsystem is used to obtain network topology information, a target dependency graph, and an initial execution plan, determine the transmission priorities of cross-platform subtasks in the initial execution plan, and send corresponding target data to each target execution platform according to each transmission priority and the network topology information.
[0011] The embodiments of the present invention have the following technical effects:
[0012] Through the resource abstraction subsystem, collect information of each resource platform and network topology information, construct a resource catalog, and construct each adaptation layer. Through the task granularity division subsystem, calculate the target granularity, and convert the initial dependency graph into a target dependency graph. Through the workflow orchestration subsystem, combine task characteristics, the resource catalog, and the target dependency graph to determine the initial execution plan. Through the data preheating management subsystem, pre-transmit each target data across platforms. After each target execution platform in the initial execution plan receives the target data through the workflow orchestration subsystem, determine the corresponding adaptation layer for each target execution platform, and submit corresponding subtasks to each target execution platform to schedule and execute the target task, realizing unified management of resources on different platforms, reducing the overhead of data transmission and task scheduling, and improving the stability of cross-domain task completion. Description of the Drawings
[0013] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0014] Figure 1 It is a schematic structural diagram of a cross-domain workflow scheduling system for dynamic resource orchestration provided by an embodiment of the present invention;
[0015] Figure 2 It is a schematic structural diagram of another cross-domain workflow scheduling system for dynamic resource orchestration provided by an embodiment of the present invention. Specific Embodiments
[0016] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present invention.
[0017] The cross-domain workflow scheduling system for dynamic resource orchestration provided by the embodiments of the present invention is mainly applicable to the situation of unifying the resources of each computing platform and performing reasonable resource orchestration and workflow scheduling for target tasks.
[0018] Figure 1 It is a schematic structural diagram of a cross-domain workflow scheduling system for dynamic resource orchestration provided by the embodiments of the present invention. Refer to Figure 1 This cross-domain workflow scheduling system for dynamic resource orchestration specifically includes: a resource abstraction subsystem 110, a task granularity division subsystem 120, a workflow orchestration subsystem 130, and a data preheating management subsystem 140.
[0019] The resource abstraction subsystem 110 is used to collect information of each resource platform and network topology information, construct a resource directory according to the information of each resource platform, and construct an adaptation layer between each task submission system type and each resource platform type.
[0020] Among them, the resource platform information is the local resource information of the platform where heterogeneous computing resources distributed in different supercomputing centers and cloud platforms are located, which can be the local resource information of resource platforms in different clusters and different domains. For example, CPU (Central Processing Unit, central processing unit), GPU (Graphics Processing Unit, graphics processing unit), memory, storage, network bandwidth, etc. The network topology information is the network topology structure and network status information formed between each resource platform. The resource directory is a directory used to integrate and statistically analyze the information of each resource platform, which can be understood as a global resource view. The task submission system type is the system type used when submitting the target task to be executed. The resource platform type is the type of different resource platforms. The adaptation layer can include task submission, monitoring, and management interfaces.
[0021] Specifically, the resource abstraction subsystem 110 can regularly collect information of each resource platform and network topology information, and integrate the information of each resource platform into a resource directory for subsequent resource allocation in combination with the target task. Moreover, an adaptation layer between different task submission system types and different resource platform types can be pre-constructed for direct invocation during task submission, monitoring, and management.
[0022] The task granularity division subsystem 120 is used to receive the target task, determine the task characteristics, each task cost, and the initial dependency graph corresponding to the target task, determine the target granularity according to each task cost, and optimize the initial dependency graph according to the target granularity to obtain the target dependency graph.
[0023] Among them, the target task is a task that requires resource allocation and execution. It can be understood that the target task is a file related to the workflow description uploaded by the user, which may include a workflow description file, a task code file, an environment configuration file, a data description file, a resource preference configuration file, etc. Task characteristics include compute-intensive, communication-intensive, and I / O (Input / Output) intensive. The task cost is the processing cost of the target task in terms of computing, communication, etc. The initial dependency graph is the task dependency graph between subtasks determined by analyzing the submitted workflow description file. The target granularity is used to describe the size of the target task divided, that is, the amount of work contained in each subtask. The target dependency graph is the task dependency graph between subtasks obtained after merging or splitting subtasks based on the target granularity.
[0024] Specifically, the task granularity division subsystem 120 receives the target task submitted by the user, analyzes the task code file in the target task to determine the task characteristics and various types of task costs, analyzes the workflow description file in the target task, and constructs an initial dependency graph. By comprehensively analyzing each task cost, the target granularity of the target task can be calculated, and it is judged whether the target granularity is within a reasonable range. If not, the initial dependency graph needs to be optimized, such as merging subtasks or splitting subtasks, and the optimized initial dependency graph is used as the target dependency graph.
[0025] It can be understood that the workflow description file is the core file for the system to understand the entire computing process. It usually adopts the JSON (JavaScript Object Notation) or YAML (YAML Ain't a Markup Language) format and usually contains task identification, task name, task description, subtask list (the subtask list contains subtask identification, subtask name, subtask description, subtask code file path, input data, output data, required computing resources such as CPU and memory, estimated execution duration, task priority, other subtasks on which the subtask depends), data entity relationship, running limit conditions, etc. The task code file is the execution code corresponding to each task that the user needs to provide. It can be various programming languages and needs to contain clear task description comments and metadata, good error handling and logging, support for configuring input / output paths through environment variables or command-line parameters, and return a clear exit code to indicate task success or failure. The environment configuration file is the environment dependency configuration that the user needs to provide for the task. The data description file usually contains data identification, data name, data type, data size, data file metadata information. The resource preference configuration file is the resource preference provided by the user to guide the system's resource allocation, which may include preferred resources, avoided resources, and performance preferences.
[0026] The workflow orchestration subsystem 130 is used to obtain task characteristics, a resource directory, and a target dependency graph, determine an initial execution plan according to the target dependency graph and the resource directory, and after each target execution platform in the initial execution plan receives target data, determine an adaptation layer corresponding to each target execution platform, and submit corresponding subtasks to each target execution platform based on each adaptation layer, so that each target execution platform jointly executes the target task.
[0027] Among them, the initial execution plan is an execution plan determined according to the target dependency graph to determine the workflow and combined with task characteristics and the resource directory for resource allocation. The target execution platform is the resource platform that needs to execute subtasks after resource allocation. The target data is the data required for task execution transmitted across platforms.
[0028] Specifically, through the workflow orchestration subsystem 130, the task characteristics and the target dependency graph determined by the task granularity division subsystem 120 can be obtained, and a new workflow can be constructed and resource allocation can be performed in combination with the resource directory constructed by the resource abstraction subsystem 110 to obtain the initial execution plan. And, after each target execution platform in the initial execution plan receives the target data, that is, when the cross-platform data preparation is completed, according to the task submission system type of the target task and each target execution platform, look up in each pre-constructed adaptation layer in the resource abstraction subsystem 110 to obtain the adaptation layer corresponding to each target execution platform respectively. Furthermore, for the target execution platform, submit the corresponding subtask to the target execution platform based on the corresponding adaptation layer, so that each target execution platform jointly executes the target task.
[0029] The data preheating management subsystem 140 is used to obtain network topology information, a target dependency graph, and an initial execution plan, determine the transmission priority of cross-platform subtasks in the initial execution plan, and send corresponding target data to each target execution platform according to each transmission priority and the network topology information.
[0030] Among them, the cross-platform subtask is a subtask that needs to switch resource platforms for execution, that is, the previous subtask is executed on one resource platform and the current subtask is executed on another resource platform, and the current subtask is the cross-platform subtask. The transmission priority is the priority of the target data required by the cross-platform subtask.
[0031] Specifically, the data preheating management subsystem 140 can obtain network topology information from the resource abstraction subsystem 110, obtain the target dependency graph from the task granularity division subsystem 120, and obtain the initial execution plan from the workflow orchestration subsystem 130. By analyzing the target dependency graph and the initial execution plan, each cross-platform subtask in the initial execution plan can be obtained, and the transmission priority can be calculated and determined for these cross-platform subtasks. Furthermore, by combining each transmission priority and the network topology information, it is determined whether parallel transmission can be performed, and the corresponding target data is sent to each target execution platform in a parallel transmission manner or a non-parallel transmission manner, so as to pre-transmit the required target data to the target execution platform before each subtask is executed, reducing the data transmission waiting time.
[0032] It can be understood that the data preheating management subsystem 140 runs after the initial execution plan is generated and determined in the workflow orchestration subsystem 130, and before the corresponding adaptation layers for each target execution platform are determined and the corresponding subtasks are submitted to each target execution platform based on each adaptation layer. The data preheating management subsystem 140 analyzes the task dependency relationship, the location of the dependent target data, and each target execution platform in the initial execution plan, and decides which data needs to be migrated or copied (target data) to perform the preheating operation.
[0033] Figure 2 It is a schematic structural diagram of another cross-domain workflow scheduling system for dynamic resource orchestration provided by an embodiment of the present invention.
[0034] As Figure 2 shown, the resource abstraction subsystem 110 includes: a resource proxy module 111, a resource directory module 112, a resource adaptation module 113, and a network topology module 114.
[0035] The resource proxy module 111 is used to collect information of each resource platform based on a first preset period.
[0036] Among them, the first preset period is a preset period for resource platform information collection and can be set according to requirements.
[0037] Specifically, the resource proxy module 111 deploys a lightweight resource proxy on each resource platform, which is responsible for collecting resource platform information based on the first preset period and regularly reporting it to the resource directory module 112.
[0038] Exemplarily, the multi-thread technology can be adopted to implement the function of collecting resource platform information. Using the timing task mechanism, the system command is called every first preset period (such as 5 minutes) to collect the resource platform information. The collected resource platform information is formatted and encapsulated into a data packet in a format such as JSON, and can be reported to the resource directory module 112 through the HTTP (Hypertext Transfer Protocol) / HTTPS (Hypertext Transfer Protocol Secure) protocol. To ensure the security of data transmission, the SSL / TLS (Secure Sockets Layer / Transport Layer Security) encrypted communication can be adopted.
[0039] The resource directory module 112 is used to receive the information of each resource platform and construct a resource directory according to the information of each resource platform.
[0040] Among them, the resource directory module 112 includes a resource query interface and a resource allocation interface.
[0041] Specifically, the resource directory module 112 maintains a global resource view, including the information of each resource platform, that is, information such as resource type, specification, and available status, and can provide a resource query interface and a resource allocation interface.
[0042] Exemplarily, a relational database (such as MySQL) is used to store the information of each resource platform and construct a global resource view. The database table design includes fields such as platform identifier, resource type, resource specification, available status, and update time. The resource query interface can support queries according to conditions such as platform identifier, resource type, and available status. The resource allocation interface ensures the atomicity of resource allocation through a transaction mechanism to avoid resource conflicts. At the same time, a caching mechanism can be introduced, such as Redis (Remote Dictionary Server), to cache the commonly used resource platform information, reduce the database query pressure, and improve the query efficiency.
[0043] The resource adaptation module 113 is used to construct an adaptation group according to each task submission system type and each resource platform type, and pre-construct an adaptation layer corresponding to each adaptation group.
[0044] Among them, an adaptation group is a combination of a task submission system type and a resource platform type.
[0045] Specifically, combine each task submission system type with each resource platform type to obtain multiple adaptation groups. Furthermore, for each adaptation group, pre-construct the task submission, monitoring, and management interfaces corresponding to the adaptation group as the corresponding adaptation layer. That is, for different types of task submission systems, such as: Slurm (Simple Linux Utility for Resource Management), PBS (Portable Batch System), Kubernetes (open-source container orchestration platform), etc., provide unified task submission, monitoring, and management interfaces.
[0046] Exemplarily, develop corresponding adaptation layers for task submission systems on different platforms. Taking Slurm as an example, use the command-line tools provided by Slurm for task submission, monitoring, and management. Package the unified task submission, monitoring, and management interfaces into Python modules, and call the implementation modules corresponding to different resource platforms through the importlib (dynamic import module and check module) dynamic import mechanism. When submitting a task, automatically call the corresponding task submission interface according to the resource platform type queried from the resource directory.
[0047] The network topology module 114 is used to obtain the network topology structure, collect the network status information corresponding to each network node in the network topology structure, and determine the network topology information according to the network topology structure and the network status information corresponding to each network node.
[0048] Specifically, obtain the network topology structure, regard each node in the network topology structure as a network node, collect the network status information of each network node, such as network bandwidth, availability, etc., write the network status information of each network node into each network node of the network topology structure to obtain the network topology information.
[0049] As Figure 2 shown, the task granularity division subsystem 120 includes: a task characteristic analysis module 121, a dependency graph construction module 122, a granularity optimization module 123, a dependency graph update module 124, and a critical path identification module 125.
[0050] The task characteristic analysis module 121 is used to receive a target task, determine the quantity of each type of operation according to the task code file in the target task, determine the task characteristics according to the quantity of each type of operation, determine the task calculation cost in the task cost according to the task code file, determine the task communication cost in the task cost according to the data description file in the target task, and determine the task switching cost in the task cost according to the historical execution time data of similar tasks corresponding to the target task.
[0051] Among them, the type operation quantity includes the computation operation quantity, the transmission operation quantity, and the I / O operation quantity. The task cost includes the task computation cost, the task communication cost, and the task switching cost. The task computation cost is used to describe the computational complexity of the target task. The task communication cost is used to describe the estimated cost of data dependencies in the target task. The task switching cost is used to describe the time cost from the start of task execution to the start of scheduling. The historical execution time data is the time from the start of task execution to the start of scheduling.
[0052] Specifically, based on the target task uploaded by the user, analyze the task code file in the target task, judge the quantity of each type of operation in the target task, and use the type with the largest quantity of type operations as the task characteristic to facilitate determining the basic resource requirements of the target task. Moreover, obtain the task computation cost in the task cost according to the analysis of the task code file, estimate the task communication cost in the task cost according to the data dependency quantity in the data description file, find similar tasks of the target task in the historical tasks, and statistically obtain the task switching overhead in the task cost according to the historical data of the similar tasks. Specifically, use the difference between the task start execution time and the task start scheduling time of the similar task as the historical execution time data, and use the average value of each historical execution time data as the task switching overhead.
[0053] Exemplarily, a method combining static code analysis and dynamic runtime monitoring is used to analyze the task characteristics. Static code analysis is to parse the syntax tree of the task code file and count the quantity of different types of operations to preliminarily judge whether the task characteristic of the target task is computation-intensive, I / O-intensive, or communication-intensive. Dynamic runtime monitoring is to use system performance monitoring tools to collect the resource usage situation during task runtime in real time during the task trial operation stage to further accurately determine the basic resource requirements of the task.
[0054] The dependency graph construction module 122 is used to receive the target task, determine each initial node, the node attributes corresponding to each initial node, each initial directed edge, and the edge attributes corresponding to each initial directed edge according to the initial workflow in the target task, and construct an initial dependency graph according to each initial node, the node attributes corresponding to each initial node, each initial directed edge, and the edge attributes corresponding to each initial directed edge.
[0055] Among them, the initial workflow is the workflow description file in the target task. The initial nodes are the sub-task nodes in the initial workflow. The node attributes include sub-task identification, task priority, required computing resources, estimated execution duration, etc. The initial directed edges are the edge structures used to describe the dependencies of each sub-task node, and the edge attributes include data dependency information, etc.
[0056] Specifically, receive the target task, analyze the initial workflow in the target task, and obtain each initial node, the node attributes corresponding to each initial node, each initial directed edge, and the edge attributes corresponding to each initial directed edge. Construct a directed acyclic graph (DAG) of the dependency relationships between subtasks with each initial node, the node attributes corresponding to each initial node, each initial directed edge, and the edge attributes corresponding to each initial directed edge, to obtain the initial dependency graph, which serves as the basis for subsequent task partitioning and workflow orchestration.
[0057] Exemplarily, based on the workflow description file, use a graph data structure (such as an adjacency list) to construct a directed acyclic graph (DAG) of the dependency relationships between tasks, that is, the initial dependency graph. During the construction process, parse the task dependency relationship information in the workflow description file, create an initial node for each subtask, and establish directed edges between the initial nodes according to the task dependency relationship information, that is, the initial directed edges. At the same time, add corresponding node attributes and edge attributes to each initial node and initial directed edge, such as subtask identifiers, task priorities, required computing resources, estimated execution durations, data dependency information, etc., for subsequent analysis and processing.
[0058] The granularity optimization module 123 is used to receive the task computing cost, task communication cost, and task switching cost, take the sum value of the task communication cost and the task switching cost as the process value, and take the quotient of the task computing cost and the process value as the target granularity.
[0059] Among them, the process value is the sum value of the task communication cost and the task switching cost.
[0060] Specifically, receive the task computing cost, task communication cost, and task switching cost determined by the task characteristic analysis module 121, sum the task communication cost and the task switching cost to obtain the process value, and divide the task computing cost by the process value to obtain the target granularity, without considering the real-time resource status here. The target granularity can be calculated according to the following formula:
[0061] G opt = C comp / (C comm + C overhead )
[0062] Among them, G opt is the target granularity, C comp is the task computing cost, C comm is the task communication cost, C overhead is the task switching cost.
[0063] The dependency graph update module 124 is configured to receive an initial dependency graph and a target granularity, determine an adjustment method for the initial dependency graph according to a preset granularity range and the target granularity, and adjust the initial dependency graph based on the adjustment method to obtain a target dependency graph.
[0064] Wherein, the preset granularity range is a preset granularity range for determining whether to adjust the initial dependency graph. The adjustment methods include a task merging method and a task splitting method. The target dependency graph is either the adjusted initial dependency graph or the initial dependency graph that does not need to be adjusted.
[0065] Specifically, receive the initial dependency graph constructed by the dependency graph construction module 122 and the target granularity determined by the granularity optimization module 123, and judge whether the target granularity is within the preset granularity range. If not, when the target granularity is greater than the preset granularity range, determine the adjustment method of the initial dependency graph as the task merging method; when the target granularity is less than the preset granularity range, determine the adjustment method of the initial dependency graph as the task splitting method. Furthermore, the initial dependency graph can be adjusted and optimized based on the determined adjustment method to obtain the target dependency graph. If the target granularity is within the preset granularity range, the initial dependency graph can be directly used as the target dependency graph.
[0066] Exemplarily, for a computationally intensive target task, if the target granularity is large and the dependencies between subtasks permit, multiple adjacent small subtasks can be merged into one large subtask to reduce task scheduling overhead. For an I / O intensive target task, if the target granularity is small, splitting a large subtask into multiple small subtasks can improve resource utilization. During the process of subtask merging and splitting, the initial dependency graph is updated to obtain the target dependency graph to ensure the correct task execution order.
[0067] Based on the above example, the dependency graph update module 124 is further configured to:
[0068] Initialize the number of adjustment times and judge the relationship between the target granularity and the preset granularity range;
[0069] In response to the target granularity being greater than the upper limit value of the preset granularity range, determine that the adjustment method of the initial dependency graph is the task merging method; determine whether the number of adjustments is equal to the preset number; in response to the number of adjustments being less than the preset number, based on the initial dependency graph, take two adjacent initial nodes with the smallest data dependency as the merging nodes, merge the merging nodes to obtain new initial nodes, update the initial dependency graph and the number of adjustments, determine the target granularity of the updated initial dependency graph, and return to execute the step of judging the relationship between the target granularity and the preset granularity range; in response to the number of adjustments being equal to the preset number, take the updated initial dependency graph as the target dependency graph;
[0070] In response to the target granularity being within the preset granularity range, determine that the initial dependency graph is the target dependency graph;
[0071] In response to the target granularity being less than the lower limit value of the preset granularity range, determine that the adjustment method of the initial dependency graph is the task splitting method; determine whether the number of adjustments is equal to the preset number; in response to the number of adjustments being less than the preset number, take the initial nodes in the initial dependency graph that contain parallel tasks as the splitting nodes, split the splitting nodes into at least two new initial nodes, update the initial dependency graph and the number of adjustments, determine the target granularity of the updated initial dependency graph, and return to execute the step of judging the relationship between the target granularity and the preset granularity range; in response to the number of adjustments being equal to the preset number, take the updated initial dependency graph as the target dependency graph.
[0072] Among them, the preset granularity range is an interval including the upper limit value and the lower limit value. The number of adjustments is the count of the current iteration of merging or splitting. The preset number is the upper limit value of the preset number of adjustments. The merging nodes are two adjacent initial nodes with the smallest data dependency. The splitting nodes are the initial nodes with parallel tasks.
[0073] Specifically, initialize the adjustment count, that is, set the adjustment count to zero. Determine the relationship between the target granularity and the preset granularity range. If the target granularity is greater than the upper limit value of the preset granularity range, it indicates that the target granularity is too large. It can be determined that the adjustment method of the initial dependency graph is the task merging method. Furthermore, determine whether the adjustment count is equal to the preset count. If the adjustment count is less than the preset count, analyze the initial dependency graph, and determine two adjacent initial nodes with the smallest data dependency from it. Take these two initial nodes as the merging nodes, merge these two merging nodes to obtain a new initial node, update the initial dependency graph, analyze and calculate the task cost corresponding to the updated initial dependency graph, and then update the target granularity. Increment the adjustment count by one. At this time, it is possible to return to the step of determining the relationship between the target granularity and the preset granularity range and enter the next iteration. If the adjustment count is equal to the preset count, it indicates that the adjustment upper limit has been reached, and directly use the updated initial dependency graph as the target dependency graph. If the target granularity is within the preset granularity range, it indicates that the target granularity already meets the requirements, and directly use the current initial dependency graph as the target dependency graph. If the target granularity is less than the lower limit value of the preset granularity range, it indicates that the target granularity is too small. Determine the adjustment method of the initial dependency graph is the task splitting method. Furthermore, determine whether the adjustment count is equal to the preset count. If the adjustment count is less than the preset count, check whether there is an initial node containing parallel tasks in the initial dependency graph. Take this initial node as the splitting node, and split the splitting node into at least two new initial nodes according to the number of parallel tasks. Update the initial dependency graph, analyze and calculate the task cost corresponding to the updated initial dependency graph, and then update the target granularity. Increment the adjustment count by one. At this time, it is possible to return to the step of determining the relationship between the target granularity and the preset granularity range and enter the next iteration. If the adjustment count is equal to the preset count, it indicates that the adjustment upper limit has been reached, and directly use the updated initial dependency graph as the target dependency graph. If the target granularity is within the preset granularity range, it indicates that the target granularity already meets the requirements, and directly use the current initial dependency graph as the target dependency graph.
[0074] Exemplarily, for a target task where the target granularity is greater than the upper limit value of the preset granularity range (such as 1.5, etc.), consider the adjustment method as the task merging method: identify adjacent subtasks in the initial dependency graph with small data dependencies, evaluate the target granularity after merging the subtask nodes. If it is within the preset granularity range, end the adjustment; if not, continue the adjustment, and the number of adjustments does not exceed the preset number. According to the adjustment, update the initial dependency graph, and update the nodes of the new subtasks and their dependencies into the initial dependency graph to obtain the target dependency graph. For a target task where the target granularity is less than the lower limit value of the preset granularity range (such as 0.5, etc.), consider the adjustment method as the task splitting method: analyze the internal structure of the task in combination with the initial dependency graph, identify the parallelizable parts, generate a splitting strategy, including splitting nodes and data flow transfer methods, evaluate the target granularity after splitting. If it is within the preset granularity range, end the adjustment; if not, continue the adjustment, and the number of adjustments does not exceed the preset number. According to the adjustment, update the initial dependency graph, replace the splitting nodes with the initial nodes of multiple subtasks, and update their dependencies.
[0075] The critical path identification module 125 is configured to receive the target dependency graph, determine the earliest start time and the latest start time corresponding to each target node in the target dependency graph according to the target dependency graph, determine the time float corresponding to each target node according to the earliest start time and the latest start time corresponding to each target node, and determine the critical path and the path weight corresponding to each target node according to the time float corresponding to each target node, and add the critical path and the path weight corresponding to each target node to the target dependency graph.
[0076] Among them, the target nodes are the subtask nodes in the target dependency graph. The earliest start time is the latest completion time of all predecessor subtasks. The latest start time is the start time of the latest subtask that does not affect the completion time of the workflow. The time float is the difference between the latest start time and the earliest start time. The critical path is the path composed of target nodes with a time float of 0. The path weight is used to describe the importance of the path.
[0077] Specifically, receive the target dependency graph in the dependency graph update module 124, analyze the target dependency graph, and calculate the earliest start time and the latest start time corresponding to each target node in the target dependency graph. For each target node, use the difference between the latest start time and the earliest start time corresponding to the target node as the time float corresponding to the target node. Use the path composed of target nodes with a time float of 0 as the critical path, and analyze each time float to calculate and determine the path weight corresponding to each target node. Moreover, add the critical path and the path weight corresponding to each target node to the target dependency graph to provide information for subsequent workflow orchestration.
[0078] Based on the above example, the critical path identification module 125 is further configured to:
[0079] For each target node, take the difference between the latest start time and the earliest start time corresponding to the target node as the time float corresponding to the target node;
[0080] Construct a critical path according to the target nodes with a time float of 0;
[0081] Determine the maximum value of the time float according to the time floats corresponding to the target nodes;
[0082] For each target node, determine the path weight corresponding to the target node according to the preset maximum weight, the maximum value of the time float, and the time float corresponding to the target node.
[0083] Wherein, the maximum value of the time float is the maximum value among the time floats corresponding to the target nodes. The preset maximum weight is the maximum value that the preset path weight can take, such as 10, etc. For example, if the value range of the path weight is 0 - 10, then the preset maximum weight is 10. The larger the path weight, the more important the subtask corresponding to the target node.
[0084] Specifically, for each target node, take the difference between the latest start time and the earliest start time corresponding to the target node as the time float corresponding to the target node. Construct a critical path from the target nodes with a time float of 0. Take the maximum value of the time floats corresponding to the target nodes as the maximum value of the time float. For each target node, establish a corresponding relationship between the preset maximum weight and the maximum value of the time float. Furthermore, convert the time float corresponding to the target node into the corresponding weight, which is the path weight corresponding to the target node.
[0085] Exemplarily, the following attributes are calculated for each target node: EST (Earliest Start Time): The latest completion time of all predecessor subtasks. EFT (Earliest Finish Time): EST plus the task execution time. LST (Latest Start Time): The start time of the latest subtask that does not affect the workflow completion time. LFT (Latest Finish Time): LST plus the execution time of the subtask. Specifically, it can be calculated in the following way: First, calculate EST and EFT (forward traversal). For the entry target nodes (target nodes without predecessor subtasks): EST = 0; For other target nodes: EST = max{EFT (of the target nodes of all predecessor subtasks)}; For all target nodes: EFT = EST + the estimated execution duration of the target node. Second, calculate LFT and LST (backward traversal). For the exit target nodes (target nodes without successor subtasks): LFT = the earliest completion time of the entire workflow (the maximum value of EFT among all exit target nodes); For other target nodes: LFT = min{LST (of the target nodes of all successor subtasks)}; For all target nodes: LST = LFT – the estimated execution duration of the target node. Calculate the time float slack = LST – EST for each target node. If the time float is 0, that is, the target node with slack = 0 is the node on the critical path. The path weight of the target node can be calculated by the following formula:
[0086]
[0087] Wherein, is the maximum value of the time float, is the time float of this target node, is the preset maximum weight, is the path weight of this target node. If slack = 0, it is the node on the critical path, and the path weight can be considered as ; If slack = slack max , it means that this task is the least critical and the path weight is 0.
[0088] Finally, mark the critical path and each path weight in the target dependency graph, and the subtask sequence of the critical path can be output.
[0089] As Figure 2 shown, the workflow orchestration subsystem 130 includes: a task priority determination module 131, a workflow construction module 132, a resource matching and allocation module 133, and an execution plan generation module 134.
[0090] The task priority determination module 131 is used to obtain the target dependency graph and the initial workflow in the target task, and determine the specified priority corresponding to each target node in the target dependency graph according to the initial workflow; for each target node, determine the task priority of the target node according to the estimated execution duration in the node attributes of the target node, the specified priority corresponding to the target node, and the path weight corresponding to the target node.
[0091] Among them, the specified priority is the priority assigned by the user in the initial workflow. The estimated execution duration is the duration expected for executing the subtasks of the target node. The task priority is used to describe the order of execution of the subtasks corresponding to each target node.
[0092] Specifically, receive the target dependency graph in the dependency graph update module 124 and the initial workflow in the target task. Since the initial workflow contains the specified priority of each subtask, the specified priority corresponding to each target node in the target dependency graph can be determined. For each target node, calculate by integrating the estimated execution duration in the node attributes of the target node, the specified priority corresponding to the target node, and the path weight corresponding to the target node, and determine the task priority of the target node, so as to comprehensively determine the execution priority of the subtasks corresponding to each target node based on multiple factors, that is, the task priority, providing a decision basis for subsequent resource allocation.
[0093] Exemplarily, the task priority can be calculated through the following formula:
[0094] Task priority = α × specified priority + β × path weight + γ × estimated execution duration score
[0095] Among them, the specified priority is from the workflow description file (initial workload) and can be normalized to 0 - 10. The estimated execution duration score is the estimated execution duration (unit: second) mapped to 0 - 10, which can be a linear normalization value based on the maximum / minimum duration. α, β, γ are preset coefficients, which can be configured according to the system policy. For example, the preset coefficient β corresponding to the path weight is relatively high, and it can be α = 0.2, β = 0.5, γ = 0.3.
[0096] Estimated execution duration score = 10×(1 - (t - t min ) / (t max - t min ))
[0097] Among them, t is the estimated execution duration of the subtask corresponding to the target node, t min is the shortest estimated execution duration (such as 10 seconds) among all subtasks to be executed, t maxis the longest estimated execution duration among all tasks to be executed (e.g., 360 seconds). The range of the estimated execution duration score: the shorter the subtask, the higher the score (close to the preset maximum value of 10), and the longer the subtask, the lower the score (close to the preset minimum value of 0), that is, the shorter the subtask, the higher the priority.
[0098] Optionally, dynamic priority adjustment can also be performed. According to the execution progress of the target task and the changes in the resource platform information, the estimated execution duration and the estimated execution duration score are recalculated regularly to calculate the new task priority.
[0099] The workflow construction module 132 is used to construct a target workflow according to the target dependency graph and the task priorities of each target node.
[0100] Among them, the target workflow is a workflow standardized according to the target dependency graph and the task priorities of each target node in accordance with the standard workflow description language.
[0101] Specifically, the standard workflow description language (such as JSON, YAML, etc.) is used to describe the workflow corresponding to the target dependency graph and the task priorities of each target node, which supports expressing task dependencies, resource requirements, data dependencies, etc.
[0102] Exemplarily, the standard workflow description language is defined and standardized to ensure the accuracy and consistency of the workflow description. When describing a scientific computing workflow, nested structures are supported, which can express complex task dependencies, resource requirements, data dependencies, and task priorities in detail. For example, for a complex task containing multiple subtasks, the execution order, dependencies, and respective resource requirements of the subtasks can be defined in the workflow description.
[0103] The resource matching and allocation module 133 is used to obtain the task characteristics and the task priorities of each target node; obtain the resource catalog, and determine the real-time resource information of each candidate resource platform according to the resource catalog; determine the target execution platform corresponding to each target node according to the task characteristics, the real-time resource information of each candidate resource platform, and the task priorities of each target node.
[0104] Among them, the candidate resource platforms are the resource platforms recorded in the resource catalog. The real-time resource information is the real-time information of the candidate resource platforms. The target execution platform is the candidate resource platform allocated for the target node.
[0105] Specifically, obtain the task characteristics analyzed by the task characteristic analysis module 121 and the task priorities of each target node determined by the task priority determination module 131, and obtain the resource directory constructed by the resource directory acquisition module 112. The real-time resource information of each candidate resource platform can be obtained from the resource directory. By comprehensively analyzing the task characteristics, the real-time resource information of each candidate resource platform, and the task priorities of each target node, a corresponding target execution platform can be allocated for each target node, so as to allocate the most suitable resource platform for each subtask in combination with the task resource requirements and the current available resource status.
[0106] Exemplarily, the real-time resource information in the resource directory obtained from the resource abstraction subsystem 110 includes the resource types, specifications, available status, etc. of each resource platform. According to the task resource requirements in the target dependency graph provided by the task granularity division subsystem 120, search for eligible resources. The matching degree between task characteristics and resource characteristics can be considered: allocate high-performance CPU / GPU resources preferentially for compute-intensive tasks; allocate high-bandwidth storage resources preferentially for I / O-intensive tasks. And resource reservation and allocation can be performed: reserve resources for subtasks with high task priorities, and submit resource allocation requests to the target execution platform through the resource abstraction subsystem 110, and maintain task queues for different resource pools (high task priorities first).
[0107] The execution plan generation module 134 is used to obtain the target execution platform and the target workflow corresponding to each target node, determine the initial execution plan; after each target execution platform in the initial execution plan receives the target data, determine the corresponding adaptation layer for each target execution platform, and submit the subtasks corresponding to the target nodes to each target execution platform based on each adaptation layer.
[0108] Among them, the initial execution plan includes a task allocation plan and an execution schedule.
[0109] Specifically, by identifying the target dependency graph and each task priority corresponding to the target workflow, and obtaining the resource allocation result, that is, the target execution platform corresponding to each target node, an initial execution plan can be generated. After each target execution platform in the initial execution plan receives the target data, it can be determined that the corresponding target execution platform can execute the corresponding task. Therefore, determine the corresponding adaptation layer for each target execution platform, and submit the subtasks corresponding to the target nodes to each target execution platform based on each adaptation layer, so that each target execution platform can execute the corresponding task.
[0110] Exemplarily, the target dependency graph is received, and combined with the resource matching and allocation results, the corresponding target execution platform and the expected execution time are determined for each task. Furthermore, a multi-objective optimization algorithm can be used to generate an initial execution plan. Taking the multi-objective optimization using NSGA-II (Nondominated Sorting Genetic Algorithm II) as an example, the specific steps are as follows: First, an objective function is constructed with the goals of minimizing the completion time, maximizing the resource utilization rate, and minimizing the data transmission cost. Then, an initial population is constructed, that is, multiple feasible scheduling schemes are randomly generated, and then fitness evaluation is carried out. The performance of each scheme on each objective is calculated respectively, non-dominated sorting is performed, the schemes are divided into different Pareto front levels, crowding degree calculation is carried out to maintain the diversity of the solution set, and a new generation of population is generated through selection, crossover, and mutation, and iterative optimization is carried out, that is, the above steps are repeated until convergence or the maximum number of iterations is reached. From the finally obtained non-dominated solution set, the most suitable scheduling scheme is selected according to the current system state and user preferences as the initial execution plan. This multi-objective optimization method can take into account multiple performance indicators at the same time and generate a balanced and efficient scheduling scheme.
[0111] Based on the above example, the workflow orchestration subsystem 130 further includes: a dynamic scheduling module 135 and a workflow mutation module 136.
[0112] The dynamic scheduling module 135 is used to determine, for each target execution platform, whether the target execution platform meets at least one rescheduling condition. If it meets, the target execution platform and the rescheduling condition triggered by the target execution platform are fed back to the execution plan generation module, so that the execution plan generation module updates the initial execution plan.
[0113] Among them, the rescheduling condition is the condition that triggers the update of the initial execution plan. For example: the task execution time deviates significantly from the expectation, the resource availability changes, the network condition changes significantly, the user adjusts the task priority, etc.
[0114] Specifically, for each target execution platform, it is monitored, and it is judged in real time whether the target execution platform meets at least one rescheduling condition. If it meets, it means that rescheduling is required. Therefore, the target execution platform and the rescheduling condition triggered by the target execution platform are fed back to the execution plan generation module, so that the execution plan generation module updates the initial execution plan, which is convenient for subsequent execution according to the updated initial execution plan.
[0115] Exemplarily, the execution status and resource status of tasks are monitored in real time. By deploying monitoring agents on task execution nodes to collect information such as task execution time and resource usage, logs and historical data can be recorded. When the target execution platform meets at least one rescheduling condition, for example, when the task execution time significantly deviates from the expectation (such as exceeding 20% of the expected time), the resource availability changes (such as the CPU usage rate exceeds 80% or the remaining memory is less than 20%, etc.), the network condition changes significantly (such as the network bandwidth is lower than the set threshold), or the user adjusts the task priority, rescheduling is triggered. During rescheduling, an event-driven scheduling algorithm is adopted. According to the type of the triggering event and the current system state (the target execution platform and the rescheduling conditions triggered by the target execution platform), the tasks with problems are recalculated based on a multi-objective optimization algorithm, and the dynamically adjusted content includes the target execution platform used for task execution and the expected execution time.
[0116] The workflow mutation module 136 is used to, during the execution of the target task, in response to a change in the algorithm of at least one target node in the target dependency graph, update the target execution platform of the corresponding target node according to the changed algorithm and the resource catalog, and feedback the algorithm and the target execution platform of the updated target node to the execution plan generation module, so that the execution plan generation module updates the initial execution plan; in response to a change in the configuration parameters of at least one target node in the target dependency graph, send new configuration parameters to the target execution platform corresponding to the target node whose configuration parameters have changed.
[0117] Specifically, during the execution of the target task according to the initial execution plan, if the algorithm of at least one target node in the target dependency graph changes, it is necessary to combine the changed algorithm and the resource catalog to reassign the corresponding target execution platform for the corresponding target node, and it is necessary to feedback the algorithm and the target execution platform of the updated target node to the execution plan generation module, so that the execution plan generation module updates the initial execution plan. During the execution of the target task according to the initial execution plan, if the configuration parameters of at least one target node in the target dependency graph change, send new configuration parameters to the target execution platform corresponding to the target node whose configuration parameters have changed, so that the target execution platform is reconfigured to facilitate the execution of the subtasks corresponding to the target node according to the new configuration parameters. The workflow mutation module 136 supports dynamically adjusting the workflow structure during task execution, such as replacing the algorithm implementation for a certain task, adjusting parameter configurations, etc.
[0118] Exemplarily, during the execution according to the initial execution plan corresponding to the target workflow, the structure of the workflow can be dynamically adjusted by modifying the constructed target dependency graph. For example, for a task to implement a replacement algorithm, according to the resource requirements and interface specifications of the new algorithm, it is necessary to re-allocate the corresponding target execution platform for this task and modify the execution code of this task; when adjusting the parameter configuration, the corresponding parameter values can be directly modified on the target nodes of the constructed target dependency graph, and relevant task nodes and their corresponding target execution platforms can be notified to reload the configuration parameters.
[0119] As Figure 2 shown, the data preheating management subsystem 140 includes: a target data determination module 141, a transmission priority determination module 142, a multipath transmission module 143, and a data recording module 144.
[0120] The target data determination module 141 is used to obtain the target dependency graph and the initial execution plan, determine the cross-platform subtasks in the initial execution plan, and use the data required by the cross-platform subtasks as the target data.
[0121] Specifically, according to the target dependency graph provided by the task granularity division subsystem 120 and the initial execution plan in the workflow orchestration subsystem 130, analyze the data dependency relationships between tasks in the target workflow. Among them, the task relationships and data sizes are in the target dependency graph, and the target execution platforms corresponding to the tasks are in the initial execution plan. Based on this, cross-platform subtasks can be identified, and the data that needs to be transmitted across platforms can be determined as the target data.
[0122] The transmission priority determination module 142 is used to determine the transmission priority of the target data for each group of target data according to the preset weight, the data size of the target data, the task priority of the cross-platform subtask corresponding to the target data, and the network bandwidth between the two target execution platforms corresponding to the target data.
[0123] Among them, the preset weight is the weight corresponding to each target data obtained by training according to the data size, task priority, network bandwidth, and corresponding real priority of the historical records of each target data.
[0124] Specifically, the network bandwidth between the two target execution platforms can be obtained through the resource abstraction subsystem 110, and moreover, the data size and task priority of each target data can be determined according to the initial execution plan, etc. Through the data size, network conditions, task execution plan, and their corresponding preset weights, weighted summation can be performed to obtain the transmission priority of the target data.
[0125] Exemplarily, the transmission priority of the target data can be determined by using a weight score based on the access frequency, and the calculation method is as follows:
[0126] P = α × data size + β × network bandwidth + γ × task priority
[0127] Wherein, P is the transmission priority of the target data, and α, β, and γ are the preset weights corresponding to the data size, network bandwidth, and task priority respectively, which are obtained by training the data size, task priority, network bandwidth, and corresponding true priority of each target data in the historical records. Exemplarily, the result range of the transmission priority is 0.0~1.0.
[0128] The multi-path transmission module 143 is configured to obtain network topology information, and based on the transmission priorities of the respective target data and the network topology information, send the corresponding target data to each target execution platform in an incremental synchronization manner.
[0129] Specifically, the network topology information can be obtained through the resource abstraction subsystem 110. According to the transmission priorities of the respective target data and the network topology information, at least one transmission path and transmission order of each target data are analyzed, and in an incremental synchronization manner, the data is transmitted to the corresponding target execution platform according to the transmission path and transmission order corresponding to each target data. Multiple transmission paths can be used to transmit the target data in parallel to improve the transmission efficiency. Moreover, when transmitting data, in an incremental synchronization manner, the data difference algorithm is used to only transmit the changed part of the data, which can reduce the amount of transmitted data.
[0130] Exemplarily, according to the network topology information, a network routing algorithm (such as the Dijkstra algorithm) is used to calculate multiple transmission paths. During the data transmission process, multi-threading technology is adopted to transmit data in parallel, and each thread is responsible for transmitting the data of one path. To avoid network congestion, a flow control mechanism is introduced to dynamically adjust the data sending rate of each thread according to the network bandwidth utilization rate. At the same time, a data verification and retransmission mechanism is adopted to ensure the integrity and accuracy of the data transmission. The version information of the data can be recorded, and when the data is updated each time, a data change log is generated to record the modified content of the data. When transmitting data, according to the version information and the data change log, only the changed part of the data is transmitted. For example, the data difference algorithm is used to calculate the data change to improve the efficiency of incremental synchronization.
[0131] The data recording module 144 is configured to record the data size, task priority, network bandwidth, and true priority of each transmitted target data during the transmission process of each target data, and update the preset weights according to the data size, task priority, network bandwidth, and true priority of each transmitted target data.
[0132] Wherein, the true priority is the priority actually used when each target data is transmitted.
[0133] Specifically, for each target data, when the data transmission process of the target data to the corresponding target execution platform is completed, record the data size, task priority, network bandwidth, and real priority of the target data. Furthermore, periodically use the data size, task priority, network bandwidth, and real priority of each stored target data to train and adjust the preset weights, and apply the trained and adjusted preset weights to the transmission priority determination module 142.
[0134] Exemplarily, record the data size (such as MB), task priority (0 - 10), network bandwidth (such as Mbps), and real priority (the actual transmission priority order adopted during scheduling, which can be represented by a decimal, with a value range of 0.0 - 1.0, and the larger the value, the higher the priority) of the target data as historical data. Use the historical data as sample data and adopt the method of regression analysis to fit each preset weight (α, β, γ). For example, the least squares method can be used to fit each preset weight. Before regression, first normalize each input variable to a unified scale (0 - 1), so that α + β + γ ≈ 1 after training. Before the initial operation, a batch of sample data (at least dozens to hundreds of sample data) can be prepared first, calculate each preset weight first, and then retrain each preset weight periodically after running. It should be noted that before training the preset weights, normalize the sample data.
[0135] Based on the above example, the cross - domain workflow scheduling system for dynamic resource orchestration further includes: a cross - domain fault - tolerance subsystem 150, as Figure 2 shown, the cross - domain fault - tolerance subsystem 150 includes: a fault detection and replacement module 151 and a critical data preservation module 152.
[0136] The fault detection and replacement module 151 is used to detect faults in each resource platform during the execution of the target task. When at least one of the target execution platforms corresponding to the target task fails, determine the replacement execution platforms corresponding to each failed target execution platform according to the information of each resource platform, and send each failed target execution platform and the corresponding replacement execution platform to the workflow orchestration subsystem to update the initial execution plan.
[0137] Among them, the replacement execution platform is a platform that can replace the target execution platform to execute the corresponding subtask.
[0138] Specifically, during the execution of the target task, a fault detection mechanism can be configured for each resource platform to perform real-time fault detection, facilitating the timely identification of platform or network faults. When at least one of the target execution platforms corresponding to the target task fails, it is necessary to replace the resource platform. Combining the information of each resource platform in the resource abstraction subsystem 110, find candidate resource platforms similar to each failed target execution platform as the corresponding alternative execution platforms. And send each failed target execution platform and the corresponding alternative execution platform to the workflow orchestration subsystem 130, so that the workflow orchestration subsystem 130 can replace each failed target execution platform with the corresponding alternative execution platform and update the initial execution plan.
[0139] Exemplarily, a distributed heartbeat detection mechanism can be deployed, and heartbeat detection agents are deployed on the nodes of each resource platform and the nodes of the network link. The heartbeat detection agent can send heartbeat messages to other nodes at regular intervals (such as every 10 seconds), and at the same time, receive heartbeat messages from other nodes. If the heartbeat message of a certain node is not received within a certain period (such as 30 seconds), it is determined that the node has failed. Furthermore, the distributed hash table technology can be used to manage the information of the heartbeat detection agent to ensure the reliability and scalability of the heartbeat detection and facilitate the timely discovery of faults in the nodes of the resource platform and the nodes of the network link. When the target execution platform fails and is unavailable, find and obtain a list of available alternative execution platforms from the resource abstraction subsystem 110. Furthermore, according to the resource requirements of the corresponding subtasks and the resource platform information of each candidate resource platform in the list of alternative execution platforms, a resource matching algorithm (such as the Hungarian algorithm) can be used to select the most suitable alternative execution platform. Then, according to the requirements of the task submission system of the alternative execution platform, regenerate the task submission script, and make the workflow orchestration subsystem 130 update the initial execution plan. If the corresponding target execution platform is marked in the target dependency graph, the task granularity division subsystem 120 can be made to update the target dependency graph.
[0140] The key data saving module 152 is used to incrementally save the data of each target node on the critical path in the target dependency graph according to the second preset period.
[0141] Among them, the second preset period is a preset period for incrementally saving the data of each target node on the critical path, which can be set according to requirements.
[0142] Specifically, according to the second preset period, periodically save the intermediate state during the execution of the subtasks corresponding to each target node on the critical path to support task recovery.
[0143] Exemplarily, the intermediate states of the subtasks corresponding to each target node on the critical path are periodically saved during the execution process, and the information is stored using a file system or a database. During the task execution process, every second preset period (such as 10 minutes) or after the subtasks corresponding to each target node on the critical path are completed, the intermediate state of the task (such as memory data, file pointers, task execution progress, etc.) is saved to a checkpoint file or a database table. To improve the saving efficiency, an incremental saving method is adopted, and only the data that has changed compared with the previous time is saved.
[0144] Based on the above example, the cross-domain fault-tolerant subsystem 150 further includes: a degraded execution module 153, configured to determine the task level corresponding to each target node according to the critical path in the target dependency graph and the task priorities corresponding to each target node; for each target node, in response to insufficient resources in the target execution platform corresponding to the target node, the degraded execution module 153 determines a degradation strategy according to the task level of the target node and the resource state in the target execution platform, and controls the degraded execution of the subtask corresponding to the target node in the target execution platform based on the degradation strategy.
[0145] Among them, the task level is used to describe the importance of the subtasks corresponding to each target node, and may include core, important, secondary, etc. The resource state includes memory state, CPU usage rate, system load, etc.
[0146] Specifically, the task levels corresponding to the target nodes in the critical path in the target dependency graph are determined as core, and the remaining target nodes are classified according to the corresponding task priorities to obtain the task levels corresponding to each target node. For each target node, when there are insufficient resources in the target execution platform corresponding to the target node, it is combined with the task level of the target node to determine whether degradation processing can be performed. If so, considering the resource state in the corresponding target execution platform, a degradation strategy is determined to control the use of the function degradation strategy to execute the subtask corresponding to the target node in the target execution platform. When the resources are severely insufficient, a degradation strategy is formulated according to the task level of the target node and the resource state in the target execution platform to support degraded execution and ensure the core functions. Degradation can reduce the calculation accuracy or reduce the calculation steps to reduce resource consumption. During the degraded execution process, the user is notified that the workflow is being degraded, and a degradation operation log is recorded for subsequent analysis and recovery.
[0147] Exemplarily, if the value range of task priority is 0 - 10, and the larger the value, the higher the priority. Then, the core tasks are those with task priority >= 7 or tasks of target nodes on the critical path, indicating that they must be completed and cannot be downgraded; important tasks are those with task priority > 3 and < 7 and not on the target nodes of the critical path, indicating that they can be moderately downgraded to ensure basic functions; secondary tasks are those with task priority <= 3, indicating that they can be significantly downgraded or postponed. The downgrading strategy can include: reducing algorithm accuracy (reducing the number of iterations, convergence accuracy, or sampling rate), reducing features (temporarily closing non-core functional modules), reducing data scale (reducing the amount of processed data, such as reducing resolution, sampling), reducing resource requirements (adjusting algorithm parameters, reducing memory and CPU requirements), etc. For example, a decision-making mechanism can be formulated: Level 1 (mild): CPU usage > 80% or memory occupancy > 70%, and a downgrading strategy of data sampling and reducing visualization frequency will be adopted; Level 2 (moderate): available memory < 2GB or obvious I / O congestion, and a downgrading strategy of data sampling, reducing visualization frequency, reducing calculation accuracy, and restricting modules will be adopted; Level 3 (severe): system load > 95% or the number of scheduling failures > N (preset number of failures), only core tasks will be retained, and all non-critical processes will be closed. This downgrading mechanism can ensure that the most core scientific computing tasks can still be completed under resource constraints, and the availability of basic functions is guaranteed.
[0148] The present invention has the following technical effects: Through the resource abstraction subsystem, information of each resource platform and network topology information are collected, a resource directory is constructed, and each adaptation layer is constructed. Through the task granularity division subsystem, the target granularity is calculated, and the initial dependency graph is converted into a target dependency graph. Through the workflow orchestration subsystem, combined with task characteristics, resource directory, and target dependency graph, an initial execution plan is determined. Through the data preheating management subsystem, each target data is pre-transferred across platforms in advance. After each target execution platform in the initial execution plan receives the target data through the workflow orchestration subsystem, the adaptation layer corresponding to each target execution platform is determined, and corresponding subtasks are submitted to each target execution platform to schedule the execution of target tasks, realizing unified management of resources on different platforms, reducing the overhead of data transmission and task scheduling, and improving the stability of cross-domain task completion.
[0149] It should be noted that the terms used in the present invention are only for describing specific embodiments and do not limit the scope of the present application. Unless the context clearly indicates an exception, words such as "a", "an", "one" and / or "the" are not specifically singular and may also include the plural. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such a process, method or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method or device comprising the said element.
[0150] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A cross-domain workflow scheduling system for dynamic resource orchestration, characterized in that Including: A resource abstraction subsystem, which is used to collect information of each resource platform and network topology information, construct a resource catalog according to the information of each resource platform, and construct an adaptation layer between each task submission system type and each resource platform type; A task granularity division subsystem, which is used to receive a target task, determine the task characteristics, each task cost and the initial dependency graph corresponding to the target task, determine the target granularity according to each task cost, and optimize the initial dependency graph according to the target granularity to obtain a target dependency graph; A workflow orchestration subsystem, which is used to obtain the task characteristics, the resource catalog and the target dependency graph, determine an initial execution plan according to the target dependency graph and the resource catalog, and after each target execution platform in the initial execution plan receives target data, determine the adaptation layer corresponding to each target execution platform, and submit corresponding subtasks to each target execution platform based on each adaptation layer, so that each target execution platform jointly executes the target task; A data preheating management subsystem, which is used to obtain the network topology information, the target dependency graph and the initial execution plan, determine the transmission priority of cross-platform subtasks in the initial execution plan, and send corresponding target data to each target execution platform according to each transmission priority and the network topology information.
2. The system according to claim 1, wherein The resource abstraction subsystem includes: A resource proxy module, which is used to collect information of each resource platform based on a first preset period; A resource catalog module, which is used to receive information of each resource platform and construct a resource catalog according to the information of each resource platform; the resource catalog module includes a resource query interface and a resource allocation interface; A resource adaptation module, which is used to construct adaptation groups according to each task submission system type and each resource platform type, and pre-construct an adaptation layer corresponding to each adaptation group for each adaptation group; A network topology module, which is used to obtain a network topology structure, collect network status information corresponding to each network node in the network topology structure, and determine network topology information according to the network topology structure and the network status information corresponding to each network node.
3. The system according to claim 1, wherein The task granularity division subsystem includes: A task characteristic analysis module, which is used to receive a target task, determine the quantity of each type of operation according to the task code file in the target task, determine task characteristics according to the quantity of each type of operation, determine the task calculation cost in the task cost according to the task code file, determine the task communication cost in the task cost according to the data description file in the target task, and determine the task switching cost in the task cost according to the historical execution time data of similar tasks corresponding to the target task; A dependency graph construction module, which is used to receive the target task, determine each initial node, the node attribute corresponding to each initial node, each initial directed edge and the edge attribute corresponding to each initial directed edge according to the initial workflow in the target task, and construct an initial dependency graph according to each initial node, the node attribute corresponding to each initial node, each initial directed edge and the edge attribute corresponding to each initial directed edge. A granularity optimization module, which is configured to receive the task computing cost, the task communication cost, and the task switching cost, use the sum of the task communication cost and the task switching cost as a process value, and use the quotient of the task computing cost and the process value as the target granularity; A dependency graph update module, which is configured to receive the initial dependency graph and the target granularity, determine an adjustment method for the initial dependency graph according to a preset granularity range and the target granularity, and adjust the initial dependency graph based on the adjustment method to obtain a target dependency graph; wherein, the adjustment method includes a task merging method and a task splitting method; A critical path identification module, which is configured to receive the target dependency graph, determine the earliest start time and the latest start time corresponding to each target node in the target dependency graph according to the target dependency graph, determine the time float corresponding to each target node according to the earliest start time and the latest start time corresponding to each target node, and determine the critical path and the path weight corresponding to each target node according to the time float corresponding to each target node, and add the critical path and the path weight corresponding to each target node to the target dependency graph.
4. The system according to claim 3, wherein The dependency graph update module is further configured to: Initialize the adjustment times and judge the relationship between the target granularity and the preset granularity range; In response to the target granularity being greater than the upper limit value of the preset granularity range, determine the adjustment method for the initial dependency graph as the task merging method; judge whether the adjustment times are equal to the preset times; In response to the adjustment times being less than the preset times, take two adjacent initial nodes with the least data dependency in the initial dependency graph as merging nodes, merge the merging nodes to obtain new initial nodes, update the initial dependency graph and the adjustment times, determine the target granularity of the updated initial dependency graph, and return to execute the step of judging the relationship between the target granularity and the preset granularity range; In response to the adjustment times being equal to the preset times, use the updated initial dependency graph as the target dependency graph; In response to the target granularity being within the preset granularity range, determine the initial dependency graph as the target dependency graph; In response to the target granularity being less than the lower limit value of the preset granularity range, determine the adjustment method for the initial dependency graph as the task splitting method; Judge whether the adjustment times are equal to the preset times; In response to the adjustment times being less than the preset times, take the initial node containing parallel tasks in the initial dependency graph as a splitting node, split the splitting node into at least two new initial nodes, update the initial dependency graph and the adjustment times, determine the target granularity of the updated initial dependency graph, and return to execute the step of judging the relationship between the target granularity and the preset granularity range; In response to the adjustment times being equal to the preset times, use the updated initial dependency graph as the target dependency graph.
5. The system according to claim 3, wherein The critical path identification module is further configured to: For each target node, the difference between the latest start time and the earliest start time corresponding to the target node is used as the time float corresponding to the target node; Based on the target nodes with a time float of 0, a critical path is constructed; Based on the time floats corresponding to the respective target nodes, the maximum value of the time floats is determined; For each target node, based on a preset maximum weight, the maximum value of the time floats, and the time float corresponding to the target node, the path weight corresponding to the target node is determined.
6. The system according to claim 1, wherein The workflow orchestration subsystem includes: A task priority determination module, configured to obtain a target dependency graph and an initial workflow in a target task, and determine the specified priority corresponding to each target node in the target dependency graph according to the initial workflow; for each target node, based on the estimated execution duration in the node attributes of the target node, the specified priority corresponding to the target node, and the path weight corresponding to the target node, determine the task priority of the target node; A workflow construction module, configured to construct a target workflow according to the target dependency graph and the task priorities of the respective target nodes; A resource matching and allocation module, configured to obtain the task characteristics and the task priorities of the respective target nodes; obtain the resource catalog, and determine the real-time resource information of each candidate resource platform according to the resource catalog; according to the task characteristics, the real-time resource information of each candidate resource platform, and the task priorities of the respective target nodes, determine the target execution platform corresponding to each target node; An execution plan generation module, configured to obtain the target execution platform corresponding to each target node and the target workflow, and determine an initial execution plan; after the target data is received by each target execution platform in the initial execution plan, determine the adaptation layer corresponding to each target execution platform, and submit the subtask corresponding to the target node corresponding to each target execution platform to each target execution platform based on each adaptation layer.
7. The system according to claim 6, wherein The workflow orchestration subsystem further includes: A dynamic scheduling module, configured to, for each target execution platform, determine whether the target execution platform meets at least one rescheduling condition, and if so, feedback the target execution platform and the rescheduling condition triggered by the target execution platform to the execution plan generation module, so that the execution plan generation module updates the initial execution plan; A workflow mutation module, configured to, during the execution of the target task, in response to a change in the algorithm of at least one target node in the target dependency graph, update the target execution platform corresponding to the corresponding target node according to the changed algorithm and the resource catalog, and feedback the algorithm and the target execution platform of the updated target node to the execution plan generation module, so that the execution plan generation module updates the initial execution plan; in response to a change in the configuration parameter of at least one target node in the target dependency graph, send new configuration parameters to the target execution platform corresponding to the target node with the changed configuration parameter.
8. The system according to claim 1, wherein The data preheating management subsystem includes: A target data determination module, configured to obtain the target dependency graph and the initial execution plan, determine cross-platform subtasks in the initial execution plan, and use the data required by the cross-platform subtasks as target data; A transmission priority determination module, configured to, for each group of target data, determine the transmission priority of the target data according to a preset weight, the data size of the target data, the task priority of the cross-platform subtask corresponding to the target data, and the network bandwidth between two target execution platforms corresponding to the target data; A multipath transmission module, configured to obtain the network topology information, and based on the transmission priority of each target data and the network topology information, send the corresponding target data to each target execution platform in an incremental synchronization manner; A data recording module, configured to record the data size, task priority, network bandwidth, and actual priority of each transmitted target data during the transmission of each target data, and update the preset weight according to the data size, task priority, network bandwidth, and actual priority of each transmitted target data; 9. The system according to claim 1, wherein It further includes a cross-domain fault tolerance subsystem, and the cross-domain fault tolerance subsystem includes: A fault detection and replacement module, configured to perform fault detection on each resource platform during the execution of the target task. When at least one of the target execution platforms corresponding to the target task fails, determine the replacement execution platform corresponding to each failed target execution platform according to each resource platform information, and send each failed target execution platform and the corresponding replacement execution platform to the workflow orchestration subsystem to update the initial execution plan; A critical data saving module, configured to incrementally save the data of each target node on the critical path in the target dependency graph according to a second preset period; 10. The system according to claim 9, characterized in that, The cross-domain fault tolerance subsystem further includes: A degraded execution module, configured to determine the task level corresponding to each target node according to the critical path in the target dependency graph and the task priority corresponding to each target node; for each target node, in response to insufficient resources in the target execution platform corresponding to the target node, determine a degradation strategy according to the task level of the target node and the resource status in the target execution platform, and control the degraded execution of the subtask corresponding to the target node in the target execution platform based on the degradation strategy.
Citation Information
Patent Citations
Data analysis method and system based on visual modeling and job arrangement scheduling
CN118819903A
Intelligent operation and maintenance method and system for managing different types of computing resources
CN119645614A