Workflow parallel computing execution method and system, terminal equipment and storage medium
Through task dependency graph construction and dynamic resource allocation optimization task decomposition, combined with lightweight communication and fault tolerance mechanism, the problems of uneven load, high communication overhead and insufficient fault tolerance in parallel computing are solved, and efficient and reliable parallel computing execution is achieved.
Patent Information
- Application Number
- CN202510150497.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-07-04
AI Technical Summary
The existing parallel computing technology has insufficient flexibility and accuracy in task decomposition strategies, resulting in uneven load allocation or computing bottlenecks, high communication overhead between nodes, and data transmission delays have become obstacles to performance improvement. The fault tolerance mechanism design is not perfect enough, making it difficult to ensure the integrity of the calculation results.
Through task dependency graph construction, resource status monitoring and dynamic allocation, lightweight communication protocols and fault tolerance mechanisms, task decomposition, resource scheduling and result integration are optimized to ensure subtask independence, uniformity and separability, reduce synchronization and communication overhead, and achieve efficient parallel execution and result verification.
Significantly improve the processing efficiency of computing tasks and system resource utilization, ensure the task is quickly recovered in the event of failure, ensure the accuracy and reliability of calculation results, and reduce communication delays and resource waste.
Smart Images

Figure CN120256083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of task execution, and particularly to a method, system, terminal device and storage medium for parallel computing execution of a workflow. Background Art
[0002] As an important part of modern computing technology, parallel computing can significantly improve the computing efficiency and resource utilization rate of a system by decomposing complex tasks into multiple subtasks and processing them in parallel among multiple nodes. Its applications cover fields such as high-performance computing, big data processing, and artificial intelligence training, and it has become the core means to solve ultra-large-scale computing problems.
[0003] Although significant progress has been made in parallel computing technology, there are still many problems in practical applications. First, the flexibility and accuracy of the task decomposition strategy are insufficient, which easily leads to uneven load distribution or computing bottlenecks. Second, the communication overhead between nodes is relatively high. Especially in large-scale distributed systems, data transmission delay has become the main obstacle to performance improvement. In addition, the design of the fault tolerance mechanism is not perfect enough, and it is difficult to recover in time and ensure the integrity of the computing results when a node fails or a task is interrupted.
[0004] Facing the above challenges, there is an urgent need for a solution that can comprehensively optimize task decomposition, resource scheduling, communication efficiency, and fault tolerance mechanism to make up for the deficiencies of the existing technology in terms of efficiency, reliability, and scalability. Summary of the Invention
[0005] Aiming at the above defects, the purpose of the present invention is to provide a method, system, terminal device and storage medium for parallel computing execution of a workflow, aiming to provide an efficient and stable solution for complex tasks through task decomposition, resource scheduling, node communication, fault tolerance mechanism, and result integration.
[0006] To achieve this purpose, the present invention adopts the following technical solutions:
[0007] A method for parallel computing execution of a workflow, the method for parallel computing execution of the workflow includes:
[0008] Obtain the task to be executed. After performing logical analysis on the task to be executed, divide the task to be executed into several subtasks according to the task type and construct a task dependency graph, and formulate an execution model according to the characteristics of the subtasks;
[0009] Allocate resources to the subtasks according to the resource status and task priority of the subtasks, and set the execution order for the subtasks based on the dependency relationships in the task dependency graph;
[0010] According to the execution mode and the execution order, the subtasks enter the parallel execution state to obtain the execution results of each subtask;
[0011] According to the original logic and data dependencies of the task to be executed, summarize, sort, and merge the execution results to obtain a summary result, and perform a verification operation on the summary result to obtain the final execution result.
[0012] Preferably, the task types include loop tasks, large data volume tasks, and multi-functional tasks.
[0013] Preferably, after dividing the task to be executed into several subtasks according to the task type, it includes:
[0014] Analyze the task to be executed to obtain the recommended granularity range and recommended quantity of the subtasks, and re-divide the subtasks based on the recommended granularity range and recommended quantity.
[0015] Preferably, the execution model includes a data parallel model, a task parallel model, and a hybrid parallel model;
[0016] The data parallel model includes applying the same operation to different data blocks;
[0017] The task parallel model includes making different subtasks execute in parallel on multiple different processors;
[0018] The hybrid parallel model includes first applying the same operation to different data blocks and making different subtasks execute in parallel on multiple different processors.
[0019] Preferably, according to the resource status and task priority of the subtasks, the resource allocation for the subtasks includes:
[0020] Use a machine learning model to predict the task characteristics and load trends of the subtasks, and perform the first resource allocation for the subtasks according to the prediction results;
[0021] Continuously collect the computing power, load status, and network latency of each subtask, construct a dynamic resource map of the subtasks, and perform the second resource allocation for the subtasks according to the dynamic resource map;
[0022] Perform the third resource allocation for the subtasks according to the running status of the subtasks when the subtasks are executing;
[0023] When performing the first, second, or third resource allocation, use a distributed coordination mechanism to optimize the allocation results of the first, second, or third resource allocation.
[0024] Preferably, when the subtasks enter the parallel execution state, it further includes:
[0025] When an error occurs during the execution of the subtasks or the subtasks fail, perform error analysis and error recording on the subtasks;
[0026] According to the error analysis and error record results, restart the subtask or allocate the subtask to a preset backup node.
[0027] Preferably, after performing a verification operation on the summary result, the final execution result includes:
[0028] Verify the upper and lower limits of the data values in the summary result according to a preset numerical range, and verify the data format and logical rules in the summary result;
[0029] Perform multi-source comparison on the statistical data in the summary result, determine whether there are conflicts in the statistical data from multiple sources, and judge the rationality of the current summary result according to the current summary result and the historical summary result;
[0030] Identify outliers in the data in the summary result according to a statistical method, and judge whether the summary result deviates from the expectation according to the distribution form of the data in the summary result.
[0031] A workflow parallel computing execution system, which is applied to the workflow parallel computing execution method as described above. The workflow parallel computing execution system includes:
[0032] A task decomposition module, configured to obtain a task to be executed, perform logical analysis on the task to be executed, divide the task to be executed into several subtasks according to the task type, construct a task dependency graph, and formulate an execution model according to the characteristics of the subtasks;
[0033] A task scheduling module, configured to perform resource allocation on the subtasks according to the resource status and task priorities of the subtasks, and set an execution order for the subtasks based on the dependency relationships in the task dependency graph;
[0034] An execution module, configured to enter the parallel execution state for the subtasks according to the execution mode and the execution order, and obtain the execution results of each subtask;
[0035] A result integration module, configured to summarize, sort, and merge the execution results according to the original logic and data dependencies of the task to be executed to obtain a summary result, and perform a verification operation on the summary result to obtain a final execution result.
[0036] A terminal device, which includes a memory, a processor, and a program stored on the memory and executable on the processor. The program is configured to implement the steps of the workflow parallel computing execution method as described above.
[0037] A storage medium stores a workflow parallel computing execution program thereon. When the workflow parallel computing execution program is executed by a processor, the steps of the workflow parallel computing execution method described above are implemented.
[0038] One of the above technical solutions has the following advantages or beneficial effects:
[0039] Through task analysis, task decomposition, resource scheduling, parallel execution, and result integration, the present invention significantly improves the processing efficiency of computing tasks and the utilization rate of system resources; in the task decomposition stage, complex computing tasks are refined into independent and uniform subtasks and optimized according to the task dependency graph, effectively reducing synchronization and communication overheads and laying a foundation for parallel computing; resource scheduling ensures the balanced distribution of tasks among different nodes through real-time monitoring, predictive analysis, and adaptive adjustment mechanisms, avoiding uneven load and resource waste; in the parallel execution stage, the parallelism and system stability of task execution are improved through lightweight communication protocols and dynamic adjustment strategies; at the same time, through a built-in fault tolerance mechanism, it is ensured that tasks can be quickly restored in case of failures, ensuring the continuity and accuracy of computing tasks, and in the result integration link, through strict consistency verification and multi-dimensional error analysis, the accuracy and reliability of the final result are ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0041] Figure 1 is a flowchart of the workflow parallel computing execution method provided by the embodiment of the present invention;
[0042] Figure 2 is a schematic structural diagram of the workflow parallel computing execution system provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The following details the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.
[0044] In the present invention, the terms "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0045] As an important part of modern computing technology, parallel computing can significantly improve the computing efficiency and resource utilization rate of the system by decomposing complex tasks into multiple subtasks and processing them in parallel among multiple nodes. Its applications cover fields such as high-performance computing, big data processing, and artificial intelligence training, becoming the core means to solve ultra-large-scale computing problems.
[0046] Although significant progress has been made in parallel computing technology, there are still many problems in practical applications. Firstly, the flexibility and accuracy of task decomposition strategies are insufficient, which easily leads to uneven load distribution or computing bottlenecks; secondly, the communication overhead between nodes is relatively high, especially in large-scale distributed systems, and the data transmission delay has become the main obstacle to performance improvement; in addition, the design of the fault tolerance mechanism is not perfect enough, and it is difficult to recover in time and ensure the integrity of the computing results when nodes fail or tasks are interrupted.
[0047] Facing the above challenges, there is an urgent need for a solution that can comprehensively optimize task decomposition, resource scheduling, communication efficiency, and fault tolerance mechanism to make up for the deficiencies of existing technologies in terms of efficiency, reliability, and scalability.
[0048] Therefore, a workflow parallel computing execution method is proposed, as Figure 1 shown in a preferred embodiment of the present invention, the workflow parallel computing execution method includes the following steps:
[0049] S1: Obtain the task to be executed. After logical analysis of the task to be executed, divide the task to be executed into several subtasks according to the task type and construct a task dependency graph, and formulate an execution model according to the characteristics of the subtasks;
[0050] This embodiment takes the scheduling system as an example. Therefore, the system mentioned below can decompose, schedule, execute, and integrate tasks. In step S1, first, the task to be executed is received, and a logical analysis of the task is performed. The goal of task analysis is to clarify the scale, complexity, and execution logic of the task. When analyzing, the data structure, operation type, and computing requirements of the task are analyzed from multiple dimensions. This analysis provides a basis for subsequent task decomposition. According to the characteristics and computing requirements of the task, the task to be executed is divided into multiple subtasks. The principles of task decomposition include: 1. Independence: Subtasks should be as independent as possible, reducing mutual dependencies to ensure that they can run in parallel on multiple processors; 2. Uniformity: The computing amount of subtasks should be as uniform as possible to avoid overloading some processors and affecting the overall efficiency; 3. Divisibility: Complex tasks need to be reasonably divided into multiple subtasks with clear logic, and ensure that each subtask can be further decomposed into smaller fine-grained units so that they can be processed in parallel. Task dependencies are presented by constructing a task dependency graph (which can be a directed acyclic graph, DAG). Nodes in the graph represent subtasks, and edges represent the dependency relationships between tasks. According to the characteristics and dependency relationships of these subtasks, a suitable execution model is designed, which may be data parallelism, task parallelism, or hybrid parallelism, etc. The design of the execution model for each subtask needs to consider how to maximize the utilization of computing resources while minimizing the synchronization and communication costs between tasks. In addition, task scheduling and resource allocation can assign fine-grained subtasks to different processors through a task scheduler, and dynamically adjust the allocation of operation units by real-time monitoring the load of each processor to avoid the problem of uneven load. During the decomposition process, processors may need to cooperate and share data. By reasonably designing the data dependencies and task mappings between tasks, the communication frequency and data transmission volume can be reduced. Methods such as barrier synchronization, locks, or queues can be used to ensure the coordination of tasks between processors is consistent.
[0051] For example, in numerical simulation, the solution of complex equations can be decomposed into small-scale matrix operations for parallel execution; in the MapReduce framework, data is decomposed into multiple partitions, processed in parallel through the Map operation, and then the results are aggregated through Reduce; for the filtering process of large-scale images, the image can be divided into blocks, each block is calculated by an independent processor, and then combined to generate a complete result.
[0052] S2: According to the resource status and task priorities of the subtasks, allocate resources to the subtasks, and set the execution order for the subtasks based on the dependency relationships in the task dependency graph;
[0053] After task decomposition, the next steps are task scheduling and resource allocation. The scheduling system needs to understand the resource requirements of each subtask, including computing power, storage space, communication bandwidth, etc. At the same time, the scheduling system will also monitor the resource status in real time, such as the current load, network latency, etc., to ensure the load balance of each processing unit and avoid overloading of some nodes or resource idleness. Resource allocation needs to be adjusted according to the priority of the tasks, and tasks on the critical path are executed first to ensure the optimization of the overall system throughput and response time. The execution order of the tasks is set according to the edges in the task dependency graph, ensuring that subtasks without dependencies are executed first, and subtasks that depend on other tasks need to wait until their prerequisite tasks are completed before they can start execution. A reasonable execution order can effectively avoid resource conflicts and task waiting, thereby improving the computing efficiency.
[0054] S3: According to the execution mode and the execution order, the subtasks enter the parallel execution state to obtain the execution results of each subtask;
[0055] In this step, all subtasks enter the actual parallel computing stage according to the predetermined execution order and parallel execution mode. Subtasks may adopt different parallel execution strategies according to their characteristics, such as data parallelism, task parallelism, or hybrid parallelism. In the data parallel mode, the same operation is applied to different data blocks simultaneously; in the task parallel mode, different task modules may run simultaneously on multiple processors to perform their respective independent functions. During the execution process, the system will monitor the status of each subtask in real time to ensure that the tasks are executed as expected and handle any possible resource contention, load imbalance, or synchronization issues.
[0056] S4: According to the original logic and data dependencies of the tasks to be executed, the execution results are summarized, sorted, and merged to obtain a summary result, and after performing a verification operation on the summary result, the final execution result is obtained.
[0057] After the parallel computing is completed, each subtask will generate its own execution results. The next task is to summarize, sort, and merge these results. The summarization process is to integrate the outputs of different subtasks into a complete result set, and this process needs to follow the original logic and data dependency relationships of the tasks.
[0058] The original logic refers to the computing process, business rules, and data processing order followed during task design. It defines how to perform the calculations and the execution steps of the tasks. Specifically, the original logic describes the core computing process of the tasks, how to decompose a complex computing task into multiple subtasks, and the order and rules for the execution of these subtasks. It can ensure the logical consistency and integrity of the computing tasks and is the basis for task decomposition and parallel execution. Through the original logic, the system can understand how to complete the tasks.
[0059] Data dependency refers to the relationship that occurs between tasks or subtasks during task execution due to the need to share or transfer data. In parallel computing, data dependency determines the execution order and parallelism between tasks, and tasks with data dependency must be executed in a specific order. Data dependency can be classified into various types, such as data flow dependency, control dependency, and data synchronization dependency. These dependency relationships affect the parallel execution ability of tasks. By analyzing data dependency, the system can reasonably understand the execution order of tasks and resource scheduling, and better summarize the results.
[0060] For data that needs to be summarized in a specific order (such as data sorted by time order or data merged according to specific rules), the system will perform sorting and merging operations to ensure the accuracy and integrity of the final result. Then, the system will conduct result verification to ensure the correctness and consistency of the data. This verification process includes various checks, such as data range check, format check, logical relationship verification, etc., to ensure that the final output result conforms to the expected business logic and data standards. If inconsistent data is found during the verification process, the system will mark and correct it, and finally generate a reliable execution result. This process can ensure the quality and reliability of task execution, and avoid system failures or result errors caused by calculation errors or data inconsistencies.
[0061] Through task analysis and task dependency graph construction, the present invention divides complex tasks into multiple subtasks to ensure the independence, uniformity, and divisibility between subtasks, thereby improving the efficiency of parallel computing. The task scheduler dynamically allocates tasks according to task dependency relationships and resource status, and preferentially executes tasks on the critical path to avoid uneven load and resource conflicts, optimizing the system throughput and response time; by designing appropriate execution models for different subtasks, it is possible to select the most suitable parallel strategy according to the characteristics of the tasks and the execution environment. The flexibility of the parallel mode can reduce the synchronization and communication overhead between tasks, ensuring efficient resource utilization and minimizing waiting time.
[0062] By reasonably designing the data dependency and task mapping between tasks, the communication frequency and data transmission volume between processors are reduced. In addition, mechanisms such as barrier synchronization, locks, or queues are used to ensure the coordination and consistency between tasks, thereby reducing the latency caused by data transmission. A strict verification mechanism is implemented during task execution to ensure the accuracy and consistency of the final result. Through the analysis of data dependency relationships, the system can efficiently perform result summarization, sorting, and merging, avoiding problems caused by calculation errors and data inconsistencies.
[0063] Preferably, the task types include loop tasks, large data volume tasks, and multi-functional tasks.
[0064] For multi-functional tasks, the multi-functional tasks can be functionally decomposed. The functional decomposition divides complex computing tasks according to their different functional modules, and each module is executed as an independent operation unit. For example, when performing matrix multiplication, the matrix can be decomposed into several sub-matrices, and each sub-matrix is calculated independently; for the data sorting task, the input data can be divided into several small pieces, and each piece of data is sorted separately, and then the final sorting result is obtained through a merging operation. In this way, different functional modules can be processed in parallel, thereby effectively improving the computing efficiency, especially when multiple processors are available.
[0065] For large data volume tasks, large data volume tasks usually need to process massive amounts of data. A single computer will face performance bottlenecks when processing large-scale data. To overcome this problem, the data can be divided into chunks, and each chunk of data can be processed independently. Taking image processing as an example, a large image can be cut into multiple small pieces, and each piece of the image is processed independently, such as image filtering or feature extraction, and then the processing results are merged; for large data computing, such as MapReduce tasks, the data can be partitioned into multiple parts and distributed to different computing nodes for parallel processing. This data chunking strategy not only helps to reduce the burden on a single computing node, but also improves the processing efficiency, especially suitable for distributed computing environments.
[0066] For tasks involving a large number of iterative computations, in a typical loop task, each iteration in the loop body is usually independent, meaning that the computations for each iteration can be executed in parallel on different processors. By splitting the computations in the loop body into multiple independent tasks, parallel processing can be achieved, significantly improving the computing efficiency. For example, in numerical simulations or machine learning training, iterative algorithms are often used. Utilizing multiple computing nodes to execute each iteration step in parallel can greatly shorten the computing time. However, since there may be dependencies among the data in the loop task, appropriate synchronization mechanisms (such as barrier synchronization) are required to ensure that the tasks are executed in the correct order to avoid conflicts in parallel computing.
[0067] Preferably, after dividing the task to be executed into several subtasks according to the task type, it includes:
[0068] Analyze the task to be executed to obtain the recommended granularity range and recommended quantity of the subtasks, and re-divide the subtasks based on the recommended granularity range and recommended quantity.
[0069] Before dividing the tasks to be executed, it is necessary to first conduct a detailed analysis of the tasks themselves. The goal of this step is to determine the "granularity range" of the subtasks based on the characteristics of the tasks (such as computational complexity, data dependence, parallelism requirements, etc.). Granularity refers to the amount of work or data processed by each subtask. The recommended granularity range should consider the specific requirements of the task. Excessive granularity will result in an overly heavy single subtask and reduce parallel efficiency; while too small a granularity may lead to overly fine task division, bringing scheduling overhead and synchronization burden between tasks. Through the analysis of the tasks, the appropriate granularity range for the task can be obtained for subsequent division decisions. In addition to the granularity range, another key parameter is the recommended number of subtasks. The recommended number of tasks is usually determined based on the scale of the task, the availability of computing resources, and the requirements for parallelism. The recommended number of subtasks should match the processing resources (such as computing cores or computing nodes) to avoid resource overload or idleness. For example, if the processing resources are 10 cores, it is recommended to divide into 10 subtasks, with each core responsible for one subtask. The recommendation of the number of tasks should also balance the division accuracy of the tasks and the scheduling overhead to ensure maximum parallelism without introducing unnecessary burdens. After obtaining the recommended granularity range and the recommended number, the subtasks can be re-divided based on these parameters. Specifically, first determine the workload of each subtask according to the recommended granularity range, and then divide the task into the corresponding number of subtasks according to the recommended number. If the recommended number is large, the subtasks will be divided more finely and the task will be more parallel; if the recommended number is small, the granularity of the subtasks will increase accordingly. In this process, it is also necessary to consider the possible data dependence relationships between tasks. If there are strong dependencies between some subtasks, ensure a reasonable execution order or appropriate data transfer during division. Finally, the performance of the divided subtasks can be verified to ensure that the divided subtasks can be efficiently executed on the expected computing resources. If the task division is improper, it may lead to resource waste, execution bottlenecks, or excessive synchronization operations. At this time, the division strategy can be further optimized by adjusting the granularity range and the number of tasks, and multiple iterations may be required to adjust the division strategy based on the actual execution situation to finally achieve the optimal execution effect.
[0070] Preferably, the execution model includes a data parallel model, a task parallel model, and a hybrid parallel model;
[0071] The data parallel model includes applying the same operation to different data blocks;
[0072] The task parallel model includes making different subtasks execute in parallel on multiple different processors;
[0073] The hybrid parallel model includes first applying the same operation to different data blocks and making different subtasks execute in parallel on multiple different processors.
[0074]
[0074] The core idea of the data parallel model is to apply the same operations to different data blocks. Each processing unit processes a subset of the dataset, and all processing units perform the same operations. The independence of the data and the repetition of the same operations make parallelization very efficient, especially when dealing with large-scale data. Taking matrix operations as an example, each row or column of a matrix can be regarded as a data block and calculated independently. The operations performed on each processing unit are the same, only the data is different. Usually, such tasks have very low inter-task dependencies, so they are suitable for the data parallel model. The advantages are the efficient utilization of multi-core or multi-processors, the reduction of computational redundancy, and the significant improvement in execution speed. Application scenarios include matrix multiplication, image processing, large-scale data analysis, etc.
[0075] The core idea of the task parallel model is to break down the overall task into multiple different subtasks. Each subtask may involve different types of operations, and these subtasks can be executed in parallel on multiple processors. Different from the data parallel model, task parallelism focuses on the diversity and heterogeneity of tasks. Each subtask may perform different computational or logical operations, and the dependencies between subtasks are relatively complex. Pipeline-style task decomposition is a typical example of task parallelism. In this model, multiple processing units are responsible for different computational stages, and the output of each stage is the input of the next stage. For example, during video encoding, one processing unit can be responsible for compressing a certain part of the video, while another processing unit is responsible for the subsequent encoding task. Its advantages are high flexibility, the ability to handle different types of computational tasks, and each subtask can be adjusted according to its own optimization. Application scenarios include complex scientific computing, data processing pipelines, and tasks with different computational modules.
[0076]
[0075] The hybrid parallel model combines the characteristics of data parallelism and task parallelism, and can utilize the advantages of both to solve complex computational problems. In this model, the data is first divided into blocks, and the same operations are applied to each data block using data parallelism. After completing data parallelism, the tasks are further broken down into different subtasks, and task parallelism is used to execute them in parallel on multiple processors. It can improve computational efficiency at multiple levels, ensuring the parallel processing of large-scale data and allowing the execution of heterogeneous tasks on different processors. Its advantage is the ability to effectively combine the advantages of the two parallel methods and improve performance in multiple dimensions, especially suitable for complex applications that require both data processing and task decomposition. For example, during deep learning training, data parallelism can parallelly process different data batches on multiple computing nodes, while task parallelism can allocate different neural network layers or computational steps to different processing units. Application scenarios include deep learning training of large-scale datasets, scientific simulations, image rendering, etc.
[0077] Preferably, resource allocation for the subtasks according to the resource status and task priorities of the subtasks includes:
[0078] Predicting the task characteristics and load trends of the subtasks using a machine learning model, and performing first resource allocation for the subtasks according to the prediction results;
[0079] Prediction and pre-scheduling are techniques based on historical data and machine learning models, used to foresee in advance the load changes and resource requirements of tasks. By analyzing historical task execution situations, node load trends, and task characteristics (such as compute-intensive, I / O-intensive, etc.), the scheduler can predict the future resource requirements of each subtask and the load change trends of nodes. This prediction can help the scheduler perform resource allocation in advance, so that the required resources are already prepared when the task arrives, avoiding delays in resource allocation and potential bottleneck problems. Machine learning models (such as time series prediction, regression analysis, etc.) can provide accurate load predictions based on the real-time characteristics and historical data of tasks, enabling the scheduler to dynamically adjust the resource allocation strategy, improving the system's response efficiency and resource utilization rate. For example, when the load of a certain node is about to increase, the scheduler will schedule more resources to that node in advance to avoid excessive waiting time for tasks.
[0080] Continuously collecting the computing power, load status, and network latency of each subtask, constructing a dynamic resource map of the subtasks, and performing second resource allocation for the subtasks according to the dynamic resource map;
[0081] Real-time monitoring and analysis are the basis of task scheduling, aiming to dynamically adjust the resource allocation strategy by continuously tracking the system status. Specifically, the scheduler needs to collect in real-time multi-dimensional operation status information such as the computing power, current load, network latency, and memory usage rate of each computing node. This information helps the system construct a dynamic resource map, describing the computing resources and task load status of all current nodes. Through continuous monitoring, the scheduler can timely identify overloaded or idle situations of nodes, ensuring that the resource allocation of tasks can accurately match the actual requirements. For example, when the load of some nodes is too heavy, the scheduler can choose to migrate some tasks to nodes with lighter loads to avoid performance bottlenecks caused by uneven loads.
[0082] Performing third resource allocation for the subtasks according to the running status of the subtasks when the subtasks are being executed;
[0083] The third resource allocation can be referred to as an adaptive adjustment mechanism. The adaptive adjustment mechanism enables the scheduling algorithm to cope with unexpected events during operation, such as node failures and sudden increases in task volume. During operation, unexpected events may occur, such as a sudden failure of a certain node or a sudden increase in load, resulting in the invalidation of the resource allocation strategy or a decline in performance. In this case, the scheduler needs to respond quickly and reallocate resources through the adaptive adjustment mechanism to ensure that the system can operate stably under high load or abnormal conditions. This can be achieved by combining load monitoring, node health status monitoring, and dynamic adjustment of task priorities. For example, if a node fails, the scheduler can immediately migrate the tasks of that node to other healthy nodes; if the execution time of a certain task exceeds the expectation, the scheduler can dynamically adjust the resource allocation of the task to ensure that the task can be completed on time.
[0084] When performing the first, second, or third resource allocation, a distributed coordination mechanism is adopted to optimize the allocation results of the first, second, or third resource allocation.
[0085] The optimization of the allocation results of the first, second, or third resource allocation by adopting the distributed coordination mechanism occurs during the allocation process. It does not only focus on a certain index (such as load balancing), but comprehensively considers multiple factors, such as load balancing, communication overhead, task completion time, etc. The goal of scheduling is to achieve global optimality through the distributed coordination mechanism and avoid the negative impact of local optimization on the overall performance. For example, simply pursuing load balancing may lead to an increase in communication overhead, and simply pursuing task completion time may lead to resource waste. Through multi-objective optimization, when allocating resources, a balance will be found among various objectives according to the priorities and requirements of tasks. Some advanced algorithms, such as genetic algorithms and particle swarm optimization (PSO), can be adopted to optimize multiple objectives simultaneously, and ultimately maximize the performance of the entire system. The distributed coordination mechanism can ensure global optimization, avoid resource contention or scheduling conflicts among multiple schedulers, and at the same time increase the system throughput.
[0086] Preferably, when the subtask enters the parallel execution state, it further includes:
[0087] When an error occurs during the execution of the subtask or the subtask fails, error analysis and error recording are performed on the subtask;
[0088] According to the error analysis and error recording results, the subtask is restarted, or the subtask is allocated to a preset standby node.
[0089] During the parallel execution of subtasks, various types of errors may occur, such as hardware failures, improper resource allocation, computational anomalies, task timeouts, network problems, etc. When the system detects an error in a subtask, error analysis is required. The goal of error analysis is to quickly locate the root cause of the failure, such as whether it is due to insufficient computational resources of the node, or due to code defects or external dependencies. Error analysis relies on logging, monitoring systems, and diagnostic tools to collect the execution status of subtasks, resource usage, error stack information, etc. After analysis, the detailed error information is recorded in the system's error log as a basis for subsequent recovery, debugging, and optimization. The content of the error record includes the error type, timestamp, computational nodes involved, error code, task input and output, etc. This information can help developers or the system itself identify problems and make corresponding adjustments.
[0090] After error analysis and recording are completed, the system needs to take recovery measures based on the analysis results. If an error occurs during the execution of a subtask, the first step in the recovery mechanism is to determine the severity and transience of the error. For example, if a task fails due to a memory overflow or an external resource access timeout, this error may be temporary, and the system can attempt to restart the subtask to resume normal execution. For transient errors, the restart operation can release resources and reschedule the task, thus avoiding repeated resource or execution problems. However, if the error is caused by a hardware failure, node failure, or a system exception that cannot be recovered for a long time, restarting may not solve the problem. At this time, the system will decide to migrate the subtask to a preset backup node for execution based on the error record. The backup node is pre-reserved computational resources in the system, with relatively independent computational capabilities, and can take over and execute the failed task. By migrating the task to the backup node, the system can avoid single-point failures and improve the stability and reliability of task execution.
[0091] Restarting and task migration are two main strategies for task recovery. For minor errors or temporary problems, first try to restart the subtask. The restart process includes cleaning up resources from the current node, reloading the task and environment, ensuring a clean state of resources, and restarting the execution of the task, which is effective for solving problems such as temporary insufficient computational resources or task timeouts. However, if error analysis indicates that there are irreparable hardware problems, high load, or network instability on the current node, task restart cannot solve the problem, and the system will start the task migration mechanism. Task migration transfers the subtask from the failed node to a preset backup node. The backup node is usually a healthy computational node that can take over the execution of the task. The selection of the backup node can be based on the principles of load balancing, resource availability, and fault isolation to ensure the smooth completion of the task.
[0092] Preferably, the final execution result obtained after performing a verification operation on the summary result includes:
[0093] Verify the upper and lower limits of the data values in the summary result according to the preset value range, and verify the data format and logical rules in the summary result;
[0094] In the summary result, first, it is necessary to check the range of each data value. To ensure that all data fall within the predetermined legal range, for example, the percentages in statistical data should be between 0% and 100%. If any result exceeds this range (such as negative numbers or percentages greater than 100%), it indicates data inconsistency. The system will automatically perform range verification and mark or correct these out-of-range data to prevent unreasonable values from affecting the final result. Second, verify the data format in the summary result. The data format should conform to the standard to ensure data unity and readability. For example, the date format needs to be consistent, usually in the specification of "YYYY-MM-DD". If some date data adopt different formats (such as "DD-MM-YYYY" or other non-standard formats), the system will identify and mark them as format errors. In addition, numerical data should not contain non-numeric characters. If letters, special symbols, etc. appear in the data, the system will regard them as inconsistent. Finally, the system will verify the data consistency according to the predefined business logic rules. For example, in financial data, if there is a relationship of "income = cost + profit", and the actual summary data does not conform to this formula, it indicates logical inconsistency between the data. Similar logical rules can include mathematical constraints between data fields, comparison of budget and actual expenditure, etc. If any logical contradiction is found, the system will issue a warning or make automatic corrections.
[0095] Perform multi-source comparison on the statistical data in the summary result to determine whether there are conflicts in the statistical data from multiple sources, and judge the rationality of the current summary result based on the current summary result and the historical summary result;
[0096] The summary result may come from multiple different data sources, so cross-source consistency checking is required. For example, if the statistical data comes from multiple systems or departments, the system will check whether there are significant differences in these data. If the differences between the data from multiple sources exceed the tolerance range (for example, the error exceeds the preset percentage), it will be marked as inconsistent data, which can help identify data transmission errors, synchronization problems, or other factors causing data conflicts.
[0097] In addition, the system will also compare the current summary result with historical data to analyze the rationality and trend of the data. For example, the growth rate or fluctuation range of a certain statistical data should be within a reasonable range. If there is a significant deviation between the current result and the historical data (such as a sudden increase or decrease), it may be an error in the data processing or summarization process. By comparing historical data, the system can timely detect anomalies and make corrections or marks.
[0098] Identify outliers in the data in the summary result according to statistical methods, and determine whether the summary result deviates from the expectation based on the distribution pattern of the data in the summary result.
[0099] Statistical methods can be used to identify outliers in data. Common methods include standard deviation analysis, box plot analysis, etc. Through these methods, outliers that are significantly different from most data points can be identified. For example, if the value of a certain piece of data is much higher or lower than other values in the dataset, and this deviation exceeds a certain multiple of the standard deviation, then this data point will be marked as an outlier, which helps to ensure the normality of the data and exclude extreme outliers caused by incorrect input or calculation.
[0100] In addition to outlier detection, the distribution pattern of the data can also be analyzed. If the data in the summary result should conform to a certain specific distribution pattern (such as normal distribution), and the actual distribution of the data deviates significantly from the expected pattern, it may indicate that there are inconsistencies in the data. For example, some statistical data may be supposed to show a symmetric distribution. If the data distribution is skewed, it may indicate errors, data loss, or input errors. The system will use distribution analysis methods to detect whether the overall form of the data is reasonable and ensure that it conforms to the expected distribution rules.
[0101] In addition, the verification of the summary result also includes duplicate detection and null value and missing value checks. Duplicate detection includes checking whether there are duplicate data or duplicate records in the result. Duplicate data or duplicate records may lead to redundancy or distortion of the integration result. Null values and missing values may cause some results to be incomplete and need to be marked as inconsistent data.
[0102] A workflow parallel computing execution system, which is applied to the workflow parallel computing execution method as described above, as Figure 2 shown, the workflow parallel computing execution system includes:
[0103] Task decomposition module 1, which is used to obtain the tasks to be executed. After performing logical analysis on the tasks to be executed, divide the tasks to be executed into several subtasks according to the task type and construct a task dependency graph, and formulate an execution model according to the characteristics of the subtasks;
[0104] Task scheduling module 2, which is used to allocate resources to the subtasks according to the resource status and task priorities of the subtasks, and set the execution order for the subtasks based on the dependency relationships in the task dependency graph;
[0105] Execution module 3, which is used to make the subtasks enter the parallel execution state according to the execution mode and the execution order, and obtain the execution results of each subtask;
[0106] The result integration module 4 is configured to summarize, sort, and merge the execution results according to the original logic and data dependencies of the task to be executed to obtain a summary result, and perform a verification operation on the summary result to obtain the final execution result.
[0107] This embodiment implements a workflow parallel computing execution method and its implementation process. Please refer to the above embodiments, and details will not be repeated here.
[0108] In addition, an embodiment of the present invention further provides a terminal device, which includes a memory, a processor, and a program stored on the memory and executable on the processor. The program is configured to implement the steps of the workflow parallel computing execution method as described above.
[0109] Since the program is configured to implement the steps of the workflow parallel computing execution method as described above, the program at least has all the beneficial effects brought by all the technical solutions of the foregoing embodiments, and details will not be repeated here.
[0110] In addition, an embodiment of the present application further provides a computer-readable storage medium, on which a workflow parallel computing execution program is stored. When the workflow parallel computing execution program is executed by a processor, it implements the steps of the workflow parallel computing execution method as described above.
[0111] Since the workflow parallel computing execution program, when executed by a processor, adopts all the technical solutions of the foregoing embodiments, it at least has all the beneficial effects brought by all the technical solutions of the foregoing embodiments, and details will not be repeated here.
[0112] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0113] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention. The scope of the present invention is defined by the claims and their equivalents.
Claims
1. A method for parallel computing execution of a workflow, characterized in that The method for parallel computing execution of the workflow includes: Obtain the task to be executed. After performing logical analysis on the task to be executed, divide the task to be executed into several subtasks according to the task type and construct a task dependency graph, and formulate an execution model according to the characteristics of the subtasks; According to the resource status and task priority of the subtasks, allocate resources to the subtasks, and set the execution order for the subtasks based on the dependency relationships in the task dependency graph; According to the execution mode and the execution order, the subtasks enter the parallel execution state to obtain the execution results of each subtask; According to the original logic and data dependency of the task to be executed, summarize, sort, and merge the execution results to obtain a summary result, and perform a verification operation on the summary result to obtain the final execution result.
2. The workflow parallel computing execution method according to claim 1, wherein The task types include loop tasks, large data volume tasks, and multi-functional tasks.
3. The workflow parallel computing execution method according to claim 1, wherein After dividing the task to be executed into several subtasks according to the task type, it includes: Analyze the task to be executed to obtain the recommended granularity range and recommended quantity of the subtasks, and re-divide the subtasks based on the recommended granularity range and recommended quantity.
4. The workflow parallel computing execution method according to claim 1, characterized in that, The execution models include a data parallel model, a task parallel model, and a hybrid parallel model; The data parallel model includes applying the same operation to different data blocks; The task parallel model includes enabling different subtasks to be executed in parallel on multiple different processors; The hybrid parallel model includes first applying the same operation to different data blocks and enabling different subtasks to be executed in parallel on multiple different processors.
5. The workflow parallel computing execution method according to claim 1, wherein According to the resource status and task priority of the subtasks, allocating resources to the subtasks includes: Use a machine learning model to predict the task characteristics and load trends of the subtasks, and perform the first resource allocation for the subtasks according to the prediction results; Continuously collect the computing power, load status, and network latency of each subtask, construct a dynamic resource map of the subtasks, and perform the second resource allocation for the subtasks according to the dynamic resource map; Perform the third resource allocation for the subtasks according to the running status of the subtasks when the subtasks are being executed; When performing the first, second, or third resource allocation, adopt a distributed coordination mechanism to optimize the allocation results of the first, second, or third resource allocation.
6. The workflow parallel computing execution method according to claim 1, wherein When the subtasks enter the parallel execution state, it further includes: When an error occurs during the execution of the subtask or the subtask fails, perform error analysis and error recording on the subtask; According to the error analysis and error recording results, restart the subtask or allocate the subtask to a preset backup node.
7. The workflow parallel computing execution method according to claim 1, characterized in that Obtaining the final execution result after performing a verification operation on the summary result includes: Verify the upper and lower limits of the data values in the summary result according to the preset numerical range, and verify the data format and logical rules in the summary result; Perform multi-source comparison on the statistical data in the summary result, judge whether there are conflicts in the statistical data from multiple sources, and judge the rationality of the current summary result according to the current summary result and the historical summary result; Identify outliers in the data of the summary result according to the statistical method, and judge whether the summary result deviates from the expectation based on the distribution pattern of the data in the summary result.
8. A workflow parallel computing execution system, which is applied to the workflow parallel computing execution method described in any one of claims 1-7, and is characterized in that, The workflow parallel computing execution system includes: A task decomposition module, configured to obtain the task to be executed, perform logical analysis on the task to be executed, divide the task to be executed into several subtasks according to the task type, construct a task dependency graph, and formulate an execution model according to the characteristics of the subtasks; A task scheduling module, configured to allocate resources to the subtasks according to the resource status and task priorities of the subtasks, and set an execution order for the subtasks based on the dependency relationships in the task dependency graph; An execution module, configured to enable the subtasks to enter a parallel execution state according to the execution mode and the execution order, and obtain the execution results of each subtask; A result integration module, configured to summarize, sort, and merge the execution results according to the original logic and data dependencies of the task to be executed to obtain a summary result, and perform a verification operation on the summary result to obtain a final execution result.
9. A terminal device, characterized in that, The terminal device includes: a memory, a processor, and a program stored on the memory and executable on the processor, the program being configured to implement the steps of the workflow parallel computing execution method according to any one of claims 1 to 7.
10. A storage medium, characterized in that, A workflow parallel computing execution program is stored on the storage medium, and when the workflow parallel computing execution program is executed by the processor, the steps of the workflow parallel computing execution method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Variable-granularity task decomposition method
CN111311072A
Task scheduling method, device and system and electronic equipment
CN112882813A
Scheduling method for high-concurrency processing of small tasks in distributed cluster
CN118394476A
Edge computing task unloading optimization method and system
CN118567851A
Mass data summarization method and system based on task chain and divide-and-conquer method
CN118672790A
Cited By
Oil and gas industry chain decision optimization method and device
CN120688703A
Oil and Gas Industry Chain Decision Optimization Methods and Devices
CN120688703B
Task decomposition arrangement and exception retry method and system for workflow engine
CN121961495A