Large-scale scientific computing workflow collaborative management and fault-tolerant system based on multiple clusters

By introducing entropy regularized optimal transmission scheduling and the GMPC model, a linkage mechanism between the task scheduling layer and the coding fault tolerance layer is established, which solves the problems of low resource utilization and low fault tolerance efficiency in multi-cluster environments, realizes optimized allocation of cross-cluster resources and adaptive fault tolerance recovery, and improves computing reliability and resource utilization.

CN121833255APending Publication Date: 2026-04-10RUNYU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing scientific workflow management systems lack unified modeling and optimized scheduling in multi-cluster environments, resulting in low utilization of computing resources, high scheduling latency, and uneven task allocation. Furthermore, existing fault-tolerant methods incur significant computational and communication overhead, making it difficult to achieve efficient recovery in multi-cluster heterogeneous computing environments.

Method used

By introducing entropy regularized optimal transmission scheduling and the GMPC model, a linkage mechanism between the task scheduling layer and the coding fault tolerance layer is established. By constructing a topological anisotropic multi-element coding structure, the redundancy rate and transportation plan are dynamically adjusted to achieve cross-cluster resource optimization allocation and adaptive fault tolerance recovery.

Benefits of technology

It achieves dynamic optimal matching of tasks and resources, reduces the imbalance of cross-cluster task allocation and communication costs, improves the global utilization of computing resources, and significantly enhances the fault tolerance performance of the system in the event of node failure and network fluctuations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833255A_ABST
    Figure CN121833255A_ABST
Patent Text Reader

Abstract

The invention discloses a large-scale scientific computing workflow collaborative management and fault-tolerant system based on multiple clusters, and the system comprises a workflow and resource modeling module which obtains task and resource feature data; the entropy regular optimal transmission scheduling module is used for calculating a transportation plan result; the GMPC coding structure construction module is used for introducing cluster dimensions into a GMPC model and constructing a topological anisotropic multi-element coding structure; the parameter linkage control module is used for adaptively adjusting a redundancy rate and a cluster dimension index; the multi-cluster distributed encoding and computing module is used for executing encoding and parallel computing; the cross-cluster fault-tolerant recovery and reconstruction module is used for executing GMPC interpolation recovery and dynamically adjusting parameters when the performance is insufficient; the global result integration and feedback updating module is used for integrating reconstruction results and updating task execution states; and the operation monitoring and data management module is used for acquiring system operation information and uniformly managing data. According to the invention, the multi-cluster collaborative scheduling and fault-tolerant efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of scientific workflow management, and more particularly to a collaborative management and fault-tolerant system for large-scale scientific computing workflows based on multiple clusters. Background Technology

[0002] Currently, large-scale scientific computing tasks typically run in multi-cluster high-performance computing environments, relying on workflow systems for task decomposition and scheduling. However, existing scientific workflow management systems mostly use a single cluster as the scheduling unit, lacking a unified modeling and optimization scheduling mechanism for heterogeneous resources across clusters. This makes it difficult to achieve dynamic matching of tasks and resources across multiple clusters, resulting in low utilization of computing resources, high scheduling latency, and uneven task allocation.

[0003] On the other hand, existing multi-cluster computing environments typically rely on static redundancy or recalculation mechanisms for fault tolerance in the event of node failure, communication interruption, or loss of some computation results. These methods incur high computational and communication overhead, lack the ability to adaptively adjust based on the system's operating state, and cannot dynamically optimize redundancy parameters or coding strategies during task execution, easily leading to system performance degradation and resource waste.

[0004] Furthermore, existing distributed coding methods are mostly designed for single data distribution scenarios and fail to fully integrate cluster topology characteristics and task allocation structure. They are difficult to balance coding efficiency and fault tolerance and recovery performance, especially in multi-cluster heterogeneous computing environments where they exhibit problems such as high recovery threshold and long data reconstruction delays. Existing technologies urgently need a system that can achieve collaborative scheduling, dynamic fault tolerance and efficient reconstruction of scientific computing workflows in multi-cluster environments to improve the collaborative utilization efficiency and fault tolerance and recovery capabilities of multi-cluster resources.

[0005] Therefore, how to provide a collaborative management and fault-tolerant system for large-scale scientific computing workflows based on multiple clusters is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a collaborative management and fault-tolerant system for large-scale scientific computing workflows based on multiple clusters. By introducing entropy-regularized optimal transmission scheduling and the GMPC model, a linkage mechanism between the task scheduling layer and the coding fault-tolerant layer is established to achieve optimized resource allocation and adaptive fault-tolerant recovery across clusters. This invention constructs a topologically anisotropic multi-element coding structure and introduces cluster-dimensional adjustment parameters, which can dynamically adjust redundancy rates and transportation plans when nodes fail or loads are uneven, achieving efficient reconstruction of computational results. It possesses advantages such as strong scheduling collaboration, high fault-tolerant efficiency, and excellent computational reliability.

[0007] A large-scale scientific computing workflow collaborative management and fault-tolerant system based on multiple clusters according to an embodiment of the present invention includes: The workflow and resource modeling module is used to acquire task and resource characteristic data and generate a resource cost matrix and a task and resource mapping set. The entropy-regular optimal transmission scheduling module is used to calculate the transportation plan results based on the resource cost matrix and the task-resource mapping set, and output the task allocation entropy and allocation variance. The GMPC coding structure construction module is used to combine resource topology and transportation plan results, introduce a cluster dimension into GMPC, construct a topological anisotropic multivariate coding structure, and provide a multidimensional coding kernel, coding number vector and evaluation point set. The parameter linkage control module is used to adaptively adjust the redundancy rate and cluster dimension index based on the task allocation entropy and allocation variance, and generate a set of GMPC coding parameters. The multi-cluster distributed coding and computing module is used to perform coding and parallel computing according to the transportation plan results and the GMPC coding parameter set, and aggregate them to form a coding result set. The cross-cluster fault-tolerant recovery and reconstruction module is used to perform GMPC interpolation recovery based on the encoding result set and the GMPC encoding parameter set. When the value is lower than the recovery threshold, the parameters are adjusted and the transportation plan is recalculated, and the reconstruction result set is output. The global results integration and feedback update module is used to integrate reconstruction results and update task execution status and cluster health information; The operation monitoring and data management module is used to collect operation and failure information and manage the data in a unified manner to support scheduling and fault tolerance.

[0008] Optionally, modules can be integrated using the following methods: Establish a workflow and multi-cluster resource model, acquire task and resource characteristic data, generate a resource cost matrix, and form a task and resource mapping set; Based on the resource cost matrix and the task-resource mapping set, the entropy regularized optimal transmission model is executed to schedule the transportation plan, and the task allocation entropy and allocation variance are output. Based on the transportation plan results, the GMPC model is called to construct a topological anisotropic multivariate coding structure. A cluster dimension is introduced into the GMPC model to establish a multidimensional coding kernel and determine the coding number vector and the evaluation point set. Based on task allocation entropy, allocation variance, encoding frequency vector and cluster dimension, the execution parameters are linked to control, and the redundancy rate and cluster dimension index are adaptively adjusted to generate a set of GMPC encoding parameters. Based on the transportation plan results and the GMPC coding parameter set, multi-cluster distributed coding and task computation are performed. The task blocks and intermediate data are GMPC coded and distributed to each cluster node according to the evaluation point set for parallel operation. The coding result set returned by the nodes is collected. Based on the set of encoding results and the set of GMPC encoding parameters, cross-cluster fault-tolerant recovery and result reconstruction are performed. The GMPC interpolation mechanism is used to complete the result recovery and generate the reconstructed result set. Based on the reconstruction result set, perform global result integration and workflow feedback updates, perform dependency verification and result merging, and update task execution status and cluster health information.

[0009] Optionally, the task and resource feature data includes task dependencies, resource topology, node performance indicators, network status parameters, and failure statistics. The resource cost matrix is ​​generated by weighted normalization calculation based on node performance indicators, network status parameters, and failure statistics. The task and resource mapping set is generated by matching analysis based on task dependencies and resource topology.

[0010] Optionally, the generation of the task allocation entropy and allocation variance includes: Based on the resource cost matrix and the task-resource mapping set, the comprehensive cost of each task on different resource nodes is calculated. The execution cost, network transmission cost and failure rate are weighted and integrated to obtain the weighted cost data between tasks and resources. Construct an entropy-regularized optimal transmission model, set entropy-regularization parameters in the model to constrain the transmission probability distribution between tasks and resources, and generate an initial transportation plan matrix by combining weighted cost data; Through iterative optimization, the transportation plan matrix is ​​updated so that the matrix's rows, columns, and constraints simultaneously satisfy the total task demand and resource load capacity limits, resulting in a converged transportation plan. Based on the transportation plan results, the task allocation entropy is obtained by calculating the information entropy of the task allocation probability distribution in the transportation plan results, and the allocation variance is obtained by calculating the variance of the task load allocation probability in the transportation plan results.

[0011] Optionally, the generation of the encoding frequency vector and the evaluation point set includes: Based on the resource topology and transportation plan results, a cluster dimension parameter is introduced into the GMPC model to establish a mapping relationship between the cluster dimension and resource topology features. We perform weighted analysis on the bandwidth, latency, and failure rate parameters of each cluster, construct a set of topological feature weights, and initialize a multidimensional coding kernel in the GMPC model. In the GMPC model, based on the set of topological feature weights and cluster dimension parameters, the dimensional distribution and weight coefficients of the multidimensional coding kernel are adaptively adjusted to generate a topological anisotropic multidimensional coding structure. By jointly analyzing the dimensional distribution of the multidimensional coding kernel and the set of topological feature weights, the coding frequency vector is determined. Based on the resource topology and transportation plan results, the set of evaluation points is selected by comprehensively evaluating the communication delay and bandwidth distribution of each cluster node.

[0012] Optionally, the generation of the GMPC encoding parameter set includes: Based on task allocation entropy, allocation variance, encoding frequency vector and cluster dimension, a parameter linkage control model is constructed to establish the parameter mapping relationship between the task scheduling layer and the GMPC encoding layer. Based on the real-time changes in task allocation entropy and allocation variance, the redundancy rate and cluster dimension index in the parameter linkage control model are adaptively adjusted. Based on the adjustment results of the encoding number vector and cluster dimension, an optimal combination parameter set of redundancy rate and cluster dimension index is generated. The optimal set of combined parameters is input into the GMPC model to calculate the GMPC encoding parameter set.

[0013] Optionally, the generation of the encoding result set includes: Based on the transportation plan results and the GMPC coding parameter set, the task partitioning scheme and intermediate data processing scheme are determined. The task partitions to be executed and the intermediate data are organized according to the coding number vector and cluster dimension in the GMPC coding parameter set. Perform GMPC encoding on task blocks and intermediate data, generate encoded data for corresponding evaluation points according to the encoding number vector and cluster dimension, and correspond one-to-one with the set of evaluation points; Based on the evaluation point set, the coded task is divided into blocks and distributed to each cluster node for parallel execution. Calculations are performed according to the node allocation information in the evaluation point set, and corresponding evaluation point feedback data is generated. Based on the evaluation point set, the data returned by each cluster node is collected, and the returned data is merged and verified for consistency according to the transportation plan results to generate a set of encoded results.

[0014] Optionally, the generation of the reconstruction result set includes: Based on the set of encoding results and the set of GMPC encoding parameters, the integrity of the results and the recoverability of the data of each cluster node are detected, and the failed nodes are identified and a set of failed nodes is generated by combining the failure statistics information. By combining the set of failed nodes, the GMPC interpolation mechanism is invoked to perform result recovery. The encoding results corresponding to the set of evaluation points are used for interpolation reconstruction to restore the calculation results corresponding to the missing nodes and generate the initial recovery results. Based on the initial recovery results and task allocation entropy and allocation variance, the recovery performance index is calculated. When the recovery performance index is detected to be lower than the recovery threshold, the redundancy rate and cluster dimension index in the GMPC coding parameter set are dynamically adjusted according to the task allocation entropy and allocation variance, and the transportation plan results are recalculated. Iterative recovery is performed based on the updated GMPC coding parameter set and transportation plan results to generate a reconstruction result set.

[0015] Optionally, the global result integration and workflow feedback update include: Dependency verification is performed based on the reconstruction result set and task dependencies to confirm the output that meets the dependency conditions. The output results that meet the dependency conditions are merged to generate an updated result; Update the task execution status and cluster health information based on the update results, and record the task completion status and resource health status of the current round. The updated results are used as input for the next round of transportation planning and GMPC coding parameter set to complete the workflow feedback update and enter the next round of collaborative management and fault tolerance processing.

[0016] The beneficial effects of this invention are: First, it achieves dynamic optimal matching between tasks and resources, effectively reducing the imbalance in cross-cluster task allocation and communication costs, and improving the global utilization of computing resources.

[0017] Secondly, a topological anisotropic multivariate coding structure is constructed using the GMPC model, and a cluster dimension adjustment mechanism is introduced to enable the coding strategy to adaptively adjust the redundancy rate according to the cluster topology and task allocation status, thereby significantly enhancing the system's fault tolerance performance in the event of node failure and network fluctuations.

[0018] Furthermore, by establishing a coupling relationship between the task scheduling layer and the coding fault tolerance layer through a parameter linkage control model, dynamic linkage optimization of redundant parameters is achieved, reducing the resource waste caused by traditional fixed redundancy schemes.

[0019] Finally, this invention achieves cyclical updates of the execution status of multi-cluster tasks through a global result feedback mechanism, enabling the system to have continuous self-learning and self-optimization capabilities, thereby achieving collaborative, efficient, fault-tolerant, reliable, and dynamically scalable operation in complex distributed scientific computing environments. Attached Figure Description

[0020] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0021] Figure 1This is an overall flowchart of a large-scale scientific computing workflow collaborative management and fault-tolerant system based on multiple clusters proposed in this invention. Figure 2 This is a schematic diagram of the entropy regular optimal transmission scheduling mechanism in this invention; Figure 3 This is a schematic diagram of the GMPC encoding structure and parameter linkage control mechanism in this invention. Detailed Implementation

[0022] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0023] refer to Figure 1-3 A large-scale scientific computing workflow collaborative management and fault-tolerant system based on multiple clusters, including: The workflow and resource modeling module is used to receive task dependencies, resource topology, node performance indicators, network status parameters and failure statistics, perform standardized processing of task and resource feature data, generate a resource cost matrix and form a task and resource mapping set to support subsequent multi-cluster scheduling and mapping calculations. The entropy regularized optimal transmission scheduling module is used to construct an entropy regularized optimal transmission model based on the resource cost matrix and the task-resource mapping set. It sets entropy regularization parameters to constrain the transmission probability distribution between tasks and resources, generates an initial transportation plan matrix by combining weighted cost data, and obtains the transportation plan result through iterative optimization. At the same time, it calculates the task allocation entropy and allocation variance to characterize the uncertainty of task allocation and the resource load balance. The GMPC coding structure construction module is used to call the GMPC model based on the resource topology and transportation plan results. It introduces cluster dimension parameters into the model, establishes the mapping relationship between cluster dimension and resource topology features, forms a topological anisotropic multivariate coding structure, generates a multidimensional coding kernel, and determines the coding number vector and evaluation point set, which are used to implement coding and fault tolerance mechanisms in subsequent multi-cluster distributed computing. The parameter linkage control module is used to construct a parameter linkage control model based on task allocation entropy, allocation variance, encoding frequency vector and cluster dimension, establish the parameter mapping relationship between the task scheduling layer and the GMPC encoding layer, and dynamically adjust the redundancy rate and cluster dimension index according to the real-time changes of task allocation entropy and allocation variance. When the task allocation entropy increases, the redundancy rate is increased to enhance fault tolerance. When the allocation variance increases, the cluster dimension index is adjusted to improve computational balance. After the parameter adjustment converges, a GMPC encoding parameter set is generated to guide subsequent distributed computing and fault tolerance recovery. The multi-cluster distributed coding and computing module is used to perform GMPC coding on task blocks and intermediate data based on the transportation plan results and the GMPC coding parameter set. It distributes the coded task blocks to each cluster node for parallel operation by combining the coding number vector and the evaluation point set. It collects the feedback results from each node and performs merging, comparison and consistency verification based on the transportation plan results and the evaluation point set to form a coding result set. The cross-cluster fault-tolerant recovery and reconstruction module is used to perform fault-tolerant recovery and result reconstruction based on the encoding result set, GMPC encoding parameter set and failure statistics. It detects failed nodes and calls the GMPC interpolation mechanism to perform result recovery. It performs interpolation reconstruction based on the encoding results corresponding to the evaluation point set, generates initial recovery results and calculates recovery performance indicators. When the recovery performance indicators are detected to be lower than the preset recovery threshold, it dynamically adjusts the redundancy rate and cluster dimension index in the GMPC encoding parameter set based on task allocation entropy and allocation variance, recalculates the transportation plan results, and outputs the final reconstruction result set to achieve multi-cluster fault tolerance and adaptive recovery. The global result integration and feedback update module is used to perform global result integration and feedback update based on the reconstruction result set and task dependency relationship. It performs dependency verification and result merging on the task execution results, updates the task execution status and cluster health information, and feeds the updated results back to the entropy regularized optimal transmission scheduling module and parameter linkage control module as input for the next round of transportation plan and GMPC coding parameter set, realizing adaptive collaborative management and fault-tolerant closed loop of multi-cluster workflow. The operation monitoring and data management module is used to collect node performance indicators, network status parameters and failure statistics in real time, store and verify the consistency of key data generated by each module, maintain the global operation status monitoring and historical data tracking of the system, and provide continuous data support and performance feedback mechanism for workflow execution, scheduling optimization and fault tolerance recovery.

[0024] In this embodiment, the modules are interconnected using the following method: Establish a workflow and multi-cluster resource model, acquire task and resource characteristic data, including task dependencies, resource topology, node performance indicators, network status parameters and failure statistics, generate a resource cost matrix and form a task and resource mapping set; Based on the resource cost matrix and the task-resource mapping set, the entropy regularized optimal transmission model is executed to schedule the transportation plan, and the task allocation entropy and allocation variance are output. Based on the resource topology and transportation plan results, the GMPC model is called to construct a topological anisotropic multivariate coding structure. A cluster dimension is introduced into the GMPC model to establish a multidimensional coding kernel and determine the coding number vector and evaluation point set. Based on task allocation entropy, allocation variance, encoding frequency vector and cluster dimension, the execution parameters are linked to control, and the redundancy rate and cluster dimension index are adaptively adjusted to generate a set of GMPC encoding parameters to guide distributed computing and fault tolerance recovery. Based on the transportation plan results and the GMPC coding parameter set, multi-cluster distributed coding and task computation are performed. The task blocks and intermediate data are GMPC coded and distributed to each cluster node according to the evaluation point set for parallel operation. The coding result set returned by the nodes is collected. Based on the encoding result set, GMPC encoding parameter set and failure statistics, cross-cluster fault-tolerant recovery and result reconstruction are performed. The GMPC interpolation mechanism is used to complete the result recovery. When the recovery threshold is insufficient, the GMPC encoding parameter set is dynamically adjusted according to the task allocation entropy and allocation variance and the transportation plan is recalculated to generate the reconstruction result set, thus realizing multi-cluster fault-tolerant recovery. Based on the reconstruction result set and task dependencies, global result integration and workflow feedback updates are performed, dependency verification and result merging are carried out, task execution status and cluster health information are updated, and the updated results are used as input for the next round of transportation plan and GMPC coding parameter set, realizing adaptive collaboration and fault-tolerant closed loop of multi-cluster workflow.

[0025] In this embodiment, the task and resource feature data includes task dependencies, resource topology, node performance indicators, network status parameters, and failure statistics. The resource cost matrix is ​​generated by weighted normalization calculation based on node performance indicators, network status parameters, and failure statistics, and is used to characterize the execution cost of each task on different resource nodes. The task and resource mapping set is generated by matching analysis based on task dependencies and resource topology, and is used to describe the candidate allocation relationship of tasks among multiple cluster resources.

[0026] In this embodiment, the generation of task allocation entropy and allocation variance includes: Based on the resource cost matrix and the task-resource mapping set, the comprehensive cost of each task on different resource nodes is calculated. The execution cost, network transmission cost and failure rate are weighted and integrated to obtain the weighted cost data between tasks and resources. The execution cost represents the computation time or energy cost of executing a task on a resource node, derived from node computing power metrics. The network transmission cost represents the communication latency, bandwidth usage, and other costs required for task data transmission between nodes, derived from network status parameters. The failure rate represents the historical failure probability of a node or link, derived from failure statistics. Construct an entropy-regularized optimal transport model, set entropy regularization parameters in the model to constrain the transport probability distribution between tasks and resources, and generate an initial transport plan matrix by combining weighted cost data, where each matrix element represents the allocation probability of a task on the corresponding resource; The entropy-regular optimal transport model includes a cost function construction unit, an entropy-regular term constraint unit, a transport plan matrix solution unit, and a convergence determination unit. The system comprises the following components: a cost function construction unit, which establishes a cost matrix between tasks and resources based on weighted cost data to characterize the comprehensive allocation cost of task execution cost, network transmission cost, and failure rate; an entropy regularization constraint unit, which introduces an entropy regularization parameter into the optimal transmission objective function to smooth the probability distribution of the transportation plan matrix, thereby avoiding over-concentration and improving the stability of task allocation; a transportation plan matrix solving unit, which updates the transportation plan matrix using an iterative optimization method based on task demand constraints and resource capacity constraints, ensuring that the allocation probabilities of each task satisfy the row and column constraints and gradually converge to the optimal transportation plan result; and a convergence determination unit, which calculates the convergence error of the transportation plan matrix and the change in the cost function, outputs the final transportation plan result when the error is below a set threshold, and generates task allocation entropy and allocation variance to reflect the uncertainty of multi-cluster task allocation and resource load balance. Through iterative optimization, the transportation plan matrix is ​​updated so that the matrix's rows, columns, and constraints simultaneously satisfy the total task demand and resource load capacity limits, resulting in a converged transportation plan. Based on the transportation plan results, the task allocation entropy is obtained by calculating the information entropy of the task allocation probability distribution in the transportation plan results. This entropy is used to characterize the degree of uncertainty in multi-cluster task allocation. At the same time, the allocation variance is obtained by calculating the variance of the task load allocation probability in the transportation plan results. This variance is used to reflect the balance of resource load allocation. The system outputs transportation plan results, task allocation entropy, and allocation variance, and uses these results as input data for subsequent GMPC model construction and parameter linkage control to achieve collaborative control between the task scheduling layer and the coding fault tolerance layer.

[0027] In this embodiment, the generation of the encoding frequency vector and the evaluation point set includes: Based on the resource topology and transportation planning results, a cluster dimension parameter is introduced into the GMPC model to establish a mapping relationship between the cluster dimension and resource topology features, which is used to characterize the topological differences between different clusters. The GMPC model includes a multidimensional coding kernel, a multidimensional coding kernel modulation structure, a topology weight adjustment unit, a cluster dimension expansion unit, and an evaluation point generation unit. The multidimensional encoding kernel is constructed based on the principle of Generalized Multivariate Polynomial Codes (GMMC). It takes task-blocked data and intermediate data as input and performs polynomial encoding mapping within a multidimensional variable space, with each variable dimension corresponding to either a data dimension or a cluster dimension. The multidimensional encoding kernel modulation structure introduces a topological anisotropic modulation mechanism, transforming bandwidth, latency, and failure rate parameters in the resource topology features into encoding weight coefficients. This forms a one-to-one mapping between topological features and encoding coefficients, enabling structural differentiation of encoding across different clusters. The topology weight adjustment unit normalizes and updates the encoding weights based on the resource topology structure and transportation plan results. To ensure coding stability and recoverability in a multi-cluster environment, the cluster dimension extension unit introduces cluster dimension parameters into the multi-dimensional coding kernel of the GMPC model. By establishing a mapping relationship between the cluster dimension and resource topology features, a joint coding space of data dimension and cluster dimension is formed, enabling distinguishable coding representation between different clusters. The evaluation point generation unit determines the coding frequency vector and generates an evaluation point set based on the transportation plan results and cluster dimension parameters, ensuring interpolation consistency and cross-cluster recovery capability of the coding results across multiple cluster nodes. The output of the GMPC model includes the coding frequency vector, evaluation point set, and cluster dimension, which are used for adaptive adjustment of redundancy rate and cluster dimension index in the parameter linkage control stage, realizing the collaborative linkage between the task scheduling layer and the coding fault tolerance layer. We perform weighted analysis on the bandwidth, latency, and failure rate parameters of each cluster, construct a set of topological feature weights, which are used to adjust the encoding weights of each dimension in the GMPC model. We also initialize a multidimensional encoding kernel in the GMPC model to represent the encoding relationship between the data dimension and the cluster dimension. In the GMPC model, based on the set of topological feature weights and cluster dimension parameters, the dimension distribution and weight coefficients of the multidimensional encoding kernel are adaptively adjusted to generate a topologically anisotropic multidimensional encoding structure, so that the encoding structure can reflect the topological differences between multiple cluster resources. By jointly analyzing the dimensional distribution of the multidimensional coding kernel and the set of topological feature weights, the coding frequency vector is determined. This vector is used to set the frequency parameters of the multidimensional coding kernel in the data dimension and the cluster dimension. Based on the resource topology and transportation plan results, the communication latency and bandwidth distribution of each cluster node are comprehensively evaluated, and a set of evaluation points is selected to complete coding and recovery in a multi-cluster environment. Output the topological anisotropic multivariate coding structure, coding number vector, evaluation point set and cluster dimension, which are used for parameter linkage control and multi-cluster distributed coding and task computation.

[0028] In this embodiment, the generation of the GMPC encoding parameter set includes: Based on task allocation entropy, allocation variance, encoding frequency vector and cluster dimension, a parameter linkage control model is constructed to establish the parameter mapping relationship between the task scheduling layer and the GMPC encoding layer, which is used to determine the initial values ​​of redundancy rate and cluster dimension index. The parameter linkage control model includes a parameter mapping construction unit, an adjustment input parsing unit, a redundancy rate and cluster dimension index calculation unit, an adaptive adjustment unit, and a convergence determination unit. The parameter mapping construction unit establishes a multi-parameter mapping relationship between the task scheduling layer and the GMPC encoding layer, using task allocation entropy, allocation variance, encoding frequency vector, and cluster dimension as input variables to form a coupled mapping between task allocation features and encoding structure parameters. The adjustment input parsing unit monitors the changing trends of task allocation entropy and allocation variance in real time, calculates task allocation uncertainty and resource load balancing indicators, and uses these as dynamic inputs to the parameter-linked control model. The redundancy rate and cluster dimension index calculation unit calculates initial values ​​for redundancy rate and cluster dimension index based on the parameter mapping relationship and adjustment inputs. The redundancy rate is used to determine the redundancy in the GMPC model. The redundant coding ratio and cluster dimension index are used to determine the expansion depth of the multidimensional coding kernel in the cluster dimension direction; the adaptive adjustment unit is used to iteratively adjust the redundancy rate and cluster dimension index based on the real-time fluctuations of task allocation entropy and allocation variance, so that the redundancy rate increases when the task allocation entropy increases to improve fault tolerance, and the cluster dimension index is automatically adjusted to optimize resource balance when the allocation variance increases; the convergence judgment unit is used to detect the change range of redundancy rate and cluster dimension index in continuous iterations. When the parameter adjustment range is lower than the set convergence threshold, the final combination of redundancy rate and cluster dimension index is output, and a GMPC coding parameter set is generated to guide distributed computing and fault tolerance recovery. Based on the real-time changes in task allocation entropy and allocation variance, the redundancy rate and cluster dimension index in the parameter linkage control model are adaptively adjusted. When the task allocation entropy increases, the redundancy rate is increased to enhance fault tolerance. When the allocation variance increases, the cluster dimension index is adjusted to improve the balance of multi-cluster computing. Based on the adjustment results of the encoding number vector and cluster dimension, an optimal combination parameter set of redundancy rate and cluster dimension index is generated to control the encoding depth and cross-cluster redundancy ratio in the GMPC model. The optimal set of combined parameters is input into the GMPC model to calculate the GMPC coding parameter set, which is used to guide multi-cluster distributed computing and fault tolerance recovery, and output to subsequent distributed coding and task computing steps.

[0029] In this embodiment, the generation of the encoding result set includes: Based on the transportation plan results and the GMPC coding parameter set, the task partitioning scheme and intermediate data processing scheme are determined. The task partitions to be executed and the intermediate data are organized according to the coding number vector and cluster dimension in the GMPC coding parameter set. Perform GMPC encoding on task blocks and intermediate data, generate encoded data for corresponding evaluation points according to the encoding number vector and cluster dimension, and correspond one-to-one with the set of evaluation points; Based on the evaluation point set, the coded task is divided into blocks and distributed to each cluster node for parallel execution. Calculations are performed according to the node allocation information in the evaluation point set, and corresponding evaluation point feedback data is generated. Based on the evaluation point set, the data returned by each cluster node is collected, and the returned data is merged and verified for consistency according to the transportation plan results to generate a set of encoded results. Output a set of encoded results for cross-cluster fault tolerance recovery and result reconstruction.

[0030] In this embodiment, the generation of the reconstruction result set includes: Based on the set of encoding results and the set of GMPC encoding parameters, the integrity of the results and the recoverability of the data of each cluster node are detected, and the failed nodes are identified and a set of failed nodes is generated by combining the failure statistics information. By combining the set of failed nodes, the GMPC interpolation mechanism is invoked to perform result recovery. The encoding results corresponding to the set of evaluation points are used for interpolation reconstruction to restore the calculation results corresponding to the missing nodes and generate the initial recovery results. The GMPC interpolation mechanism is as follows: Based on the evaluation point set, available coding results are located; the minimum recovery requirement is determined according to the GMPC coding parameter set, and a subset of coding results that meet the recovery threshold is selected as the recovery input; data alignment and deduplication are performed on the recovery input; an interpolation equation system is constructed according to the order of the evaluation point set, and a numerical stability strategy is used for solving, including sorting and scaling the evaluation points to obtain the calculation results corresponding to the missing nodes and form the initial recovery result; consistency verification is performed on the initial recovery result, matching and verifying the initial recovery result with the task allocation relationship in the transportation plan result, and performing back-substitution verification against the known results in the coding result set; when consistency or accuracy does not meet the requirements, additional available coding results are introduced from the evaluation point set for incremental solving until the recovery threshold and accuracy threshold are met; the initial recovery result is output, and data items used for subsequent recovery performance index calculations are recorded. Based on the initial recovery results and task allocation entropy and allocation variance, a recovery performance index is obtained by calculating the consistency measure between the recovery results and the original transportation plan results. When the recovery performance index is detected to be lower than the recovery threshold, the redundancy rate and cluster dimension index in the GMPC coding parameter set are dynamically adjusted according to the task allocation entropy and allocation variance, and the transportation plan results are recalculated. Iterative recovery is performed based on the updated GMPC coding parameter set and transportation plan results to generate a reconstruction result set for multi-cluster fault-tolerant recovery and subsequent global result integration.

[0031] In this embodiment, the global result integration and workflow feedback update include: Dependency verification is performed based on the reconstruction result set and task dependencies to confirm the output that meets the dependency conditions for integration processing. The output results that meet the dependency conditions are merged to generate an updated result; Update the task execution status and cluster health information based on the update results, and record the task completion status and resource health status of the current round. The updated results are used as input for the next round of transportation planning and GMPC coding parameter set to complete the workflow feedback update and enter the next round of collaborative management and fault tolerance processing.

[0032] Example 1: To verify the feasibility of this invention in practice, it was applied to a large-scale climate model simulation task. This task involves multi-temporal and spatiotemporal resolution meteorological data analysis and coupled modeling, comprising approximately 28,000 computational tasks and a data size of approximately 45TB. The operating environment consists of four heterogeneous computing clusters with 640, 512, 384, and 256 nodes respectively. Inter-node communication latency ranges from 0.4 to 1.2 ms, and some nodes experience random failures. Traditional multi-cluster scheduling systems are prone to problems such as uneven task allocation, cross-cluster communication conflicts, and delayed fault recovery in such tasks, leading to decreased overall computational efficiency and increased task latency.

[0033] After deployment in this environment, the system of this invention completes unified modeling of task dependencies, resource topology, node performance, and failure statistics through the workflow and resource modeling module, generating a resource cost matrix and a task mapping set. Task allocation optimization is performed in the entropy-regular optimal transmission scheduling module, reducing communication costs by constraining the transmission probability distribution, and outputting transportation plan results and key indicators such as task allocation entropy and allocation variance. Subsequently, the system calls the GMPC model to establish a topological anisotropic multivariate coding structure, introducing cluster dimension parameters to achieve multi-cluster collaborative coding. The parameter linkage control module dynamically adjusts the redundancy rate and cluster dimension index during operation, achieving collaborative optimization between the task scheduling layer and the coding fault tolerance layer. The multi-cluster distributed coding and computation module completes task segmentation, coding, and distribution, with each cluster node performing parallel computation and transmitting result data. After detecting partial node failure, the cross-cluster fault tolerance recovery and reconstruction module automatically initiates the GMPC interpolation mechanism for recovery, and completes the integration of computation results and status updates in the global result integration and feedback update module.

[0034] Table 1. Performance comparison of different methods in multi-cluster scientific computing tasks.

[0035] As shown in Table 1, the overall performance of the system of this invention in multi-cluster scientific computing tasks is significantly better than that of traditional scheduling systems. The task completion rate increased from 86.2% to 92.4%, indicating that the task allocation is more balanced and the cross-cluster collaboration efficiency is higher; the average scheduling latency decreased from 142.7 seconds to 121.5 seconds, indicating that the entropy regularization optimal transmission mechanism effectively reduced the communication and transmission costs of cross-cluster tasks; the average node utilization increased from 78.6% to 89.1%, reflecting the dynamic matching capability of multi-cluster resource allocation; the average fault tolerance recovery time decreased from 56.3 seconds to 42.1 seconds, and the recovery success rate increased from 83.5% to 89.2%, indicating that the GMPC interpolation mechanism can quickly complete the result reconstruction under the condition of node failure; the task result accuracy increased from 86.8% to 91.3%, indicating that the system maintained high computational consistency during fault tolerance reconstruction; the resource load variance decreased from 0.27 to 0.18, and the number of scheduling optimization convergence rounds decreased from 9 to 7, indicating that the overall system operation is more stable and the scheduling process converges faster.

[0036] The performance improvement stems from the introduction of entropy-regularized optimal transmission scheduling and GMPC encoding mechanisms in a multi-cluster environment, which simultaneously optimizes task allocation and data redundancy strategies. Through entropy regularization constraints, tasks can achieve dynamic balanced distribution across different clusters according to the cost matrix, reducing communication bottlenecks. With the introduction of cluster-dimensional parameters into the GMPC model, the encoding process can automatically adjust the redundancy rate based on resource topology and task load, improving the system's adaptive fault tolerance performance. Furthermore, the parameter-linked control model establishes a feedback relationship between task allocation entropy and encoding parameters, enabling continuous optimization of the system in multiple rounds of scheduling, achieving a simultaneous improvement in task execution efficiency and fault recovery capabilities.

[0037] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A collaborative management and fault-tolerant system for large-scale scientific computing workflows based on multiple clusters, characterized in that: include: The workflow and resource modeling module is used to acquire task and resource characteristic data and generate a resource cost matrix and a task and resource mapping set. The entropy-regular optimal transmission scheduling module is used to calculate the transportation plan results based on the resource cost matrix and the task-resource mapping set, and output the task allocation entropy and allocation variance. The GMPC coding structure construction module is used to combine resource topology and transportation plan results, introduce a cluster dimension into GMPC, construct a topological anisotropic multivariate coding structure, and provide a multidimensional coding kernel, coding number vector and evaluation point set. The parameter linkage control module is used to adaptively adjust the redundancy rate and cluster dimension index based on the task allocation entropy and allocation variance, and generate a set of GMPC coding parameters. The multi-cluster distributed coding and computing module is used to perform coding and parallel computing according to the transportation plan results and the GMPC coding parameter set, and aggregate them to form a coding result set. The cross-cluster fault-tolerant recovery and reconstruction module is used to perform GMPC interpolation recovery based on the encoding result set and the GMPC encoding parameter set. When the value is lower than the recovery threshold, the parameters are adjusted and the transportation plan is recalculated, and the reconstruction result set is output. The global results integration and feedback update module is used to integrate reconstruction results and update task execution status and cluster health information; The operation monitoring and data management module is used to collect operation and failure information and manage the data in a unified manner to support scheduling and fault tolerance.

2. The large-scale scientific computing workflow collaborative management and fault-tolerant system based on multi-cluster architecture according to claim 1, characterized in that, The modules are connected in the following way: Establish a workflow and multi-cluster resource model, acquire task and resource characteristic data, generate a resource cost matrix, and form a task and resource mapping set; Based on the resource cost matrix and the task-resource mapping set, the entropy regularized optimal transmission model is executed to schedule the transportation plan, and the task allocation entropy and allocation variance are output. Based on the transportation plan results, the GMPC model is called to construct a topological anisotropic multivariate coding structure. A cluster dimension is introduced into the GMPC model to establish a multidimensional coding kernel and determine the coding number vector and the evaluation point set. Based on task allocation entropy, allocation variance, encoding frequency vector and cluster dimension, the execution parameters are linked to control, and the redundancy rate and cluster dimension index are adaptively adjusted to generate a set of GMPC encoding parameters. Based on the transportation plan results and the GMPC coding parameter set, multi-cluster distributed coding and task computation are performed. The task blocks and intermediate data are GMPC coded and distributed to each cluster node according to the evaluation point set for parallel operation. The coding result set returned by the nodes is collected. Based on the set of encoding results and the set of GMPC encoding parameters, cross-cluster fault-tolerant recovery and result reconstruction are performed. The GMPC interpolation mechanism is used to complete the result recovery and generate the reconstructed result set. Based on the reconstruction result set, perform global result integration and workflow feedback updates, perform dependency verification and result merging, and update task execution status and cluster health information.

3. The large-scale scientific computing workflow collaborative management and fault-tolerant system based on multi-cluster architecture according to claim 2, characterized in that, The task and resource feature data includes task dependencies, resource topology, node performance indicators, network status parameters, and failure statistics. The resource cost matrix is ​​generated by weighted normalization calculation based on node performance indicators, network status parameters, and failure statistics. The task and resource mapping set is generated by matching analysis based on task dependencies and resource topology.

4. The large-scale scientific computing workflow collaborative management and fault-tolerant system based on multi-cluster architecture according to claim 2, characterized in that, The generation of the task allocation entropy and allocation variance includes: Based on the resource cost matrix and the task-resource mapping set, the comprehensive cost of each task on different resource nodes is calculated. The execution cost, network transmission cost and failure rate are weighted and integrated to obtain the weighted cost data between tasks and resources. Construct an entropy-regularized optimal transmission model, set entropy-regularization parameters in the model to constrain the transmission probability distribution between tasks and resources, and generate an initial transportation plan matrix by combining weighted cost data; Through iterative optimization, the transportation plan matrix is ​​updated so that the matrix's rows, columns, and constraints simultaneously satisfy the total task demand and resource load capacity limits, resulting in a converged transportation plan. Based on the transportation plan results, the task allocation entropy is obtained by calculating the information entropy of the task allocation probability distribution in the transportation plan results, and the allocation variance is obtained by calculating the variance of the task load allocation probability in the transportation plan results.

5. The large-scale scientific computing workflow collaborative management and fault-tolerant system based on multi-cluster architecture according to claim 2, characterized in that, The generation of the encoding frequency vector and the evaluation point set includes: Based on the resource topology and transportation plan results, a cluster dimension parameter is introduced into the GMPC model to establish a mapping relationship between the cluster dimension and resource topology features. We perform weighted analysis on the bandwidth, latency, and failure rate parameters of each cluster, construct a set of topological feature weights, and initialize a multidimensional coding kernel in the GMPC model. In the GMPC model, based on the set of topological feature weights and cluster dimension parameters, the dimensional distribution and weight coefficients of the multidimensional coding kernel are adaptively adjusted to generate a topological anisotropic multidimensional coding structure. By jointly analyzing the dimensional distribution of the multidimensional coding kernel and the set of topological feature weights, the coding frequency vector is determined. Based on the resource topology and transportation plan results, the set of evaluation points is selected by comprehensively evaluating the communication delay and bandwidth distribution of each cluster node.

6. The large-scale scientific computing workflow collaborative management and fault-tolerant system based on multiple clusters according to claim 2, characterized in that, The generation of the GMPC encoding parameter set includes: Based on task allocation entropy, allocation variance, encoding frequency vector and cluster dimension, a parameter linkage control model is constructed to establish the parameter mapping relationship between the task scheduling layer and the GMPC encoding layer. Based on the real-time changes in task allocation entropy and allocation variance, the redundancy rate and cluster dimension index in the parameter linkage control model are adaptively adjusted. Based on the adjustment results of the encoding number vector and cluster dimension, an optimal combination parameter set of redundancy rate and cluster dimension index is generated. The optimal set of combined parameters is input into the GMPC model to calculate the GMPC encoding parameter set.

7. The large-scale scientific computing workflow collaborative management and fault-tolerant system based on multiple clusters according to claim 2, characterized in that, The generation of the encoding result set includes: Based on the transportation plan results and the GMPC coding parameter set, the task partitioning scheme and intermediate data processing scheme are determined. The task partitions to be executed and the intermediate data are organized according to the coding number vector and cluster dimension in the GMPC coding parameter set. Perform GMPC encoding on task blocks and intermediate data, generate encoded data for corresponding evaluation points according to the encoding number vector and cluster dimension, and correspond one-to-one with the set of evaluation points; Based on the evaluation point set, the coded task is divided into blocks and distributed to each cluster node for parallel execution. Calculations are performed according to the node allocation information in the evaluation point set, and corresponding evaluation point feedback data is generated. Based on the evaluation point set, the data returned by each cluster node is collected, and the returned data is merged and verified for consistency according to the transportation plan results to generate a set of encoded results.

8. The large-scale scientific computing workflow collaborative management and fault-tolerant system based on multiple clusters according to claim 2, characterized in that, The generation of the reconstruction result set includes: Based on the set of encoding results and the set of GMPC encoding parameters, the integrity of the results and the recoverability of the data of each cluster node are detected, and the failed nodes are identified and a set of failed nodes is generated by combining the failure statistics information. By combining the set of failed nodes, the GMPC interpolation mechanism is invoked to perform result recovery. The encoding results corresponding to the set of evaluation points are used for interpolation reconstruction to restore the calculation results corresponding to the missing nodes and generate the initial recovery results. Based on the initial recovery results and task allocation entropy and allocation variance, the recovery performance index is calculated. When the recovery performance index is detected to be lower than the recovery threshold, the redundancy rate and cluster dimension index in the GMPC coding parameter set are dynamically adjusted according to the task allocation entropy and allocation variance, and the transportation plan results are recalculated. Iterative recovery is performed based on the updated GMPC coding parameter set and transportation plan results to generate a reconstruction result set.

9. The large-scale scientific computing workflow collaborative management and fault-tolerant system based on multiple clusters according to claim 2, characterized in that, The global result integration and workflow feedback update include: Dependency verification is performed based on the reconstruction result set and task dependencies to confirm the output that meets the dependency conditions. The output results that meet the dependency conditions are merged to generate an updated result; Update the task execution status and cluster health information based on the update results, and record the task completion status and resource health status of the current round. The updated results are used as input for the next round of transportation planning and GMPC coding parameter set to complete the workflow feedback update and enter the next round of collaborative management and fault tolerance processing.