Distributed computing method and system combined with coding fault tolerance
By adopting coding fault tolerance methods and load balancing technology in distributed computing, the inaccuracy and system instability caused by computing node errors are solved, the computing efficiency and resource utilization are improved, and the accuracy of the calculation results and system stability are ensured.
Patent Information
- Application Number
- CN202410042648.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-11
AI Technical Summary
In distributed computing, the calculation results in inaccuracy and system instability caused by errors and interruptions of computing nodes, especially in quantum chemocomputing, existing redundant computing methods lead to waste of resources and inefficiency.
The encoding fault tolerance method is adopted to perform redundant calculations on important tasks by generating encoding matrix and decoding matrix, and combining static and dynamic load balancing to ensure the accuracy of the calculation results and system stability.
It improves the fault tolerance and computing efficiency of distributed computing, reduces waste of computing resources, and ensures the accuracy of calculation results and the robustness of the system.
Smart Images

Figure CN120295827A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of high-performance computing, distributed parallel computing, and high-throughput computing; it relates to a distributed computing method and system combined with coding fault tolerance. Specifically, in distributed computing scenarios such as a computing system where the target task is decomposable or high-throughput computing tasks, by introducing coding theory into distributed computing and encoding the computing results of distributed nodes, the computing errors of a certain number of nodes will not affect the final result, thereby enhancing the computing fault tolerance of the entire system and improving the efficiency and stability of distributed computing. Background Art
[0002] Distributed computing technology integrates computing resources, provides high-performance scientific computing capabilities, and provides a solid foundation for completing complex scientific computing. With the wide application of distributed computing in new technological trends such as high-throughput, big data, and artificial intelligence, ensuring the correctness of the computing results of each node in distributed computing has become increasingly important. The accuracy of the node computing results directly affects the overall correctness and reliability of the system, and is crucial for ensuring the accuracy and credibility of scientific computing, data analysis, and other complex tasks. By having the computing nodes perform additional computations and encoding the computing results of each node, the fault tolerance and stability of the entire system can be effectively improved.
[0003] For example, in quantum chemistry calculations, many computing tasks are deployed on a distributed computing system for computing acceleration. However, in the actual process of quantum chemistry calculations, various interruptions, errors, and other abnormal situations are inevitable. Even after excluding subjective factors (such as structural problems, parameter settings, etc.), there will still be computing anomalies caused by the electrical performance of computing devices, initial value selection, random parameters, convergence algorithms, etc. For example, in high-throughput calculations based on the popular density functional theory (DFT), there are often a very small number of calculations showing unreasonable convergence trends (such as the energy curve jumping up and down), which may lead to convergence to incorrect results or even interruption of the computing task. Therefore, it is necessary to encode the computing results through coded computing to improve the fault tolerance and robustness of the system. In scenarios where the frequency of computing anomalies is relatively high, coded computing can not only enhance the stability of the system, but also improve the overall efficiency of the system, and avoid the situation of continuous repeated computing due to computing errors.
[0004] In recent years, researchers have applied coding theory to the field of distributed computing and proposed Coded Distributed Computing (CDC). By leveraging the flexibility of coding and through redundant computing, on the one hand, it can reduce the time and load required for communication, and on the other hand, it can alleviate problems caused by straggler nodes or computational errors through coding. In terms of alleviating the straggler delay in distributed computing, Sun et al. proposed the Hierarchical Short-Dot (HSD) coded distributed computing scheme based on hierarchical coding. Through hierarchical coding, straggler nodes are made to undertake a small amount of computational tasks, thus reducing the time required for computation while avoiding waste of computing resources. In the cross-field of distributed computing and artificial intelligence, Tandon et al. proposed the Gradient Coding scheme, which utilizes the additional computation and storage of worker nodes, enabling popular distributed machine learning processes to tolerate partially randomly straggling or computationally erroneous nodes, greatly improving the stability of distributed machine learning computations.
[0005] It should be noted that the approach of enhancing system stability through redundant computing will also result in a relatively low utilization rate of computing resources. Using additional resources for redundant computing in a stable system will undoubtedly cause waste of computing resources. In the field of quantum computing, the probabilities of computational errors for different computational tasks vary. A possible improvement scheme is to perform redundant computing for some specific tasks, which can improve the utilization efficiency of computing power as much as possible. Summary of the Invention
[0006] To further improve the fault tolerance and computational stability of the system, the present invention proposes a distributed computing method and system combining coding fault tolerance, which has been applied to distributed and high-throughput calculations of first principles, mainly used in the fields of quantum chemistry and quantum physics. The tasks include single-point calculations for given structures, configuration optimization calculations, wave function solving, and property calculations. In the field of quantum chemistry, when calculating the properties of large molecular systems, the molecules are usually segmented, and then these segmented molecules are calculated separately by a distributed cluster. For these large quantum tasks to be calculated, each task can use 1 to n computing resources (such as computing nodes). For these numerous subtasks, redundant calculations can be performed on some important tasks according to custom rules and encoded. For example, by marking the binding site region of drug molecule - protein or detecting whether it is a material system containing specific elements (such as rare earth elements), important tasks are selected for redundant calculations. These selected subtasks are divided into subtasks to be encoded that need to perform redundant calculations, and the remaining subtasks are used as subtasks to be calculated only once. The subtasks to be encoded participate in static load balancing together with the subtasks that do not need to be encoded according to the encoding matrix generated by the set redundancy and the number of subtasks to be encoded. It should be noted that in this patent, "redundancy" represents the number of nodes or straggler nodes allowed to make mistakes in distributed computing. For example, when the redundancy is 1, each task that needs to perform redundant calculations will have 1 identical copy task. Static load balancing allocation can be based on prior experience or use machine learning means to batch predict the computing time required for these subtasks, and then allocate them through a planning scheme such as the Greedy algorithm to obtain a task list. In the task list, the subtasks to be encoded and their copies have the highest priority and are all at the front of the task list. After the static load planning is completed, for the tasks at the back of the task list (excluding the tasks that need to be encoded), a dynamic load balancing scheme is introduced to maximize the utilization of the computing power of the computing cluster. After the calculation is completed, each computing node processing the subtasks to be encoded encodes the calculation results of the subtasks to be encoded according to the encoding matrix. Then, after the encoding results are transmitted to the master node at each computing node, the results are decoded according to the decoding matrix generated by the encoding matrix to obtain the final calculation results. In the final calculation results, if there are no calculation errors, several identical calculation results will be obtained, and this result is the correct result; if there are nodes with less than the redundancy making calculation errors, some error results and several identical correct results will be obtained, and the total number of all calculation results is the combination number of all computing nodes minus the redundancy. During this process, calculation errors or delays in the number of nodes below the set redundancy will not affect the final calculation results, and the stability and robustness of the entire system are greatly guaranteed.
[0007] The technical solution of the present invention is as follows:
[0008] A distributed computing method combining coding fault tolerance, the steps of which include:
[0009] 1) The input module performs rationality analysis and format conversion on the received batch computing tasks to be calculated, and then generates input files and sends them to the prediction module;
[0010] 2) The prediction module predicts the computing time required for each computing task and sends it to the coding module;
[0011] 3) The coding module selects several computing tasks as the sub-tasks to be encoded that need to perform redundant computing according to the computing time of each computing task, and takes the remaining computing tasks as the sub-tasks to be calculated only once; then generates a coding matrix and a decoding matrix according to the set redundancy and the number of sub-tasks to be encoded; then groups the sub-tasks to be encoded according to the coding matrix to generate grouped files, and plans and performs load balancing on the encoded computing tasks and the sub-tasks to be calculated only once to obtain a static planning result; then sends the static planning result and the coding matrix to the computing module; sends the decoding matrix and the grouped files to the decoding module;
[0012] 4) The computing module distributes each computing task to the corresponding computing node for calculation according to the static planning result, and sends the coding matrix to the corresponding computing node; each computing node processing the sub-tasks to be encoded encodes the calculation result according to the coding matrix and sends it to the decoding module, and the computing node processing the sub-tasks to be calculated only once directly sends its calculation result to the decoding module;
[0013] 5) The decoding module decodes the calculation results received from each computing node according to the grouped files and the decoding matrix,
[0014] to obtain the calculation result of the batch computing tasks to be calculated.
[0015] Further, the method for the encoding module to generate the encoding matrix is as follows: generate a random matrix H with s rows and n columns according to the redundancy s and the number of computing nodes n, and operate on the column with index n - 1 of the random matrix H to make it equal to the opposite of the sum of all columns before this column; initialize an n - order matrix B with all elements being zero; then generate the elements of matrix B in a loop. The elements on the diagonal position of matrix B are 1. In the i - th row, j is an increasing index sequence with a length of s starting from i + 1 and an increment of 1. The element B(i, j) in the i - th row of matrix B and with the column index in the sequence j is equal to the matrix - H(:, j(1:s)) divided by the vector H(:, j(0)), where j(1:s) refers to all elements in the sequence j from the element with index 1 to the element with index s, and the matrix - H(:, j(1:s)) is a matrix composed of all columns in the - H matrix with indexes in j(1:s), the number of columns is equal to the number of elements in j(1:s), and the vector H(:, j(0)) is the j(0) - th column in matrix H, and j(0) is the first element of the sequence j.
[0016] Further, the method for the encoding module to generate the decoding matrix is as follows: take the combination number of arbitrarily selecting s elements from n different elements as f, and initialize a matrix A with f rows and n columns with all elements being zero; then generate all combinations of selecting n - s elements from 0 to n - 1 and store them in matrix I. Each row in matrix I represents a combination; then traverse matrix I row - by - row. Let i represent the current row number and ii represent the current combination, that is, I(i). Initialize a vector a with all elements being zero and a length of n. Update the vector a(ii) = ones(1, k) / B(ii, :), where ones(1, k) represents a row vector with a length of k and all elements being 1, B(ii, :) is a matrix composed of all rows in matrix B with the row index in ii, a(ii) is all elements in vector a with the index in ii, and the division " / " is element - by - element division. The i - th row of matrix A is the vector a.
[0017] Further, perform dimensionality reduction on the encoding matrix and the decoding matrix respectively; group the encoded computing tasks according to the dimensionality - reduced encoding matrix, and then perform planning and load balancing with the subtasks that are only calculated once.
[0018] Further, use the greedy algorithm for planning and load balancing, and continuously plan the computing task with the longest time consumption to the computing node with the least computing tasks.
[0019] Further, in the static planning result, the subtasks to be encoded and their copies are set to the highest priority, and a dynamic load - balancing scheme is introduced for the subtasks that are only calculated once in the static planning result for planning and load balancing.
[0020] A distributed computing system combined with coding error tolerance, characterized by comprising an input module, a prediction module, an encoding module, a computing module and a decoding module; wherein
[0021] The input module is used to perform rationality analysis and format conversion on the received batch computing tasks to be calculated, and then generate input files and send them to the prediction module;
[0022] The prediction module is used to predict the computing time required for each computing task and send it to the encoding module;
[0023] The encoding module is used to take several computing tasks as sub-tasks to be encoded that need to perform redundant computing according to the computing time of each computing task, and take the remaining computing tasks as sub-tasks that are only calculated once; then generate an encoding matrix and a decoding matrix according to the set redundancy and the number of sub-tasks to be encoded; then group the sub-tasks to be encoded according to the encoding matrix to generate grouped files, and perform planning and load balancing on the encoded computing tasks and the sub-tasks that are only calculated once to obtain a static planning result; then send the static planning result and the encoding matrix to the computing module; send the decoding matrix and the grouped files to the decoding module;
[0024] The computing module is used to allocate each computing task to the corresponding computing node for calculation according to the static planning result, and send the encoding matrix to the corresponding computing node; each computing node processing the sub-tasks to be encoded encodes the calculation result according to the encoding matrix and sends it to the decoding module, and the computing node processing the sub-tasks that are only calculated once directly sends its calculation result to the decoding module;
[0025] The decoding module decodes the calculation results received from each computing node according to the grouped files and the decoding matrix to obtain the calculation result of the batch computing tasks to be calculated.
[0026] The advantages of the present invention are as follows:
[0027] 1. In the scenario of distributed computing, it can significantly improve the error tolerance of computing tasks and avoid data verification and secondary calculation caused by individual errors in computing tasks, such as computing errors and node failures.
[0028] 2. In the scenario of distributed computing, it can improve the computing efficiency and reliability of the overall task.
[0029] 3. In the scenario where the error rate of computing nodes is relatively high, it can significantly improve the overall computing efficiency of the system, and only some sub-tasks need to be encoded, reducing the waste of computing resources. Description of the Drawings
[0030] Figure 1 It is the overall architecture of the system of the present invention.
[0031] Figure 2 This is a schematic diagram of grouping and matrix dimensionality reduction in the system of the present invention;
[0032] (a) Before matrix dimensionality reduction, (b) After matrix dimensionality reduction.
[0033] Figure 3 This is the process in which the encoding subtask and the non-encoding subtask jointly participate in load planning in the system of the present invention.
[0034] Figure 4 This is the decoding process of the system of the present invention under different conditions;
[0035] (a) Simple encoding, (b) Multiplexing encoding (i.e., encoding after matrix dimensionality reduction), (c) When encoding-non-encoding hybrid. Detailed implementation manner
[0036] The present invention will be further described in detail below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0037] This solution can be divided into five main modules: an input module, a prediction module, an encoding module, a calculation module, and a decoding module. The overall process architecture is as shown in the appendix Figure 1 as shown. Each module is briefly described as follows:
[0038] (1) Input module
[0039] This module is responsible for receiving batch calculation tasks to be calculated. This module can also perform functions such as rationality analysis of each calculation task, format conversion, or input file generation. Finally, the input module passes this information to the prediction module. The calculation task is the input file required for the calculation; for quantum chemistry calculations, the calculation task may include the structure of the molecule and the calculation method. For the rationality analysis of quantum chemistry calculations, specifically including: 1. The rationality of atomic bonding in the calculation task, and hydrogen supplementation operations can be performed if it is unreasonable; 2. Check the charge and spin multiplicity of the molecular structure in the calculation task to be calculated, and correction can be performed if it is unreasonable; 3. Repeated check of the atomic coordinate positions in the calculation task, if there are repetitions, the calculation task is cleared from the task list and an output prompt is made. After the above processing, the result is that an input file of a reasonable calculation task can be obtained, avoiding many error reports caused by the above factors; at the same time, the reasonable input file also ensures that the prediction module can give a reliable calculation time prediction result.
[0040] (2) Prediction module
[0041] This module is mainly responsible for predicting the computing time required for each computing task. For this part, we can use the computing time prediction module (ACSOmega 6_2001_2021) developed independently by us, which is based on cheminformatics and various machine learning models.
[0042] (3) Encoding module
[0043] This module mainly generates an encoding matrix and a decoding matrix based on the computing time data provided by the prediction module, and groups, plans, and performs load balancing on the encoded subtasks and the subtasks that do not require encoding. First, according to a function set by a custom rule, many subtasks are divided into two parts: subtasks that require encoding and subtasks that do not require encoding. Then, an encoding matrix and a decoding matrix are generated based on the set redundancy and the number of subtasks to be encoded. The generation methods of the encoding matrix and the decoding matrix are as shown in Algorithm 1 and Algorithm 2. It should be noted that the generated encoding matrix and decoding matrix are static. According to the algorithm, when the redundancy, the number of subtasks to be encoded, and the number of nodes are determined, the encoding matrix and the decoding matrix can be generated and determined, and the encoding matrix and the decoding matrix correspond one by one.
[0044] The algorithms for generating the encoding matrix and the decoding matrix will be introduced below:
[0045] First, assume that the redundancy is s and the number of nodes is n. In the following algorithms, matrix indices start from 0, and for convenience of description, all coordinates greater than or equal to n will be regarded as mod n. For example, if a coordinate of (n, n + 2) appears, the actual coordinate represented is (0, 2).
[0046] Algorithm 1:
[0047] First, generate a random matrix H with s rows and n columns, and operate on the column with index n - 1 of matrix H, making it equal to the opposite of the sum of all columns before this column. Then initialize a zero matrix B of order n, and then loop to generate the elements of matrix B. The elements on the diagonal position of matrix B are 1. In the i-th row, let j be an increasing index sequence of length s starting from i + 1 with an increment of 1. It should be noted that the elements in the index sequence j in each row are all different, and all elements in the sequence j will be modulo n to ensure that all elements in the sequence j are between 0 and n - 1, that is, all elements in the sequence j are indices. B(i, j) is equal to the matrix -H(:, j(1:s)) divided by the vector H(:, j(0)), where j(1:s) refers to all elements in the sequence j from the element with index 1 to the element with index s, and the matrix -H(:, j(1:s)) is a matrix composed of all columns in the -H matrix with indices in j(1:s), and the number of columns is equal to the number of elements in j(1:s), and the vector H(:, j(0)) is the j(0)-th column in matrix H, and j(0) is the first element of the sequence j.
[0048] Algorithm 2:
[0049] f is the number of combinations of choosing s elements from n different elements.
[0050] Create a zero matrix A with f rows and n columns, then generate all possible combinations of choosing n - s elements from 0 to n - 1 and store them in matrix I. Each row in matrix I represents a possible combination.
[0051] Traverse matrix I row by row. Let i represent the current row number and ii represent the current combination, that is, the combination I(i) of the i-th row in matrix I. Initialize a zero vector a with length n. Vector a(ii) = ones(1,k) / B(ii,:), where ones(1,k) represents a row vector with length k and all elements being 1, and B(ii,:) is the matrix composed of all rows in matrix B whose row indices are in ii. a(ii) refers to all elements in vector a with indices in ii. Here, the division is element-wise division in Matlab. The i-th row of matrix A is then a.
[0052] In addition, to avoid excessive computing power consumption due to large matrices during encoding and decoding, we respectively reduced the dimensions of the encoding matrix and the decoding matrix, as Figure 2 shown. According to the redundancy, copies with the number of sub-tasks to be encoded equal to the redundancy will be generated. The sub-tasks to be encoded and their copies are grouped according to the reduced-dimension encoding matrix. Tasks in the same group include multiple sub-tasks and their corresponding redundant tasks, and tasks in the same group will be encoded by the same reduced-dimension encoding matrix. After obtaining the grouped files, all sub-tasks that need to be encoded and those that do not need to be encoded participate in static load balancing together. The greedy algorithm is used during planning, continuously planning the task with the longest execution time to the computing node with the least computing tasks, so as to make the expected load on each computing unit (computing node) as close as possible, as Figure 3 shown.
[0053] (4) Calculation module
[0054] The computing module receives the static planning result (i.e., the task list) transmitted by the encoding module, and distributes the tasks to each computing node according to the result of the static planning. Each computing node will receive multiple computing tasks and perform calculations accordingly. During the calculation process, the computing module can dynamically adjust several computing tasks with lower rankings in the static planning or tasks with calculation times less than the set time length to idle computing nodes, that is, to achieve task scheduling that combines static and dynamic load balancing. To efficiently implement real-time task scheduling that combines static and dynamic aspects, we have independently developed a distributed computing engine based on the Julia high-performance computing language (J. Comput. Chem. 44_1174_2023). Compared with the default queuing scheduling systems such as SLURM and LSF, this computing engine can be regarded as a dedicated computing scheduling system with static and dynamic job scheduling capabilities.
[0055] (5) Decoding module
[0056] The decoding module receives the grouped file and the decoding matrix transmitted by the encoding module on the master node, and receives the calculation results of each computing node in the computing module, as Figure 4 shown. In the grouped file, the numbers of the encoded subtasks in the same group are placed on the same line separated by spaces. Each group can be found through the grouped file. When decoding, it is done in groups. The computing tasks contained in each group can be recombined into a matrix in order. Performing matrix operations on the decoding matrix and the matrices recombined from each group can obtain the final calculation result. When a computing error occurs in a node below the redundancy level, the correct calculation result can still be obtained.
[0057] The machine learning-assisted load balancing method and its implementation system described in the present invention can be implemented by the Python and Julia computer programming languages. In the current implementation, we use the Julia language to implement the input module; the prediction module uses the computer time prediction module based on the Python language developed by the research group in the early stage (ACSOmega_6_2001_2021); the encoding module is also programmed in the Python language based on the data of the prediction module; the computing module is programmed in the Julia language, mainly applying the distributed computing function of the Julia language, and is compatible with the task scheduling systems of the cluster (such as the tested SLURM, LSF, etc., J. Comput. Chem. 44_1174_2023). The decoding module is programmed in the Julia language.
[0058] Although specific embodiments of the present invention are disclosed for illustrative purposes, which are intended to help understand the content of the present invention and implement it accordingly, those skilled in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the best embodiments, and the scope of protection required by the present invention shall be defined by the scope defined in the claims.
Claims
1. A distributed computing method combined with encoding fault tolerance, the steps of which include: 1) The input module performs rationality analysis and format conversion on the received batch computing tasks to be calculated, and then generates input files and sends them to the prediction module; 2) The prediction module predicts the computing time required for each computing task and sends it to the encoding module; 3) The encoding module selects several computing tasks as the sub-tasks to be encoded that need to perform redundant computing according to the computing time of each computing task, and regards the remaining computing tasks as the sub-tasks to be calculated only once; then generates an encoding matrix and a decoding matrix according to the set redundancy and the number of sub-tasks to be encoded; Then, group the sub-tasks to be encoded according to the encoding matrix to generate grouped files, and perform planning and load balancing on the encoded computing tasks and the sub-tasks to be calculated only once to obtain a static planning result; then send the static planning result and the encoding matrix to the computing module; send the decoding matrix and the grouped files to the decoding module; 4) The computing module distributes each computing task to the corresponding computing node for calculation according to the static planning result, and sends the encoding matrix to the corresponding computing node; each computing node processing the sub-tasks to be encoded encodes the calculation result according to the encoding matrix and sends it to the decoding module, and the computing node processing the sub-tasks to be calculated only once directly sends its calculation result to the decoding module; 5) The decoding module decodes the calculation results received from each computing node according to the grouped files and the decoding matrix, to obtain the calculation result of the batch computing tasks to be calculated.
2. The method according to claim 1, wherein The method for the encoding module to generate the encoding matrix is: generate a random matrix H with s rows and n columns according to the redundancy s and the number of computing nodes n, and operate on the column with index n - 1 of the random matrix H to make it equal to the opposite of the sum of all columns before this column; initialize an n-order matrix B to zero; then loop to generate the elements of matrix B, and the elements on the diagonal of matrix B are 1; set an increasing index sequence j with a length of s starting from i + 1 in the i-th row, and the increment is 1; in the i-th row of matrix B and the column index in the sequence j, B(i,j) is equal to the matrix -H(:,j(1:s)) divided by the vector H(:,j(0)), where j(1:s) refers to all elements from the element with index 1 to the element with index s in the sequence j, and the matrix -H(:,j(1:s)) is a matrix composed of all columns in the -H matrix with indices in j(1:s), and the number of columns is equal to the number of elements in j(1:s), and the vector H(:,j(0)) is the j(0)-th column in matrix H, and j(0) is the first element of the sequence j.
3. The method according to claim 2, wherein The method for the encoding module to generate the decoding matrix is as follows: Take the combination number of arbitrarily selecting s elements from n different elements as f, and initialize a matrix A of f rows and n columns with all elements being zero. Then generate all combinations of selecting n - s elements from 0 to n - 1 and store them in matrix I, where each row in matrix I represents a combination. Then traverse matrix I row by row. Let i represent the current row number and ii represent the current combination, that is, I(i). Initialize a vector a of length n with all elements being zero, and update vector a(ii) = ones(1,k) / B(ii,:), where ones(1,k) represents a row vector of length k with all elements being 1, B(ii,:) is the matrix composed of all rows in matrix B whose row indices are in ii, a(ii) are all elements in vector a whose indices are in ii, and the division " / " is element-wise division, so that the i-th row of matrix A is vector a.
4. The method according to claim 1 or 2 or 3, characterized in that, Reduce the dimensions of the encoding matrix and the decoding matrix respectively; group the encoded computing tasks according to the reduced encoding matrix, and then perform planning and load balancing with the subtasks that are only calculated once.
5. The method according to claim 1 or 2 or 3, characterized in that, Use the greedy algorithm for planning and load balancing, and continuously plan the computing tasks with the longest time consumption to the computing node that undertakes the fewest computing tasks.
6. The method according to claim 1 or 2 or 3, characterized in that, In the static planning result, the subtasks to be encoded and their copies are set to the highest priority, and a dynamic load balancing scheme is introduced for the subtasks that are only calculated once in the static planning result for planning and load balancing.
7. A distributed computing system incorporating coding fault tolerance, characterized in that, It includes an input module, a prediction module, an encoding module, a computing module, and a decoding module; among them The input module is used to perform rationality analysis and format conversion on the received batch of computing tasks to be calculated, and then generate an input file and send it to the prediction module; The prediction module is used to predict the computing time required for each computing task and send it to the encoding module; The encoding module is used to regard several computing tasks as the subtasks to be encoded that need to perform redundant calculations according to the computing time of each computing task, and regard the remaining computing tasks as the subtasks that are only calculated once; then generate an encoding matrix and a decoding matrix according to the set redundancy and the number of subtasks to be encoded; Then group the subtasks to be encoded according to the encoding matrix to generate a grouped file, and perform planning and load balancing on the encoded computing tasks and the subtasks that are only calculated once to obtain a static planning result; then send the static planning result and the encoding matrix to the computing module; Send the decoding matrix and the grouped file to the decoding module; The computing module is used to allocate each computing task to the corresponding computing node for calculation according to the static planning result, and send the encoding matrix to the corresponding computing node; each computing node processing the subtasks to be encoded encodes the calculation result according to the encoding matrix and sends it to the decoding module, and each computing node processing the subtasks that are only calculated once directly sends its calculation result to the decoding module; The decoding module decodes the calculation results received from each computing node according to the grouped file and the decoding matrix to obtain the calculation result of the batch of computing tasks to be calculated.