Elastic coding calculation parameter optimization method based on LLM bayesian optimization
Patent Information
- Application Number
- CN202610903519.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]目前恢复阈值 k 完全依靠工程经验人工选定,无法适配节点随机上下线的动态集群环境,难以平衡存储成本与任务重构失败惩罚成本,造成分布式系统综合运行成本偏高、容错性能无法达到最优,是现有弹性编码计算落地亟需解决的技术痛点
[0032]依托集群历史节点运行数据作为输入数据源:利用历史节点在线数量统计数据表征集群真实抢占波动规律,区别于人工经验选型,参数选型贴合集群实际运行工况,解决经验选型脱离现场节点动态变化的缺陷。
Smart Images

Figure CN122817048A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of distributed computing and elastic coding computing technology, and in particular to an elastic coding computing parameter optimization method based on LLM Bayesian optimization. Background Technology
[0002] With the large-scale deployment of business applications, distributed computing has become the mainstream architecture for parallel processing of massive tasks. Service providers such as Amazon EC2 and Microsoft Azure have launched preemptive idle low-priority computing node resources. These resources are inexpensive, but the nodes have strong dynamic preemption characteristics: nodes are forced offline when high-priority tasks arrive, and nodes are randomly reconnected to the cluster when system resources are idle.
[0003] Existing distributed computing frameworks generally employ a stop-and-restart fault tolerance mechanism. When a node is preempted, all tasks are paused, waiting for the node to recover or be rescheduled. Frequent node offlines can easily cause prolonged computational stagnation. To address this deficiency, the industry has proposed an elastic coding computing scheme based on MDS (Maximum Distance Divisible Code). This scheme relies on coding redundancy to achieve task fault tolerance in dynamic node scenarios. Its core configuration parameter is the recovery threshold k: k is the minimum number of online nodes required to reconstruct the task result; the smaller the value of k, the stronger the system's fault tolerance capability, but the higher the coding storage overhead; the larger the value of k, the lower the storage cost, but the higher the probability of reconstruction failure due to insufficient online nodes.
[0004] Currently, the recovery threshold k is entirely determined manually based on engineering experience, which cannot adapt to the dynamic cluster environment where nodes randomly go online and offline. It is difficult to balance storage costs and task reconstruction failure penalty costs, resulting in high overall operating costs of distributed systems and failure to achieve optimal fault tolerance performance. This is a technical pain point that urgently needs to be addressed in the implementation of existing elastic coding computing. Summary of the Invention
[0005] To overcome the shortcomings of the prior art, the purpose of this invention is to provide a method for optimizing elastic coding computation parameters based on LLM Bayesian optimization.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] This invention provides a method for optimizing elastic coding computation parameters based on LLM Bayesian optimization, including:
[0008] S1. Input steps: Collect historical task execution data of the target distributed cluster in multiple rounds, count the number of online nodes that have not been preempted in each round of task execution, and summarize them to form a historical node dataset that represents the dynamic preemption pattern of cluster nodes; during task execution, nodes randomly go offline according to the inherent preemption rules of the cluster, offline nodes randomly reconnect to the cluster, and changes in the number of online nodes are marked as new round of task data statistics.
[0009] S2, First-level processing steps: Define a system comprehensive cost function with the recovery threshold as the independent variable. The comprehensive cost function consists of a weighted sum of storage cost and task reconstruction failure penalty cost. Set a fixed trade-off coefficient to adjust the proportion of the two types of costs, and determine minimizing the comprehensive cost as the optimization objective of the recovery threshold. The storage cost is determined by the ratio of the total number of nodes in the cluster to the recovery threshold. The reconstruction failure penalty cost is obtained by counting the total number of task rounds in which the number of online nodes is less than the current recovery threshold, based on the historical node dataset obtained in S1.
[0010] S3, Secondary Processing Steps: Within the legal range of the recovery threshold, select several initial candidate recovery thresholds using equidistant sampling. Substitute each threshold into the comprehensive cost function of S2 and combine it with the historical node dataset to calculate the corresponding comprehensive cost. Pair the candidate recovery thresholds with their corresponding comprehensive costs one by one to construct the initial sampling dataset.
[0011] S4. Three-level processing steps: Based on the initial sampled dataset, a Gaussian process surrogate model is built, and a kernel function is used to characterize the degree of numerical correlation between any two different recovery thresholds; a large language model is introduced to read the distribution characteristics of cost data in the initial sampled dataset, and the squared exponential kernel function or the Matern3 / 2 kernel function is adaptively selected according to the distribution difference of the cost data being smooth and continuous or abruptly changing. The large language model relies on the numerical range of the cost data and the length of the recovery threshold interval corresponding to the cost abrupt change to adaptively solve the signal variance hyperparameter and length scale hyperparameter of the kernel function to complete the parameter initialization of the Gaussian process surrogate model.
[0012] S5, Level 4 Processing Steps: Based on the current sampled dataset, construct the sample point covariance matrix, the covariance vector between the unsampled candidate recovery threshold and the sampled points, and the sample cost vector. Introduce a regularization term to correct the covariance matrix and avoid matrix singularity. Based on the Gaussian process posterior derivation formula, solve for the posterior mean and posterior variance of the cost corresponding to each unsampled candidate recovery threshold. The posterior mean represents the expected value of the predicted cost corresponding to the candidate threshold, and the posterior variance represents the uncertainty of the prediction result.
[0013] S6. Fifth-level processing steps: Extract the minimum comprehensive cost in the current sampled dataset as the benchmark optimal cost. Combine the posterior mean, posterior standard deviation, standard normal cumulative distribution function and probability density function corresponding to the candidate recovery threshold to construct the expected improvement collection function. The output value of the collection function represents the expected improvement of the cost after further sampling of the corresponding candidate recovery threshold.
[0014] S7, Level 6 processing steps: Traverse all unsampled candidate recovery thresholds and select the candidate parameter with the largest expected value of the acquisition function as the next sampling recovery threshold to be measured in this round;
[0015] S8, Level 7 processing steps: Substitute the next sampling recovery threshold obtained from the screening into the S2 comprehensive cost function, calculate the actual comprehensive cost in combination with the historical node dataset, supplement the new sample of this set of parameters-cost into the sampling dataset, and complete the sampling dataset update;
[0016] S9, Level 8 processing steps: The large language model reads the updated sampled dataset, reselects the kernel function type and optimizes the kernel function hyperparameters based on the updated data distribution characteristics, and recalculates the posterior mean and posterior variance of all unsampled candidate recovery thresholds based on the updated Gaussian process surrogate model.
[0017] S10, Iteration Judgment Step: Iterate through steps S6 to S9 until the preset iteration termination condition is met. The termination condition includes the number of iterations reaching the preset upper limit, or the deviation between the comprehensive cost of the sampled samples and the theoretical optimal cost falling within the preset error range.
[0018] S11. Output steps: After the iteration terminates, retrieve all sample data in the final sampled dataset, select the recovery threshold corresponding to the sample with the lowest comprehensive cost as the global optimal recovery threshold and output it outward. Configure the optimal recovery threshold to the elastic coding distributed computing system to optimize the task fault tolerance parameters.
[0019] Furthermore, in S2, the expression for the comprehensive cost function is as follows: In the formula, As a weighting factor, Where n is the storage cost, k is the total number of nodes in the cluster, and k is the recovery threshold. To reconstruct the failure penalty cost, T represents the total number of rounds in the historical task. Let be the number of online nodes in the i-th round of tasks. For indicator functions, when The value is 1 when it is active and 0 otherwise.
[0020] Furthermore, in S4: when the cost curve is smooth and continuous, the quadratic exponential kernel function is selected; when the cost curve has abrupt inflection points, the Matern3 / 2 kernel function is selected.
[0021] ,
[0022] in, For signal variance hyperparameters, For length scale hyperparameters, Let be any two feasible recovery thresholds.
[0023] Furthermore, the formulas for calculating the posterior mean and posterior variance in S5 are as follows:
[0024]
[0025] in, for identity matrix of order 1, regularization term This is used to prevent the covariance matrix A strange problem has occurred.
[0026] Furthermore, S6 aims to improve the expression of the acquisition function:
[0027] In the formula, The minimum overall cost within the current dataset. For the posterior standard deviation, The standard normal cumulative distribution function is... It is the standard normal probability density function.
[0028] Furthermore, the dynamic rules for nodes in S1 are as follows: idle online nodes are preempted and taken offline according to a preset offline probability, and offline nodes rejoin the cluster to participate in computation according to a preset online probability.
[0029] Furthermore, the equidistant sampling described in S3 involves selecting a specified number of candidate k values at uniform intervals within a continuous integer range formed by the upper and lower limits of the recovery threshold.
[0030] Furthermore, it is applied to the elastic coding distributed computing architecture based on MDS maximum distance separable codes, where the optimal recovery threshold is used to constrain the number of original data splits and the minimum number of online nodes required for the reconstruction of coding subtasks.
[0031] Compared with the prior art, the technical solution disclosed in this invention has the following beneficial effects:
[0032] Using historical node operation data as the input data source: the statistical data of the number of online nodes in history is used to characterize the actual preemption fluctuation pattern of the cluster. Unlike manual experience-based selection, the parameter selection is close to the actual operating conditions of the cluster, which solves the problem of experience-based selection being detached from the dynamic changes of on-site nodes.
[0033] Construct a weighted total cost function that integrates storage costs and reconstruction failure penalty costs: By flexibly adjusting the weights of the two types of costs through a trade-off coefficient, an objective evaluation standard for the merits of the recovery threshold k is quantified, thereby achieving a quantitative balance between storage overhead and fault tolerance loss.
[0034] Initial equidistant sampling establishes the original sample dataset: a small amount of initial sampling is sufficient to complete the model initialization, avoiding the massive computational overhead caused by traversing all parameters.
[0035] LLM adaptively selects the Gaussian process kernel function and configures hyperparameters based on the distribution characteristics of the sample data: it automatically switches between the squared exponential kernel and the Matern3 / 2 kernel according to the smoothing / abrupt characteristics of the cost curve, and simultaneously optimizes the signal variance and length scale hyperparameters. Compared with Bayesian optimization with a fixed kernel function, it improves the fitting accuracy of the Gaussian process to the nonlinear cost function and optimizes the prediction accuracy of the surrogate model.
[0036] Gaussian process to solve for the posterior mean and variance of unsampled thresholds: quantifies the expected cost and prediction uncertainty of candidate parameters, taking into account both cost and the value of sampling exploration.
[0037] Based on EI, the sampling point is iteratively optimized by improving the acquisition function: each iteration selects only the recovery threshold with the best expected return, which approaches the global optimum with far fewer iterations than a full traversal, greatly reducing the simulation computation of parameter optimization.
[0038] After iterative convergence, the optimal recovery threshold is output: the k value with the lowest overall cost is automatically obtained, balancing storage costs and the risk of task reconstruction failure, reducing the overall operation and maintenance cost of the distributed cluster, and improving the stability of elastic coding computation in a preemptive dynamic node environment.
[0039] In summary, compared with the traditional manual experience-based selection and full-parameter traversal optimization scheme, this invention can accurately balance the losses of storage and fault tolerance in dynamic preemptive node cluster scenarios, and significantly reduce the number of calculations for optimal parameter search, thereby improving the overall operating efficiency of elastic coding distributed computing. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a diagram illustrating the total system cost under different recovery thresholds provided in embodiments of the present invention.
[0042] Figure 2 A diagram illustrating the Bayesian optimization process provided in this embodiment of the invention;
[0043] Figure 3 This is a schematic diagram of the process for optimizing elastic coding computation parameters based on LLM Bayesian optimization, as provided in an embodiment of the present invention. Detailed Implementation
[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0046] like Figures 1-3 As shown, this embodiment of the invention provides a method for optimizing elastic coding computation parameters based on LLM Bayesian optimization, including:
[0047] S1. The input steps can be summarized as collecting historical cluster operation data, specifically including: collecting historical cluster data and counting the number of nodes that were not preempted in each round of tasks. To obtain historical datasets ,in The number of tasks is denoted by _____. In each round of computation, a computing node may leave due to preemption by a high-priority task, or it may rejoin the computation after becoming idle. When the number of unpreempted nodes changes, the proposed elastic coding computation method adjusts the algorithm based on the current number of unpreempted nodes. The computing tasks on each node are redistributed, but the overall number of tasks to be computed remains unchanged, and this process is regarded as entering a new round of task execution.
[0048] S2, the first-level processing step, can be summarized as constructing a quantized total cost function for the recovery threshold, specifically including: for any recovery threshold Based on the computational redundancy and reconstruction computation methods in the elastic coding computation paradigm, a cost function is defined. The function is determined by storage cost. Penalty cost for failed reconstruction of computation results Composition, in which Storage cost and penalty costs The trade-off coefficient. This refers to the number of nodes that were not preempted in a given round of tasks. At that time, the task reconstruction failed. The value is 1 if it is true, and 0 otherwise. The optimization objective is to minimize the cost function. .
[0049] S3, the second-level processing step, can be summarized as generating the initial sample dataset through equidistant sampling, specifically including: equidistant sampling selection. Recovery threshold For each recovery threshold Based on the above historical dataset Calculate the corresponding total system cost using the cost function defined in S2. To obtain the initial sampling dataset ,in It is the first One sampling point.
[0050] S4, the third-level processing steps, can be summarized as LLM-driven Gaussian process surrogate model construction and kernel function adaptive configuration, specifically including: constructing the Gaussian process surrogate model, and restoring the threshold... Cost function of elastically encoded distributed computing The nonlinear mapping relationship between them is fitted to a Gaussian random process. This process uses a kernel function... As a covariance function of the prior distribution, it is used to evaluate any two feasible recovery thresholds. The degree of correlation between them. Among them, to improve the surrogate model's response to the cost function... To improve the fitting ability, a large language model is introduced. Based on the distribution characteristics of each point in the sampled dataset D (e.g., cost function) The kernel function is selected based on factors such as the degree of volatility and smoothness of the cost function. Smooth and continuous change, Use a squared exponential kernel function; if there are obvious transition or abrupt change regions... use Kernel function, i.e.
[0051]
[0052] Among them, large language model Based on the sampled dataset Medium cost function value The range and the corresponding recovery threshold when costs change significantly. The interval length is given, and the signal variance is given. and length scale parameters The appropriate value.
[0053] S5, the fourth-level processing step, can be summarized as solving the posterior statistic of the unsampled candidate threshold using a Gaussian process, specifically including: quantifying the cost function... At each unsampled point The expected value and uncertainty of the prediction at a given point, the Gaussian process surrogate model based on the covariance function of the prior distribution and the current sampled dataset Derive the cost function At unsampled points posterior mean at and posterior variance Specifically, when the total number of sampled points is At that time, for a given unsampled point Calculate the covariance function value between it and each of the sampled points. The covariance vector is constructed. , Simultaneously, a covariance matrix is constructed based on the covariance among the sampled points. , and cost vector , Based on this, the posterior mean is calculated using the following formula. and posterior variance ,
[0054]
[0055]
[0056] in, for identity matrix of order 1, regularization term This is used to prevent the covariance matrix A strange problem has occurred.
[0057] S6, the five-stage processing steps, can be summarized as constructing the EI to improve the acquisition function, specifically including: based on the minimum cost function value in the current sampled dataset D. Each unsampled point posterior mean at and posterior standard deviation (It is the posterior variance) (the value after taking the square root), and the cumulative distribution function of the standard normal distribution. and density function Constructing the desired improvement Acquisition function Used to quantize at each unsampled point The cost function value at that point is relative to the currently known optimal cost. The expected improvement is expressed as follows:
[0058]
[0059] S7, the sixth-level processing steps, can be summarized as selecting the optimal sampling point for the next round based on the acquisition function. Specifically, this includes: maximizing the acquisition function, selecting the optimal sampling point that makes... The maximum recovery threshold is used as the next sampling point. This allows us to approach the global minimum of the cost function C with as few attempts as possible.
[0060] S8, the seven-level processing steps, can be summarized as new sample testing and dataset updating, specifically including: based on historical datasets and the cost function defined by S2 ,calculate Corresponding cost function value and new samples Add to sampling dataset .
[0061] S9, the eight-level processing steps, can be summarized as LLM dynamic reconfiguration of Gaussian process model hyperparameters, specifically including: the large language model based on the updated sampled dataset. Choose a suitable kernel function Give the signal variance and length scale parameters The appropriate value is determined, and the cost function is recalculated using a Gaussian process surrogate model. At each unsampled point posterior mean at and posterior variance .
[0062] S10, the iterative determination step, can be summarized as loop and termination determination, specifically including: repeatedly executing steps S6-S9, and stopping the iteration when the termination condition is met; wherein, the termination condition can be that the number of iterations reaches a preset value; or it can be that, given that the minimum cost function value is known, the cost function value corresponding to at least one sampling point has approached the minimum cost function value within the preset value range.
[0063] S11, the output step, can be summarized as filtering the globally optimal recovery threshold, specifically including: Finally, from the sampled dataset... Select the recovery threshold with the lowest cost function value. As the globally optimal solution, it is substituted into the calculation.
[0064] The above steps will now be simulated using a cluster containing several computing nodes.
[0065] 1. Assume a total number of nodes Number of tasks The probability of a node being preempted The probability of a preempted node being re-added to the calculation. ,parameter , , Preset value for the number of iterations , and Initially set to 1.
[0066] 2. None of the initial cluster nodes were preempted, i.e. For each round of tasks, unclaimed nodes are... The probability of being preempted, and the preempted node is... The probability of a node rejoining the cluster for computation is determined. The number of nodes that were not preempted at the end of each task round is counted to obtain the historical dataset. .
[0067] 3. Select 5 recovery thresholds using interval sampling. For each recovery threshold, the historical dataset is used. and the cost function defined by S2 Calculate the corresponding total system cost. To obtain the initial sampling dataset .
[0068] 4. Construct a Gaussian process surrogate model and set the recovery threshold. Cost function of elastically encoded distributed computing The nonlinear mapping relationship between them is fitted to a Gaussian random process.
[0069] 5. Introduce a large language model By sampling datasets Choose a suitable kernel function and signal variance and length scale parameters The value of .
[0070] 6. Calculate other unsampled points using a Gaussian process surrogate model. Corresponding posterior mean and posterior variance Then by formula Calculate its expected improvement .
[0071] 7. Choose the option with the highest expected improvement. Calculate the corresponding total system cost. and new samples Add to sampling dataset .
[0072] 8. Repeat steps 5-7. When the preset number of iterations (5) is reached, select the sampling dataset. Total system cost Minimum recovery threshold As the globally optimal solution.
[0073] The technical effects of the present invention will be further explained below with reference to simulation experiments.
[0074] Simulation Result Analysis:
[0075] Depend on Figure 1 It can be seen that when At that time, storage costs With recovery threshold The decrease in [amount] rises rapidly and is a major component of total cost; when At that time, the penalty cost of task refactoring failure With recovery threshold The increase gradually rises, while storage costs... continuously decreasing, therefore Gradually becoming a major component of total cost; when At that time, the total system cost Reaching the minimum value, i.e., the optimal recovery threshold. .
[0076] Depend on Figure 2 It can be seen that Bayesian optimization found the global optimal solution for this simulation in the 4th iteration. Furthermore, the entire process requires only 10 simulation calculations (5 initial samplings and 5 iterative samplings). If an iterative method is used to find the optimal recovery threshold, 50 simulation calculations are required. Therefore, the method of this invention improves the efficiency of selecting the optimal recovery threshold k.
[0077] The basic principles of the present invention have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in the present invention are merely examples and not limitations, and should not be considered as essential features of each embodiment of the present invention. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the present invention to the necessity of employing the aforementioned specific details.
[0078] The block diagrams of devices, apparatuses, devices, and systems involved in this invention are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0079] It should also be noted that in the apparatus, device, and method of the present invention, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of the present invention.
[0080] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the invention. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the invention. Therefore, the invention is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0081] It should be understood that the qualifying terms "first", "second", "third", "fourth", "fifth" and "sixth" used in the description of the embodiments of the present invention are only used to more clearly illustrate the technical solutions and are not intended to limit the scope of protection of the present invention.
[0082] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the invention to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A method for optimizing elastic coding computational parameters based on LLM Bayesian optimization, characterized in that, include: S1. Input steps: Collect historical task running data of the target distributed cluster in multiple rounds, count the number of online nodes that have not been preempted in each round of task running, and summarize them to form a historical node dataset that represents the dynamic preemption pattern of cluster nodes. During task execution, nodes randomly go offline according to the cluster's inherent preemption rules, and offline nodes randomly reconnect to the cluster. Changes in the number of online nodes are marked as a new round of task data statistics. S2, First-level processing steps: Define a system comprehensive cost function with the recovery threshold as the independent variable. The comprehensive cost function consists of a weighted sum of storage cost and task reconstruction failure penalty cost. Set a fixed trade-off coefficient to adjust the proportion of the two types of costs, and determine minimizing the comprehensive cost as the optimization objective of the recovery threshold. The storage cost is determined by the ratio of the total number of nodes in the cluster to the recovery threshold. The reconstruction failure penalty cost is obtained by counting the total number of task rounds in which the number of online nodes is less than the current recovery threshold, based on the historical node dataset obtained in S1. S3, Secondary Processing Steps: Within the legal range of the recovery threshold, select several initial candidate recovery thresholds using equidistant sampling. Substitute each threshold into the comprehensive cost function of S2 and combine it with the historical node dataset to calculate the corresponding comprehensive cost. Pair the candidate recovery thresholds with their corresponding comprehensive costs one by one to construct the initial sampling dataset. S4. Three-level processing steps: Based on the initial sampled dataset, a Gaussian process surrogate model is built, and a kernel function is used to characterize the degree of numerical correlation between any two different recovery thresholds; a large language model is introduced to read the distribution characteristics of cost data in the initial sampled dataset, and the squared exponential kernel function or the Matern3 / 2 kernel function is adaptively selected according to the distribution difference of the cost data being smooth and continuous or abruptly changing. The large language model relies on the numerical range of the cost data and the length of the recovery threshold interval corresponding to the cost abrupt change to adaptively solve the signal variance hyperparameter and length scale hyperparameter of the kernel function to complete the parameter initialization of the Gaussian process surrogate model. S5, Level 4 processing steps: Based on the current sampled dataset, construct the sample point covariance matrix, the covariance vector between the unsampled candidate recovery threshold and the sampled points, and the sample cost vector. Introduce a regularization term to correct the covariance matrix and avoid matrix singularity. Based on the posterior derivation formula of Gaussian process, the posterior mean and posterior variance of the cost corresponding to each unsampled candidate recovery threshold are solved. The posterior mean represents the expected value of the predicted cost corresponding to the candidate threshold, and the posterior variance represents the uncertainty of the prediction result. S6. Fifth-level processing steps: Extract the minimum comprehensive cost in the current sampled dataset as the benchmark optimal cost. Combine the posterior mean, posterior standard deviation, standard normal cumulative distribution function and probability density function corresponding to the candidate recovery threshold to construct the expected improvement collection function. The output value of the collection function represents the expected improvement of the cost after further sampling of the corresponding candidate recovery threshold. S7, Level 6 processing steps: Traverse all unsampled candidate recovery thresholds and select the candidate parameter with the largest expected value of the acquisition function as the next sampling recovery threshold to be measured in this round; S8, Level 7 processing steps: Substitute the next sampling recovery threshold obtained from the screening into the S2 comprehensive cost function, calculate the actual comprehensive cost in combination with the historical node dataset, supplement the new sample of this set of parameters-cost into the sampling dataset, and complete the sampling dataset update; S9, Level 8 processing steps: The large language model reads the updated sampled dataset, reselects the kernel function type and optimizes the kernel function hyperparameters based on the updated data distribution characteristics, and recalculates the posterior mean and posterior variance of all unsampled candidate recovery thresholds based on the updated Gaussian process surrogate model. S10, Iteration Judgment Step: Iterate through steps S6 to S9 until the preset iteration termination condition is met. The termination condition includes the number of iterations reaching the preset upper limit, or the deviation between the comprehensive cost of the sampled samples and the theoretical optimal cost falling within the preset error range. S11. Output steps: After the iteration terminates, retrieve all sample data in the final sampled dataset, select the recovery threshold corresponding to the sample with the lowest comprehensive cost as the global optimal recovery threshold and output it outward. Configure the optimal recovery threshold to the elastic coding distributed computing system to optimize the task fault tolerance parameters.
2. The method according to claim 1, characterized in that, In S2, the expression for the comprehensive cost function is: In the formula, As a weighting factor, Where n is the storage cost, k is the total number of nodes in the cluster, and k is the recovery threshold. To reconstruct the failure penalty cost, T represents the total number of rounds in the historical task. Let be the number of online nodes in the i-th round of tasks. For indicator functions, when The value is 1 when it is active and 0 otherwise.
3. The method according to claim 1, characterized in that, In S4: When the cost curve is smooth and continuous, the quadratic exponential kernel function is selected; when the cost curve has abrupt inflection points, the Matern3 / 2 kernel function is selected. , in, For signal variance hyperparameters, For length scale hyperparameters, Let be any two feasible recovery thresholds.
4. The method according to claim 1, characterized in that, Formulas for calculating the posterior mean and posterior variance in S5: in, for identity matrix of order 1, regularization term This is used to prevent the covariance matrix A strange problem has occurred.
5. The method according to claim 1, characterized in that, The desired improvement to the acquisition function expression in S6 is as follows: In the formula, The minimum overall cost within the current dataset. For the posterior standard deviation, The standard normal cumulative distribution function is... It is the standard normal probability density function.
6. The method according to claim 1, characterized in that, The dynamic rules for nodes in S1 are as follows: idle online nodes are preempted and taken offline according to a preset offline probability, and offline nodes rejoin the cluster to participate in computation according to a preset online probability.
7. The method according to claim 1, characterized in that, S3 The equidistant sampling refers to selecting a specified number of candidate k values at uniform intervals within a continuous integer interval formed by the upper and lower limits of the recovery threshold.
8. The method according to claim 1, characterized in that, It is applied to a resilient coding distributed computing architecture based on MDS (Maximum Distance Divisible Code), where the optimal recovery threshold is used to constrain the number of original data splits and the minimum number of online nodes required for the reconstruction of coding subtasks.