A Modeling Method of Pointwise Weighted Hybrid Surrogate Model Based on Hybrid Measure
By constructing a point-by-point weighted hybrid agent model based on mixed measurements, using LOO cross-validation and GMSE error to determine the local and global benchmark models, combining the crowding distance and local error measurement to calculate the weight coefficients, the problems of insufficient accuracy and high calculation cost in the local area are solved, and efficient global and local predictions are achieved.
Patent Information
- Application Number
- CN202411165716.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-08-23
AI Technical Summary
The existing hybrid proxy models have insufficient prediction accuracy in local areas and are highly computationally cost-effectively, especially in high-dimensional problems, and the global fitting accuracy is poor, so they cannot effectively adjust weights based on local and global errors.
The training samples were obtained by using the Latin hypercube sampling method, the PRS, RBF, KRG, and SVR sub-agent models were constructed, and the local and global benchmark models were determined through LOO cross-validation and GMSE error. The weight coefficient was calculated based on the crowding distance and local error measurement, and the point-by-point weighted hybrid agent model based on mixed measurements was constructed.
It improves the local and global prediction accuracy of the hybrid proxy model, reduces the computational cost, is suitable for low-dimensional and high-dimensional problems, and ensures accurate prediction of the model in different regions.
Smart Images

Figure CN119272606B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the construction of an approximate relationship between the input and output of a mechanical structure, and particularly to a modeling method of a point-by-point weighted hybrid surrogate model based on a hybrid measure. Background Art
[0002] With the progress of mathematics and computing technology, computer simulation technology has shown significant advantages in the field of mechanical design and has gradually replaced traditional experimental methods, providing strong technical support for mechanical design optimization. However, computer simulation technology is not perfect. Its computational cost is often high, especially when dealing with complex high-fidelity engineering simulation models, the computational cost will increase sharply, and the design cycle will also be lengthened, making it difficult to rely solely on computer simulation for design optimization.
[0003] To solve the problem of high computational cost of computational simulation experiments, a mechanical design method using a surrogate model to replace the expensive computational simulation model has emerged. This method not only effectively reduces the computational cost but also significantly shortens the design cycle to a large extent. Compared with the traditional computational simulation model, the surrogate model exists in a pure mathematical form and can quickly and accurately predict and simulate results, providing more efficient and reliable technical support for mechanical design optimization. However, due to the complexity of actual engineering problems, how to select an appropriate surrogate model before design optimization remains a difficult problem in the application field of surrogate models.
[0004] The hybrid surrogate model forms a new surrogate model by carefully screening and combining multiple surrogate model methods. A large number of studies have shown that using a hybrid surrogate model to replace a single surrogate model for modeling can effectively avoid the risk of the surrogate model method selection strategy caused by the differences in factor surrogate model methods. This not only significantly improves the overall modeling accuracy of the model but also enhances the robustness of the model. However, the performance of the hybrid surrogate model is not always better than that of the optimal sub-surrogate model that composes it. In addition, most hybrid models are global average hybrid surrogate models, and their weight coefficients are set as constants, which means that once the weight coefficients are determined, they remain unchanged throughout the design interval. This method of global fixed weight coefficients cannot reveal the prediction accuracy of each component surrogate model in the local region, so accurate predictions may not be provided in some local regions. That is, the global average hybrid surrogate model mostly only focuses on the global error of the sample set to ensure the global accuracy of the surrogate model. This places certain requirements on the distribution of the selected sample points. When the sample points are sparse or some sample points are far from the main part, the overall prediction accuracy will be greatly affected due to the neglect of local errors.
[0005] To improve the local prediction accuracy of the hybrid surrogate model, researchers have proposed a hybrid surrogate model based on the variation of weight coefficients with local accuracy, called the pointwise weighted hybrid surrogate model. This self-adaptive intelligent hybrid surrogate model method can flexibly adjust the weight coefficients according to the requirements of different regions, thereby improving the prediction accuracy. However, they may have poor global fitting accuracy. At the same time, most of the current adaptive weight coefficients require optimization methods to assist in the search, especially for high-dimensional problems, and the computational cost is relatively high.
[0006] Therefore, for complex engineering application problems, finding a modeling method for the pointwise weighted hybrid surrogate model that is applicable to both low-dimensional and high-dimensional problems, comprehensively considering the global error and local error, ensuring that its weight coefficients can guarantee the model has sufficient global accuracy and local accuracy, and a low computational cost, is an important technical problem that needs to be solved in this field currently. Summary of the Invention
[0007] The purpose of the present invention is to provide a modeling method for a pointwise weighted hybrid surrogate model based on a hybrid measure in view of the deficiencies of the prior art, and solve the problems of low accuracy, difficult calculation and high cost of weight coefficients of the existing hybrid surrogate model.
[0008] The technical solution adopted by the present invention to achieve the above purpose is as follows:
[0009] A modeling method for a pointwise weighted hybrid surrogate model based on a hybrid measure, which includes the following steps:
[0010] Step 1: Use the Latin hypercube sampling method to extract training samples in the variable space, and obtain the true response value through simulation analysis;
[0011] Step 2: Use the initial sample points and the corresponding true values to construct four sub-surrogate models of PRS, RBF, KRG, and SVR respectively and perform hyperparameter optimization;
[0012] Step 3: Determine the local benchmark model and the global benchmark model for each sample point according to the LOO cross-validation error and the GMSE error, and determine the global weight coefficient;
[0013] Step 4: Determine the sample density based on the crowding distance;
[0014] Step 5: Determine the local accuracy coefficient based on the local error measure;
[0015] Step 6: Determine the local weight coefficients of each sub-surrogate model;
[0016] Step 7: Determine the preliminary hybrid surrogate model based on the local error measure;
[0017] Step 8: Determine the point - by - point weighted hybrid surrogate model based on the hybrid measure.
[0018] The specific steps of Step 1 are as follows:
[0019] Use the Latin - hypercube sampling method to extract training samples in the original variable space and form the initial training point set \(x^{(j)}\) (\(j = 1,\cdots,N\)); \(N\) represents the number of training samples. (j) (\(j = 1,\cdots,N\)); \(N\) represents the number of training samples;
[0020] Conduct simulation analysis on the training samples to obtain the true response value \(g(x^{(j)})\), and form the database \(DB[x^{(j)}|g(x^{(j)})]\) (\(j = 1,\cdots,N\)); where \(x=(x_1,\cdots,x_D)\) is a \(D\) - dimensional input variable. (j) ) and form the database \(DB[x^{(j)}|g(x^{(j)})]\) (\(j = 1,\cdots,N\)); where \(x=(x_1,\cdots,x_D)\) is a \(D\) - dimensional input variable. (j) |g(x^{(j)})](\(j = 1,\cdots,N\)); where \(x=(x_1,\cdots,x_D)\) is a \(D\) - dimensional input variable. (j) )](\(j = 1,\cdots,N\)); where \(x=(x_1,\cdots,x_D)\) is a \(D\) - dimensional input variable. D ) is a \(D\) - dimensional input variable.
[0021] The specific steps of constructing the PRS sub - surrogate model by using the initial sample points and the corresponding true values in Step 2 are as follows:
[0022] Step 2.1.1: Determine the data dimension \(D\), select the basis function as a complete quadratic polynomial, and conduct multiple - polynomial regression analysis on the training sample points and their true response values in the database \(DB\) according to the least - squares method to determine the regression function.
[0023] Step 2.1.2: Conduct variance and error statistical analysis on the regression function to determine the PRS model, and its basic form is:
[0024]
[0025] where \(y_1(x)\) is the prediction function of the objective function in the PRS model, \(x_i\) (\(i = 1,\cdots,D\)) is the \(i\) - th component of the \(D\) - dimensional variable, and \(\beta\) is the unknown weight coefficient of the prediction function. i (\(i = 1,\cdots,D\)) is the \(i\) - th component of the \(D\) - dimensional variable, and \(\beta\) is the unknown weight coefficient of the prediction function.
[0026] The specific steps of constructing the RBF sub - surrogate model and optimizing its hyperparameters by using the initial sample points and the corresponding true values in Step 2 are as follows:
[0027] Step 2.2.1: Determine the data dimension \(D\), select the kernel function type as the Gaussian function, and its function is as follows:
[0028]
[0029] where \(r\) is the distance norm between the position point and the sample point, and the Euclidean distance \(\vert\vert x - x^{(j)}\vert\vert\) is used, i ||, is a common distance kernel function, \(\sigma\) is the bias - expansion constant, \(\sigma\gt0\), and is used to control the shape of the kernel function.
[0030] Step 2.2.2: Use the radial function to linearly weight the sample points to determine the RBF model, whose basic form is:
[0031]
[0032] where y2(x) is the prediction function of the objective function in the RBF model, and x i (i = 1, …, D) is the i-th component of the D-dimensional variable, is the distance kernel function, m is the number of kernel functions, and λ is the linear combination weight coefficient; r is the distance norm between the position point and the center point (sample), and the Euclidean distance ||x - x i || is adopted;
[0033] The construction of the KRG sub surrogate model by using the initial sample points and the corresponding true values described in Step 2 specifically includes:
[0034] Step 2.3.1: Determine the data dimension D, select the regression function of the KRG model as a quadratic polynomial model, and the relevant model function as the most typical Gaussian function;
[0035] Step 2.3.2: Define the value of theta and the thresholds lob and upb, where theta represents the hyperparameter of the selected relevant function;
[0036] Step 2.3.3: Train the KRG model with the training sample points and their true response values in the database DB, and its basic form:
[0037]
[0038] where y3(x) is the prediction function of the objective function in the KRG model, and β i is the regression coefficient, f i (x) is the regression function, and Z l (x) is a stationary random process with a mean of 0 and a variance of σ 2 ;
[0039] The construction of the SVR sub surrogate model by using the initial sample points and the corresponding true values described in Step 2 and the hyperparameter optimization specifically include:
[0040] Step 2.4.1: Determine the data dimension D, and select e-SVR with the kernel function type as the RBF radial basis function;
[0041] Step 2.4.2: Define the values of the loss function c and the gamma function g in the kernel function. c and g represent the selected hyperparameters;
[0042] Step 2.4.3: Train the SVR model with the training sample points and their true response values in the database DB, and its basic form:
[0043]
[0044] Among them, y4(x) is the prediction function of the objective function in the SVR model, and k(x i T x) = θ(x i T )θ(x j ) is the kernel function.
[0045] Step 2.4.4: Initialize the particle swarm. The positions of the particles in the particle swarm are the values of the loss function c and the gamma function g in the kernel function (hyperparameters), and the velocities are the change directions and rates of the hyperparameters. Determine the number of particles in the particle swarm, as well as the initial positions and initial velocities of each particle. The optimization objective is to minimize the cross-validation error of the SVR sub-agent model, and set the convergence condition in the optimization process. After completing the particle swarm optimization, obtain the optimal values of the loss function c and the gamma function g in the kernel function, and reconstruct the SVR sub-agent model.
[0046] Step 3 for determining the local benchmark model for each sample point according to the LOO cross-validation error specifically includes:
[0047] Step 3.1.1: For any training sample point x (j) (j = 1,..., N), respectively construct four sub-agent models for LOO cross-validation, and obtain its LOO cross-validation error error. The expression is as follows:
[0048]
[0049] Among them represents the LOO cross-validation error of the k-th sub-agent model at the j-th sample point, and y (j) is the true response value of the model at this sample point, is the predicted response value of the k-th sub-agent model at this sample point.
[0050] Step 3.1.2: For each sample point x (j) (j = 1,..., N), select the sub-agent model with the smallest LOO cross-validation error error as the local benchmark model μ(x (i) );
[0051] Step 3 for determining the global benchmark model and the global weight coefficient according to the LOO generalized mean square cross-validation error specifically includes:
[0052] Step 3.2.1: Conduct LOO cross-validation on the four sub-agent models again, and obtain the GMSE error. The expression is as follows:
[0053] Among them, GMSE k represents the GMSE error of the k-th sub-agent model, and y (j) is the true response value of the model at the j-th sample point, is the predicted response value of the k-th sub-agent model at this sample point, and N represents the number of training points.
[0054] Step 3.2.2: Sort all the sub-agent models from low to high according to the GMSE error results. The sub-agent model with the smallest GMSE error and the highest global accuracy is selected as the global benchmark model f g (x), and at the same time calculate the global weight coefficient, and its expression is as follows:
[0055]
[0056] Among them, w g is the global weight coefficient, where GMSE1 is the smallest GMSE value and GMSE2 is the second smallest GMSE value.
[0057] The determination of sample density based on crowding distance described in Step 4 specifically includes:
[0058] Step 4.1: For any prediction point x (i) , calculate the sum D (j) of its Euclidean distances from all training sample points x i , which is defined as the crowding distance, and its expression is as follows:
[0059]
[0060] Among them, D i represents the crowding distance of the i-th prediction point, using the Euclidean distance, x (j) represents the j-th training point, x (i) represents the i-th prediction point, and N is the number of training points.
[0061] Step 4.2: Calculate the sample density ρ i at any prediction point, and its expression is as follows:
[0062]
[0063] Step 4.3: Normalize the sample density to obtain the normalized sample density Nρ i at any prediction point, and its expression is as follows:
[0064]
[0065] The determination of the local precision coefficient based on the local error measure described in step 5 specifically includes:
[0066] For any prediction point x (i) , calculate the local precision coefficient of each sub-agent model at this point, and its expression is as follows:
[0067] Where P k (x (i) ) represents the local precision coefficient of the k-th sub-agent model at the i-th prediction point, and y k (x (i) ) represents the predicted response value of the k-th sub-agent model at the i-th prediction point, and μ(x (i) ) represents the predicted response value of the local benchmark model at the i-th prediction point. The sensitivity parameter σ 2 The calculation formula is as follows:
[0068]
[0069] Where σ1(x ( i ) ) represents the sensitivity parameter at the i-th prediction point, and Nρ i is the normalized sample density at the i-th prediction point.
[0070] The determination of the local weight coefficient of each sub-agent model described in step 6 specifically includes:
[0071] For any prediction point x (i) , by normalizing the local precision coefficients of each sub-agent model in the sub-model library, the local weight coefficients corresponding to each sub-agent model can be calculated, and the expression is as follows:
[0072]
[0073] Where w localk (x (i) ) represents the local weight coefficient of the sub-agent model k at the i-th prediction point, and P k (x (i) ) represents the local precision coefficient of the sub-agent model k at the i-th prediction point, and n s represents the number of sub-agent models in the model library.
[0074] The determination of the preliminary hybrid surrogate model based on the local error measure described in step 7 specifically includes:
[0075] Determine the preliminary hybrid surrogate model based on the local error measure through each sub-agent model and the local weight coefficient, and its expression is as follows:
[0076]
[0077] Among them, f PHSM represents the preliminary hybrid proxy model, and y k represents the sub-proxy model k, and w local represents the local weight coefficient of the sub-proxy model k, and n s represents the number of sub-proxy models in the model library.
[0078] Determining the point-by-point weighted hybrid proxy model based on the hybrid measure described in step 8 specifically includes:
[0079] Based on the preliminary hybrid proxy model f based on the local error measure PHSM , the global benchmark model f g , and the global weight coefficient w g , determine the point-by-point weighted hybrid proxy model f PWHSMHM , and its expression is as follows:
[0080] f PWHSMHM = w g f g +(1 - w g )f PHSM
[0081] Among them, f PWHSMHM represents the point-by-point weighted hybrid proxy model based on the hybrid measure, f g represents the global benchmark model, and f PHSM represents the preliminary hybrid proxy model.
[0082] Compared with the prior art, the present invention has the following beneficial effects:
[0083] (1), The present invention uses the LOO cross-validation error as the selection basis for the local benchmark model, and calculates the local weight coefficients of each sub-proxy model based on the local benchmark models of each sample point, ensuring that each prediction point should be predicted by its corresponding local sub-proxy model with the best prediction ability as much as possible, thereby improving the local prediction accuracy of the point-by-point weighted hybrid proxy model, and there is no need for optimization calculation, and the time cost is low.
[0084] (2), The present invention uses the GMSE error as the selection basis for the global benchmark model and the calculation basis for the global weight coefficient, thereby ensuring the global prediction accuracy of the point-by-point weighted hybrid proxy model.
[0085] (3), The hybrid measure of the present invention integrates the global error measure and the local error measure, thereby ensuring the overall prediction accuracy of the point-by-point weighted hybrid proxy model.
[0086] In summary, the present invention has advantages such as simple logic and effectively improving the overall prediction accuracy of the hybrid surrogate model, and has high practical value and popularization value in the technical field of hybrid surrogate model construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 is a schematic flowchart of the method of the present invention.
[0088] Figure 2 is the box plot of R in the modeling performance indexes of the point-by-point weighted hybrid surrogate model based on the hybrid measure and four sub-surrogate models for 40 standard test functions in the specific application example of the present invention. 2 of the present invention.
[0089] Figure 3 is the comparison chart of the mean value of R and the number of R 2 >0.8 in the modeling performance indexes of the point-by-point weighted hybrid surrogate model based on the hybrid measure and four sub-surrogate models for 40 standard test functions in the specific application example of the present invention. 2 >0.8 in the specific application example of the present invention.
[0090] Figure 4 is the comparison chart of the mean value of NRMSE in the modeling performance indexes of the point-by-point weighted hybrid surrogate model based on the hybrid measure and four sub-surrogate models for 40 standard test functions in the specific application example of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0091] Next, with reference to the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0092] In this embodiment, the term "and / or" only describes the associated relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone.
[0093] The terms "first" and "second" in the description and claims of this embodiment are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first target object and the second target object are used to distinguish different target objects, rather than to describe a specific order of the target objects.
[0094] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.
[0095] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" refers to two or more. For example, a plurality of processing units refers to two or more processing units; a plurality of systems refers to two or more systems.
[0096] As Figures 1-4 shown, the embodiments of the present invention propose a modeling method for a point - by - point weighted hybrid surrogate model based on a hybrid measure. As Figure 1 shown, in this embodiment, 40 standard test functions from low - dimension (1 - D) to high - dimension (16 - D) are modeled. The method of the present invention is used to determine the point - by - point weighted hybrid surrogate model based on the hybrid measure and verify its accuracy, so as to understand the specific modeling process and technical advantages of the disclosed technical method. The specific steps are as follows:
[0097] Step 1: Use the Latin hypercube sampling method to extract training samples in the original variable space and form an initial training point set \(x^{(j)}\) \((j = 1,\ldots,N)\); \(N\) represents the number of training samples, \(N = 10D\), and \(D\) is the dimension of the input variable. (j) (j = 1,…,N); N represents the number of training samples, N = 10D, D is the dimension of the input variable;
[0098] Perform simulation analysis on the training samples to obtain the true response value \(g(x^{(j)})\), and form a database \(DB[x^{(j)}|g(x^{(j)})]\) \((j = 1,\ldots,N)\); where \(x=(x_1,\ldots,x_D)\) is a D - dimensional input variable. (j) )], forming the database DB[x (j) |g(x (j) )](j = 1,…,N); where x = (x1,…,x D ) is the D - dimensional input variable.
[0099] The 40 standard test functions are shown in Table 1:
[0100] Table 1 40 standard test functions
[0101]
[0102]
[0103]
[0104] Step 2: Use the initial sample points and the corresponding true values to construct a PRS sub - surrogate model in matlab using the fitnlm toolbox, specifically including:
[0105] Step 2.1.1: Determine the data dimension D, select the basis function as a complete quadratic polynomial, and perform multivariate polynomial regression analysis on the training sample points and their true response values in the database DB according to the least squares method to determine the regression function;
[0106] Step 2.1.2: Conduct variance and error statistical analysis on the regression function to determine the PRS model, and its basic form:
[0107]
[0108] where y1(x) is the prediction function of the objective function in the PRS model, and x i (i = 1, …, D) is the i-th component of the D-dimensional variable, and β is the unknown weight coefficient of the prediction function;
[0109] Using the initial sample points and the corresponding true values, construct an RBF sub-agent model in matlab using the RBF neural network toolbox, and optimize the hyperparameters using the particle swarm algorithm, specifically including:
[0110] Step 2.2.1: Determine the data dimension D, select the kernel function type as the Gaussian function, and its function is as follows:
[0111]
[0112] where r is the distance norm between the position point and the sample point, and the Euclidean distance ||x - x i || is used, which is a common distance kernel function, and σ is the bias expansion constant, σ > 0, used to control the shape of the kernel function.
[0113] Step 2.2.2: Linearly weight the sample points using the radial function to determine the RBF model, and its basic form is:
[0114]
[0115] where y2(x) is the prediction function of the objective function in the RBF model, and x i (i = 1, …, D) is the i-th component of the D-dimensional variable, is the distance kernel function, m is the number of kernel functions, and λ is the linear combination weight coefficient; r is the distance norm between the position point and the center point (sample), and the Euclidean distance ||x - x i ||;
[0116] Using the initial sample points and the corresponding true values, construct a KRG sub-agent model in matlab using the DACE toolbox, specifically including:
[0117] Step 2.3.1: Determine the data dimension D, select the regression function of the KRG model as a quadratic polynomial model, and the relevant model function is the most typical Gaussian function;
[0118] Step 2.3.2: Define the value of theta and the thresholds lob and upb, where theta represents the hyperparameter of the selected relevant function;
[0119] Step 2.3.3: Train the KRG model on the training sample points and their true response values in the database DB, and its basic form:
[0120]
[0121] Among them, y3(x) is the prediction function of the target function in the KRG model, and β i is the regression coefficient, and f i (x) is the regression function, and Z l (x) is a stationary random process with a mean of 0 and a variance of σ 2 ;
[0122] Using the initial sample points and the corresponding true values, construct an SVR sub-agent model using the SVR support vector regression toolbox in matlab, and optimize the hyperparameters using the particle swarm algorithm, specifically including:
[0123] Step 2.4.1: Determine the data dimension D, and select e-SVR with the kernel function type as the RBF radial basis function;
[0124] Step 2.4.2: Define the values of the loss function c and the gamma function g in the kernel function. c and g represent the selected hyperparameters;
[0125] Step 2.4.3: Train the SVR model on the training sample points and their true response values in the database DB, and its basic form:
[0126]
[0127] Among them, y4(x) is the prediction function of the target function in the SVR model, and k(x i T x) = θ(x i T )θ(x j ) is the kernel function.
[0128] Step 2.4.4: Initialize the particle swarm. The positions of the particles in the particle swarm are the values of the loss function c and the gamma function g in the kernel function (hyperparameters), and the velocities are the change directions and rates of the hyperparameters. Determine the number of particles in the particle swarm, as well as the initial positions and initial velocities of each particle. The optimization objective is to minimize the cross-validation error of the SVR sub-agent model, and set the convergence condition in the optimization process. After completing the particle swarm optimization, obtain the optimal values of the loss function c and the gamma function g in the kernel function, and reconstruct the SVR sub-agent model.
[0129] In this embodiment, the number of particles in the particle swarm is set to 10, and the preset convergence condition is that the number of optimization times reaches the preset maximum number of iterations of 100 times.
[0130] Step 3: Determine the local benchmark model for each sample point according to the LOO cross-validation error, specifically including:
[0131] Step 3.1.1: For any training sample point x (j) (j = 1,…, N), respectively construct four sub-agent models for LOO cross-validation, and obtain its LOO cross-validation error error. The expression is as follows:
[0132]
[0133] Where represents the LOO cross-validation error of the k-th sub-agent model at the j-th sample point, y (j) is the true response value of the model at this sample point, is the predicted response value of the k-th sub-agent model at this sample point.
[0134] Step 3.1.2: For each sample point x (j) (j = 1,…, N), select the sub-agent model with the smallest LOO cross-validation error error as the local benchmark model μ(x (i) );
[0135] Determine the global benchmark model and the global weight coefficient according to the LOO generalized mean square cross-validation error, specifically including:
[0136] Step 3.2.1: Conduct LOO cross-validation on the four sub-agent models again, and obtain the GMSE error. The expression is as follows:
[0137] Where GMSE k represents the GMSE error of the k-th sub-agent model, y (j) is the true response value of the model at the j-th sample point, is the predicted response value of the k-th sub-agent model at the sample point, and N represents the number of training points.
[0138] Step 3.2.2: Sort all sub-agent models from low to high according to the GMSE error results. The sub-agent model with the smallest GMSE error and the highest global accuracy is selected as the global benchmark model f g (x), and at the same time calculate the global weight coefficient, and its expression is as follows:
[0139]
[0140] where w g is the global weight coefficient, where GMSE1 is the smallest GMSE value and GMSE2 is the second smallest GMSE value.
[0141] Step 4: Determine the sample density based on the crowding distance, specifically including:
[0142] Step 4.1: For any prediction point x (i) , calculate the sum D (j) of its Euclidean distances from all training sample points x i , which is defined as the crowding distance, and its expression is as follows:
[0143]
[0144] where D i represents the crowding distance of the i-th prediction point, using the Euclidean distance, x (j) represents the j-th training point, x (i) represents the i-th prediction point, and N is the number of training points.
[0145] Step 4.2: Calculate the sample density ρ i at any prediction point, and its expression is as follows:
[0146]
[0147] Step 4.3: Normalize the sample density to obtain the normalized sample density Nρ i at any prediction point, and its expression is as follows:
[0148]
[0149] Step 5: Obtain the local accuracy coefficient based on local uncertainty, specifically including:
[0150] For any prediction point x (i) , calculate the local accuracy coefficient of each sub-agent model at this point, and its expression is as follows:
[0151] Among them, P k (x (i) ) represents the local accuracy coefficient of the k-th sub-agent model at the i-th prediction point, and y k (x (i) ) represents the predicted response value of the k-th sub-agent model at the i-th prediction point. μ(x (i) ) represents the predicted response value of the local benchmark model at the i-th prediction point. The sensitivity parameter σ 2 The calculation formula is as follows:
[0152]
[0153] Among them, σ1(x (i) ) represents the sensitivity parameter at the i-th prediction point, and Nρ i is the normalized sample density at the i-th prediction point.
[0154] Step 6: Determine the local weight coefficients of each sub-agent model, specifically including:
[0155] For any prediction point x (i) , by normalizing the local accuracy coefficients of each sub-agent model in the sub-model library, the local weight coefficients corresponding to each sub-agent model can be calculated. The expression is as follows:
[0156]
[0157] Among them, w localk (x (i) ) represents the local weight coefficient of sub-agent model k at the i-th prediction point, and P k (x (i) ) represents the local accuracy coefficient of sub-agent model k at the i-th prediction point, and n s represents the number of sub-agent models in the model library.
[0158] Step 7: Determine the preliminary hybrid surrogate model based on the local error measure, specifically including:
[0159] Through each sub-agent model and the local weight coefficients, determine the preliminary hybrid surrogate model based on the local error measure. Its expression is as follows:
[0160]
[0161] Among them, f PHSM represents the preliminary hybrid surrogate model, y k represents sub-agent model k, w local represents the local weight coefficient of sub-agent model k, and n s represents the number of sub-agent models in the model library.
[0162] Step 8: Determine the point - by - point weighted hybrid surrogate model based on the hybrid measure, specifically including:
[0163] According to the preliminary hybrid surrogate model \(f\) based on the local error measure PHSM , the global benchmark model \(f\) g and the global weight coefficient \(w\) g , determine the point - by - point weighted hybrid surrogate model \(f\) PWHSMHM based on the hybrid measure, and its expression is as follows:
[0164] \(f\) PWHSMHM \(=\) \(w\) g \(f\) g +(1 - \(w\) g )\(f\) PHSM
[0165] where \(f\) PWHSMHM represents the point - by - point weighted hybrid surrogate model based on the hybrid measure, \(f\) g represents the global benchmark model, and \(f\) PHSM represents the preliminary hybrid surrogate model.
[0166] Step 9: Conduct the accuracy verification of the surrogate model, specifically including:
[0167] Use the four sub - surrogate models and the point - by - point weighted hybrid surrogate model to fit the above 40 standard test functions, then randomly generate 500D test points through Latin hypercube sampling, and use the established surrogate models to predict the test points. Their prediction performances are evaluated by calculating the coefficient of determination (\(R\) 2 ) and the normalized root - mean - square error (NRMSE) of 500D random test points, and the expressions are as follows:
[0168]
[0169] where \(N\) test is the number of samples in the test set, \(y\) i is the true value of the \(i\) - th test point, is the predicted value of the \(i\) - th test point. Tables 2 - 5 and Figures 2-4 give the test results of the point - by - point weighted hybrid surrogate model based on the hybrid measure (PWHSMHM) using the proposed method for 40 standard test functions, as well as the comparison with the four sub - surrogate models.
[0170] The closer the \(R\) 2 value is to 1, the smaller the overall error between the surrogate model and the true model, and the higher the approximation degree. According to Tables 2 - 3 and Figures 2-4 it can be seen that for 40 standard test functions, the approximation degree of the point - by - point weighted hybrid surrogate model based on the hybrid measure (PWHSMHM) to the true model is generally higher than that of the four sub - surrogate models.
[0171] The smaller the NRMSE value, the higher the fitting accuracy of the model. According to Tables 4 - 5 and Figure 4 it can be seen that for 40 standard test functions, the point - wise weighted hybrid surrogate model based on hybrid measure (PWHSMHM) has a higher overall fitting accuracy for the true model than the four sub - surrogate models.
[0172] R 2 and the test results of NRMSE are consistent, which proves that the fitting accuracy and robustness of the point - wise weighted hybrid surrogate model based on hybrid measure (PWHSMHM) are better than those of the classical sub - surrogate model methods. It can be seen that considering the integrated error measure of global accuracy and local accuracy is of great significance for improving the prediction accuracy of the hybrid surrogate model. It should be noted that for some surrogate models, especially for high - dimensional problems, due to their inability to achieve approximate fitting, the 2 calculated value of R is negative, but this negative value only represents the failure of fitting and has no actual statistical significance, and it will affect the analysis of the 2 R mean. Here, all R values less than 0 are set to 0, representing the failure of fitting. 2
[0173] Table 2 Test Results: R 2 Index
[0174]
[0175] Table 3 Statistical Results of R 2 Index Statistical Results
[0176]
[0177] Table 4 Test Results: NRMSE Index
[0178]
[0179]
[0180] Table 5 Statistical Results of NRMSE Index
[0181]
[0182] Finally, it should be noted that the above - mentioned embodiments are only used for exemplifying and explaining the present invention, and are not intended to limit the present invention within the scope of the described embodiments. In addition, those skilled in the art can understand that the present invention is not limited to the above - mentioned embodiments, and more variations and modifications can be made according to the teachings of the present invention, and these variations and modifications all fall within the scope of protection required by the present invention.
Claims
1. A modeling method for a point - by - point weighted hybrid surrogate model based on a hybrid measure, characterized in that, It includes the following steps: Step 1: Use the Latin hypercube sampling method to extract training samples in the variable space, and obtain the true response values through simulation analysis; Step 2: Use the initial sample points and the corresponding true response values to construct PRS, RBF, KRG, and SVR surrogate models respectively and perform hyperparameter optimization; Step 3: Determine the local benchmark model and the global benchmark model for each sample point according to the LOO cross-validation error and the GMSE error, and determine the global weight coefficient; The specific content of Step 3 includes: Step 3.1.1: For any training sample point x (j) , where j = 1, …, N, construct four sub surrogate models respectively for leave-one-out cross-validation, and obtain its leave-one-out cross-validation error error, and the expression is as follows: Among them represents the LOO cross - validation error of the k - th surrogate model at the j - th sample point, and y (j) is the true response value of the model at this sample point, and is the predicted response value of the k - th surrogate model at this sample point; Step 3.1.2: For each sample point x (j) , where j = 1, …, N, select the surrogate model with the smallest leave-one-out cross-validation error error as the local reference model μ(x (i) ); Step 3.2.1: Perform LOO cross-validation on the four surrogate models again to obtain the GMSE error, and the expression is as follows: where GMSE k represents the GMSE error of the k-th sub surrogate model, and y (j) is the true response value of the model at the j-th sample point, is the predicted response value of the k-th sub surrogate model at this sample point, and N represents the number of training points; Step 3.2.2: Sort all sub-agent models from low to high according to the GMSE error results. The sub-agent model with the smallest GMSE error and the highest global accuracy is selected as the global benchmark model f g (x). At the same time, calculate the global weight coefficient, and its expression is as follows: where w g is the global weight coefficient, where GMSE1 is the minimum GMSE value and GMSE2 is the second minimum GMSE value. Step 4: Determine the sample density based on the crowding distance; Step 5: Determine the local accuracy coefficient based on the local error measure; The specific content of Step 5 includes: For any prediction point x (i) , calculate the local accuracy coefficients of each sub-agent model at this point, and its expression is as follows: where P k (x (i) ) represents the local accuracy coefficient of the k-th surrogate model at the i-th prediction point, y k (x (i) ) represents the predicted response value of the k-th surrogate model at the i-th prediction point, μ(x (i) ) represents the predicted response value of the local benchmark model at the i-th prediction point, and the sensitivity parameter σ 2 is calculated as follows: where σ1(x (i) ) represents the sensitivity parameter at the i-th prediction point, and Nρ i is the normalized sample density at the i-th prediction point; Step 6: Determine the local weight coefficient of each surrogate model; Step 7: Determine the preliminary hybrid surrogate model based on the local error measure; Step 8: Determine the point-by-point weighted hybrid surrogate model based on the hybrid measure. The specific content of Step 8 includes: Based on the preliminary hybrid surrogate model \(f\) of the local error measure PHSM , the global benchmark model \(f\) g , and the global weight coefficient \(w\) g , determine the point - by - point weighted hybrid surrogate model \(f\) PWHSMHM based on the hybrid measure, and its expression is as follows: f PWHSMHM = w g f g +(1 - w g )f PHSM where f PWHSMHM represents a point - by - point weighted hybrid surrogate model based on a hybrid measure, f g represents a global benchmark model, f PHSM represents a preliminary hybrid surrogate model.
2. The modeling method of the point-by-point weighted hybrid surrogate model based on the hybrid measure according to claim 1, wherein, The specific content of Step 1 includes: Training samples are drawn in the original variable space using the Latin hypercube sampling method, and an initial training point set \(x\) is formed, where \(j = 1,\ldots,N\); \(N\) represents the number of training samples; (j) , where \(j = 1,\ldots,N\); \(N\) represents the number of training samples; Perform simulation analysis on the training samples to obtain the true response value g(x (j) ), and form the database DB[x (j) |g(x (j) )]; where j = 1, …, N, and x = (x1, …, x D ) is the D-dimensional input variable.
3. The modeling method of the point-by-point weighted hybrid surrogate model based on the hybrid measure according to claim 2, wherein The specific content of Step 2 includes: Construct the PRS surrogate model: Step 2.1.1: Determine the data dimension D, select the basis function as a complete quadratic polynomial, and perform multivariate polynomial regression analysis on the training sample points and their true response values in the database DB according to the least squares method to determine the regression function; Step 2.1.2: Perform variance and error statistical analysis on the regression function to determine the PRS model, and its basic form: where y1(x) is the prediction function of the objective function in the PRS model, and x i is the i-th component of the D-dimensional variable, where i = 1, …, D, and β is the unknown weight coefficient of the prediction function; Construct the RBF surrogate model: Step 2.2.1: Determine the data dimension D, select the kernel function type as the Gaussian function, and its function is as follows: where r is the distance norm between the position point and the sample point, and the Euclidean distance ||x - x i || is used, which is a common distance kernel function, σ is the partial expansion constant, σ > 0, and is used to control the shape of the kernel function; Step 2.2.2: Linearly weight the sample points using the radial function to determine the RBF model, and its basic form is: Among them, y2(x) is the prediction function of the objective function in the RBF model, and x i is the i-th component of the D-dimensional variable, where i = 1, …, D, is the distance kernel function, m is the number of kernel functions, and λ is the linear combination weight coefficient; r is the distance norm between the position point and the center point, and the Euclidean distance ||x - x i || is adopted; Construct the KRG surrogate model: Step 2.3.1: Determine the data dimension D, select the regression function of the KRG model as the quadratic polynomial model, and the related model function is the Gaussian function; Step 2.3.2: Define the value of theta and the thresholds lob and upb, where theta represents the hyperparameter of the selected related function; Step 2.3.3: Train the KRG model on the training sample points and their true response values in the database DB, and its basic form: Among them, y3(x) is the prediction function of the objective function in the KRG model, and β i is the regression coefficient, f i (x) is the regression function, and Z l (x) is a stationary random process with a mean of 0 and a variance of σ 2 ; Construct the SVR surrogate model: Step 2.4.1: Determine the data dimension D, select e-SVR, and the kernel function type as the RBF radial basis function; Step 2.4.2: Define the values of the loss function c and the gamma function g in the kernel function. c and g represent the selected hyperparameters; Step 2.4.3: Train the SVR model on the training sample points and their true response values in the database DB, and its basic form: Among them, y4(x) is the prediction function of the objective function in the SVR model, and k(x i T x) = θ(x i T )θ(x j ) is the kernel function; Step 2.4.4: Initialize the particle swarm. The positions of the particles in the particle swarm are the values of the loss function c and the gamma function g in the kernel function. c and g represent the selected hyperparameters, and the velocity is the change direction and rate of the hyperparameters. Determine the number of particles in the particle swarm, as well as the initial positions and initial velocities of each particle. The optimization objective is to minimize the cross-validation error of the SVR sub-agent model, and set the convergence condition in the optimization process. After completing the particle swarm optimization, obtain the optimal values of the loss function c and the gamma function g in the kernel function, and reconstruct the SVR sub-agent model.
4. The modeling method of the point - by - point weighted hybrid surrogate model based on the hybrid measure according to claim 3, characterized in that, The specific steps of step 4 include: Step 4.1: For any prediction point x (i) , calculate the sum D (j) of its Euclidean distances from all training sample points x i , which is defined as the crowding distance, and its expression is as follows: Among them, D i represents the crowding distance of the i-th prediction point, using the Euclidean distance, x (j) represents the j-th training point, x (i) represents the i-th prediction point, and N is the number of training points; Step 4.2: Calculate the sample density ρ at any prediction point i , and its expression is as follows: Step 4.3: Normalize the sample density to obtain the normalized sample density Nρ at any prediction point i , and its expression is as follows:
5. The modeling method of the point-by-point weighted hybrid surrogate model based on the hybrid measure according to claim 4, wherein The specific steps of step 6 include: For any prediction point x (i) , by normalizing the local accuracy coefficients of each sub-agent model in the sub-model library, the corresponding local weight coefficients of each sub-agent model are calculated, and the expression is as follows: where w localk (x (i) ) represents the local weight coefficient of the sub-agent model k at the i-th prediction point, P k (x (i) ) represents the local precision coefficient of the sub-agent model k at the i-th prediction point, n s represents the number of sub-agent models in the model library.
6. The modeling method of the point-by-point weighted hybrid surrogate model based on the hybrid measure according to claim 5, characterized in that The specific steps of step 7 include: Determine the preliminary hybrid agent model based on the local error measure through each sub-agent model and the local weight coefficient, and its expression is as follows: where f PHSM represents the preliminary hybrid proxy model, y k represents the sub - proxy model k, w local represents the local weight coefficient of the sub - proxy model k, n s represents the number of sub - proxy models in the model library.