Kriging large model adaptive training method and system for face modeling
By adopting convergence standards based on failure probability error and adaptive learning strategies in face modeling, the training process of the ALK model is optimized, and the problem of excessive number of functional function calls in the existing technology is solved, and the computing efficiency and robustness are improved.
Patent Information
- Application Number
- CN202410606333.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-01-16
AI Technical Summary
The existing technology calls functional functions too many times in face modeling, resulting in increased computational cost and time, affecting the efficiency of the Kriging model.
The convergence standard based on failure probability error is adopted, and the training process of the ALK model is optimized by dynamically adjusting the training set and introducing adaptive learning strategies, and the training process of the ALK model is reduced to reduce the number of unnecessary calculations and calls of functional functions.
It improves computing efficiency, reduces computing costs, enhances the system's robustness to uncertainty and noise, and realizes the efficiency of face model topology model.
Smart Images

Figure CN119027575B_ABST
Abstract
Description
[0001] This application is a divisional application with application number CN202410060290.2, application date January 16, 2024, and invention name “An adaptive training method and system for a vertical model for face modeling”. Technical Field
[0002] The present invention relates to the technical field of film and television 3D modeling, specifically to the technical field of intelligent generation of Kriging model for face modeling, and in particular to a Kriging large model adaptive training method and system for face modeling. Background Art
[0003] For 3D modeling technology in movies, TV or games, the human face is the most expressive part of the human body, and the surface of the human face has highly complex geometric shapes and very rich color and texture information. Therefore, whether it is movie design, game design or CG design, the precision of face modeling provides a lot of options and expansion for the effect limit and lens language of the work. However, even though computer graphics technology is relatively mature, it is still a very challenging task to build a realistic 3D model of the human face even with manual modeling technology; and the current automated topology modeling and UV unfolding (mapping) technology is developing steadily (such as the FaceBuilder technology of Blender software), and its main principles are as follows: Figure 2 As shown, it reads the training point data of the target graphics based on a specific model and outputs the actual three-dimensional model in the three-dimensional modeling software, and the modeling difficulty is also relatively difficult.
[0004] Currently, Kriging is one of the optional technologies for the above-mentioned automatic modeling and topological face 3D modeling algorithms. [2] . The Kriging model is also known as the spatial autocovariance optimal interpolation algorithm. It is a modeling method for geological environment proposed by French geomathematician Georges Matheron. From the algorithm level, it is an optimal interpolation strategy. The Kriging model not only considers the relative position of the observation point (feature point) and the estimated point (non-feature point), but also considers the relative position relationship between each observation point. Therefore, this method has a high degree of approximation, strong extrapolation ability, and a wide range of applications. At the same time, the Kriging model has become a relatively mature driving model technology for face automatic topological modeling technology. [1-2] , and can even provide a model basis for subsequent artificial intelligence video rendering technology. Its principle is also relatively intuitive after abstraction: if a face is laid flat, its three-dimensional structure and topological features are actually no different from a "geological structure."
[0005] The newly developed Monte Carlo model based on active learning Kriging model (AK-MCS) [3] Compared with the traditional method of mechanically arranging training points (e.g. Figure 3 AK-MCS can "actively" select training points close to the limit state surface to automatically construct the Kriging model, so it can greatly improve the efficiency of model training. However, the estimation of the small failure probability of the Kriging model is a challenging problem. [4-7] ; To solve this problem, the Active Learning Kriging (ALK) model was developed, but it usually requires a large number of candidate samples, and each time a sample is selected, the Kriging model needs to predict a large number of candidate points, which makes the learning process very time-consuming. To solve this problem, a variety of strategies combining the ALK model and advanced sampling methods have been proposed in the prior art.
[0006] For example, ECHARD et al. [4] The AK-IS method is proposed by combining the ALK model with the IS. The basic idea is to first use FORM to find the most likely failure point (MPP) of the functional function, and then generate an important sampling sample pool around the MPP point to establish the ALK model. However, FORM can only find a single MPP point. For a functional function with multiple MPP points, AK-IS will underestimate the failure probability. To make up for this deficiency, CADINI et al. [5] The Meta-AK-IS2 method is proposed to obtain multiple MPP points. [6] It is proposed to use adaptive sampling domain to identify all the most likely failure regions of the functional function, and then use the samples in the failure domain to construct the important sampling function through kernel density estimation, and generate a candidate sample pool for active learning. This method is named ALK-KDE-IS.
[0007] The latest technology [7-8] , proposed a strategy to search for multiple MPP points using a multi-objective optimization algorithm, namely ALK-EMO-IS and ALK-MCS. The basic idea is to first construct an approximate proxy model of the limit state function, and then identify multiple MPP points of the proxy limit state function through an evolutionary multi-peak optimization algorithm, and establish an important sampling function around these MPP points to generate an important sample pool for updating the ALK model.
[0008] Although these traditional technologies have made some progress in dealing with the problem of small failure probability, and the technology of face modeling has achieved certain beneficial effects, there are still some technical problems in general. One of them is the excessive number of calls to the function, mainly because the traditional convergence criteria are usually too strict and conservative, resulting in the need to continue to call the function for more calculations after reaching a certain accuracy requirement; research shows that [9] ,When the accuracy of the failure probability of the Kriging model reaches a certain level, even if a large number of training points are added, the effect of improving the accuracy of the failure probability can be ignored.
[0009] Therefore, the phenomenon of calling the function too many times will undoubtedly increase the computing cost and time. This will lead to a significant reduction in the computing resources and efficiency required for the Kriging model to generate a face model, which will in turn affect the efficiency of the subsequent artificial intelligence video rendering stage. To this end, the present invention proposes a Kriging large model adaptive training method and system for face modeling.
[0010] The citations in the present invention are as follows:
[0011] [1] Wang Jie, Wang Zhaoqi, et al. Facial animation system on mobile phone platform. The First Intelligent CAD and Digital Entertainment Academic Conference [R], 2004: 95-102.
[0012] [2] Zhao Na. Research on 3D face modeling technology based on photos[D]. Yanshan University: 2007(02).
[0013] [3] ECHARD B, GAYTON N, LEMAIRE M. AK-MCS: An active learning reliability method combining Kriging and Monte Carlo simulation[J]. Structural Safety, 2011, 33(2): 145-154.
[0014] [4]ECHARD B,GAYTON N,LEMAIRE M,et al. A combined importance sampling and kriging reliability method for small failure probabilities with time-demanding numerical models[J]. Reliability Engineering & System Safety, 2013, 111: 232-240.
[0015] [5]CADINI F,SANTOS F,ZIO E. An improved adaptive Kriging-basedimportance technique for sampling multiple failure regions of low probability[J]. Reliability Engineering & System Safety,2014,131:109-117.
[0016] [6]YANG X,LIU Y,MI C,et al. Active learning Kriging model combiningwith kernel-density-estimation-based importance sampling method for theestimation of low failure probability[J]. Journal of Mechanical Design,2018,140(5).
[0017] [7]YANG X,CHENG X. Active learning method combining Kriging model andmultimodal-optimization-based importance sampling for the estimation of smallfailure probability[J]. International Journal for Numerical Methods inEngineering,2020,121(21):4843-4864.
[0018] [8]ECHARD B,GAYTON N,LEMAIRE M. AK-MCS:An active learning reliabilitymethod combining Kriging and Monte Carlo simulation[J].Structural Safety,2011,33(2):145-154.
[0019] [9]WANG Z, SHAFIEEZADEH A. ESC: an efficient error-based stopping criterion for kriging-based reliability analysis methods[J]. Structural and Multidisciplinary Optimization, 2019, 59: 1621-1637. Summary of the invention
[0020] In view of this, the embodiment of the present invention hopes to provide a Kriging large model adaptive training method and system for face modeling to solve or alleviate the technical problems existing in the prior art, that is, how to provide an intelligent mechanism for the training of the ALK model, that is, an intelligent convergence and judgment mechanism, so as to optimize the efficiency and provide at least a beneficial option for this. The technical solution of the embodiment of the present invention is implemented as follows:
[0021] First, the adaptive training method of the Kriging large model for face modeling:
[0022] 1. Overview:
[0023] For the vertical ALK model used in the field of face modeling, when evaluating the accuracy of the Kriging model in the ALK model, it is necessary to focus on the accuracy of its estimated failure probability, rather than just the accuracy of the symbol prediction. This issue involves convergence and error control in reliability analysis. Traditional convergence criteria are usually based on mathematical convergence proofs to ensure the accuracy and stability of the calculation results. However, in traditional technologies, these criteria usually use the function value error or its statistical characteristics as the convergence basis, rather than directly judging the convergence of the failure probability error.
[0024] Therefore, in order to solve this problem, it is necessary to further design a technical solution based on the convergence criterion of failure probability error. The criterion of this technical solution can directly measure the accuracy of failure probability estimation and terminate the learning process in time when the target accuracy requirement is reached. By reducing unnecessary calculations and the number of function calls, the computational efficiency can be improved and the cost can be reduced.
[0025] 2. Technology leadership:
[0026] 2.1 ALK model:
[0027] The existing ALK model consists of two parts: a (constant) regression process and a random process:
[0028] ;
[0029] is the regression process part of the ALK model, representing the global average response of the function g(x), and β is only expressed as a constant in 2.1; z(x) is a stationary Gaussian random process, which reflects the deviation between the local response and the global response. For example, given a set of DoE (Design of Experiment, test scale, i.e., sample point set), where m is the total number of training points:
[0030] ;
[0031] For DoE, the response value output by the performance function g(x) will follow a Gaussian distribution with a predicted mean Defined as:
[0032] ;
[0033] Prediction variance of DoE for:
[0034] ;
[0035] Where T is the transformation matrix, are the known training points x and the unknown training points The correlation function vector between is the autocorrelation matrix composed of training points, , , , r is the Kriging hyperparameter in the ALK model, and its value (including other terms) is detailed in the literature. [7-8] .
[0036] 2.2 ALK convergence rules and technical improvements:
[0037] The convergence criteria of traditional ALK models are defined based on the numerical value of the learning function. [8] The convergence criteria of the ALK-MCS are defined as: ;
[0038] When the ALK model is trained on the face training set, the accuracy of the failure probability of the MPP point is the core goal of the ultimate concern, not just the accuracy of the symbol prediction. Although the accuracy of the symbol prediction can ensure an accurate estimate of the failure probability, this causes the established Kriging model to be too focused on accurately predicting the failure probability at the macro level. Studies have shown that [9] ,When the accuracy of the failure probability of the Kriging model reaches a certain level, even if a large number of training points are added, the effect of improving the accuracy of the failure probability is not significant.
[0039] Therefore, the applicant believes that when evaluating the accuracy of the Kriging model, attention should be paid to the accuracy of its estimated failure probability, rather than just the accuracy of symbol prediction. Therefore, it can be clearly seen that the starting point for the improvement of this technical solution is: how to formulate a convergence standard based on the failure probability accuracy based on the evaluated failure probability error, and then derive the error convergence criterion for active learning, so that the Kriging model automatically stops the learning process after reaching the predetermined target accuracy, avoiding the addition of unnecessary training points; and optimizing the training efficiency of the ALK model based on this starting point.
[0040] (III) Technical content:
[0041] Considering the failure probability error of the evaluation, we can first formulate an optimal auxiliary probability density function (PDF) to demonstrate the phenomenon of its error; PDF includes a random variable x, and PDF describes the probability properties of x. The probability density function is defined as p(x), which is a non-negative value, that is, p(x) ≥ 0, and the integral value is 1 within the range of its possible values, that is:
[0042] ;
[0043] The meaning of the above formula is used to describe the probability of a random variable in any interval, so it can be further expressed as:
[0044] ;
[0045] However, the above formula has no practical significance because its denominator contains the unknown failure probability PF. But it indirectly reminds us that we might as well use the failure domain samples to approximately construct a function similar to the failure probability PF and take it into consideration, and construct an approximately optimal auxiliary probability density function p(x)' based on the above formula. The auxiliary probability density function p(x)' is used to form a candidate sample pool S, and it is regenerated with the update of the Kriging model. In this process, the function value s of each sample in the candidate sample pool S is calculated, and the upper bound of the failure probability error evaluated by the current Kriging model is obtained; in this way, the target failure error threshold is formulated, and the convergence of the Kriging model can be made precise. At this point, the context of the present technical solution is already very clear, that is, the corresponding operations are performed according to the following steps S1~S4.
[0046] 3.1 Step S1, generate initial training points:
[0047] N uniformly distributed initial sample points are extracted from the random variable space σ. This process requires the execution of the Halton Sequence algorithm for sampling. The extracted initial sample points are uniformly substituted into the functional function g(x). The response values of these sample points are first calculated, and then they are aggregated to form an initial training set T. Then, the Kriging model is established or updated based on the current training set T.
[0048] 3.1.1 Step S100, Halton sequence sampling algorithm:
[0049] The Halton sequence is a low-discrepancy sequence used for uniform sampling in high-dimensional space. In one dimension, the sample value of the i-th element of the Halton sequence is for:
[0050] ;
[0051] is the cardinality of uniform distribution, k is the maximum value of the total element i, is an integer uniformly distributed in the range [0, b-1].
[0052] Similarly, for multi-dimensional space, each dimension uses a different prime number as the base to ensure better uniformity. For example, taking two dimensions as an example, the extracted value of the i-th element of the Halton sequence is for:
[0053] ;
[0054] The same is true for other dimensions.
[0055] 3.1.2 Step S101, obtaining uniformly distributed initial sample points:
[0056] Each element of the Halton sequence is considered as a component of a random variable. If there is a k-dimensional space, then the i-th element of the Halton sequence It can be expressed as:
[0057] ;
[0058] The form of the above elements represents the value of the i-th element in the j-th dimension. This value is obtained by sampling the Halton sequence using different prime numbers as the base. In order to make the Halton sequence evenly distributed in each dimension, different prime numbers can be selected as the base, for example ;
[0059] Specifically, the i-th sample point It is expressed as:
[0060] ;
[0061] Then, these sample points Substitute into the performance function g(x), the form of the performance function g(x) is given in Section 2.1 above, and because the ALK model is available, its operation form and the corresponding response value are obtained It can be simplified as:
[0062] ;
[0063] 3.1.3 Step S102, forming an initial training set T:
[0064] The sample points obtained And the corresponding function response value Combine to form the initial training set T:
[0065] ;
[0066] is the sample point corresponding to The response value output by the function g(x).
[0067] 3.1.4 Step S103, establishing the Kriging model required for the ALK model based on the current training set T:
[0068] Step S103 will exemplarily construct a standard constant Kriging model:
[0069] Through the N known sample points in the training set T , and the value The linear combination of The unknown quantity at , so that the variance of the estimation error is minimized.
[0070] Therefore, we can assume that Z(x) is a second-order stationary random function, then the position The estimated value of the point is:
[0071] ;
[0072] is the estimated coefficient;
[0073] The estimated error is recorded as: ,in is the true value of the estimated point. Since these true values are unknown, they can be regarded as random variables, then the mean square error corresponding to the random variable r is for:
[0074] ;
[0075] E(r) represents the mathematical expectation (i.e. mean) of the random variable r;
[0076] In order to minimize the mean square error, the estimated coefficients The constraints also need to be met:
[0077] ;
[0078] Then, based on the Lagrange multiplier method, when the variance of the estimation error is minimized, N sample points are input ,in is the i-th sample point (same for the response value), is the jth sample point (same for the response value), is the initial sample point (same for the response value), and the equation system of the Kriging model is obtained:
[0079] ;
[0080] µ is the Lagrange multiplier,
[0081] Represents the correlation between two positions; of course, the equation group can also be converted into a matrix form, and its form is not strictly limited.
[0082] At the same time, the above equations must also follow the Gaussian distribution described in 2.1, and its form is:
[0083] ;
[0084] Where μ is the mean and exp is the exponential function. If it is an exponential function of a one-dimensional Gaussian distribution, then , where σ is the standard deviation of the Gaussian distribution.
[0085] Combined with literature [2] Based on the above expression, the properties of the exemplary Kriging model can be summarized as follows: For identification, the topological surface passes through the sample point , rather than approximating the sample points When the Kriging model estimates a sample point When , the weight of the known point is 1, and the others are 0. If the data is completely random, that is, unrelated, then for a sample point The estimated value of is the sum of all points, which is the concept of "Nugget Effect" of Kriging model. [2]This is equivalent to all weights being 1 / n and the correlation being 0. When there is no such effect, the closer the sample point is to the estimated point, the The greater the weight of .
[0086] It is understandable that the above-mentioned Kriging model is merely an example of introducing an existing Kriging model into the training set T. It can also be conventionally transformed through several different forms of Kriging models.
[0087] 3.2 Step S2, construct the limit state function proxy model and generate the candidate sample pool S:
[0088] The Gaussian mixture algorithm is executed to obtain the auxiliary probability density function p(x)' of the Kriging model, and the failure domain set G is combined with the auxiliary probability density function p(x)' to generate n random variable samples. , thus forming a candidate sample pool S. The candidate sample pool S will be used to select the training points for the next update, and each step will be regenerated with the update of the Kriging model.
[0089] 3.2.1 Step S200, performing a Gaussian mixture algorithm on the auxiliary probability density function p(x)':
[0090] Gaussian mixture model (GMM) is a probabilistic model algorithm used to cluster and estimate density of data. In this step, the auxiliary probability density function p(x)' is in the form of:
[0091] ;
[0092] in, are the weights of the mixture components, is the probability density function of the multidimensional Gaussian distribution, is the mean vector, is the covariance matrix and K is the number of components in the mixture model.
[0093] 3.2.2 Get the invalid domain set G:
[0094] Assume that the failure probability of the Kriging model for the input variable x is estimated as , the confidence interval is [L(x), U(x)], where L(x) is the lower bound and U(x) is the upper bound. At the same time, a threshold T1 is set to represent the threshold of the failure probability. Then the acquisition mechanism of the failure domain is:
[0095] ;
[0096] That is, if If it exceeds the threshold T1, a value of 1 is assigned, x is determined to belong to the failure domain, and is added to the failure domain set G; otherwise, a value of 0 is assigned, i.e., it is not considered.
[0097] The failure domain set G includes several failure domain samples :
[0098] ;
[0099] n is the failure domain sample The total number of .
[0100] Through such a mechanism, each time step S2 is executed, the failure domain samples are dynamically determined according to the prediction results and confidence intervals of the current Kriging model. The failure domain samples obtained in this way can be used for subsequent GMM or other method modeling.
[0101] Failure probability Modeling and analysis of probability distributions around the performance function g(x). Because the performance function g(x) defines the relationship between system performance and design requirements. Failure probability The way to obtain can be:
[0102] ;
[0103] Failure probability It is the probability that the corresponding value output by the functional function g(x) is less than or equal to zero.
[0104] You can also use numerical integration or Monte Carlo simulation to obtain .
[0105] 3.2.3 Step S201, using the failure domain set G combined with the auxiliary probability density function p(x)' to generate n random variable samples :
[0106] First, according to the content of 3.2.1, use the parameters of the Gaussian mixture model to generate samples. That is, randomly select a component from the mixture model according to the weight of each component, and then generate a sample from the multidimensional Gaussian distribution corresponding to the failure domain set G. This process can be expressed as:
[0107] ;
[0108] The symbol “~” represents the statistical concept of “…distributed in…”.
[0109] Considering each failure domain sample point in the random variable sample , randomly select a sample from the failure domain set G, expressed as , and according to the auxiliary probability density function Generate random variables:
[0110] ;
[0111] 3.2.4 Step S202, generate candidate sample pool S:
[0112] Repeat step S201 of 3.2.3 to generate n random variable samples , forming a candidate sample pool S:
[0113] ;
[0114] 3.3 Step S3, identify the next updated training point in the candidate sample pool S:
[0115] In this stage, we need to first calculate the function value s of each sample in the candidate sample pool S, and then select the smallest function value s* and its corresponding failure domain sample point in the function value s. , as the training point for the next update.
[0116] 3.3.1 Step S300, calculate the function value of each sample in the candidate sample pool S :
[0117] First, draw each random variable sample from the candidate sample pool S As input variables; for the technical solution of the present invention, the priority is the difference between system performance and design requirements, so a limit state function G(y) is introduced, and its input is the vector of variables:
[0118] ;
[0119] , is the same as the i-th input variable Related regulatory factors. is applied to the ith input variable The limit state function represents the difference between system performance and design requirements, and a negative value indicates that the system is within the design requirements, while a zero value indicates that the system is on the boundary of the design requirements.
[0120] 3.3.2 Step S301, select the minimum sample point in the function value s as the training point for the next update:
[0121] The function value of each sample output by the limit state function G(y) in 3.3.1 , aggregate it and select the smallest value s∗, and then find the corresponding failure domain sample point .
[0122] In this process, by calculating the function value of each sample in the candidate sample pool and then selecting the sample point with the minimum function value, it is ensured that the difference between system performance and design requirements is considered when selecting the training point for the next update. Such a selection strategy can guide the update of the model in active learning, making the model pay more attention to the area near the failure domain.
[0123] 3.4 Step S4, determine the convergence of the learning process:
[0124] Based on the minimum function value s* and its corresponding failure domain sample point , and the mean prediction provided by the current ALK model , calculate the upper bound of the failure probability error evaluated by the current Kriging model. If the upper bound is less than the target failure error threshold, stop the learning process; otherwise, calculate the upper bound confidence interval and return to S2 to update the Kriging model.
[0125] 3.4.1 Step S400, calculate the upper bound of the failure probability error:
[0126] Get the predicted mean based on 2.1 , this upper bound on the error can be formulated based on the uncertainty of statistical inference and model prediction:
[0127] ;
[0128] in, is to use the failure domain sample points The calculated failure probability, is the failure probability calculated using the ALK model predicted mean.
[0129] 3.4.2 Step S401, determine whether the learning stop condition is met:
[0130] ;
[0131] T2 is the target failure error threshold;
[0132] If the upper bound of the failure probability error If the error is less than the target failure error threshold T2, the learning stops. Otherwise, the learning continues.
[0133] 3.4.3 Step S402, calculate the upper confidence interval:
[0134] If learning is not stopped, the upper confidence interval of the failure probability error can be calculated to perform the next round of Kriging model update. This process involves further normal distribution algorithm and model update. It specifically includes steps S4020~S4023.
[0135] 3.4.3.1 Step S4020, Calculate the Standard Deviation of Failure Probability Error :
[0136] ;
[0137] is the ith failure probability error, is the average value of the total failure probability error; n is the total number of samples;
[0138] Step S4021, calculate the confidence interval radius Ma: ;
[0139] Z is the “Z-score” in the standard normal distribution, corresponding to the chosen confidence level α (taken as 95%);
[0140] Ma can be regarded as a marginal value, which refers to the radius of the confidence interval, indicating the distance between the upper and lower limits of the confidence interval and the mean. It represents the uncertainty range of the estimated failure probability error within the confidence level. If Ma is larger, it means that the confidence interval is wider and the uncertainty of the estimated failure probability error is greater. Conversely, if Ma is smaller, it means that the uncertainty of the estimated failure probability error is smaller.
[0141] Step S4022, calculate the new confidence interval CI:
[0142] ;
[0143] Once a new confidence interval CI of the failure probability error is obtained, a decision can be made based on this interval to determine whether to stop learning or continue updating the Kriging model.
[0144] 3.4.3.2 Step S4023, learning strategy plan:
[0145] According to 3.2.2, the confidence interval is [L(x), U(x)], where L(x) is the lower bound and U(x) is the upper bound. The above confidence interval is adjusted according to the new confidence interval CI obtained in 3.4.3.1. The strategy selected by this technical solution is:
[0146] 1) If the confidence interval is wide, it means that the estimate of the failure probability is less certain, and a more conservative learning strategy can be adopted, that is, increasing the number of samples or adjusting the learning step size;
[0147] 2) If the confidence interval is narrow, it means that the estimate of the failure probability is more reliable, and a more aggressive learning strategy can be adopted, that is, reducing the number of samples or accelerating learning.
[0148] The expression of the above strategy is: ;
[0149] γ is the original learning step size or learning rate, and γ' is the updated learning step size or learning rate. f(CI) is an adjustment function for the confidence interval CI:
[0150] ;
[0151] w(CI) is the width of the confidence interval, where w(CI)=U(x)−L(x). v is the adjustment parameter used to control the magnitude of the adjustment.
[0152] This scheme makes w(CI) larger and f(CI) smaller when the confidence interval width is larger, so γ' will decrease accordingly, making learning more conservative. When the confidence interval width is smaller, f(CI) is larger, and γ' will increase accordingly, making model learning more aggressive.
[0153] 3.5 Summary of intelligent convergence and judgment mechanism:
[0154] (1) Active learning: Use active learning strategies, especially the Active Learning Kriging (ALK) method based on the Kriging model. Active learning allows the model to autonomously choose where to obtain more training points, so that it can focus on areas near the failure domain and optimize the model more intelligently.
[0155] (2) Dynamically adjust the training set: During the training process, the training set is dynamically adjusted. Using the prediction information of the current model, new sample points that may be useful near the failure domain are selected to continuously improve the prediction performance of the model and reduce the number of calls to the function. This helps avoid spending too much computing resources in irrelevant areas.
[0156] (3) Termination condition based on failure probability error: A termination condition based on failure probability error is introduced, that is, when the Kriging model's estimation accuracy of the failure probability reaches the target level, the learning process is stopped. This can avoid adding unnecessary training points when the model is already sufficiently accurate, thus reducing computational costs.
[0157] (4) Considering the upper confidence interval of the failure probability error: In the learning process, we not only focus on the failure probability predicted by the model, but also introduce the upper confidence interval of the failure probability error. This strategy controls the uncertainty in the learning process and stops learning in time when the accuracy meets the requirements, thereby reducing the computational cost.
[0158] 3.6 The second technical solution:
[0159] After section 3.4.3.2, another technical solution can be introduced to connect to step S4, and an adaptive adjustment mechanism (referred to as step S5 in the present invention) can be introduced for the learning of the overall ALK model. The mechanism is implemented by the DS evidence theory algorithm (Dempster Shafer, DS), which considers information from different channels, aggregates and performs mapping and forms a correction factor ψ, and uses the correction factor ψ to perform adaptive correction on the auxiliary probability density function p(x)' in step S200.
[0160] 3.6.1 Step S500, obtaining evidence A and evidence B:
[0161] First, the candidate sample pool S is formed by collecting the samples in step S202 at the current time step t as evidence A (set):
[0162] ;
[0163] Then, the candidate sample pool S' is formed by collecting the samples in step S202 at the last time step t-1 as evidence B (set):
[0164] ;
[0165] The above time step assignment mechanism can be linked to the learning rate and step size of the Kriging model; that is, evidence A is obtained at the current learning rate and step size (based on the candidate sample pool S), and evidence B (based on the candidate sample pool S) is read at the previous learning rate and step size.
[0166] 3.6.2 Step S501, execute Dempster's combination principle:
[0167] Dempster's Combination Principle is a mathematical principle in DS evidence theory. Its principle is to solve the set of evidence A and evidence B, and output the set of combined confidence distribution Bel(C) and the combined uncertainty distribution Pl(C); where C is the possible hypothesis or state space.
[0168] Step S5020, merging trust allocation Bel(C):
[0169] ;
[0170] Ai∩B=∅ means that evidence A and evidence B have an intersection, and Ai∩B=∅ means that evidence A and evidence B do not have an intersection. Similarly, Ai∩C=∅ and Bi∩C=∅ mean that evidence A and evidence B have an intersection with hypothesis set C. Bel(C) represents the confidence in hypothesis set C (a value in the range [0,1]). By combining the Bel(C) of different evidences, a more comprehensive and reliable assessment of the confidence in hypothesis set C can be obtained. This helps to comprehensively consider the information from different evidences and improve the confidence level for different states.
[0171] Step S5021, combined uncertainty allocation Pl(C):
[0172] ;
[0173] In S5020~S5021, ∅ represents the empty set; and are a subset of the evidence A and the evidence B respectively; W is the weight; C is the possible set;
[0174] Pl(C) represents the uncertainty of the hypothesis set C (a value in the range [0,1]), which is complementary to Bel(C). By combining Pl(C) of different evidences, a more comprehensive and reliable assessment of the uncertainty of the hypothesis set C can be obtained. Combined with Bel(C), it can provide a more comprehensive understanding, including the confidence in a certain state and the consideration of the uncertainty of that confidence.
[0175] 3.6.3 Step S502, obtain the joint trust function Bel(A∪B):
[0176] ;
[0177] (Pl(A)∩Pl(B)) is the uncertainty distribution of the intersection of the evidence A and the evidence B respectively.
[0178] The numerator in the formula is the intersection of evidence A and evidence B, which is normalized by multiplying the uncertainty distribution and trust of the other party and then dividing by 1-Pl(A∩B) to obtain the joint trust function. This joint trust function can be used to comprehensively consider the relationship between different evidences and provide a more comprehensive trust evaluation for a state or hypothesis.
[0179] Among them, the specific method of obtaining (Pl(A)∩Pl(B)) is:
[0180] ;
[0181] and are subsets of evidence A and evidence B respectively. The numerator is the sum of the evidences wi supporting A∩B, and 1 minus this sum gives Pl(A∩B), which indicates the degree of uncertainty about A∩B.
[0182] in, Represents a subset of evidence A and a subset of evidence B In other words, it is able to find a subset and a subset , and is not zero under the correction of weight W, that is, the two subsets have an intersection. Responsible for finding some specific elements in hypothesis set C that are supported by both evidence A and evidence B. It is used to indicate the support for each element in hypothesis set C. It considers the support of two pieces of evidence for the possible set and combines these supports through Dempster's combination principle. The higher the value of this function, the stronger the support of the two pieces of evidence for an element in the possible set, and therefore the more confident the conclusion about the element is.
[0183] 3.6.4 Step S503, mapping the joint trust function Bel(A∪B) to an interval value between [0,1] as a correction factor ψ through a sigmoid function:
[0184] ;
[0185] e is the base of natural logarithm (approximately equal to 2.718); the input of sigmoid function is the output of Bel(A∪B). The characteristic of sigmoid function is that its output range is between [0,1] regardless of the input, and the change of input value is relatively smooth, so it is used to map real numbers to probability range.
[0186] In other words, the size of the correction factor ψ will be affected by Bel(A∪B), thereby adjusting the state transition function in the system. The output value of the sigmoid function approaches 1 when Bel(A∪B) is large, and approaches 0 when Bel(A∪B) is small. Therefore, the size of the correction factor ψ reflects the size of Bel(A∪B). This mapping preserves the causal relationship, that is, the change of ψ is consistent with the change of Bel(A∪B). Through this correction, the system can introduce a certain uncertainty correction when considering the credibility of historical evidence. This design helps the system to be more adaptive and robust when facing new data and changes.
[0187] 3.6.5 Step S504, modify the auxiliary probability density function p(x)':
[0188] Based on the content of the previous step S200, it can be known that the auxiliary probability density function p(x)' is in the form of:
[0189] ;
[0190] in, are the weights of the mixture components, is the probability density function of the multidimensional Gaussian distribution, is the mean vector, is the covariance matrix and K is the number of components in the mixture model.
[0191] Step S505 needs to correct the auxiliary probability density function p(x)' to obtain a new auxiliary probability density function p(x)'', and the new auxiliary probability density function p(x)'' has undergone adaptive correction:
[0192] ;
[0193] The modification process of p(x)'' is actually a weighted adjustment of the original function. The purpose of this adjustment is to introduce dependence on the trust of the modification factor ψ, so as to adaptively modify the model according to the current evidence (reflected by ψ).
[0194] In the second aspect, a Kriging large model adaptive training system for face modeling: the system includes a processor and a register connected to the processor, the register stores program instructions, and when the program instructions are executed by the processor, the processor executes the Kriging model adaptive training method as described above.
[0195] Compared with the prior art, the present invention has the following beneficial effects:
[0196] 1. Intelligent learning and convergence mechanism: By adopting active learning and Kriging model, the present invention can dynamically select the next training point in the learning process, so as to more intelligently approach the limit state surface. This helps to avoid adding training points in unnecessary areas and improve the efficiency of model learning. Moreover, by dynamically adjusting according to the confidence interval, the technology can enhance the robustness of the system to uncertainty and noise. This means that even when there is a certain error or uncertainty in the estimation, the Kriging model can still provide reliable face model topology results.
[0197] 2. Accurately estimate the failure probability: The technology of the present invention focuses on the accuracy of the failure probability estimation, not just the accuracy of the symbol prediction. The convergence mechanism of the present invention can ensure that the Kriging model reaches the predetermined target accuracy when estimating the failure probability, thereby improving the reliability of the model and achieving high efficiency in topological modeling of the face model.
[0198] 3. Reduce computational cost and time: By intelligently selecting the next training point and introducing a convergence criterion based on the failure probability error, the present invention is expected to reduce unnecessary calculations and the number of calls to functional functions. This helps to reduce computational costs and the time of the learning process, thereby improving overall efficiency. At the same time, since calculations and updates are only performed when necessary, this method can make more efficient use of computing resources.
[0199] 4. Avoid excessive calls to function: Traditional technologies have the problem of calling function too many times, which leads to increased computational costs. The technology of the present invention introduces innovative convergence criteria to terminate the learning process in time, avoid excessive calls to function, improve computational efficiency, and thus improve the efficiency of the Kriging model for face modeling topology.
[0200] 5. Adaptive adjustment of learning strategy: Using the information of confidence interval, the technology of the present invention adopts the method of dynamically adjusting learning parameters so that the learning strategy can be adaptive at different learning stages. This helps to achieve better performance at different stages of model training. By dynamically adjusting learning parameters (such as step size or learning rate) based on the current confidence interval, the system exhibits adaptive characteristics. This means that in different estimation uncertainty scenarios, the system can automatically adjust its behavior to ensure a good balance between estimation accuracy and computational efficiency.
[0201] 6. Adaptive adjustment mechanism based on historical environmental factors: The present invention can further introduce DS evidence theory, consider and apply the different environments encountered by the Kriging model during the training process, and form a conversion adjustment strategy for weight distribution, so that the Kriging model can consider the conflict and uncertainty between historical factors and current factors during the learning process, and make autonomous adjustments and corrections, which helps the system to be more adaptive and robust when facing new data and changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0202] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0203] Figure 1 It is a schematic diagram of the method flow of the present invention;
[0204] Figure 2 Schematic diagram of the basic process of automatic modeling of traditional faces;
[0205] Figure 3 A schematic diagram of the traditional arrangement of training points;
[0206] Figure 4 It is a schematic diagram of the method flow of the connection between step S5 and other sub-steps of the present invention;
[0207] Figure 5 A schematic diagram of a facial topology structure for demonstrating generation of a human face topology model in the present invention;
[0208] Figure 6 A schematic diagram of a topological structure of specific attributes in the face modeling direction obtained in advance by the functional function g(x) of the present invention and the initial training set T;
[0209] Figure 7 A screenshot of a step length of the cyclic update of the topological structure of face modeling in step S3 of the present invention;
[0210] Figure 8 A cutout in one step of the topological structure cyclic update of face modeling in step S3 of the present invention (with Figure 7 The number of faces is different); DETAILED DESCRIPTION
[0211] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below;
[0212] Embodiment 1: Figure 1 As shown, for the existing Kriging model perpendicular to the face modeling direction, it should first perform step S1, which specifically includes sub-steps S100 to S103. That is, generate initial training points:
[0213] Extract N uniformly distributed initial sample points from the random variable space σ. This process requires the execution of the Halton sequence sampling algorithm for sampling; this step is the starting point of the entire modeling process. Its goal is to generate an initial training set to provide a starting point for the subsequent Kriging model construction.
[0214] The Halton sequence sampling algorithm uniformly substitutes the extracted initial sample points into the function g(x), calculates the response values of these initial sample points, and aggregates them to form the initial training set T; the Kriging model is established or updated based on the current initial training set T. The Kriging model estimates the attributes of unknown locations by considering the spatial relationship between sample points. In face modeling, this means that the model will learn the variability of the face's geometric structure and topological information and be able to make predictions. The final topological result of the face is as follows: Figure 5 shown.
[0215] Specifically, regarding the sampling of the random variable space σ: the random variable space is the parameter range that describes the face modeling. By executing the Halton sequence sampling algorithm, a set of initial sample points can be extracted from this space.
[0216] Specifically, regarding the functional function g(x): Substitute the extracted initial sample points into the functional function g(x), which describes the geometric shape and topological features of the face. The calculation of this function depends on the extracted initial sample points. By substituting these points into g(x), the specific properties of each sample point in the direction of face modeling are obtained. By calculating the response value of the functional function, the specific properties of each sample point in the direction of face modeling are obtained. The functional function g(x) consists of two parts: a (constant) regression process and a random process:
[0217] ;
[0218] It is the regression process part of the ALK model, which represents the global average response of the function g(x), β is a constant, and z(x) is a stationary Gaussian random process, which reflects the deviation between the local response and the global response. This part introduces randomness to take into account the local differences of the sample points.
[0219] By substituting the extracted initial sample points into the definition of g(x), the response value of each sample point in the face modeling direction is calculated, which includes the contribution of the global average response and local randomness.
[0220] Furthermore, the functional function g(x) includes:
[0221] 1) Fusion of global and local information: The design of g(x) fully considers the global average response and local differences, enabling the model to capture both the overall shape and local features of the face.
[0222] 2) Introduction of adaptability and randomness: The introduction of the random process z(x) makes the model more adaptable and able to handle the variability of faces in different regions, rather than just the global trend.
[0223] Furthermore, the function g(x) and the initial training set T pre-obtain specific attributes in the direction of face modeling. Figure 6 As an example, only the topological structure at one training point is shown; the C language program executed is:
[0224] #include<stdio.h>
[0225] #include<gsl / gsl_matrix.h>
[0226] #include<gsl / gsl_blas.h>
[0227] #include<gsl / gsl_rng.h>
[0228] / / Kriging model regression coefficient beta
[0229] #define N_FEATURES 10
[0230] double beta[N_FEATURES] = {0.5, 0.3, -0.2, 0.7, -0.4, 0.1, 0.2, -0.5,0.6, -0.3};
[0231] / / Sample data of initial training set T
[0232] #define N_SAMPLES 100
[0233] double training_set[N_SAMPLES][N_FEATURES]; / / Each sample has 10 features
[0234] / / Define the Gaussian random number generator for the random process z(x)
[0235] gsl_rng *rng;
[0236] / / Initialize the Kriging model
[0237] void initialize_kriging_model() {
[0238] / / Initialize the random number generator
[0239] rng = gsl_rng_alloc(gsl_rng_default);
[0240] / / Set the initial training set and other parameters
[0241] }
[0242] / / Calculate the function g(x)
[0243] double calculate_g_function(double *sample) {
[0244] / / f(x)^T·beta
[0245] double regression_part = 0.0;
[0246] gsl_vector_view beta_view = gsl_vector_view_array(beta, N_FEATURES);
[0247] gsl_vector_view sample_view = gsl_vector_view_array(sample, N_FEATURES);
[0248] / / Calculation of f(x)^T·beta
[0249] gsl_blas_ddot(&sample_view.vector, &beta_view.vector, ®resion_part);
[0250] / / z(x) is a normally distributed random number
[0251] double random_part = gsl_ran_gaussian(rng, 0.1); / / standard deviation is 0.1
[0252] / / Calculate the function g(x)
[0253] double g_x = regression_part + random_part;
[0254] return g_x;
[0255] }
[0256] / / Main function
[0257] int main() {
[0258] / / Initialize the Kriging model
[0259] initialize_kriging_model();
[0260] / / Traverse the initial training set T and calculate the value of the function g(x)
[0261] for (int i = 0; i < N_SAMPLES; ++i) {
[0262] double g_value = calculate_g_function(training_set[i]);
[0263] printf("Sample %d: g(x) = %f\n", i + 1, g_value);
[0264] }
[0265] / / Release the random number generator to generate a face topology model
[0266] gsl_rng_free(rng);
[0267] return 0;
[0268] }
[0269] The above program mainly involves two core functions, namely initialize_kriging_model and calculate_g_function.
[0270] 1) initialize_kriging_model: This function is used to initialize the Kriging model, including setting the random number generator and other necessary parameters. gsl_rng_alloc(gsl_rng_default) is used to initialize the random number generator provided by GSL.
[0271] 2) calculate_g_function: This function is used to calculate the value of the function g(x), where g(x) consists of a regression process and a random process. The regression process uses a linear model, where f(x) is the sample feature vector and β is the regression coefficient. The random process simulates a stationary Gaussian random process, assuming a normally distributed random number. Finally, the value of g(x) is calculated and returned.
[0272] Specifically, regarding the formation of the initial training set T: the calculated sample points and their response values are used to form the initial training set T. This set will contain initial information about the direction of face modeling.
[0273] In this embodiment, regarding step S100, the Halton sequence sampling algorithm: the Halton sequence is a low-discrepancy sequence used for uniform sampling of high-dimensional space. For this technical solution, one-dimensional space and multi-dimensional space need to be considered:
[0274] 1) In one dimension, the extracted value of the i-th element of the Halton sequence for:
[0275] ;
[0276] is the cardinality of uniform distribution, k is the maximum value of the total element i, is an integer uniformly distributed in the range [0, b-1]
[0277] 2) Similarly, in multidimensional space, in order to ensure better uniformity, each dimension uses a different prime number as the base. Taking two-dimensional space as an example, the extracted value of the i-th element of the Halton sequence is for:
[0278] ;
[0279] Indicates that each dimension has a different cardinality ,and is Integers uniformly distributed in the range.
[0280] In this embodiment, regarding step S101, uniformly distributed initial sample points are obtained: this step executes the process of obtaining uniformly distributed initial sample points, using the elements of the Halton sequence as the components of the random variable, ensuring uniform distribution of the sample points in each dimension:
[0281] Specifically, each element of the Halton sequence is taken as a component of the random variable. If there is a k-dimensional space, then the i-th element of the Halton sequence It can be expressed as:
[0282] ;
[0283] N is the total number of dimensions of the k-dimensional space; the above elements are of the form represents the value of the i-th element in the j-th dimension. This value is obtained by sampling the Halton sequence using different prime numbers as the base. In order to make the Halton sequence evenly distributed in each dimension, different prime numbers can be selected as the base, for example .
[0284] It can be understood that by sampling the Halton sequence, uniformly distributed sample points can be obtained to ensure uniformity in each dimension. In order to make the Halton sequence uniformly distributed in each dimension, different prime numbers are selected as the bases. Different prime numbers are used for different dimensions to ensure the differences in the distribution of sample points in each dimension, thereby improving the comprehensiveness of the sample points. Through this process, the uniform distribution of the initial sample points is ensured, which provides representative input data for the establishment of the subsequent Kriging model, and helps to fit the face modeling more comprehensively and accurately. This also provides initial information for the direction of face topology modeling.
[0285] Specifically, the i-th sample point It is expressed as:
[0286] ;
[0287] Then, these sample points Substituting into the utility function g(x) described above, where x is represented by an n-dimensional vector (where the constant vector β should match the dimension of f(x)):
[0288] ;
[0289] This expression can be used to calculate each sample point Corresponding function , and obtain the specific properties of the functional function in the direction of face modeling. This also provides a data basis for the subsequent Kriging model establishment.
[0290] It can be understood that the above expression is only used to substitute the constant form of the Kriging model in the existing ALK model; if the ALK model adopts other forms of the Kriging model, it can be adjusted accordingly by using a similar substitution form as above; for example, using the following C language framework:
[0291] #include<stdio.h>
[0292] #include<stdlib.h>
[0293] / / A library that comes with a Kriging model, which includes initialization, training, and prediction functions
[0294] #include "kriging_library.h"
[0295] / / Define the length of the constant vector
[0296] #define DIMENSION 3
[0297] int main() {
[0298] / / Declare and initialize the Kriging model
[0299] kriging_model model;
[0300] initialize_kriging_model(&model, DIMENSION);
[0301] / / Generate initial training points, here use Halton sequence
[0302] int num_samples = 10;
[0303] double** initial_samples = generate_halton_sequence(num_samples,DIMENSION);
[0304] / / Bring in sample points
[0305] printf("Initial Samples:\n");
[0306] for (int i = 0; i < num_samples; ++i) {
[0307] printf("Sample %d: [", i + 1);
[0308] for (int j = 0; j < DIMENSION; ++j) {
[0309] printf("%lf", initial_samples[i][j]);
[0310] if (j < DIMENSION - 1) {
[0311] printf(", ");
[0312] }
[0313] }
[0314] printf("]\n");
[0315] }
[0316] / / Substitute the Kriging model for training
[0317] for (int i = 0; i < num_samples; ++i) {
[0318] double response = calculate_function_response(initial_samples[i]);
[0319] train_kriging_model(&model, initial_samples[i], response);
[0320] }
[0321] / / New test point
[0322] double test_point[DIMENSION] = {0.5, 0.5, 0.5};
[0323] / / Use Kriging model for prediction
[0324] double predicted_response = predict_kriging_model(&model, test_point);
[0325] / / Execute prediction results and topology
[0326] printf("Predicted Response at Test Point: %lf\n", predicted_response);
[0327] / / Release memory
[0328] for (int i = 0; i < num_samples; ++i) {
[0329] free(initial_samples[i]);
[0330] }
[0331] free(initial_samples);
[0332] return 0;
[0333] }
[0334] The principle of the above program is to store any type of kriging model in the library "kriging_library.h" and use the Halton sequence to generate initial training points. The Halton sequence is a low-discrepancy sequence used to generate uniformly distributed sample points in multidimensional space. These sample points are stored in initial_samples. Then the initial training points are traversed, and the calculate_function_response function is called for each point to obtain the response value of the function function, and then the train_kriging_model function is called to use the sample points and their response values to train the Kriging model. This step is the process of building a Kriging model. Then a new test point test_point is generated, and then the predict_kriging_model function is called to predict the response value of the point using the trained Kriging model. At the end of the program, the memory allocated for the sample points and the Kriging model is released.
[0335] In this embodiment, regarding step S102, an initial training set T is formed:
[0336] The obtained sample points and the corresponding functional response values are combined to form the initial training set T:
[0337] ;
[0338] is the sample point corresponding to The response value output by the function g(x).
[0339] Specifically, for each sample point , whose response value has been calculated by the function g(x) This is calculated by calling the functional function, which describes the geometric shape and topological features of the face. The initial training set T now contains sample points and their corresponding functional function response values, which will be used to train the Kriging model. The goal of this model is to capture the spatial relationship between sample points in order to interpolate or predict unknown points.
[0340] Specifically, each sample point Generated by Halton sequence, the value of each dimension represents a position in the face space. Therefore, the set of sample points T reflects the positions evenly distributed in the face topology. Response value Contains the geometric shape and topological features of the face at the corresponding position. Therefore, each element in T The description of the face topology is saved. T is then used as training data to build the Kriging model. The Kriging model learns the structural information of the face topology by interpolating the spatial relationship between sample points. The purpose of model training is to infer the features of unknown positions in the face space based on known sample points. Ultimately, the Kriging model can predict unknown points in the face space. In this way, the face topology information saved by the initial training set T can be revealed in the entire face area, thus forming a comprehensive face model.
[0341] In this embodiment, regarding step S103, a standard constant Kriging model is constructed as an example:
[0342] First of all, it needs to be made clear that the Kriging model is a modeling method based on spatial interpolation, which aims to estimate the values of unknown locations and provide quantification of the uncertainty of the estimate. The basic idea of the Kriging model is to estimate the function values of unknown points by linearly combining the function values of known sample points, where the estimated coefficients need to meet certain constraints.
[0343] Therefore, the N known sample points in the training set T can be , and the value The linear combination of The unknown quantity at , so that the variance of the estimation error is minimized. Therefore, let Z(x) be a second-order stationary random function, with position The estimated value of the point is:
[0344] ;
[0345] is the estimated coefficient;
[0346] The estimated error is recorded as: ,in is the true value of the estimated point. Since these true values are unknown, the two are regarded as random variables, and the mean square error corresponding to the random variable r is for:
[0347] ;
[0348] E(r) represents the mathematical expectation (i.e. mean) of the random variable r;
[0349] In order to minimize the mean square error, the estimated coefficients The constraints also need to be met:
[0350] ;
[0351] In order to minimize the variance of the estimation error, a constrained optimization problem is established by introducing the Lagrange multiplier method. The constraints include: based on the Lagrange multiplier method, when the variance of the estimation error is minimized, input N sample points ,in is the i-th sample point (same for the response value), is the jth sample point (same for the response value), is the initial sample point (same for the response value);
[0352] That is, the sum of the estimated coefficients is 1, and the correlation between sample points satisfies the correlation function:
[0353]
[0354] Based on the constrained optimization problem, we obtain a and Lagrange multipliers. The equations include the requirements for the correlation between the estimated coefficients and the sample points. Finally, the equations for the Kriging model are obtained:
[0355] ;
[0356] µ is the Lagrange multiplier, Represents the correlation between two positions; of course, the equation group can also be converted into a matrix form, and its form is not strictly limited.
[0357] Further, we impose Gaussian distribution constraints on the equations to ensure that the Kriging model satisfies the properties of the Gaussian process. This step introduces the Gaussian kernel function to express the correlation between sample points. Its form is:
[0358] ;
[0359] Where μ is the mean and exp is the exponential function. If it is an exponential function of a one-dimensional Gaussian distribution, then , where σ is the standard deviation of the Gaussian distribution.
[0360] The properties of the Kriging model can be summarized as follows: For identification, the topological surface passes through the sample point , rather than approximating the sample points When the Kriging model estimates a sample point When , the weight of the known point is 1, and the others are 0. That is, the Kriging model estimates each sample point not by simply approximating the point, but by passing through a surface at the point.
[0361] If the data is completely random, that is, unrelated, then for a sample point The estimated value is the sum of all points, which is the concept of "NuggetEffect" of Kriging model. This is equivalent to all weights being 1 / n and the correlation being 0. When there is no such effect, the closer the sample point is to the estimated point, the smaller the sample point will be. The larger the weight of . This means that the model relies entirely on known sample points when estimating, and has zero influence on other points.
[0362] When there is no Nugget Effect, the closer the sample point is to the estimated point, the The larger the weight of . This reflects the sensitivity of the Kriging model to the distance of sample points. In the absence of "NuggetEffect", the model emphasizes the contribution of neighboring points to the estimation.
[0363] Specifically, the Kriging model C language form of the above step S103 is:
[0364] #include<stdio.h>
[0365] #include<math.h>
[0366] / / The structure represents the sample points
[0367] typedef struct {
[0368] double x; / / x coordinate of the sample point
[0369] double response; / / Functional response value of the sample point
[0370] } SamplePoint;
[0371] / / Training set T
[0372] SamplePoint trainingSet[] = {
[0373] {1.0, 10.0},
[0374] {2.0, 20.0},
[0375] / / Other sample points
[0376] };
[0377] / / Kriging model parameters
[0378] double lambda
[100] ; / / estimated coefficients
[0379] double mu; / / Lagrange multiplier
[0380] double nuggetEffect; / / Nugget Effect
[0381] / / Kriging model calculation function
[0382] double krigingModel(double x0) {
[0383] double estimate = 0.0;
[0384] / / Calculate the estimated value
[0385] for (int i = 0; i < sizeof(trainingSet) / sizeof(trainingSet[0]);i++) {
[0386] double distance = fabs(trainingSet[i].x - x0);
[0387] double correlation = exp(-mu * distance * distance);
[0388] estimate += lambda[i] * correlation;
[0389] }
[0390] / / Add Nugget Effect
[0391] estimate += nuggetEffect;
[0392] return estimate;
[0393] }
[0394] int main() {
[0395] / / Use Kriging model for face topology modeling
[0396] double x0 = 3.0; / / point position
[0397] double faceTopology = krigingModel(x0);
[0398] / / Output 3D model
[0399] printf("Estimated face topology at x=%.2f: %.2f\n", x0,faceTopology);
[0400] return 0;
[0401] }
[0402] The principle of the above program is: use the structure SamplePoint to represent the sample point, which contains the x-coordinate of the sample point and the response value of the function. Then declare a training set trainingSet, which contains a series of known sample points (that is, T mentioned above). Then declare the parameters of the Kriging model, including the estimated coefficient lambda, the Lagrange multiplier mu, and the NuggetEffect. Then the krigingModel function uses the Kriging model to calculate the corresponding face topology estimate based on the input position x0. This function calculates the weight and correlation for each sample point in the training set and finally obtains the estimated value. In the main function, set the parameters of the Kriging model and estimate the face topology at the specified position by calling the krigingModel function.
[0403] It should be noted that the above-mentioned Kriging model is only an example of introducing an existing constant form of Kriging model into the training set T. It can also be conventionally transformed through several different forms of Kriging models.
[0404] In this embodiment, regarding step S2, a limit state function proxy model is constructed to generate a candidate sample pool S; it specifically includes sub-steps S200 to S202:
[0405] Execute the Gaussian mixture algorithm, which is a method for modeling complex probability distributions. It uses the weighted sum of multiple Gaussian distributions to approximate the true probability distribution. Through this algorithm, the auxiliary probability density function p(x)' of the Kriging model can be obtained, which describes the probability distribution of sample points in the random variable space. And use the failure domain set G combined with the auxiliary probability density function p(x)' to generate n random variable samples , thus forming a candidate sample pool S. The candidate sample pool S will be used to select the training points for the next update, and each step will be regenerated with the update of the Kriging model. In this way, the candidate sample pool S will be updated in each iteration to better guide the selection of the next training point.
[0406] Specifically, the failure domain set G is obtained in the early stage, which contains a series of sample points related to failure. In this step, N random variable samples are generated by combining the failure domain set G with the auxiliary probability density function p(x)' These sample points represent possible locations in the failure domain.
[0407] In this embodiment, regarding step S200, a Gaussian mixture algorithm is performed on the auxiliary probability density function p(x)':
[0408] Gaussian mixture model (GMM) is a probability model algorithm, which is usually used for data clustering and density estimation. It assumes that the observed data is composed of several potential Gaussian distributions, each of which is called a mixture component. The core idea of GMM is to regard the observed data as the weighted sum of each Gaussian distribution. In this step, the auxiliary probability density function p(x)' is in the form of:
[0409] ;
[0410] in, are the weights of the mixture components, is the probability density function of the multidimensional Gaussian distribution, is the mean vector, is the covariance matrix and K is the number of components in the mixture model.
[0411] Specifically, the execution process of the Gaussian mixture algorithm is:
[0412] P1. Initialization: Randomly initialize the mean of each Gaussian distribution , covariance matrix and weight .
[0413] P2, GMM's Expectation-Maximization (EM) iteration: The model parameters are estimated through the iterative EM algorithm. In the E step, the posterior probability of each sample belonging to each component is calculated, and then in the M step, the model parameters are updated according to these posterior probabilities.
[0414] P3, Convergence criterion: Repeat the EM steps until the model parameters converge or the predetermined number of iterations is reached.
[0415] P4. Application of auxiliary probability density function p(x)': Get the weight of each sample point in the auxiliary probability density function, which represents the contribution of the sample point in different components. Such weight information can be used in subsequent steps, such as generating a candidate sample pool S.
[0416] Specifically, is the weight of the mixture component. In the Gaussian mixture model (GMM), it is the relative importance of each mixture component. Its assignment needs to meet the following three conditions:
[0417] 1) Non-negativity: All weights must be non-negative, that is: ;
[0418] 2) The sum is one: ;
[0419] 3) Assuming there are K mixed components, the weights under uniform distribution can be assigned according to the following formula: ;
[0420] This ensures that the relative importance of each mixture component is the same. Of course, it is understandable that if there is prior information indicating that some components should be more important, those skilled in the art can make adjustments according to their subjective initiative based on the characteristics of the problem.
[0421] Specifically, is the probability density function of the multidimensional Gaussian distribution, which is in the form of:
[0422] ;
[0423] x is a D-dimensional column vector (D is the dimension), μ is a D-dimensional column vector, μ i represents the mean vector, is the covariance matrix of D*D, is the determinant of the covariance matrix. This formula describes the multidimensional Gaussian distribution under a given mean μ and covariance matrix Σ. For each input vector The probability density of .
[0424] Specifically, is the mean vector, which is the mean of the i-th component in the Gaussian mixture model, expressed as a D-dimensional column vector. The specific assignment is obtained through training data, that is, the mean of the sample points corresponding to the i-th component. Taking the training set T mentioned above as an example, the mean of the i-th component Defined as:
[0425] ;
[0426] is the number of sample points assigned to the i-th component. The calculation of this mean vector is the average value of the coordinate components of all sample points on the corresponding component.
[0427] Specifically, the choice of K is a hyperparameter, and the optimal K is selected through some conventional evaluation indicators (such as Akaike Information Criterion (AIC) or Bayesian Information Criterion (BIC)). These criteria take into account the degree of fit and complexity of the model, helping to find a balance and avoid overfitting.
[0428] In this embodiment, after executing step S200, it is also necessary to obtain the failure domain set G:
[0429] Assume that the failure probability of the Kriging model for the input variable x is estimated as , the confidence interval is [L(x), U(x)], where L(x) is the lower bound and U(x) is the upper bound. At the same time, a threshold T1 is set to represent the threshold of the failure probability. The acquisition mechanism of the failure domain is:
[0430] ;
[0431] That is, if If it exceeds the threshold T1, a value of 1 is assigned, x is determined to belong to the failure domain, and is added to the failure domain set G; otherwise, it is not considered.
[0432] Specifically, the threshold T1 is assigned using cross validation and ROC curve. The steps include:
[0433] P1. Calculate the True Positive Rate (TPR) and False Positive Rate (FPR) according to different thresholds and draw the ROC curve:
[0434] ;
[0435] ;
[0436] True Positives (TP): The number of positive examples correctly predicted by the model;
[0437] False Negatives (FN): The number of positive examples that the model incorrectly predicted as negative examples;
[0438] False Positives (FP): The number of negative examples that the model incorrectly predicts as positive examples;
[0439] True Negatives (TN): The number of negative examples correctly predicted by the model.
[0440] These values can be calculated by comparing the model's predictions to the actual labels. Use a confusion matrix calculation tool in Python, such as the confusion_matrix function in scikit-learn, to get these values:
[0441] from sklearn.metrics import confusion_matrix
[0442] # y_true is the actual label, y_pred is the model prediction result
[0443] y_true = [1, 0, 1, 1, 0, 1, 0, 0, 1]
[0444] y_pred = [1, 0, 1, 0, 0, 1, 1, 0, 1]
[0445] # Calculate the confusion matrix
[0446] conf_matrix = confusion_matrix(y_true, y_pred)
[0447] # Get TP, FN, FP, TN from the confusion matrix
[0448] TP = conf_matrix[1, 1]
[0449] FN = conf_matrix[1, 0]
[0450] FP = conf_matrix[0, 1]
[0451] TN = conf_matrix[0, 0]
[0452] P2. Use the confusion matrix calculation tool in Python to plot the working point on the ROC curve so that the true positive rate is maximized and the false positive rate is minimized as the threshold T1:
[0453] from sklearn.metrics import roc_curve, auc
[0454] # Calculate ROC curve
[0455] fpr, tpr, thresholds = roc_curve(y_true, y_prob)
[0456] # Calculate the area under the curve (AUC)
[0457] roc_auc = auc(fpr, tpr)
[0458] # Choose the best threshold
[0459] best_threshold_index = np.argmax(tpr - fpr)
[0460] best_threshold = thresholds[best_threshold_index]
[0461] # best_threshold selected threshold T1
[0462] Specifically, the failure domain set G is a set of all sample points that are judged to be failed. Each time, the prediction results and confidence intervals of the current Kriging model are used to make a judgment, and the sample points that meet the conditions are added to the failure domain set G. It includes several failure domain samples. :
[0463] ;
[0464] n is the failure domain sample The total number of .
[0465] Through such a mechanism, each time step S2 is executed, the failure domain samples are dynamically determined according to the prediction results and confidence intervals of the current Kriging model. The failure domain samples obtained in this way can be used for subsequent GMM or other method modeling.
[0466] Failure probability Modeling and analysis of probability distributions around the performance function g(x), because the performance function g(x) defines the relationship between system performance and design requirements.
[0467] Specifically, regarding the failure probability The way to obtain can be:
[0468] ;
[0469] Failure probability It is the probability that the corresponding value output by the function g(x) is less than or equal to zero. The obtained failure domain samples can be used for subsequent GMM (Gaussian mixture model) or other modeling methods. This provides sample data on the failure state of the system for the subsequent modeling process, which helps to more accurately capture the relationship between system performance and design requirements.
[0470] In this embodiment, regarding step S201, the failure domain set G is combined with the auxiliary probability density function p(x)' to generate n random variable samples. :
[0471] First, according to the content of step S200, the parameters of the Gaussian mixture model are used to generate samples. That is, a component is randomly selected from the mixture model according to the weight of each component, and then a sample is generated from the multidimensional Gaussian distribution corresponding to the failure domain set G. Each generated sample Both represent a sample sampled from a mixed model. This process can be expressed as:
[0472] ;
[0473] The symbol “~” represents the statistical concept of “…distributed in…”.
[0474] Specifically, considering each failure domain sample point in the random variable sample , randomly select a sample from the failure domain set G, expressed as , and according to the auxiliary probability density function Generate random variables, that is, generate random variables related to the failure domain according to the auxiliary probability density function of the failure domain sample:
[0475] ;
[0476] It can be understood that the core of this step is to combine the samples of the Gaussian mixture model with the samples of the failure domain set, and to update the auxiliary probability density function of the current Kriging model by generating random variables related to the failure domain. Such an update strategy helps to more accurately capture the probability distribution of system failures, so that training points can be selected more specifically in the next round of updating the Kriging model.
[0477] In this embodiment, regarding step S202, a candidate sample pool S is generated:
[0478] Repeat step S201 to generate n random variable samples , forming a candidate sample pool S:
[0479] ;
[0480] In this process, the Gaussian mixture model parameters are used to simulate the uncertainty of system behavior, and the failure domain set G provides sample points under failure conditions. By repeatedly executing S201, N new random variable samples are obtained to form a candidate sample pool S. Each sample in the candidate sample pool S They are all generated by comprehensively considering system uncertainty and failure domain information. This process can be executed using the following python program:
[0481] # Parameter settings
[0482] N = 100 # The number of random variable samples generated
[0483] parameters = get_parameters() # Get the parameters of the hybrid model and failure domain set
[0484] # Loop to generate samples
[0485] candidate_samples = []
[0486] for _ in range(N):
[0487] # Randomly select a failure domain sample
[0488] x_fail = choose_random_fail_sample(parameters['G'])
[0489] # Generate random variables based on auxiliary probability density function
[0490] y_sample = generate_random_variable(x_fail, parameters['p(x)\''])
[0491] # Add to candidate sample pool
[0492] candidate_samples.append(y_sample)
[0493] # Candidate sample pool S
[0494] S = candidate_samples
[0495] In the above program, get_parameters() and choose_random_fail_sample() are functions used to obtain parameters and select failure domain samples, while generate_random_variable() is a function used to generate random variables based on the auxiliary probability density function. The goal of this program is to generate N qualified random variable samples to form a candidate sample pool S.
[0496] It is understandable that N new random variable samples are generated through multiple simulations. In this way, the samples in the candidate sample pool S will better reflect the uncertainty and failure of the current system. The generation of candidate samples is to select the most informative samples as the training points for the next update in the next step to optimize the performance of the Kriging model. The entire process is continuously iterated to gradually improve the prediction accuracy of the Kriging model and the understanding of system behavior.
[0497] In this embodiment, step S3, identifying the training point to be updated next time in the candidate sample pool S, includes two sub-steps S300-S301:
[0498] 1) Step S300, calculate the function value of each sample in the candidate sample pool S :
[0499] First, draw each random variable sample from the candidate sample pool S As input variables; for this technical solution, the priority is the difference between system performance and design requirements, so a limit state function G (y) is introduced; its input is the vector of variables.
[0500] ;
[0501] , is the same as the i-th input variable Related regulatory factors.
[0502] The limit state function is used to represent the difference between system performance and design requirements. A negative value indicates that the system is within the design requirements, and a zero value indicates that the system is on the boundary of the design requirements.
[0503] 2) Step S301, select the minimum sample point in the function value s as the training point for the next update:
[0504] The function value of each sample output by the limit state function G(y) based on S300 , aggregate it and select the smallest value s∗, and then find the corresponding failure domain sample point .
[0505] This selection strategy ensures that the difference between system performance and design requirements is taken into account when selecting the training point for the next update. By calculating the function value of each sample in the candidate sample pool and then selecting the sample point with the minimum function value, this strategy guides the update of the model in active learning, making the model pay more attention to the area near the failure domain. This method helps the model fit the structure of the failure domain more accurately and improves the prediction performance of the model.
[0506] For details, please refer to Figure 5~Figure 8 , the loop form of step S3 is:
[0507] P1. Generate n random variable samples using the failure domain set G: Generate samples based on the parameters of the Gaussian mixture model. Randomly select failure domain sample points from the failure domain set G to generate random variable samples.
[0508] P2. Generate a candidate sample pool S: Repeat step S201 to generate n random variable samples to form a candidate sample pool S.
[0509] P3. Calculate the function value S of each sample in the candidate sample pool S: Use the limit state function G(y) to calculate the function value of each sample.
[0510] P4. Select the minimum sample point s∗ in the function value S as the training point for the next update: Select the sample point with the minimum function value and find the corresponding failure domain sample point .
[0511] In this process, the topological structure of the face (taking a training point as an example) changes from Figure 6 gradually became Figure 5 The change in this process randomly selects a step size in the form of Figure 7 If the size of the topology is taken into consideration, it can also be adjusted as follows Figure 8 The form shown.
[0512] Specifically, , is the same as the i-th input variable The values of the two are standardized or normalized so that different input variables have similar ranges in magnitude:
[0513] ;
[0514] is the input variable The mean of is the input variable The standard deviation of is the normalized input.
[0515] For the adjustment factor and The value of can be 0 or 1, so that the input variable Normalized to a range with a mean of 0 and a standard deviation of 1.
[0516] In this embodiment, regarding step S4, the convergence of the learning process is determined based on the minimum function value s* and its corresponding failure domain sample point. , and the mean prediction provided by the current ALK model , calculate the upper bound of the failure probability error evaluated by the current Kriging model. If the upper bound is less than the target failure error threshold, stop the learning process; otherwise, calculate the upper bound confidence interval and return to step S2 to update the Kriging model. Step S4 specifically includes sub-steps S400 to S402;
[0517] In this embodiment, regarding step S400, the upper bound of the failure probability error is calculated:
[0518] Get the predicted mean , represents the predicted probability of system failure given the input variable x. This is an important output in the Kriging model. Calculate the failure probability predicted by the model This failure probability represents the probability of system failure calculated by the predicted mean of the Kriging model; the upper bound of the error can be formulated based on the uncertainty of statistical inference and model prediction:
[0519] ;
[0520] in, is to use the failure domain sample points The calculated failure probability, is the failure probability calculated using the ALK model predicted mean.
[0521] Specifically, It represents the error of the predicted mean using the Kriging model relative to the failure probability calculated by the actual failure domain sample point. The calculation of this upper bound of the error is based on the difference between the two calculation methods, which measures the accuracy of the model's prediction relative to the actual situation. In layman's terms, it reflects the following information:
[0522] 1) If It is smaller, indicating that the predicted mean of the Kriging model is more accurate and has a smaller difference with the actual failure probability.
[0523] 2) If A larger value means that the prediction performance of the Kriging model is poor in some areas, and model updates or other adjustments need to be considered.
[0524] Specifically, the predicted mean You can execute the following MATLAB program to obtain:
[0525] % Sample point X and corresponding response value Y
[0526] X = [x1; x2; ...; xn]; % Sample points of input variables
[0527] Y = [y1; y2; ...; yn]; % corresponding response value
[0528] % Kriging model
[0529] theta = [1, 1]; % Hyperparameters of the Kriging model
[0530] model = fitrgp(X, Y, 'KernelFunction', 'squaredexponential', 'KernelParameters', theta);
[0531] % New input variable x_new, used for prediction
[0532] x_new = [new_x1; new_x2; ...; new_xn]; % New input variables
[0533] % Use Kriging model for prediction
[0534] [y_pred, y_pred_sd] = predict(model, x_new);
[0535] % Output predicted mean
[0536] disp('Predicted mean:');
[0537] disp(y_pred);
[0538] The above program uses Matlab's fitrgp function to build a GaussianProcessRegression model, in which the squaredexponential kernel function is selected. Then, the predict function is used to predict the new input variable and obtain the predicted mean y_pred and standard deviation y_pred_sd.
[0539] Specifically, is the failure probability calculated using the ALK model prediction mean. For the Gaussian process model (a form of the Kriging model), the failure probability can be calculated using the prediction mean and standard deviation. The predicted mean is mu and the standard deviation is sigma, then the failure probability can be calculated by the Cumulative Distribution Function (CDF) of the standard normal distribution:
[0540] ;
[0541] Where Φ represents the cumulative distribution function of the standard normal distribution, is the ALK model predicted mean, and σ is the standard deviation of the predicted mean.
[0542] Furthermore, the cumulative distribution function Φ of the standard normal distribution is:
[0543] ;
[0544] This formula describes the cumulative probability that a random variable in a standard normal distribution is less than or equal to z. Where erf is the error function:
[0545] ;
[0546] e is the base of natural logarithm, which is approximately equal to 2.71828; dt represents the integral variable, that is, the small change in the integral process. t represents the integral variable, that is, the independent variable in the integral process. In Gaussian integral, t represents the upper limit of the integral. erf(x) represents the Gaussian integral from 0 to x.
[0547] In this embodiment, regarding step S401, it is determined whether the stop learning condition is met:
[0548] ;
[0549] T2 is the target failure error threshold;
[0550] If the upper bound of the failure probability error If the error is less than the target failure error threshold T2, the learning stops. Otherwise, the learning continues. This stop learning condition is set to ensure that the learning process can automatically stop after reaching a certain level of accuracy, avoiding unnecessary calculations and improving calculation efficiency.
[0551] Specifically, regarding the target failure error threshold T2, this embodiment provides some specific assignment schemes:
[0552] Scheme 1: Conservative method, T2=0.05;
[0553] In some highly critical and risk-sensitive applications, a conservative approach is desirable to ensure that the error in the probability of failure is very small. Therefore, a relatively small T2 is set to ensure that the system reliability is within the design requirements.
[0554] Solution 2: Balanced method, T2=0.1
[0555] In some applications, a balance needs to be found between system performance and cost. Selecting a moderate T2 can better consider cost-effectiveness while ensuring system reliability.
[0556] Solution 3: Adaptive solution: as shown in the full text of Example 4.
[0557] In this embodiment, regarding step S402, the upper confidence interval is calculated: if the Kriging model has not stopped learning, the upper confidence interval of the failure probability error can be calculated to perform the next round of Kriging model update. This process involves further normal distribution algorithm and model update. It includes sub-steps S4020~S4023.
[0558] In this embodiment, regarding step S4020, the standard deviation of the failure probability error is calculated. :
[0559] ;
[0560] is the ith failure probability error, is the average value of the total failure probability error; n is the total number of samples;
[0561] Specifically, the main purpose of step S4020 is to calculate the standard deviation of the failure probability error , the significance of this step is to evaluate the dispersion of the failure probability error. Specifically, the standard deviation is a statistical indicator that measures the degree of deviation of each data in a data set from the mean value. The standard deviation reflects the dispersion of the failure probability error relative to its mean value. A larger standard deviation means that the error value is more dispersed relative to the mean value, indicating that the difference between the system performance and the design requirements is larger. The smaller the standard deviation, the more stable the distribution of the failure probability error, and the more credible the prediction of the model. Conversely, a larger standard deviation may indicate that the prediction of the model is subject to greater uncertainty. At the same time, the information of the standard deviation can be used to optimize the Kriging model, especially when selecting the next round of training points for model update, more attention can be paid to areas with larger failure probability errors to improve the prediction accuracy of the model.
[0562] Therefore, step S4020 helps to gain a deeper understanding of the distribution characteristics of the failure probability error and provides valuable information for subsequent model updating and optimization.
[0563] Specifically, as shown in step S400, the i-th failure probability error The way to obtain is:
[0564] ;
[0565] Therefore, the total failure probability error average The calculation method is:
[0566] ;
[0567] n is the total number of samples, is the sum of the failure probability errors. By calculating each failure probability error and taking its average, we get the average value of the total failure probability error. .
[0568] In this embodiment, regarding step S4021, the confidence interval radius Ma is calculated:
[0569] ;
[0570] is the standard deviation; Z is the “Z-score” in the standard normal distribution, corresponding to the chosen confidence level α (taken as 95%);
[0571] Ma can be regarded as a marginal value, which refers to the radius of the confidence interval, indicating the distance between the upper and lower limits of the confidence interval and the mean. It represents the uncertainty range of the estimated failure probability error within the confidence level. If Ma is larger, it means that the confidence interval is wider and the uncertainty of the estimated failure probability error is greater. Conversely, if Ma is smaller, it means that the uncertainty of the estimated failure probability error is smaller.
[0572] In this embodiment, regarding step S4022, a new confidence interval CI is calculated:
[0573] ;
[0574] Once a new confidence interval CI of the failure probability error is obtained, a decision can be made based on this interval to determine whether to stop learning or continue updating the Kriging model.
[0575] In this embodiment, regarding step S4023, the learning strategy scheme:
[0576] According to the content of "Failure Domain Set G" above, the confidence interval is [L(x), U(x)], where L(x) is the lower bound and U(x) is the upper bound. The above confidence interval is adjusted according to the new confidence interval CI obtained in 3.4.3.1. The strategy selected by this technical solution is: if the confidence interval is wide, it means that the estimate of the failure probability is less certain, and a more conservative learning strategy can be adopted, that is, increasing the number of samples or adjusting the learning step size. If the confidence interval is narrow, it means that the estimate of the failure probability is more reliable, and a more aggressive learning strategy can be adopted, that is, reducing the number of samples or accelerating learning; and the expression of the above strategy is:
[0577] ;
[0578] γ is the original learning step size or learning rate, and γ' is the updated learning step size or learning rate. f(CI) is an adjustment function for the confidence interval CI:
[0579] ;
[0580] w(CI) is the width of the confidence interval, where w(CI)=U(x)−L(x). v is an adjustment parameter used to control the amplitude of the adjustment. The amplitude is subjective, so technicians in this field need to exercise their subjective initiative to make adjustments when actually implementing the adjustment.
[0581] This scheme makes w(CI) larger and f(CI) smaller when the confidence interval width is larger, so γ' will decrease accordingly, making learning more conservative. When the confidence interval width is smaller, f(CI) is larger, and γ' will increase accordingly, making model learning more aggressive.
[0582] In summary, this embodiment adopts an active learning strategy to dynamically select training samples according to the prediction results and confidence intervals of the current model, so that the model can adapt to changes in different regions more flexibly. Such a strategy may make the model pay more attention to the difference between system performance and design requirements during training, and improve the modeling accuracy of the model near the failure domain. Although the traditional iCE method can achieve similar technical effects, this embodiment does not recommend the use of the traditional iCE method. Because the iCE method is based on an integrated contraction and expansion method, it achieves global optimization by gradually improving the model. In contrast, the ALK model adopts an active learning strategy to dynamically select training samples according to the prediction results and confidence intervals of the current model, emphasizing the modeling of failure probability. Therefore, the technical solution of this embodiment and the iCE method have essential differences in learning strategies and cannot be replaced. At the same time, the iCE method focuses more on the correction of global optimization problems, while the ALK model focuses more on the modeling of failure probability, which is suitable for scenarios where system reliability and performance need to be evaluated. Therefore, the implementation of the iCE method in face modeling cannot be implemented.
[0583] Moreover, for this embodiment, while paying attention to the failure domain, it emphasizes the modeling of the failure probability, which is suitable for scenarios where the reliability and performance of the system need to be evaluated. Compared with the global optimization problem, this technology pays more attention to the behavior of the system within the failure domain, and the wiring in the face topology has a strong directional strategy. At the same time, this technology seeks to reduce the number of calls to functional functions while ensuring the accuracy of the failure probability. By dynamically adjusting the learning step size, adjusting the learning strategy according to the confidence interval, etc., while improving the efficiency of the model, the failure probability can still be accurately modeled. The width of the confidence interval is used as an adjustment parameter in the learning strategy. By adjusting the width of the confidence interval, the learning step size can be flexibly controlled, making the model more flexible and adaptable to the estimation of the failure probability.
[0584] Embodiment 2: This embodiment is based on the content of embodiment 1. In obtaining the failure domain set G, a Monte Carlo simulation algorithm is used to obtain :
[0585] P1. Using the Monte Carlo method, the integral expression on the failure domain is defined as :
[0586] ;
[0587] G1 is the sample value within the failure domain, and G2 is the total sample value.
[0588] P2. Use Simpson's Law:
[0589] ;
[0590] f(a) and f(b) are the function values of the integrand f(x) at the two end points of the integration interval [a,b]. a and b are the endpoints of the integration interval. dx is a small increment of the interval, used to indicate the width of the integration region.
[0591] Specifically, this method estimates the integral by fitting a quadratic polynomial of the integrand f(x) in each small interval. In actual calculation, dx is obtained by evenly dividing the integral interval into several small intervals, calculating the integral value in each small interval, and then accumulating the estimated value of the entire integral area.
[0592] P3. Discretization: Discretize the integral region into small regions and perform numerical integration on each small region. Specifically, the interval [a, b] can be equally divided into N small regions, where a and b are the endpoints of the integral interval:
[0593] Where n is an even number. Let the width of each cell be: ;
[0594] P4, calculate the integral: perform Simpson's rule of P2 again for each small area, that is, perform numerical integration, and accumulate the results to get the final value:
[0595] P400, for the first small interval:
[0596] ;
[0597] For the second interval:
[0598] ;
[0599] And so on, for the i-th small interval:
[0600] ;
[0601] P401. Apply Simpson's rule: For each small interval [x2i,x2i+2], calculate the approximate integral value Area2i of Simpson's rule:
[0602] ;
[0603] P402, accumulate the integral value of each small area: add up the integral values of all small areas to get the final estimated integral value:
[0604] ;
[0605] The Python execution program used in all contents of this embodiment is:
[0606] def numerical_integration_simpson_rule(a, b, N):
[0607] delta_x = (b - a) / N
[0608] integral_result = 0.0
[0609] for i in range(N):
[0610] x_i = a + i * delta_x
[0611] x_ip1 = a + (i + 1) * delta_x
[0612] # Simpson's Law
[0613] integral_result += (delta_x / 6) * (
[0614] calculate_function_value(x_i) +
[0615] 4 * calculate_function_value((x_i + x_ip1) / 2) +
[0616] calculate_function_value(x_ip1) )
[0618] return integral_result
[0619] # Calculate P_f(x)
[0620] def calculate_failure_probability():
[0621] a, b = define_integration_interval() # Determine the integration interval
[0622] N = define_discretization_steps() # Define the number of discretization steps
[0623] integral_result = numerical_integration_simpson_rule(a, b, N)
[0624] P_f_x = integral_result
[0625] return P_f_x
[0626] In the above program, a and b represent the endpoints of the integration interval, and delta_x represents the width of each small interval, i.e., dx. calculate_function_value(x) is a function used to calculate the function value of the integrand at a given point x.
[0627] In this embodiment, numerical integration and Simpson's method are used to process complex probability density functions or cumulative distribution functions that are difficult to solve by analytical methods. For the complex system model provided by this embodiment, especially in high-dimensional space, these methods provide an effective way to estimate the failure probability.
[0628] Because multi-dimensional integration is usually involved in face topology modeling, numerical integration methods can handle this complexity. Simpson's rule is a method in numerical integration that is applicable to one-dimensional and multi-dimensional integration. It estimates the integral value by using polynomial interpolation. These methods are very flexible in dealing with practical problems because they do not depend on the specific form of the function, but approximate the integral value by discretizing the integral region. At the same time, Simpson's rule performs well in approximating the integral of continuous functions and can adaptively adjust the division of intervals.
[0629] Embodiment 3: Based on Embodiment 1, this embodiment can introduce another technical solution to connect to step S4, and introduce an adaptive adjustment mechanism (referred to as step S5) for the learning of the overall ALK model, such as Figure 4 As shown: This mechanism is implemented by the DS evidence theory algorithm (Dempster Shafer, DS), which considers information from different channels, aggregates and performs mapping and forms a correction factor ψ, through which the auxiliary probability density function p(x)' in step S200 is adaptively corrected. It includes sub-steps S500~S504.
[0630] In this embodiment, regarding step S500, evidence A and evidence B are obtained:
[0631] First, the candidate sample pool S is formed by collecting the samples in step S202 at the current time step t as evidence A (set):
[0632] ;
[0633] Then, the candidate sample pool S' is formed by collecting the samples in step S202 at the last time step t-1 as evidence B (set):
[0634] ;
[0635] Specifically, this step clarifies the division of two time steps, namely the current time step t and the previous time step t-1. Such a division helps capture the state changes of the system at different time points. Through the collection of candidate sample pools, two sets of evidence A and B are formed. These samples represent the system states at different time points. This step provides a basis for subsequent evidence fusion. By comparing the candidate sample pools of two time steps, we can understand the dynamic changes of the system over time and provide more information for model learning. In this scheme, by sampling candidate sample pools of continuous time steps, it supports the continuous modeling of the dynamic changes of the system, which helps to capture the changing trend of the system more accurately.
[0636] In this embodiment, regarding step S501, Dempster's combination principle is implemented:
[0637] The effect of Dempster's Combination Principle is that it solves the set of evidence A and evidence B and outputs the set of combined confidence assignments Bel(C) and the combined uncertainty assignments Pl(C); where C is the space of possible hypotheses or states.
[0638] In this embodiment, regarding step S5020, the combined trust distribution Bel(C) is:
[0639] ;
[0640] In the above formula, all the numerators are weighted sums of the evidences with intersections.
[0641] In the above formula, all denominators are weighted sums of non-overlapping evidence.
[0642] Ai∩B=∅ means that evidence A and evidence B have an intersection, and Ai∩B=∅ means that evidence A and evidence B do not have an intersection. Similarly, Ai∩C=∅ and Bi∩C=∅ mean that evidence A and evidence B have an intersection with hypothesis set C. Bel(C) represents the confidence in hypothesis set C (a value in the range [0,1]). By combining the Bel(C) of different evidences, a more comprehensive and reliable assessment of the confidence in hypothesis set C can be obtained. This helps to comprehensively consider the information from different evidences and improve the confidence level for different states.
[0643] Specifically: Dempster's combination principle is the core of DS evidence theory, which is used to combine different evidences to obtain the confidence for different hypotheses. By combining the confidence from evidence A and evidence B, the degree of support of the two sets of evidence for different hypotheses can be considered more comprehensively. The combined uncertainty distribution Pl(C) reflects the uncertainty of the combined evidence for each hypothesis, which helps to understand the uncertainty of the system state more comprehensively. The contribution of different evidence to the combined confidence can be adjusted through weights, so that more credible or reliable evidence plays a greater role in the combination.
[0644] In this embodiment, regarding step S5021, the combined uncertainty allocation Pl(C):
[0645] ;
[0646] In S5020~S5021, ∅ represents the empty set; and are a subset of the evidence A and the evidence B respectively; W is the weight (which can be averaged, that is, the weight of each subset is the same, but the sum is 1); C is the "hypothesis set";
[0647] Pl(C) represents the uncertainty of the hypothesis set C (a value in the range [0,1]), which is complementary to Bel(C). By combining Pl(C) of different evidences, a more comprehensive and reliable assessment of the uncertainty of the hypothesis set C can be obtained. Combined with Bel(C), it can provide a more comprehensive understanding, including the confidence in a certain state and the consideration of the uncertainty of that confidence.
[0648] Specifically, regarding the relationship of Bel(C):
[0649] (1) Complementary relationship: Pl(C) and Bel(C) are complementary concepts in Dempster's combination principle. If Bel(C) represents the confidence in a certain state, then Pl(C) represents the uncertainty of this confidence.
[0650] (2) Comprehensive understanding: When used in combination with Bel(C), it can provide a more comprehensive understanding. Bel(C) focuses on the degree of certainty of a state, while Pl(C) focuses on the degree of uncertainty about this certainty. The combination of the two can provide a more comprehensive understanding of the confidence and uncertainty of the system state.
[0651] It can be understood that Pl(C) is a measure of uncertainty for the hypothesis set C. By calculating the weights of different evidences, it reflects the uncertainty of the combined evidence for each hypothesis. Complementary to Bel(C), the two together constitute a complete expression of Dempster's combination principle, which can more comprehensively describe the cognition of the system state. At the same time, Bel(C) focuses on the degree of certainty of a certain state, and Pl(C) focuses on the degree of uncertainty of this certainty. Together, they provide a comprehensive assessment of the confidence and uncertainty of the system state.
[0652] In this embodiment, regarding step S502, a joint trust function Bel(A∪B) is obtained:
[0653] ;
[0654] (Pl(A)∩Pl(B)) is the uncertainty distribution of the intersection of the evidence A and the evidence B respectively.
[0655] The numerator in the formula is the intersection of evidence A and evidence B, which is normalized by multiplying the uncertainty distribution and trust of the other party and then dividing by 1-Pl(A∩B) to obtain the joint trust function. This joint trust function can be used to comprehensively consider the relationship between different evidences and provide a more comprehensive trust evaluation for a state or hypothesis.
[0656] Among them, Bel(A) and Bel(B): respectively represent the trust of evidence A and evidence B for a certain state or hypothesis. This is the trust of the system state evaluated based on different evidences in the Kriging model.
[0657] Among them, about Pl(A) and Pl(B): they represent the uncertainty of evidence A and evidence B for a certain state or hypothesis. In Dempster's combination principle, the measure of uncertainty.
[0658] Among them, Pl(A∩B): represents the uncertainty of the intersection of evidence A and evidence B. In Dempster's combination principle, the uncertainty of the intersection is the uncertainty of the part that exists in common between two pieces of evidence.
[0659] Specifically, the numerator of the above expression represents the trust that takes into account the uncertainty of the other evidence. The sum of these two parts represents the combined trust between the two evidences. The denominator is subtracted from 1 for normalization to prevent repeated calculation of the uncertainty of the intersection. By dividing this part, the normalized joint trust function is obtained.
[0660] Specifically, the specific method of obtaining (Pl(A)∩Pl(B)) is:
[0661] ;
[0662] and are subsets of evidence A and evidence B respectively. The numerator is the sum of the evidences wi supporting A∩B, and 1 minus this sum gives Pl(A∩B), which indicates the degree of uncertainty about A∩B.
[0663] in, Represents a subset of evidence A and a subset of evidence B In other words, it is able to find a subset and a subset , and is not zero under the correction of weight W, that is, the two subsets have an intersection. Responsible for finding some specific elements in hypothesis set C that are supported by both evidence A and evidence B. It is used to indicate the support for each element in hypothesis set C. It considers the support of two pieces of evidence for the possible set and combines these supports through Dempster's combination principle. The higher the value of this function, the stronger the support of the two pieces of evidence for an element in the possible set, and therefore the more confident the conclusion about the element is.
[0664] In this embodiment, regarding step S503, the joint trust function Bel(A∪B) is mapped to an interval value between [0,1] as the correction factor ψ through a sigmoid function:
[0665] ;
[0666] e is the base of natural logarithm (approximately equal to 2.718); the input of sigmoid function is the output of Bel(A∪B). The characteristic of sigmoid function is that its output range is between [0,1] regardless of the input, and the change of input value is relatively smooth, so it is used to map real numbers to probability range.
[0667] The size of the correction factor ψ is affected by Bel(A∪B), thereby adjusting the state transition function in the system. The output value of the sigmoid function approaches 1 when Bel(A∪B) is large, and approaches 0 when Bel(A∪B) is small. Therefore, the size of the correction factor ψ reflects the size of Bel(A∪B). This mapping preserves the causal relationship, that is, the change of ψ is consistent with the change of Bel(A∪B).
[0668] Specifically, by introducing ψ, the system introduces a certain uncertainty correction when considering the credibility of historical evidence. This design allows the system to adapt to different situations more flexibly, especially when facing new data and changes, the system can be more adaptive and robust. Through step S503, the system introduces uncertainty correction for historical evidence, and maps the joint trust function to the [0,1] interval through the sigmoid function, which improves the adaptability and robustness of the system and provides a more flexible and reliable foundation for the Kriging model to perform face topology modeling.
[0669] In this embodiment, regarding step S504, the auxiliary probability density function p(x)' is modified:
[0670] Based on the content of step S200 in the first embodiment, it can be known that the auxiliary probability density function p(x)' is in the form of:
[0671] ;
[0672] in, are the weights of the mixture components, is the probability density function of the multidimensional Gaussian distribution, is the mean vector, is the covariance matrix and K is the number of components in the mixture model.
[0673] Step S505 needs to correct the auxiliary probability density function p(x)' to obtain a new auxiliary probability density function p(x)'', and the new auxiliary probability density function p(x)'' has undergone adaptive correction:
[0674] ;
[0675] The process of correcting p(x)'' by the correction factor ψ is actually a weighted adjustment of the original function. The purpose of this adjustment is to introduce dependence on the trust of the correction factor ψ, so as to adaptively correct the model according to the current evidence (reflected by ψ). Because the correction factor ψ is obtained by mapping the joint trust function Bel(A∪B) through the sigmoid function. ψ reflects the credibility of historical evidence, and by introducing dependence on ψ, adaptive correction of current evidence is achieved.
[0676] Specifically, the key to the correction is to introduce the correction factor ψ into the weight of each component In , the dependence on the confidence of the correction factor is realized. This means that the contribution of each component in the correction process is regulated by ψ, making the correction more flexible and adaptive. The core of adaptability is to adjust the model according to the current evidence, rather than using static model parameters. By introducing ψ, the model can adapt to different evidence more flexibly, improving the dynamics of the model.
[0677] Specifically, regarding adaptability and robustness: Adaptive correction makes the model more adaptable and robust in the face of new data and changes. The correction process reflects the system's trust in historical evidence, allowing the system to adjust more flexibly in different situations.
[0678] Embodiment 4: Based on Embodiment 3, in this embodiment, after the correction factor ψ is obtained in step S503, step S504 is not performed, but the correction factor ψ is used as the target failure error threshold T2 as described in Embodiment 1:
[0679] ;
[0680] That is, ψ is complemented to obtain the target failure error threshold T2. The reason for complementation is:
[0681] (1) The size of ψ reflects the system's trust in historical evidence. A smaller ψ indicates that the system has a lower trust in historical evidence, while a larger ψ indicates that the system has a higher trust in historical evidence. T2, obtained by complementing ψ, reflects a trend opposite to the trust, that is, when the system has a lower trust in historical evidence, T2 is larger and the system is more likely to accept new evidence.
[0682] (2) Controlling the conservatism of learning stop: Through complementation, the design of T2 makes the system more conservative when the trust is low and has higher requirements for target failure. Such a design helps the system learn more cautiously when the trust in historical evidence is low, so as to avoid over-reliance on new evidence and maintain the robustness of the system.
[0683] (3) Introducing the target failure threshold: T2 is designed to be related to the reverse trend of ψ, so that the system can control the stop of learning by adjusting T2 at different trust levels, thereby effectively introducing the target failure error threshold. The setting of this threshold affects the system's acceptance of new evidence, thereby adjusting the conservatism of learning.
[0684] In this solution, ψ is directly used as the target failure error threshold T2. This means that the system considers the target to be invalid after reaching a certain trust level ψ. At the same time, ψ is the trust level calculated by comprehensive calculation of historical evidence, so T2 is related to the system's trust level in historical evidence. In layman's terms:
[0685] When ψ is small, that is, the system has low trust in historical evidence, T2 is large and the system is more receptive to new evidence.
[0686] When ψ is larger, that is, when the system has a higher degree of trust in historical evidence, T2 is smaller, the system is more conservative, and has a higher requirement for the failure of the target.
[0687] Specifically: By using ψ directly as the value of T2, the system exhibits different learning and decision-making behaviors at different trust levels. When ψ is small, T2 is large, and the system is more receptive to new evidence, which increases the system's adaptability and helps to adapt to environmental changes more flexibly. As ψ increases, T2 decreases, and the system's trust in historical evidence increases, which makes the system more conservative. The decrease in T2 reflects the system's higher requirements for target failure, thereby increasing its sensitivity to historical evidence. The system updates the model more cautiously to ensure that the high trust in historical evidence can be reflected in the high requirements for target failure judgment.
[0688] It is understandable that this strategy can flexibly adjust the speed of model learning: the target failure error threshold T2 controls the learning stop condition, and reflects the system's trust in historical evidence through ψ. This design allows for faster learning and faster adaptation to new situations when the system trust is low. When the system trust is high, by reducing T2, the learning speed is slowed down, and new evidence is treated more conservatively, reducing the risk of uncertainty. Similarly, since T2 is adjusted according to ψ, it can be adjusted according to specific application scenarios and system requirements. This allows the system to respond flexibly in different application scenarios and better adapt to various complex situations and data changes.
[0689] All the above embodiments only express the implementation methods of the relevant practical applications of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be based on the attached claims.
Claims
1. A Kriging large model adaptive training method for face modeling, comprising an ALK model, wherein the ALK model comprises a Kriging model, wherein the ALK model constructs a regression process and a random process of the Kriging model; characterized in that: The steps include: S1, generate initial training points: extract N uniformly distributed initial sample points from the random variable space σ, substitute them uniformly into the function g(x), calculate the response values of these initial sample points, and form an initial training set T after aggregation; establish or update the Kriging model based on the current training set T; In said S1, the following steps are included: S100, the extraction is performed by the Halton sequence sampling algorithm: Create a Halton sequence whose extracted value of the i-th element for: ; is the cardinality of uniform distribution, k is the maximum value of the total element i, is an integer uniformly distributed in the range [0, b-1]; S101, obtain uniformly distributed initial sample points: take each element of the Halton sequence as a component of the random variable; the i-th element of the Halton sequence It is expressed as: ; Represents the value of the i-th element in the j-th dimension; The i-th sample point for: ; Then all the sample points Substitute into the functional function g(x); S2, construct the limit state function proxy model and generate the candidate sample pool S: execute the Gaussian mixture algorithm to obtain the auxiliary probability density function p(x)' of the Kriging model, and use the failure domain set G to combine with the auxiliary probability density function p(x)' to generate n random variable samples , thereby forming a candidate sample pool S for selecting training points for the next update, and each step is regenerated as the Kriging model is updated; S3, identify the next updated training point in the candidate sample pool S: calculate the function value s of each sample in the candidate sample pool S, select the smallest function value s* and its corresponding failure domain sample point in the function value s , as the training point for the next update; S4, determine the convergence of the learning process: based on the minimum function value s* and its corresponding failure domain sample point , calculate the upper bound of the failure probability error evaluated by the current Kriging model, and if the upper bound is less than the target failure error threshold, stop the learning process; Otherwise, the upper confidence interval is calculated and the process returns to S2 to update the Kriging model. S400, calculate the upper bound of the failure probability error : ; in, is the predicted mean, is to use the failure domain sample points The calculated failure probability, is the failure probability calculated using the ALK model predicted mean; S401, determine whether the learning stop condition is met: ; T2 is the target failure error threshold; if the failure probability error upper bound If it is less than the target failure error threshold T2, then stop learning, otherwise continue learning; S402, calculate the upper confidence interval: If learning is not stopped, the upper confidence interval of the failure probability error is calculated again to perform the next round of Kriging model update; S4020, calculate the standard deviation of the error in the probability of failure : ; is the ith failure probability error, is the average value of the total failure probability error; n is the total number of samples; S4021, calculate the confidence interval radius Ma: ; Z is the "Z-score" from the standard normal distribution, corresponding to the chosen confidence level α; S4022, calculate the new confidence interval CI: ; After obtaining a new confidence interval CI of the failure probability error, a decision is made to determine whether to stop learning or continue updating the Kriging model; S5, by considering the information from different channels, summarizing and mapping, and forming a correction factor ψ, the auxiliary probability density function p(x)' in S2 is adaptively corrected by the correction factor ψ.
2. The adaptive training method according to claim 1, characterized in that: In the S1, it also includes: S102, forming the initial training set T: ; is the sample point corresponding to The response value output by the functional function g(x); S103: Establish or update the Kriging model based on the current training set T.
3. The adaptive training method according to claim 1, characterized in that: In said S2, it includes: S200, performing a Gaussian mixture algorithm on the auxiliary probability density function p(x)': ; in, are the weights of the mixture components, is the probability density function of the multidimensional Gaussian distribution, μ i is the mean vector, is the covariance matrix, K is the number of components of the mixed model; S201, using the failure domain set G combined with the auxiliary probability density function p(x)' to generate n random variable samples ,include: First, generate a sample from the multidimensional Gaussian distribution corresponding to the failure domain set G: ; The symbol "~" means "...distributed in..."; Then randomly select a sample from the failure domain set G , and according to the auxiliary probability density function Generate random variables: ; Then generate n random variable samples , forming a candidate sample pool S: 。 4. The adaptive training method according to claim 3, characterized in that: In the S2, it also includes: S202, generating a candidate sample pool S: Repeat S201 to generate n random variable samples , forming a candidate sample pool S: 。 5. The adaptive training method according to claim 1, characterized in that: In said S3, it includes: S300, calculate the function value of each sample in the candidate sample pool S , including the limit state function G(y), whose input is the vector of variables; each random variable sample is extracted from the candidate sample pool S As input variables: ; , is the same as the i-th input variable Related regulatory factors; is applied to the ith input variable The limit state function of S301, select the minimum sample point in the function value s as the training point for the next update: The function value of each sample output based on the limit state function G(y) , after aggregation, select the minimum function value s*, and then find the corresponding failure domain sample point .
6. The adaptive training method according to claim 5, characterized in that: In said S5, it includes: S500, obtaining evidence A and evidence B: collecting S2 at the current time step t to form a candidate sample pool S, as evidence A, collecting the S2 at the previous time step t-1 to form a candidate sample pool S', as evidence B; S501, execute Dempster's combination principle: solve the set of evidence A and evidence B, and output the set of combined confidence distribution Bel(C) and the combined uncertainty distribution Pl(C); S502, obtaining a joint trust function Bel(A∪B); S503, mapping the joint trust function Bel(A∪B) to an interval value between [0,1] as the correction factor ψ through a sigmoid function; S504, modify the auxiliary probability density function p(x)': 。 7. The adaptive training method according to claim 6, characterized in that: In the S500, it includes: ; ; In the S501, it includes: S5020, merge trust allocation Bel(C): ; S5021, combined uncertainty allocation Pl(C): ; In S5020~S5021, C is the "hypothesis set"; A i ∩B=∅ means that evidence A and evidence B have an intersection, A i ∩B=∅ means that evidence A and evidence B have no intersection, A i ∩C = ∅ and B i ∩C=∅ means that evidence A and evidence B have an intersection with hypothesis set C; Bel(C) represents the confidence in hypothesis set C; ∅ represents the empty set; and are a subset of the evidence A and the evidence B respectively; W is the weight; In the S502, it includes: ; (Pl(A)∩Pl(B)) is the uncertainty distribution of the intersection of the evidence A and the evidence B; In the step 503, the sigmoid function includes: ; Here, e is the base of natural logarithms.
8. The adaptive training method according to claim 7, characterized in that: In S4, the updating method is: S4023, Learning Strategy Program: ; γ is the original learning step size or learning rate, γ' is the updated learning step size or learning rate; f(CI) is the adjustment function of the confidence interval CI: ; w(CI) is the width of the confidence interval, where w(CI)=U(x)−L(x), L(x) is the lower bound, and U(x) is the upper bound; v is the adjustment parameter used to control the magnitude of the adjustment.
9. Kriging large model adaptive training system for face modeling, characterized by: The system includes a processor and a register connected to the processor, wherein program instructions are stored in the register, and when the program instructions are executed by the processor, the processor executes the adaptive training method as described in any one of claims 1-8.
Citation Information
Patent Citations
A fluid flow regulator for intravenous feeding devic.
ES1000233U
Iteration interpolation method based on face triangle mesh adaptive subdivision and Gauss wavelet
CN105678252A
Structural reliability analysis method based on self-adaptive agent model
CN107563067A