Sample screening method and system based on weight clustering and dynamic interval reduction
Through weight clustering and dynamic interval reduction methods, the waste of computing resources and insufficient accuracy of sample screening in complex mechanical systems are solved, efficient and accurate sample screening is achieved, and the efficiency and accuracy of reliability analysis are improved.
Patent Information
- Application Number
- CN202510374612.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
AI Technical Summary
In the reliability analysis of complex mechanical systems, the sample screening method has the problems of wasted computing resources and insufficient screening accuracy, especially in high-dimensional spaces, which are highly complex and susceptible to noise interference.
Weight clustering method and dynamic interval reduction method are used to predict the upper and lower thresholds of the limit state function through the proxy model, combined with the sample removal rate and reduction rate, the filter interval boundaries are dynamically adjusted to achieve efficient screening and iterative optimization of samples.
Effectively compress the size of candidate samples, reduce redundant calculations, improve the density of the boundary area of the sample points approaching the limit state plane, enhance the proxy model's characterization ability of the limit state plane, and improve iteration efficiency and accuracy.
Smart Images

Figure CN120296376A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sample screening methods in mechanical reliability analysis methods, and particularly relates to a sample screening method and system based on weight clustering and dynamic interval reduction. Background Art
[0002] As the core carrier in fields such as aerospace, automobile manufacturing, and energy equipment, the reliability of complex mechanical systems directly affects the operating safety of equipment and the performance throughout the life cycle. With the evolution of modern industrial equipment towards integration and intelligence, the multi-physical field coupling effect and multi-failure mode interaction inside the system are becoming increasingly significant. Traditional experience-driven design methods have been difficult to meet the requirements of high reliability. It is urgent to construct a reliability analysis system based on probability statistics to achieve the coordinated optimization of safety and economy by quantitatively evaluating the boundary distribution characteristics of the limit state function. In this process, accurate and efficient reliability analysis can not only reduce the cost of experimental verification but also be the key support technology for preventing major accidents and enhancing the market competitiveness of core equipment.
[0003] As the core link of reliability analysis, the effectiveness of sample screening directly determines the utilization efficiency of computing resources and the characterization accuracy of the limit state surface. In the reliability analysis of complex mechanical systems, the limit state function often involves non-linear finite element analysis or high-fidelity multidisciplinary simulation, and the single calculation can take several hours or even several days. Facing the massive candidate samples generated by the Monte Carlo method, how to quickly lock the key samples that significantly contribute to the failure probability through an intelligent screening mechanism, while ensuring the sample density in the boundary region and eliminating redundant calculations, has become an important breakthrough point for improving the engineering applicability of reliability analysis. An efficient sample screening strategy can not only shorten the iteration cycle but also enhance the proxy model's ability to capture complex failure boundaries through sample space optimization.
[0004] In the prior art, when patent CN107038303A adopts a two-layer experimental design, the secondary sample screening completely depends on the prediction results of the initial proxy model. When the model is insufficiently fitted in the limit state region, it is easy to cause the samples to deviate from the true critical boundary. Moreover, after generating tens of thousands of candidate samples by uniform sampling in the second stage, only the first k samples are mechanically screened, resulting in waste of computing resources. Patent CN114077776A introduces an active learning function, but the full-distance calculation in the high-dimensional space leads to a sharp increase in time complexity, and it depends on a fixed error threshold (0.05 - 0.1) to screen samples, there are risks of noise interference and error accumulation. Summary of the Invention
[0005] Objective of the Invention: In order to overcome the deficiencies in the prior art, the present invention provides a sample screening method and system based on weighted clustering and dynamic interval reduction, which has high iteration efficiency, can effectively screen out sample points near the limit state surface, and has a fast convergence speed.
[0006] Technical Solution: To achieve the above objective, the sample screening method based on weighted clustering and dynamic interval reduction of the present invention includes:
[0007] Generating an initial sample pool through the Monte Carlo method;
[0008] Based on the initial sample pool, performing multiple rounds of iterative screening using a sample screening process to obtain a screened sample pool;
[0009] The sample screening process includes:
[0010] Obtaining uniformly distributed candidate samples from the current sample pool using the weighted clustering method;
[0011] Using the interval reduction method to predict the upper threshold of the limit state function in the sample pool through a surrogate model and the lower threshold where j is the iteration round; the interval reduction method is based on the following rules:
[0012]
[0013] When j≥3,
[0014] where, is the predicted value of the surrogate model for the candidate sample x x = [x1, x2,..., x n ; k is a coefficient less than 1;
[0015] Reduction rate
[0016] In the formula, β min represents the minimum value of the reduction rate; β max represents the maximum value of the reduction rate; κ determines the change speed of the function; τ represents the rejection rate of the samples; τ c is the function center position parameter;
[0017] Selecting the candidate samples whose predicted values of the limit state function are between the upper threshold and the lower threshold to update the sample pool.
[0018] Furthermore, the step of obtaining uniformly distributed candidate samples from the current sample pool using the weighted clustering method includes:
[0019] Calculate the weight coefficient based on the joint probability distribution of each sample point in the current sample pool within the neighborhood of ±10% standard deviation of its respective dimensional variables;
[0020] Perform weighted sampling on the sample pool according to the weight coefficient to generate a weighted sample set with a uniform distribution;
[0021] Perform K-means clustering on the weighted sample set and select candidate samples with a uniform distribution from the weighted sample set.
[0022] Furthermore, the calculating the weight coefficient based on the joint probability distribution of each sample point in the current sample pool within the neighborhood of ±10% standard deviation of its respective dimensional variables includes:
[0023] Calculate the occurrence probability P(x) of the sample point within the neighborhood of ±10% standard deviation of its respective dimensional variables, specifically:
[0024]
[0025] Where: is the standard deviation of x; f(x j ) represents the probability distribution function of the random variable x j ; represents the cumulative distribution function of the random variable x j ;
[0026] Calculate the weight coefficient of the i-th sample x (i) :
[0027] Where α represents the weight smoothing adjustment parameter.
[0028] Furthermore, the performing weighted sampling on the sample pool according to the weight coefficient to generate a weighted sample set with a uniform distribution includes:
[0029] Normalize the weight coefficient to obtain the normalized weight coefficient
[0030] Calculate the eigenvalue of each sample point Where u (i) is a random number, whose value is greater than 0 and less than 1;
[0031] Select a specific number of sample points with the largest eigenvalues to obtain the weighted sample set.
[0032] Furthermore, select the corresponding number of sample points with a proportion of 1% from the current sample pool to form the weighted sample set.
[0033] Furthermore, when the change rate of the number of sample points in the sample pool obtained after the implementation of adjacent two rounds of the sample screening process is less than 5%, it is determined to converge and the iteration is stopped.
[0034] A sample screening system based on weight clustering and dynamic interval reduction, the system includes:
[0035] A generation module, which is used to generate an initial sample pool through the Monte Carlo method;
[0036] A screening module, which is used to perform multiple rounds of iterative screening based on the initial sample pool by using a sample screening process to obtain a screened sample pool;
[0037] The screening module includes:
[0038] A first screening unit, which is used to obtain uniformly distributed candidate samples from the current sample pool by using the weight clustering method;
[0039] An interval reduction unit, which is used to predict the upper threshold of the limit state function in the sample pool by using the interval reduction method and the lower threshold where j is the number of iteration rounds; the interval reduction method is carried out based on the following rules:
[0040]
[0041] When j≥3,
[0042] where, is the predicted value of the surrogate model for the candidate sample x , x = [x1, x2,..., x n ; k is a coefficient less than 1;
[0043] Reduction rate
[0044] In the formula, β min represents the minimum value of the reduction rate; β max represents the maximum value of the reduction rate; κ determines the change speed of the function; τ represents the rejection rate of the sample; τ c is the function center position parameter;
[0045] A second screening unit, which is used to select the candidate samples whose predicted values of the limit state function are between the upper threshold and the lower threshold, and update the sample pool.
[0046] Beneficial effects: The sample screening method and system based on weight clustering and dynamic interval reduction of the present invention have the following beneficial effects:
[0047] (1) In the method and system of the present invention, the weight clustering method is combined to dynamically screen candidate samples. The upper and lower thresholds of the limit state function predicted by the surrogate model are used for interval reduction. By introducing a dynamic reduction rate function linked to the sample rejection rate, the screening interval boundary is adaptively adjusted. While ensuring the screening accuracy, the scale of candidate samples is effectively compressed, the redundant calculation amount is reduced, and efficient iterative optimization is achieved. The retained sample points gradually approach the boundary region of the limit state surface, the sample density in the boundary region is increased while redundant samples are reduced, and the geometric relationship between the sample distribution after iteration and the limit state surface is continuously optimized through the threshold linkage mechanism.
[0048] (2) In each round of iteration, the sample distribution in the high and low probability regions is balanced using the weight coefficient, so that the weighted sample set gathers around the limit state surface. The representation ability of the surrogate model for the boundary of the limit state surface is enhanced through spatial equalization, and the sample points gradually converge to the adjacent region of the limit state surface during the iteration process.
[0049] (3) By adjusting the exponential factor α to optimize the calculation of the weight coefficient, the excessive aggregation of samples in the high probability region is suppressed during the iteration process, and the weight of samples in the low probability region is dynamically increased, so that the sample points after each round of screening are closer to the edge distribution of the limit state surface, and the sensitivity of the surrogate model to sparse samples at the boundary of the limit state surface is enhanced. Description of the Drawings
[0050] Figure 1 is a schematic flow diagram of the sample screening method based on weight clustering and dynamic interval reduction;
[0051] Figure 2 is a schematic flow diagram of the sample screening process;
[0052] Figure 3 is a schematic composition diagram of the sample screening system based on weight clustering and dynamic interval reduction;
[0053] Figure 4 is a schematic composition diagram of the screening module. Detailed Embodiments
[0054] The present invention will be further described in detail below with reference to the accompanying drawings.
[0055] As Figure 1 shown, the sample screening method based on weight clustering and dynamic interval reduction includes the following steps S101 - S102:
[0056] Step S101, generating an initial sample pool through the Monte Carlo method;
[0057] Step S102, based on the initial sample pool, performing multiple rounds of iterative screening using the sample screening process to obtain the screened sample pool;
[0058] As Figure 2 shown, the sample screening process described in the above step S102 includes the following steps S201 - S203:
[0059] Step S201, obtaining uniformly distributed candidate samples from the current sample pool by using the weighted clustering method;
[0060] In this step, in the first round of iteration, the current sample pool is the initial sample pool, and in subsequent iteration rounds, the current sample pool is the sample pool updated after the previous round of iteration. In the first round, the candidate samples selected from the initial sample pool by using the weighted clustering method are used to construct the surrogate model.
[0061] Step S202, predicting the upper threshold of the limit state function in the sample pool through the surrogate model by using the interval reduction method and the lower threshold where j is the iteration round; the interval reduction method is carried out based on the following rules:
[0062]
[0063] When j ≥ 3,
[0064] where, is the predicted value of the surrogate model for the candidate sample x , x = [x1, x2,..., x n , where x1, x2,..., x n are random variables; k is a coefficient less than 1, and in this embodiment, k = 0.4;
[0065] Reduction rate
[0066] In the formula, β min represents the minimum value of the reduction rate; β max represents the maximum value of the reduction rate; κ determines the change speed of the function; τ represents the sample rejection rate, which is the ratio of the number of rejected samples to the total number of samples in the sample pool after the previous iteration; τ c is the function center position parameter, and in this embodiment, τ c = 0.5;
[0067] In this step, in each round, the surrogate model has been updated, and the surrogate model used is the latest updated surrogate model.
[0068] Step S203, selecting the candidate samples whose predicted values of the limit state function are between the upper threshold and the lower threshold to update the sample pool.
[0069] In the above method, the candidate samples are dynamically screened by combining the weighted clustering method, the interval reduction is carried out by using the upper and lower thresholds of the limit state function predicted by the surrogate model, and the screening interval boundary is adaptively adjusted by introducing a dynamic reduction rate function linked to the sample rejection rate, effectively compressing the scale of candidate samples while ensuring the screening accuracy, reducing the redundant calculation amount, realizing efficient iterative optimization, making the retained sample points gradually approach the boundary region of the limit state surface, increasing the sample density in the boundary region while reducing redundant samples, and ensuring the continuous optimization of the geometric relationship between the sample distribution after iteration and the limit state surface through the threshold linkage mechanism.
[0070] Further, the step of obtaining uniformly distributed candidate samples from the current sample pool by using the weighted clustering method in the above step S201 includes the following steps S301 - S303:
[0071] Step S301, calculating the weight coefficient based on the joint probability distribution of each sample point in the current sample pool within the ±10% standard deviation neighborhood of its respective dimensional variables;
[0072] Step S302, performing weighted sampling on the sample pool according to the weight coefficient to generate a uniformly distributed weighted sample set;
[0073] Step S303, performing K - means clustering on the weighted sample set and selecting uniformly distributed candidate samples from the weighted sample set.
[0074] In the above steps S301 - S303, the weight coefficient is generated by calculating the joint probability distribution of the sample points in the neighborhood of each dimension. By combining the methods of weighted sampling and K - means clustering, while retaining the samples in the high - probability region, the coverage ability of the samples in the low - probability region is enhanced, and the distribution uniformity of the candidate sample set is improved through spatial equalization processing. In each round of iteration, the weight coefficient is used to balance the sample distribution in the high - and low - probability regions, making the weighted sample set gather around the limit state surface, enhancing the representation ability of the surrogate model for the boundary of the limit state surface through spatial equalization, and the sample points gradually converge to the neighborhood region of the limit state surface during the iteration process.
[0075] Further, the step of calculating the weight coefficient based on the joint probability distribution of each sample point in the current sample pool within the ±10% standard deviation neighborhood of its respective dimensional variables in the above step S301 includes the following steps S401 - S402:
[0076] Step S401, calculating the occurrence probability P(x) of the sample point within the ±10% standard deviation neighborhood of its respective dimensional variables, specifically:
[0077]
[0078] Where: is the standard deviation of x; f(x j) represents the probability distribution function of the random variable x j , represents the cumulative distribution function of the random variable x j ;
[0079] Step S402, calculate the weight coefficient of the i-th sample x (i) :
[0080] where α represents the weight smoothing adjustment parameter; the setting of this parameter is designed to balance the sampling probability distributions in different regions. By adjusting the exponential factor in the weight calculation, it alleviates the weight polarization phenomenon caused by uneven distribution of the probability density function. When the adjustment factor takes a relatively small value, the algorithm tends to increase the sampling weights in low-probability regions and enhance the representation ability of the edge regions; conversely, a larger value will strengthen the sample screening advantage in high-probability regions. Through experimental verification, the optimal working range of this adjustment factor is determined to be in the range of 0.5 to 1.0 in this method, so as to achieve the collaborative optimization of the probability density distribution and the spatial coverage range, and effectively improve the global representation ability of the sample set.
[0081] In the above steps, the reciprocal of the joint probability in the ±10% standard deviation neighborhood is used as the weight calculation benchmark. By adjusting the value of the exponential factor α, the weight distribution between high- and low-probability regions is balanced, the sample aggregation phenomenon caused by probability density differences is suppressed, the integrity of spatial coverage and the probability distribution characteristics are taken into account during the weight smoothing adjustment process, and the screening weights of samples in the edge regions are enhanced to improve the ability to capture the boundary of the limit state function.
[0082] Furthermore, the weighted sampling of the sample pool according to the weight coefficient in step S302 above to generate a uniformly distributed weighted sample set includes the following steps S401 - S403:
[0083] Step S401, normalize the weight coefficient to obtain the normalized weight coefficient
[0084] Step S402, calculate the eigenvalue of each sample point where u (i) is a random number, and its value is greater than 0 and less than 1;
[0085] Step S403, select a specific number of sample points with the largest eigenvalues to obtain the weighted sample set.
[0086] In the above steps, the probability proportional selection strategy is used to ensure the reasonable retention probability of samples with different weights. By repeating the sampling operation, the distribution density of high-weight samples is strengthened, and a weighted sample set that takes into account both spatial uniformity and probability distribution characteristics is formed, providing a data basis for subsequent clustering.
[0087] Further, in the above step S403, a corresponding number of the sample points accounting for 1% are selected from the current sample pool to form the weighted sample set.
[0088] Further, after the implementation of two adjacent rounds of the sample screening process, when the change rate of the number of sample points in the obtained sample pool is less than 5%, it is determined to converge, and the iteration is stopped.
[0089] A sample screening system 400 based on weight clustering and dynamic interval reduction. The sample screening system 400 may include or be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the present invention and implement the above sample screening method. The program modules referred to in the embodiments of the present invention refer to a series of computer program instruction segments that can complete specific functions, and are more suitable for describing the execution process of the sample screening method in the storage medium than the program itself. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 3 shown, the sample screening system 400 includes:
[0090] A generation module 410, which is used to generate an initial sample pool by the Monte Carlo method;
[0091] A screening module 420, which is used to perform multiple rounds of iterative screening on the basis of the initial sample pool by using the sample screening process to obtain a screened sample pool;
[0092] As Figure 4 shown, the screening module 420 includes:
[0093] A first screening unit 421, which is used to obtain uniformly distributed candidate samples from the current sample pool by using the weight clustering method;
[0094] An interval reduction unit 422, which is used to predict the upper threshold of the limit state function in the sample pool by using the interval reduction method through a surrogate model and the lower threshold where j is the iteration round; the interval reduction method is based on the following rules:
[0095]
[0096] When j≥3,
[0097] where, is the predicted value of the surrogate model for the candidate sample x , x = [x1, x2,..., x n , where x1, x2,..., x n are random variables; k is a coefficient less than 1, and in this embodiment, k = 0.4;
[0098] Reduction rate
[0099] In the formula, β min represents the minimum value of the reduction rate; β max represents the maximum value of the reduction rate; κ determines the change speed of the function; τ represents the rejection rate of samples, which is the ratio of the number of rejected samples to the total number of samples in the sample pool after the previous iteration; τ c is the function center position parameter. In this embodiment, τ c = 0.5;
[0100] In each round, the surrogate model has been updated, and the surrogate model used is the latest updated surrogate model.
[0101] The second screening unit 423 is used to select the candidate samples whose predicted values of the limit state function are between the upper threshold and the lower threshold, and update the sample pool.
[0102] Other contents of implementing the above sample screening method based on the sample screening system have been introduced in detail in the previous embodiments. For corresponding contents in the previous embodiments, reference can be made, and details are not described here again.
[0103] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A sample screening method based on weight clustering and dynamic interval reduction, which includes: Generating an initial sample pool through the Monte Carlo method; Based on the initial sample pool, using a sample screening process for multiple rounds of iterative screening to obtain a screened sample pool; It is characterized in that: the sample screening process includes: Using the weight clustering method to obtain uniformly distributed candidate samples from the current sample pool; Predicting the upper threshold of the limit state function in the sample pool through a surrogate model using an interval reduction method and the lower threshold where j is the iteration round; the interval reduction method is carried out based on the following rules: When j ≥ 3, Among them, is the predicted value of the proxy model for the candidate sample x , x = [x1, x2,..., x n ; k is a coefficient less than 1; Reduction rate where β min represents the minimum value of the reduction rate; β max represents the maximum value of the reduction rate; κ determines the change speed of the function; τ represents the rejection rate of the samples; τ c is the function center position parameter; Select the candidate samples whose predicted values of the limit state function are between the upper threshold value and the lower threshold value to update the sample pool.
2. The sample screening method based on weight clustering and dynamic interval reduction according to claim 1, wherein The step of using the weight clustering method to obtain uniformly distributed candidate samples from the current sample pool includes: Calculating the weight coefficient based on the joint probability distribution of each sample point in the current sample pool within the ±10% standard deviation neighborhood of its respective dimensional variables; Performing weighted sampling on the sample pool according to the weight coefficient to generate a uniformly distributed weighted sample set; Performing K-means clustering on the weighted sample set and selecting uniformly distributed candidate samples from the weighted sample set.
3. The sample screening method based on weight clustering and dynamic range reduction according to claim 2, wherein The step of calculating the weight coefficient based on the joint probability distribution of each sample point in the current sample pool within the ±10% standard deviation neighborhood of its respective dimensional variables includes: Calculating the occurrence probability P(x) of the sample point within the ±10% standard deviation neighborhood of its respective dimensional variables, specifically: Wherein: is the standard deviation of x; f(x j ) represents the probability distribution function of the random variable x j ; represents the cumulative distribution function of the random variable x j ; Calculate the weight coefficient of the i-th sample x (i) : Among them, α represents the weight smoothing adjustment parameter.
4. The sample screening method based on weight clustering and dynamic interval reduction according to claim 2, wherein The step of performing weighted sampling on the sample pool according to the weight coefficient to generate a uniformly distributed weighted sample set includes: Normalize the weight coefficient to obtain the normalized weight coefficient wherein Calculate the eigenvalue of each of the sample points where u (i) is a random number, whose value is greater than 0 and less than 1; Selecting a specific number of sample points with the largest eigenvalues to obtain the weighted sample set.
5. The sample screening method based on weight clustering and dynamic interval reduction according to claim 4, characterized in that Selecting 1% of the corresponding number of sample points from the current sample pool to form the weighted sample set.
6. The sample screening method based on weight clustering and dynamic interval reduction according to claim 1, characterized in that When the change rate of the number of sample points in the sample pool obtained after the implementation of the sample screening process in two adjacent rounds is less than 5%, it is determined to converge and the iteration stops.
7. A sample screening system based on weight clustering and dynamic interval reduction, the system includes: A generating module, which is used to generate an initial sample pool through the Monte Carlo method; A screening module, which is used to perform multiple rounds of iterative screening based on the initial sample pool using the sample screening process to obtain a screened sample pool; It is characterized in that the screening module includes: A first screening unit, which is used to obtain uniformly distributed candidate samples from the current sample pool using the weight clustering method; An interval reduction unit, which is used to predict the upper threshold of the limit state function in the sample pool through a surrogate model by using an interval reduction method and the lower threshold where j is the iteration round; the interval reduction method is carried out based on the following rules: When j ≥ 3, wherein, is the predicted value of the surrogate model for the candidate sample x , x = [x1, x2,..., x n ; k is a coefficient less than 1; Reduction rate where β min represents the minimum value of the reduction rate; β max represents the maximum value of the reduction rate; κ determines the change speed of the function; τ represents the rejection rate of the samples; τ c is the function center position parameter; A second screening unit, which is used to select the candidate samples whose predicted values of the limit state function are between the upper threshold and the lower threshold and update the sample pool.
Citation Information
Patent Citations
Double-layer experiment design method based on agent model and applied to mechanical reliability analysis and design
CN107038303A
Structural reliability robust optimization design method based on active learning agent model
CN114077776A
Cited By
Time-varying reliability optimization design method and system for helical gear
CN121744556A