Discrete element simulation repose angle analysis data cleaning method
Through the combination of Q-Q graph and kernel density estimation, outliers in discrete element simulation angle of rest data are screened out, which solves the problem of inaccurate data screening in the prior art, and improves the reliability of data and the accuracy of the model.
Patent Information
- Application Number
- CN202510272045.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-04
AI Technical Summary
Existing data screening methods are difficult to accurately identify and remove non-normal distributed data points in discrete element simulation angle of rest data, resulting in insufficient accuracy and reliability of the response surface model.
The method of Q-Q graph analysis combined with kernel density estimation is used to filter outliers by calculating the cumulative distribution function of the residual absolute value, including data preprocessing, theoretical quantile calculation, reference line function construction, residual absolute value calculation and kernel density estimation, and the abnormal data threshold is determined and eliminated.
Improve the accuracy and reliability of the data, reduce the error screening rate and missed screening rate, and provide more reliable parameter basis for subsequent statistical analysis and modeling.
Smart Images

Figure CN120256813A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a data screening method, and in particular to a data cleaning method for analyzing the angle of repose in discrete element simulation, belonging to the technical fields of experimental data processing and physical property measurement of granular materials. Background Art
[0002] The angle of repose is an important parameter for measuring the fluidity and packing characteristics of granular materials. In the calibration of discrete element method (DEM) parameters for granular materials, it is often used as the core response variable to characterize the particle flow characteristics. The response surface method (RSM) realizes parameter calibration by constructing a mathematical model between process parameters and the angle of repose. This method generally assumes that the response variable follows a normal distribution and there is no systematic measurement error. If there are outliers or non-normal distributions in the angle of repose data, it may lead to distortion of the regression model, and further result in inaccurate parameter calibration results. Therefore, the quality of the discrete element simulation angle of repose data directly determines the reliability of the response surface method results. In order to construct an accurate regression model and optimize the result parameters of parameter calibration, it is necessary to check whether the input discrete element simulation angle of repose data conforms to the normal distribution before performing the response surface design, and clean out the abnormal data that may be measurement errors or experimental errors.
[0003] Traditional normality test methods, such as the Shapiro-Wilk test and the Kolmogorov-Smirnov test, although they can judge whether the data conforms to the normal distribution, it is difficult to accurately screen out non-normal data points in the case of mixed abnormal data. The traditional Quantile-Quantile Plot (abbreviation: Q-Q plot) can judge the normality of data by observing whether the data is distributed along the reference line, but for the screening of non-normal distributed data in the dataset, it can only be distinguished by manual visual inspection. The manual judgment has strong subjectivity, and the differences in the experience of operators will lead to relatively high false screening rates and missed screening rates. Therefore, a data cleaning method for analyzing the angle of repose in discrete element simulation is needed, which can more accurately screen out non-normal distributed data points while testing the normality of the data, and improve the accuracy and reliability of data screening. Summary of the Invention
[0004] Aiming at the deficiencies of the existing data screening methods, the present invention proposes a data cleaning method for analyzing the angle of repose in discrete element simulation, which can more accurately screen out the data points that may have systematic measurement errors in the simulation angle of repose dataset to be subjected to the response surface design, improve the accuracy and reliability of the data, and provide a more reliable parameter basis for the subsequent response surface design.
[0005] The data cleaning method for analyzing the angle of repose in discrete element simulation includes the following steps: (1) Preprocessing of the discrete element simulation angle of repose dataset. First, duplicate data in the dataset is removed, and then the remaining sample data is sorted from smallest to largest to obtain the sample dataset u = {u1, u2, … u i , …, u n}, where i = 1~n; (2) Calculate the theoretical quantile Z (i) of each sample data. The calculation formula is: , where Φ -1 is the inverse function of the cumulative distribution function Φ(x) of the standard normal distribution; (3) Construct the reference line function y (i) , and the formula is: y (i) = σZ (i) + μ, where y (i) is the theoretical value of the i-th position that follows the normal distribution N(μ, σ 2 ), σ is the standard deviation of the sample data, , μ is the mean of the sample data, ; (4) Calculate the absolute value Δ i of the residual between each sample data and the corresponding theoretical value. The calculation formula is: Δ i = |y (i) - u i |; (5) Calculate the probability density function of the absolute value of the residual Δ i through kernel density estimation. The calculation formula is: ; where x is the continuous variable of the absolute value of the residual, , f(x) is the probability density function, K is the kernel function, and the Gaussian kernel function is selected, and h is the bandwidth; (6) Calculate the cumulative distribution function F(x) of the absolute value of the residual. The formula is , where Δ q is the threshold for determining outliers. If Δ i ≥ Δ q , then the original data u i is considered as outlier data.
[0006] Advantageous technical effects of the present invention: By combining Q-Q plot analysis and residual analysis calculations, and strictly calculating the threshold through the cumulative distribution function of the absolute value of the residual to screen outliers, it can efficiently and accurately identify and remove the non-normal distribution data points mixed in, avoiding the problems of high false screening rate and missed screening rate caused by subjective judgment in traditional methods, significantly improving the accuracy and reliability of the data, and providing a more reliable parameter basis for subsequent statistical analysis and modeling. Description of the Drawings
[0007] Figure 1 is the flowchart of the data cleaning method for the angle of repose analysis in discrete element simulation of the present invention; Figure 2 is the Q-Q plot of the cleaning effect of the sample data obtained by using the Gaussian kernel function in the present invention; Figure 3 is the residual distribution and KDE curve graph of the sample data in the present invention. Detailed implementation manners
[0008] Combined with the attached Figures 1-3 , illustrate the steps and specific operations of the present invention.
[0009] In view of the situation that the discrete element simulation angle of repose data set to be subjected to response surface design is mixed with data that may have systematic measurement errors, the present invention proposes a data cleaning method for the angle of repose analysis in discrete element simulation, which uses the Q-Q plot to test the normality of the data, calculates the residuals and combines the kernel density estimation method to accurately screen out the abnormal data points.
[0010] The data cleaning method for the angle of repose analysis in discrete element simulation includes the following steps: (1) Pretreatment of the discrete element simulation angle of repose data set. First, eliminate the duplicate data in the data set, and then sort the remaining sample data from small to large to obtain the sample data set u = {u1, u2,... u i ,..., u n}), where i = 1~n; (2) Calculate the theoretical quantile Z (i) of each sample data, and the calculation formula is: , where Φ -1 is the inverse function of the cumulative distribution function Φ(x) of the standard normal distribution; (3) Construct the reference line function y (i) , and the formula is: y (i) = σZ (i) + μ, where y (i) is the theoretical value of the i-th position obeying the normal distribution N(μ, σ 2 ), σ is the standard deviation of the sample data, , μ is the mean of the sample data, ; (4) Calculate the absolute value Δ i of the residual between each sample data and the corresponding theoretical value, and the calculation formula is: Δ i = |y (i) - u i |; (5) Calculate the probability density function of the absolute value of the residual Δ i through kernel density estimation, and the calculation formula is: ; In the formula, x is a continuous variable of the absolute value of the residual, , f(x) is the probability density function, K is the kernel function, and the Gaussian kernel function is selected , h is the bandwidth; (6) Calculate the cumulative distribution function F(x) of the absolute value of the residual, the formula is: , where Δ q is the threshold for determining abnormal values, if Δ i ≥Δ q , then it is considered that u i is abnormal data.
[0011] In the discrete element simulation software, different contact parameter combinations of particles are adjusted to conduct repose angle experiments, and the repose angle parameters of different accumulation bodies are measured. The following group of repose angles (38 in total) are used as examples to explain the operation process of this method in detail: {37.69,38.56,35.4,35.03,35.01,35.7,36.88,31.88,36.81,41.13,38.53,36.92,33.25,32.82,37.43,36.34,36.43,38.21,33.09,34.20,33.24,34.49,25.23,24.56,24.35,23.25,24.31,22.63,36.19,35.04,36.17,31.71,34.42,35.69,30.94,32.47,31.65,30.89} Specifically, the discrete element simulation repose angle analysis data cleaning process of the present invention is as follows: 1. Preprocessing of the discrete element simulation repose angle data set Preprocessing of the discrete element simulation repose angle data set is one of the key steps in the entire data screening method. Its purpose is to convert the original data into a form that is convenient for subsequent processing. Data preprocessing mainly includes the following two steps: (1) Eliminate duplicate data: By checking the duplicate values in the data set, duplicate data points are eliminated to ensure that each data point is unique. Since there are no duplicate values in the example angle of repose data set, proceed directly to the next step.
[0012] (2) Sorting: Sort the remaining sample data in ascending order to obtain the sample data set: u = {22.63, 23.25, 24.31, 24.35, 24.56, 25.23, 30.89, 30.94, 31.65, 31.71, 31.88, 32.47, 32.82, 33.09, 33.24, 33.25, 34.20, 34.42, 34.49, 35.01, 35.03, 35.04, 35.40, 35.69, 35.70, 36.17, 36.19, 36.34, 36.43, 36.81, 36.88, 36.92, 37.43, 37.69, 38.21, 38.53, 38.56, 41.13}.
[0013] The preprocessing of the data set is implemented through the numpy library and built-in functions in the Python program code. Specifically, the set function in Python is used to remove duplicate values from the sample data set, the len function is used to calculate the length n of the processed data set, the array function is used to convert the processed data set into a numpy array, and the sort function is used to sort the array to obtain the rank statistics of each data. The sorted data is convenient for subsequent calculation of theoretical quantiles and construction of reference line functions, making the data processing process more standardized.
[0014] 2. Calculate the theoretical quantile For the sorted sample data, calculate the theoretical quantile Z of each sample data (i) . The theoretical quantile is calculated based on the rank of the data and the inverse function of the cumulative distribution function of the normal distribution (standard normal quantile function). The specific calculation formula is:
[0015] In the formula, Φ -1 is the inverse function of the cumulative distribution function Φ(x) of the standard normal distribution; i is the rank of each sample data.
[0016] This formula converts the rank of the sample data into the corresponding normal distribution quantile, which is the basis for constructing the Q-Q plot and subsequent residual analysis.
[0017] It is implemented through the numpy and scipy libraries and built-in functions in the Python program code. Specifically, the arange function is used to generate an integer sequence i from 1 to n (including n), and (i - 0.375) / (n + 0.25) is named as the probability estimate p. This is based on Blom's quantile estimation method, which can better handle small sample data and make the calculation of theoretical quantiles more accurate. Then, the theoretical quantile of each rank is calculated through the norm.ppf(p) function.
[0018] The theoretical quantile set Z corresponding to the sample data set u is as follows: Z = {-2.14, -1.72, -1.49, -1.31, -1.17, -1.05, -0.94, -0.84, -0.75, -0.67, -0.59, -0.51, -0.44, -0.37, -0.30, -0.23, -0.16, -0.1, -0.03, 0.03, 0.10, 0.16, 0.23, 0.30, 0.37, 0.44, 0.51, 0.59, 0.67, 0.75, 0.84, 0.94, 1.05, 1.17, 1.31, 1.49, 1.72, 2.14} 3. Construct the reference line function The reference line in the Q-Q plot is an important basis for distinguishing whether the sample data conforms to the normal distribution. The reference line function y (i) = σZ (i) + μ reflects the linear relationship between the theoretical values of the normal distribution and the sample data. The construction of the reference line function mainly consists of two steps: (1) Calculate the standard deviation σ of the sample data. The calculation formula is: ; (2) Calculate the mean μ of the sample data. The calculation formula is: .
[0019] For the example angle of repose data set u, it is calculated that σ = 4.69 and μ = 33.38. That is, the reference line function is y (i) = 4.69Z (i) + 33.38.
[0020] The above steps are implemented through the numpy code library in the Python program code. Specifically, the std(data, ddof = 1) function in the numpy library is used to calculate the standard deviation of the data set u. The ddof = 1 parameter is used for Bessel correction to obtain an unbiased estimate that more accurately reflects the dispersion degree of the sample data. The mean(data) function is used to calculate its mean. The reference line function converts the theoretical quantile Z (i) into the theoretical value y 2 of the normal distribution N(μ, σ (i) ) that the sample data follows, which is used for subsequent comparison with the sample data and residual calculation.
[0021] 4. Calculate the absolute value of the residual Calculate the residual between the sample data and the reference line function, and take the absolute value to obtain the absolute value of the residual Δ i = |y (i) - u i|; The absolute value of the residual reflects the degree of difference between the sample data and the theoretical value of the normal distribution. A larger residual value may indicate that the corresponding sample data point deviates far from the normal distribution and may be an outlier of non-normal distribution.
[0022] This step is implemented through the built-in functions in the Python program code. The theoretical quantile Z of order i calculated in step 2 is brought into step 3 to obtain the theoretical value y (i) , bring y (i) and the sample data set u of rank i i are subtracted, and the absolute value is obtained through the abs function.
[0023] The results of the absolute values of the residuals of the example angle of repose data set u are as follows: Δ = {0.72, 2.05, 2.10, 2.87, 3.33, 3.23, 1.93, 1.52, 1.81, 1.47, 1.26, 1.50, 1.50, 1.44, 1.26, 0.95, 1.59, 1.50, 1.26, 1.47, 1.19, 0.88, 0.93, 0.90, 0.59, 0.72, 0.40, 0.19, 0.10, 0.11, 0.46, 0.88, 0.88, 1.19, 1.33, 1.83, 2.91, 2.28} 5. Calculate the probability density function of the absolute value of the residual through kernel density estimation Kernel density estimation (KDE) is a non-parametric method for estimating the probability density function of a random variable. KDE does not assume that the data follows a specific distribution form, but is directly constructed based on the sample data itself. The calculation formula is:
[0024] In the formula, x is the continuous variable of the absolute value of the residual, , f(x) is the probability density function, K is the kernel function, and h is the bandwidth.
[0025] The main steps to construct the probability density function through KDE include the following: (1) Select an appropriate kernel function. The main role of the kernel function K is to define the influence range and shape around each data point. Selecting an appropriate kernel function has an important impact on the accuracy and efficiency of density estimation.
[0026] The commonly used kernel functions include the following 4 categories: 1) Gaussian Kernel:
[0027] 2) Uniform Kernel:
[0028] 3) Triangular Kernel:
[0029] 4) Epanechnikov Kernel:
[0030] The II in the above-mentioned uniform kernel, triangular kernel, and Epanechnikov kernel function expressions is an indicator function.
[0031] Among them, the Gaussian kernel function has good robustness to outliers and can reduce the interference of local noise through smoothing; the uniform kernel is suitable for data distributions close to uniform distributions and is vulnerable to outliers for complex distributions; the triangular kernel is generally suitable for data with medium smoothing requirements, balancing local sensitivity and global stability; the Epanechnikov kernel is generally applicable to data distributions that are relatively steep and require highlighting local density changes.
[0032] The results of the four kernel functions in screening outliers in the example angle of repose dataset are as follows: The Gaussian kernel function screened out the outliers: 24.35, 24.56, 25.23, 38.56; the uniform kernel function screened out the outlier: 24.56; the triangular kernel function did not screen out any outliers; the Epanechnikov kernel function screened out the outliers: 24.56, 25.23.
[0033] Through experimental comparison, it is found that the uniform kernel function, triangular kernel function, and Epanechnikov kernel function have a relatively high rate of missing outliers, and the Gaussian kernel function has the best effect in screening outliers for probability density function estimation. For example, Figure 2 Figure (24) is the Q-Q plot of the screening results using the Gaussian kernel function. Among them, the black dots are normal points, the white dots are the screened outliers, and the dashed line is the reference line. It can be seen that the screened outliers are all points relatively far from the reference line. Figure 3 Figure (26) is the probability density curve graph and the histogram of the absolute value of residuals obtained by selecting the Gaussian kernel function for kernel density estimation. Among them, the dashed line represents the calculated outlier threshold, and the u corresponding to the absolute value of the residuals on the right side of the dashed line i are all considered outliers.
[0034] (2) Select an appropriate bandwidth. The parameter bandwidth h determines the width of the kernel function, thus affecting the smoothness of the density estimation. The selection of the bandwidth directly affects the bias and variance of the estimation result. Too large a bandwidth will cause the density estimation to be too smooth and mask the details of the data; while too small a bandwidth will cause the estimation result to be too fluctuating and capture the noise in the data. This method uses the improved Silverman's rule to calculate the bandwidth, and the calculation formula is:
[0035] In the formula, σ* is the standard deviation of Δ i , , μ* represents the mean of Δ i , , IQR is the sample interquartile range, IQR = Q3 - Q1, Q3 represents the value in Δ i that is greater than 75% of the data and less than 25% of the data, Q1 represents the value that is greater than 25% of the data and less than 75% of the data, and n is the number of samples.
[0036] For the residual absolute values of the example dataset, the calculation results are: μ* = 1.38, σ* = 0.79, IQR = 0.87, n = 38, h = 0.28 (3) Calculate the probability density function. According to the selected kernel function and the calculated bandwidth, using the mathematical formula of KDE, the values of all kernel functions at the evaluation points are weighted and averaged to obtain the density estimation value for each evaluation point x.
[0037] (4) Correct the boundary effect to avoid truncation error at the boundary. The boundary effect refers to the inaccurate situation that occurs in kernel density estimation when data points are close to the data boundary. This is because the influence range of the kernel function is truncated outside the data boundary, resulting in the estimated values near the boundary being affected by fewer kernel functions and thus deviating from the true density distribution. To solve the boundary effect problem of kernel density estimation, virtual data points can be introduced outside the data boundary by mirroring the data to correct the boundary effect and improve the accuracy of kernel density estimation.
[0038] (5) Visualize the estimated probability density function.
[0039] The above steps are implemented through the numpy, matplotlib, scipy, sklearn, and matplotlib code libraries and built-in functions in the Python program code. Specifically, first, the Δ i calculated in step 4 is converted into a NumPy array through the np.array function for convenient subsequent data operations. Since KDE requires the input data to be a two-dimensional array, then the one-dimensional array is converted into a two-dimensional array through reshape(-1,1). Next, the bandwidth is calculated. First, the 75th percentile and 25th percentile values of Δ i are subtracted to obtain the IQR through the np.percentile function, the standard deviation σ* of Δ i is calculated through the np.std function, the number of samples n of Δ i is calculated according to the len function, the minimum value of σ* and IQR / 1.34 is obtained through np.min, and then according to the formula Calculate the appropriate bandwidth. Next, call the KernelDensity function in the sklearn library to set the selected kernel type and bandwidth and name it kde for later use, and obtain the required probability density function through kde.fit(). To avoid large truncation errors at the boundaries, use the mirror reflection method to mirror the data at the boundaries, that is, symmetrically copy the data points outside the boundaries, repeat the above steps to obtain the probability density function of the mirrored data, add the two probability density functions and take the corresponding part of the original data to obtain the corrected probability density function. Next, visualize the estimated density function and draw the KDE curve through the plt function in the matplotlib library.
[0040] 6. Calculate the threshold for determining outliers and filter out the outlier data Calculate the cumulative distribution function F(x) of the absolute value of the residuals. The formula is , where Δ q is the threshold for determining outliers. If Δ i ≥Δ q , then it is considered that u i in the original data is outlier data.
[0041] Implement this step through the scipy code library in the Python program code. First, integrate the probability density function through the quad function to obtain the cumulative distribution function F(x), and then use the bisection method to obtain the threshold of the outlier at the significance level of α = 0.05, that is, the value of Δ q when F(x) = 0.95. Then determine np.abs(residuals)≥Δ q to generate a Boolean Mask for marking which data points are outliers, and obtain the outliers by returning the positions of all outliers in the data through np.where.
[0042] When the significance level is α = 0.05, according to the Gaussian kernel function, the Δ q of the absolute value of the residuals of the example data set is calculated to be 2.63. Therefore, 24.35, 24.56, 25.23, and 38.56 in the sample data set u corresponding to 2.87, 3.33, 3.23, and 2.91 in the set of absolute values of the residuals may be the results with systematic errors, and parallel experiments need to be carried out on the experiments corresponding to these several data.
[0043] To verify the advantages of this method in the cleaning of simulated angle of repose data, multiple personnel were recruited to visually identify the "outlier points" on the Q-Q plot together. The results showed that the "outlier point" results given by different personnel were not the same, and the results were highly subjective, with the situations of false screening and missed screening. This method obtains the threshold of the absolute value of the residual corresponding to the outlier through rigorous calculations, and screens the outliers in a quantitative manner, which is not only more objective, but also significantly improves the accuracy of the screening results.
Claims
1. A data cleaning method for analyzing the angle of repose in discrete element simulation, which is used to identify and eliminate abnormal data caused by measurement errors or experimental errors in the dataset, including the following steps: (1) Preprocessing the discrete element simulation angle of repose dataset; (2) Calculating the theoretical quantiles of each sample data; (3) Constructing a reference line function; (4) Calculating the absolute value of the residual between each sample data and the corresponding theoretical value; (5) Calculating the probability density function of the absolute value of the residual through kernel density estimation; (6) Calculating the cumulative distribution function of the absolute value of the residual to obtain the threshold of outliers and screening abnormal data.
2. The discrete element simulation angle of repose analysis data cleaning method according to claim 1, wherein The steps of step (1) for preprocessing the discrete element simulation angle of repose dataset are to first eliminate duplicate data in the dataset, and then sort the remaining sample data from smallest to largest to obtain a sample dataset.
3. The discrete element simulation repose angle analysis data cleaning method according to claim 1, wherein The calculation formula for the theoretical quantile of each sample data in step (2) is as follows: , where Z (i) is the theoretical quantile of the i-th sample data, n is the number of samples in the sample dataset, and Φ -1 is the inverse function of the cumulative distribution function Φ(x) of the standard normal distribution.
4. The discrete element simulation angle of repose analysis data cleaning method according to claim 1, wherein The formula of the reference line function in step (3) is: y (i) =σZ (i) +μ, where y (i) is the reference line function, which is the theoretical value of the i-th order that follows the normal distribution N(μ, σ 2 ), σ is the standard deviation of the sample data, , μ is the mean of the sample data, , and Z (i) is the theoretical quantile of the i-th sample data.
5. The discrete element simulation repose angle analysis data cleaning method according to claim 1, characterized in that The calculation formula for the absolute value of the residual in step (4) is: Δ i = |y (i) - u i |, where Δ i is the absolute value of the residual of the i-th sample rod data, y (i) is the theoretical value of the reference line function at the i-th position, and u i is the i-th sample data in the sample rod data set u.
6. The discrete element simulation repose angle analysis data cleaning method according to claim 1, wherein The calculation formula of the probability density function in step (5) is as follows: ; where x is a continuous variable of the absolute value of the residual, , Δ i is the absolute value of the residual of the i-th sample data, n is the number of samples in the sample dataset, f(x) is the probability density function, K is the kernel function, and the Gaussian kernel function is selected, and h is the bandwidth.
7. The discrete element simulation repose angle analysis data cleaning method according to claim 1, wherein The formula for calculating the cumulative distribution function of the absolute value of the residual in step (6) is , where F(x) is the cumulative distribution function of the absolute value of the residual, and Δ q is the threshold for determining outliers. If Δ i ≥Δ q , where i, q ∈ 1~n and n is the number of samples in the sample dataset, then it is considered that u i in the original data is abnormal data.