Automatic Correction and Distribution Fitting Method for Small-Batch Error Data
By using the Anderson–Darling inspection and cyclic estimation of small batch error data automatic correction and distribution fitting methods in the field of aeronautical manufacturing, the problem of small batch error data correction and distribution fitting is solved, and the data is quickly and efficiently corrected, which is suitable for small batch production processes.
Patent Information
- Application Number
- CN202210876577.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-25
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-07-25
AI Technical Summary
It is difficult for the prior art to accurately perform automatic correction and distribution fitting of small batch error data, especially when the error data may comply with truncated normal distribution, gamma distribution or t distribution.
Automatic correction and distribution fitting method of small batch error data based on Anderson–Darling test and cyclic estimation is used. By constructing Anderson–Darling test statistics under different continuous distributions, the statistical distribution type of error data is determined, and the data is automatically corrected using loop estimation until each data is optimally corrected or the p-value reaches convergence.
It realizes rapid and effective automatic correction of small batch error data, accurately obtains the statistical distribution characteristics of the data, and is suitable for various small batch production processes in the aviation manufacturing field, making up for data deviations caused by improper measurement methods or operator replacement.
Smart Images

Figure CN115455359B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of aviation production, and particularly relates to an automatic correction and distribution fitting method for small-batch error data based on Anderson–Darling test and iterative estimation. Background Art
[0002] As the top of the industrial system, the aviation industry has very strict control over product quality. The distribution characteristics of the error data between the measured value and the theoretical value of the product appearance reflect the quality information of the manufacturing process and are the basic basis for realizing statistical process control, optimization, and production management. However, due to factors such as improper measurement methods and operator replacement, the recorded values of error data often deviate from the true values, which especially brings greater challenges to the statistical distribution inference of small-batch error data. Therefore, it is of great significance to accurately correct small-batch error data automatically and analyze the statistical distribution characteristics of the error data.
[0003] Currently, most of the literature, such as "Application of Multivariate Statistical Process Control in the Reverse Flotation Production Process", "State Assessment of Wind Turbine Gearboxes by Fusing SCADA Data", "Research and Practice on Online Monitoring and Statistical Process Control in the Levelling Process", and "Comparative Analysis of Data Preprocessing Methods Based on Typical Datasets", mainly disclose that before performing statistical process control, abnormal data will be removed based on the quartile method, that is, the data beyond the upper and lower quartiles will be removed from the sample data. This method is applicable to the case of large samples. For small-batch processes, the further reduction of the sample size is not conducive to the fitting of the distribution. A more reasonable method is to find the true values of the data through data correction methods, so as to accurately obtain the statistical distribution characteristics of the data. However, the current data preprocessing methods are only limited to data normalization, standardization, and normalization. Among them, the normalization methods include Box-Cox transformation and Johnson transformation, etc., which can improve the normality and symmetry of the data; the standardization and normalization methods aim to make the data dimensionless through mathematical operations. These methods are only applicable to the case of normal distribution, while the actual manufacturing error data may also follow truncated normal distribution, gamma distribution, or t distribution, etc. Summary of the Invention
[0004] To overcome the problems existing in the above-mentioned prior art, the present invention proposes an automatic correction and distribution fitting method for small-batch error data based on Anderson–Darling test and cyclic estimation, constructs Anderson–Darling test statistics under different continuous distributions (normal distribution, truncated normal distribution, gamma distribution, t-distribution), determines the statistical distribution type of error data according to the p-values of the test statistics under different distributions, and based on this distribution type, randomly divides the data set into a historical set and an observation set by using cyclic estimation, selects the correction method with the highest overall distribution fitting p-value for the data in the observation set, and cyclically selects different data as the observation set until each data has obtained the optimal correction or the p-value reaches convergence, thus completing the automatic correction of the data.
[0005] To implement the above invention, the following technical solutions are provided:
[0006] An automatic correction and distribution fitting method for small-batch error data,
[0007] Specifically, it includes the following steps:
[0008] Step 1: Read the annual error data of the same feature of a certain small-batch production product from the record table;
[0009] Step 2: Remove the abnormal data from the error data to obtain the initial data set D = {x i , i = 1, …, n};
[0010] Step 3: Construct Anderson–Darling test statistics under four continuous distributions: normal distribution, truncated normal distribution, gamma distribution, and t-distribution;
[0011] Step 4: According to the limiting distribution of the Anderson–Darling test statistic A 2 , construct the p-values of the Anderson–Darling test statistics under four continuous distributions: normal distribution, truncated normal distribution, gamma distribution, and t-distribution, and determine the statistical distribution type of the error data by comparing the p-values. The larger the p-value, the higher the goodness of fit of the distribution, that is, the determined distribution type is j * = max j=1,2,3,4 p j ;
[0012] Step 5: Based on the obtained distribution type j * , perform automatic correction on each data by using cyclic estimation; pre-give a compensation value δ, randomly shuffle the data set D, and divide it into a historical set D 1 and an observation set D 2 ; stipulate the correction strategy of the data, and perform continuous iteration to obtain the final corrected data;
[0013] Step 6: Repeat Step 5 with different compensation values δ to find the optimal compensation value; under this compensation value, obtain the optimally corrected dataset D′ and use maximum likelihood estimation to solve for the parameters of distribution j * .
[0014] Further, in Step 3, to measure the goodness of fit between the true data distribution and the theoretical distribution, Anderson–Darling test statistics are constructed for four continuous distributions: normal distribution, truncated normal distribution, gamma distribution, and t-distribution as follows:
[0015]
[0016] where: represents the Anderson–Darling test statistic for the j-th hypothesized distribution, used to measure the gap between the hypothesized distribution and the true data distribution, the smaller it is, the closer the true data is to the hypothesized distribution, n is the number of samples, and F D (x) is the distribution function of the samples;
[0017] The normal distribution, truncated normal distribution, gamma distribution, and t-distribution are the four distributions that best fit the distribution type of error data in the aviation manufacturing field, and F j (x) is the theoretical distribution function of the j-th hypothesized distribution:
[0018]
[0019] where: Γ represents the gamma function, and μ, σ, a, b, α, β, v represent distribution coefficients related to the distribution.
[0020] Further, the specific method of Step 4 is as follows:
[0021] According to the limiting distribution of the Anderson–Darling test statistic A 2 , construct the p-values for the four distributions through j = 1, 2, 3, 4 as follows:
[0022]
[0023] p j represents the p-value of the Anderson–Darling test statistic for the j-th hypothesized distribution. The p-value is a probability value between 0 and 1, which qualitatively and intuitively represents the goodness of fit between the true data distribution and the theoretical distribution. The larger the p-value, the smaller it is, the higher the goodness of fit of the distribution. Determine the statistical distribution type of the error data by comparing the p-values, that is, the determined distribution type is j * = maxj=1,2,3,4 p j 。
[0024] Furthermore, step 5 is based on the obtained distribution type j * , and uses cyclic estimation to automatically correct each data item.
[0025] Still further, step 5 specifically includes the following steps:
[0026] Step 501: Predetermine a compensation value δ, and it is recommended that this compensation value be set to an integer multiple of the data recording accuracy;
[0027] Step 502: Randomly shuffle the data set D after the (r - 1)-th cycle correction r-1 , and divide it into a historical set and an observation set according to a ratio of 8:2 and
[0028] Step 503: There are three data correction strategies: subtract the compensation value remain unchanged and add the compensation value For each data item in the observation set, combine this data item with the historical set to form a new data set, and calculate the p-values of the three correction strategies under the distribution type j * , and denote them as Select the correction method with the highest p-value to correct x. For example, if then Repeat this step for all other data items in the observation set, and finally obtain the corrected observation set
[0029] Step 504: Denote the data set after the r-th cycle correction as and its p-value as p r ;
[0030] Step 505: Compare the magnitudes of p r and p r-1 . If p r - p r-1 < 0.001, then end the correction; otherwise, let r = r + 1, and repeat steps 502 - 504.
[0031] Compared with the prior art, the advantages of the present invention are as follows:
[0032] This method can automatically search for the optimal correction value, is applicable to various small-batch production processes in the aviation manufacturing field, can quickly and effectively compensate for data deviations caused by improper measurement methods or operator replacements, make the data set closer to the true distribution of the data, and at the same time provide a basis for subsequent statistical process control and the construction of control charts. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a flowchart for automatic correction process.
[0034] Figure 2 It is the fitting situation of the original data under four distributions.
[0035] Figure 3 It is the gamma distribution p-value of the corrected dataset under different compensation values.
[0036] Figure 4 It is the gamma distribution fitting situation of the corrected dataset under the optimal compensation value (δ = 0.006).
[0037] Figure 5 It is the maximum likelihood estimation of the gamma distribution parameters. Detailed implementation manners
[0038] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are for explaining the present invention rather than limiting the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] The following describes the specific implementation methods of the present invention with reference to the accompanying drawings and examples. The present invention is not limited to this embodiment.
[0040] Embodiment 1
[0041] As Figure 1 shown, the automatic correction and distribution fitting method for small-batch error data
[0042] specifically includes the following steps:
[0043] Step 1: Read the annual error data of the same feature of a certain small-batch production product from the record table;
[0044] Step 2: Use the prior information of the production process to clear the abnormal data in the error data, and clear the data that does not meet the technical requirements of the product characteristics (the data corresponding to unqualified products) to obtain the initial dataset D = {x i , i = 1,..., n};
[0045] Step 3: According to the basic knowledge of statistics, many random variables follow a normal distribution, such as measurement errors, product weights, human heights, etc. Therefore, it is generally defaulted in the production site that the error data follows a normal distribution, and the data is analyzed under the background of the normal distribution. However, due to factors such as improper measurement methods and operator replacements, the recorded values of the error data often deviate from the true values, and the presented data may not necessarily follow a normal distribution. After actual verification of the on-site data, the error data is most likely to follow one of the four distributions: normal distribution, truncated normal distribution, gamma distribution, and t-distribution. In order to more accurately determine the actual distribution of the data, Anderson–Darling test statistics are constructed under four continuous distributions: normal distribution, truncated normal distribution, gamma distribution, and t-distribution;
[0046] Step 4: Based on the limiting distribution of the Anderson–Darling test statistic A 2 , the p-values of the Anderson–Darling test statistics under four continuous distributions: normal distribution, truncated normal distribution, gamma distribution, and t-distribution are constructed. By comparing the magnitudes of the p-values, the statistical distribution type of the error data is determined. The larger the p-value, the higher the goodness of fit of the distribution, that is, the determined distribution type is j * = max j=1,2,3,4 p j ;
[0047] Step 5: Based on the obtained distribution type j * , cyclic estimation is used to automatically correct each data; a compensation value δ is given in advance, and the data set D is randomly shuffled and divided into a historical set D 1 and an observation set D 2 ; the correction strategy of the data is specified, and continuous iteration is performed to obtain the final corrected data;
[0048] Step 6: Repeat Step 5 with different compensation values δ to find the optimal compensation value; under this compensation value, the optimal corrected data set D′ is obtained and the parameters of the distribution j * are solved using maximum likelihood estimation.
[0049] Furthermore, in the above Step 3, in order to measure the goodness of fit between the true data distribution and the theoretical distribution, the Anderson–Darling test statistics under four continuous distributions: normal distribution, truncated normal distribution, gamma distribution, and t-distribution are constructed as follows:
[0050]
[0051] In the formula: represents the Anderson–Darling test statistic of the j-th hypothesized distribution, which is used to measure the gap between the hypothesized distribution and the true data distribution, The smaller it is, the closer the real data is to the assumed distribution. Here, n is the number of samples, and F D (x) is the distribution function of the samples;
[0052] The normal distribution, truncated normal distribution, gamma distribution, and t-distribution are the four distributions that best fit the distribution type of error data in the field of aviation manufacturing. F j (x) is the theoretical distribution function of the j-th assumed distribution:
[0053]
[0054] In the formula: Γ represents the gamma function, and μ, σ, a, b, α, β, v represent the distribution coefficients related to the distribution.
[0055] Further, the specific method of step 4 is as follows:
[0056] According to the limiting distribution of the Anderson–Darling test statistic A 2 and through j = 1, 2, 3, 4, construct the p-values of the four distributions as follows:
[0057]
[0058] p j represents the p-value of the Anderson–Darling test statistic of the j-th assumed distribution. The p-value is a probability value between 0 and 1, which qualitatively and intuitively represents the goodness of fit between the real data distribution and the theoretical distribution. The larger the p-value indicates the smaller it is, the higher the goodness of fit of the distribution. By comparing the p-values, determine the statistical distribution type of the error data, that is, the determined distribution type is j * = max j=1,2,3,4 p j .
[0059] Further, step 5 is based on the obtained distribution type j * , and uses cyclic estimation to automatically correct each data. This method can quickly and effectively compensate for the deviation caused by improper measurement methods or operator replacement, making the data set closer to the real distribution of the data. And this method is applicable to various small-batch production processes in the field of aviation manufacturing, and can accurately perform automatic correction on small-batch error data, while the existing methods are more applicable to the case of large samples.
[0060] Still further, step 5 specifically includes the following steps:
[0061] Step 501: Predetermine a compensation value δ, and it is recommended that this compensation value be set as an integer multiple of the data recording accuracy;
[0062] Step 502: Randomly shuffle the dataset D corrected in the (r - 1)-th loop, and divide it into a historical set and an observation set according to the ratio of 8:2. r-1 Randomly shuffle it and divide it into a historical set and an observation set
[0063] Step 503: There are three data correction strategies: subtracting the compensation value, remaining unchanged, and adding the compensation value. For each data in the observation set, combine this data with the historical set to form a new dataset, and calculate the p-values of the three correction strategies under the distribution type j, denoted as * respectively. Select the correction method with the highest p-value to correct x. For example, if then Then Repeat this step for all other data in the observation set, and finally obtain the corrected observation set.
[0064] Step 504: Denote the dataset corrected in the r-th loop as and its p-value as p. r ;
[0065] Step 505: Compare the magnitudes of p r and p r-1 . If p r - p r-1 < 0.001, then end the correction; otherwise, let r = r + 1, and repeat Steps 502 - 504.
[0066] Example 2
[0067] An automatic correction and distribution fitting method for small-batch error data based on Anderson–Darling test and cyclic estimation constructs Anderson–Darling test statistics under four distributions (normal distribution, truncated normal distribution, gamma distribution, t-distribution) based on the error data distribution type (normal distribution, truncated normal distribution, gamma distribution, t-distribution) that best fits the small-batch production process in the aviation manufacturing field. Through the goodness-of-fit test, the distribution type of the data is determined. Cyclic estimation is used to automatically correct each data, which can quickly and effectively compensate for data deviations caused by improper measurement methods or operator replacements, making the dataset closer to the true distribution of the data, and at the same time providing a basis for subsequent statistical process control and the construction of control charts.
[0068] The Anderson–Darling test statistic was constructed under the error data distribution type (normal distribution, truncated normal distribution, gamma distribution, t distribution) that best fits the small batch production process in the aviation manufacturing field. After actual verification with field data, the error data is most likely to obey one of the four distributions: normal distribution, truncated normal distribution, gamma distribution, and t distribution, rather than blindly assuming that the error data obeys the normal distribution and performing analysis based on basic statistical knowledge. The reason is that the recorded values of the error data at the production site may often deviate from the true value due to factors such as improper measurement methods and operator changes, and the presented data may not necessarily obey the normal distribution.
[0069] The method is applicable to small batch production processes in the field of aviation manufacturing. The true value of the data is found through the data correction method of cyclic estimation, so as to accurately obtain the statistical distribution characteristics of the data. At present, most of the literature will eliminate abnormal data based on the quartile method, that is, data exceeding the upper and lower quartiles are eliminated from the sample data. This method is suitable for large samples. For various small batch processes in the field of aviation manufacturing, further reduction of sample size is not conducive to the fitting of the distribution. A more reasonable method is to find the true value of the data through the data correction method, so as to accurately obtain the statistical distribution characteristics of the data.
[0070] The cyclic estimation is used to automatically correct each data. A compensation value δ is given in advance, and the data set D is randomly shuffled and divided into the historical set D 1 and observation set D 2 Three correction strategies for data are specified, and the final correction data is obtained through continuous iteration. This method can quickly and effectively compensate for the deviation caused by improper measurement methods or operator changes, making the data set closer to the real distribution of the data.
[0071] A method for automatic correction and distribution fitting of small batch error data, the process framework is as follows Figure 1 As shown, the following steps are included:
[0072] Step 1: Read the annual error data of the same feature of a small batch production product from the record table;
[0073] Step 2: Use the prior information of the production process to remove abnormal data from the error data and obtain the initial data set D = {x i ,i=1,…,n}. For example, to avoid scrapping of parts, some processing procedures will only allow for dimensional surplus, that is, the error data must be positive. In this case, very few records with negative numbers can be deleted;
[0074] Step 3: Construct the Anderson–Darling test statistics under four continuous distributions: normal distribution, truncated normal distribution, gamma distribution, and t distribution as follows:
[0075]
[0076] Where: The Anderson–Darling test statistic represents the jth hypothesized distribution, which measures the gap between the hypothesized distribution and the true distribution of the data. n is the number of samples, and F D (x) is the distribution function of the sample, F j (x) is the theoretical distribution function of the jth hypothesis distribution as follows:
[0077]
[0078]
[0079] Where: Γ represents the gamma function, μ, σ, a, b, α, β, v represent the distribution coefficients related to the distribution.
[0080] Step 4: According to the Anderson–Darling test quantity A 2 The limiting distribution of j=1,2,3,4 The p-values of the four distributions are calculated as follows:
[0081]
[0082] The larger the p-value is, the better the distribution fit is. By comparing the p-values, the statistical distribution type of the error data can be determined.
[0083] Step 5: Based on the distribution type, each data is automatically corrected using cyclic estimation. This method can quickly and effectively compensate for deviations caused by improper measurement methods or operator changes, making the data set closer to the true distribution of the data.
[0084] The step 5 comprises the following steps:
[0085] Step 501: A compensation value δ is given in advance. It is recommended that the compensation value be set to an integer multiple of the data recording accuracy.
[0086] Step 502: The dataset D after the r-1th cycle correction r-1 Randomly shuffle and divide it into historical sets in a ratio of 8:2 and observation set
[0087] Step 503: There are three data correction strategies: minus compensation value Keep it the same And plus compensation value For each data in the observation set, merge the data with the historical set into a new data set and calculate the distribution type j. * The p-values of the following three correction strategies are denoted as Select the correction method with the highest p value to correct x. For example, if but Repeat this step for all other data in the observation set, and finally obtain the corrected observation set
[0088] Step 504: The data set after the rth cycle correction is recorded as Its p value is p r .
[0089] Step 505: Compare p r With p r-1 The size of p r -p r-1 <0.001, the calibration is terminated; otherwise, r=r+1 is set and steps 502-504 are repeated.
[0090] Step 6: Set different compensation values δ and repeat step 5 to find the optimal compensation value. Under this compensation value, the optimal corrected data set d′ is obtained and the distribution parameters are solved using maximum likelihood estimation.
[0091] Case analysis:
[0092] The measured error data of the quality outer circle features of a certain aviation enterprise in Chengdu, Sichuan Province in the past year are selected. The small batch data set has a total of 60. Since the outer circle error cannot be a negative number, the two negative numbers are first eliminated to obtain the eliminated data set d = {x i ,i=1,…,58}, as shown in Table 1 below.
[0093] Table 1 Original dataset D
[0094] -0.010 -0.004 0 0 0 0 0 0 0.001 0.001 0.001 0.001 0.001 0.002 0.002 0.002 0.002 0.002 0.002 0.002 0.002 0.003 0.003 0.003 0.003 0.004 0.004 0.004 0.004 0.004 0.005 0.005 0.005 0.005 0.005 0.005 0.006 0.007 0.007 0.009 0.010 0.010 0.010 0.010 0.010 0.010 0.011 0.012 0.012 0.013 0.015 0.015 0.015 0.015 0.016 0.020 0.020 0.020 0.020 0.030
[0095] Four continuous distributions, normal distribution, truncated normal distribution, gamma distribution, and t distribution, are used to fit the distribution of D. The results are as follows: Figure 2 As shown in Table 2, the Anderson–Darling test statistics under the four distributions are constructed and their p values are calculated, as shown in Table 2 below, where the p values corresponding to the gamma distribution are 3 =0.326, so the data set is considered to obey the gamma distribution.
[0096] Table 2 P-values of Anderson–Darling test statistics under four distributions
[0097]
[0098] Based on the gamma distribution, the compensation value δ=0.0001a, a=1,…,10 is set according to the multiple of 0.0001 of the accuracy of the error data. For each compensation value, the data is automatically corrected using cyclic estimation. That is, the data set is randomly divided into a historical set and an observation set, and the correction method with the highest p-value of the overall distribution fit is selected for the data in the observation set. Different data are cyclically selected as the observation set until each data is optimally corrected or the p-value converges, completing the automatic correction of the data. Calculate the p-value of the gamma distribution of the corrected data set under different compensation values, such as Figure 3 As shown in Figure 3, the p value is relatively stable when the compensation value is between 0.002 and 0.006, and the p value is the highest when δ=0.006, so the optimal compensation value is 0.006. Based on the compensation value, the corrected data set D′ is obtained as shown in Table 3. According to statistics, a total of 4 samples implement the plus compensation value strategy, a total of 20 samples implement the unchanged strategy, and a total of 24 samples implement the minus compensation value strategy.
[0099] Table 3 Corrected data set D′ when δ=0.006
[0100] Delete Delete 0 0 0 0.001 0.001 0.001 0.002 0.002 0.002 0.002 0.003 0.003 0.003 0.004 0.004 0.004 0.004 0.005 0.005 0.005 0.005 0.005 0.005 0.006 0.006 0.006 0.006 0.007 0.007 0.007 0.007 0.007 0.008 0.009 0.009 0.010 0.010 0.010 0.010 0.010 0.010 0.010 0.011 0.012 0.013 0.015 0.015 0.015 0.015 0.015 0.015 0.017 0.017 0.020 0.020 0.025 0.025 0.035
[0101] After correction, the data set D ′ Fitting the gamma distribution, the results are as follows Figure 4 As shown in , its p value increases from the original 0.326 to 0.9133. Finally, the maximum likelihood estimation method is used to estimate the gamma distribution parameters. The results are as follows Figure 5 As shown, D′={x′|x′~Gamma(1.257,0.0063)}, this statistical distribution can be used for subsequent statistical process control and quality monitoring.
Claims
1. Automatic correction and distribution fitting method for small-batch error data, Characterized in that: Specifically includes the following steps: Step 1: Read the annual error data of the same feature of a small batch of products from the record table; Step 2: Clear the abnormal data from the error data to obtain the initial dataset D = [x i , = 1, …,}; Step 3: Constructed Anderson–Darling test statistics under four continuous distributions: normal distribution, truncated normal distribution, gamma distribution, and t-distribution; Step 4: Based on the limiting distribution of the Anderson–Darling test statistic A 2 the p-values of the Anderson–Darling test statistic under four continuous distributions, namely the normal distribution, truncated normal distribution, gamma distribution, and t-distribution, are constructed. By comparing the magnitudes of the p-values, the statistical distribution type of the error data is determined. The larger the p-value, the higher the goodness-of-fit of the distribution, that is, the determined distribution type is j * = max j=1,2,3,4 p j ; Step 5: Based on the obtained distribution type j * , perform automatic correction on each data using cyclic estimation; pre-give a compensation value δ, and randomly shuffle the dataset D and divide it into a historical set D 1 and an observation set D 2 ; stipulate the correction strategy of the data, perform continuous iteration to obtain the final corrected data; Step 6: Repeat Step 5 with different compensation values δ to find the optimal compensation value; under this compensation value, obtain the optimally corrected dataset D′ and solve for the parameters of distribution j using maximum likelihood estimation * ; The specific steps of said Step 5 are as follows: Step 501: Predetermine a compensation value δ, and it is recommended that this compensation value be set as an integer multiple of the data recording accuracy; Step 502: Randomly shuffle the dataset D corrected in the (r - 1)-th loop, and divide it into a historical set and an observation set according to the ratio of 8:2 r-1 Step 503: There are three data correction strategies: subtracting the compensation value remaining unchanged and adding the compensation value For each data in the observation set, combine this data with the historical set to form a new data set, and calculate the p-values of the three correction strategies under distribution type j * which are respectively denoted as Select the correction method with the highest p-value to correct x, and repeat this step for all other data in the observation set. Finally, obtain the corrected observation set Step 504: Denote the dataset after the r-th cycle correction as and its p-value as p r ; Step 505: Compare p r with p r-1 . If p r - p r-1 < 0.001, end the calibration; otherwise, let r = r + 1 and repeat steps 502 - 504.
2. The automatic correction and distribution fitting method for small-batch error data according to claim 1, Characterized in that: In said Step 3, in order to measure the goodness of fit between the true data distribution and the theoretical distribution, the Anderson–Darling test statistics under four continuous distributions: normal distribution, truncated normal distribution, gamma distribution, and t-distribution are constructed as follows: Wherein: represents the Anderson–Darling test statistic for the j-th hypothesized distribution, which is used to measure the gap between the hypothesized distribution and the true distribution of the data, the smaller value indicates that the true data is closer to the hypothesized distribution, n is the number of samples, and F D (x) is the distribution function of the samples; The normal distribution, truncated normal distribution, gamma distribution, and t-distribution are the four distributions that best fit the distribution type of error data in the aviation manufacturing field. F j (x) is the theoretical distribution function of the j-th hypothesized distribution: In the formula: Γ represents the gamma function, and μ, σ, a, b, α, β, v represent distribution coefficients related to the distribution.
3. The automatic correction and distribution fitting method for small-batch error data according to claim 2, Characterized in that: The specific method of said Step 4 is as follows: According to the limiting distribution of the Anderson–Darling test statistic A 2 , the p-values for the four distributions are constructed as follows by : p j The p-value of the Anderson–Darling test statistic representing the j-th hypothesized distribution. The p-value is a probability value between 0 and 1, which qualitatively and intuitively represents the goodness of fit between the true data distribution and the theoretical distribution. The larger the p-value, the smaller it is, and the higher the goodness of fit of the distribution. By comparing the p-values, the statistical distribution type of the error data is determined, that is, the determined distribution type is j * = max j=1,2,3,4 p j .
4. The automatic correction and distribution fitting method for small-batch error data according to claim 3, Characterized in that: The said step 5 is based on the obtained distribution type j * , and automatic correction is performed on each data by using cyclic estimation.
Citation Information
Patent Citations
Likelihood ratio test error detection method
CN101894215A
Method and apparatus for evaluating data and implementing training based on the evaluation of the data
US20040015329A1