A method for batch statistical analysis of precision data based on data processing
By constructing a normal distribution curve and transforming the data, the precision data of non-normally distributed data were filtered and optimized, achieving more accurate batch statistics, solving the problem of abnormal results in existing technologies, and improving the reliability of data analysis.
Patent Information
- Application Number
- CN202511194650.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-08-26
AI Technical Summary
When processing non-normally distributed precision data, existing technologies are prone to batch statistical results being affected by extreme values, leading to low accuracy and abnormal results.
By constructing a normal distribution curve and dividing it into two sub-curves, analyzing its characteristics, filtering data sets that do not conform to the normal distribution, and using data transformation methods to make them approximate the normal distribution, a histogram curve is constructed for accurate analysis, and the optimal batch statistical method is adaptively obtained.
It improves the accuracy and robustness of batch statistics, reduces sensitivity to extreme values, and provides a more reliable basis for decision-making.
Smart Images

Figure CN120763455B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method for batch statistical analysis of high-precision data based on data processing. Background Technology
[0002] Precision data is a collection of data used to quantify the closeness between measurement results and true values in systems, equipment, and other processes. Its core objective is to assess and control the error range, ensuring the reliability of the data results. Batch statistical analysis of precision data can reveal hidden patterns within the data, providing a scientific basis for quality control, system optimization, and decision-making. Batch statistical processing methods typically need to be highly efficient, capable of summarizing the entire dataset using a single statistic. For example, the mean and standard deviation can fully describe the distribution of the current data set without the need for clustering or complex processing. Simultaneously, batch statistics enable rapid and concise anomaly detection and handling of precision data; for example, by using the mean... "(like (Principle) can quickly identify outliers, thereby enabling accurate and effective batch statistics of data.
[0003] For existing technologies, precision data that follows a normal distribution is theoretically the most suitable for the above-mentioned batch statistical processing. However, precision data that does not follow a normal distribution may lead to distorted results when processed by the above-mentioned batch statistical processing due to violations of statistical assumptions or obscuring of internal structure. That is, the calculated statistical results may be affected by extreme values and become abnormal, making it impossible to guarantee the batch statistical effect. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method for batch statistical analysis of precision data based on data processing, in order to solve the problems of low accuracy and abnormal results of precision data under batch statistical processing.
[0005] This invention provides a method for batch statistical analysis of precision data based on data processing, the method comprising the following steps:
[0006] Obtain the set of precision data to be analyzed in batches;
[0007] A normal distribution curve for the precision data set is constructed based on the mean and standard deviation of the data set. The normal distribution curve is divided into two sub-curves based on the axis of symmetry. Based on the changing trends of the two sub-curves, a coarse analysis of the normal distribution characteristics of the normal distribution curve is performed to obtain the first normal similarity.
[0008] If the first normal similarity is less than the preset coarse similarity threshold, then at least two data transformation methods are used to transform the precision data set to obtain multiple transformed data sets. The skewness of each transformed data set is calculated, and the optimal transformed data set is obtained from all transformed data sets based on all skewnesses.
[0009] A histogram curve of the optimal transformed data set and a target normal distribution curve are constructed. The horizontal axis of the histogram curve represents the transformed data, and the vertical axis represents the frequency of the transformed data. The histogram curve is used to perform a fine analysis of the normal distribution characteristics of the target normal distribution curve to obtain a second normal similarity. Based on the second normal similarity, an optimal batch statistical method is adaptively obtained, and the optimal batch statistical method is used to perform batch statistics on the optimal transformed data set.
[0010] Preferably, the step of performing a coarse analysis of the normal distribution characteristics of the normal distribution curve based on the changing trends of the two sub-curves to obtain a first normal similarity includes:
[0011] The slope of each sub-curve is obtained, and the sum of the slopes is obtained. The sum of the slopes is then normalized inversely using an exponential function with the natural constant as the base, to obtain the mirror feature value.
[0012] The difference between the maximum and minimum values on each sub-curve is obtained, and the absolute value of the difference is obtained. The absolute value of the difference is then inversely normalized using an exponential function with the natural constant as the base, to obtain the similarity value of the curve fluctuation amplitude.
[0013] The first normal similarity is obtained by weighted summation of the mirror feature value and the curve fluctuation amplitude similarity value.
[0014] Preferably, after obtaining the first normal similarity, the method further includes:
[0015] If the first normal similarity is greater than or equal to the preset coarse similarity threshold, then the mean and standard deviation are used to perform batch statistics on the precision data set.
[0016] Preferably, the data transformation methods include: logarithmic transformation, square root transformation, Yeo-Johnson transformation, and exponential transformation.
[0017] Preferably, the step of obtaining the optimal transformed data set from all transformed data sets based on all skewnesses includes:
[0018] Calculate the absolute value of the difference between each skewness and the constant 0, and combine the transformed datasets of the skewness corresponding to the smallest absolute value of the difference into the optimal transformed dataset.
[0019] Preferably, the step of performing a fine analysis of the normal distribution characteristics of the target normal distribution curve using the histogram curve to obtain the second normal similarity includes:
[0020] Calculate the difference in the ordinate between the i-th data point in the histogram curve and the i-th data point in the target normal distribution curve. Based on the difference in the ordinate of each data point in the histogram curve, obtain the corresponding mean square error. Use an exponential function with the natural constant as the base to inversely normalize the mean square error to obtain the curve overlap.
[0021] Obtain the mean value corresponding to the target normal distribution curve. and standard deviation ,statistics The number of data within the set is calculated as a percentage of the total number of data in the optimal transformed data set. The absolute value of the difference between the percentage and the theoretical percentage under the standard normal distribution is obtained. The absolute value of the difference is then inversely normalized using an exponential function with the natural constant as the base, to obtain the standard normal distribution similarity.
[0022] The second normal similarity is obtained by weighted summation of the curve overlap and the standard normal distribution similarity.
[0023] Preferably, the step of adaptively obtaining the optimal batch statistics method based on the second normal similarity includes:
[0024] If the second normal similarity is greater than or equal to the preset precision similarity threshold, the optimal batch statistical method includes the mean and standard deviation; if the second normal similarity is less than the preset precision similarity threshold, the optimal batch statistical method includes the median, trimmed mean, and interquartile range.
[0025] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows:
[0026] This invention obtains a precision data set to be batch statistically analyzed; constructs a normal distribution curve for the precision data set based on the data mean and standard deviation of the precision data set; divides the normal distribution curve into two sub-curves based on the axis of symmetry; performs a coarse analysis of the normal distribution characteristics of the normal distribution curve based on the changing trends of the two sub-curves to obtain a first normal similarity; if the first normal similarity is less than a preset coarse similarity threshold, then uses at least two data transformation methods to transform the precision data set to obtain multiple transformed data sets; calculates the skewness of each transformed data set; and obtains the optimal transformed data set from all transformed data sets based on all skewnesses; constructs a histogram curve of the optimal transformed data set and a target normal distribution curve, where the horizontal axis of the histogram curve represents the transformed data and the vertical axis represents the frequency of the transformed data; performs a fine analysis of the normal distribution characteristics of the target normal distribution curve using the histogram curve to obtain a second normal similarity; and adaptively obtains the optimal batch statistical method based on the second normal similarity, and performs batch statistical analysis on the optimal transformed data set using the optimal batch statistical method. The process involves obtaining a first normal similarity to filter out precision data sets that do not conform to the characteristics of a normal distribution. Then, by performing optimal data transformation on these precision data sets that do not conform to the characteristics of a normal distribution, the data is made to be closer to a normal distribution or to better match existing batch statistical methods. Therefore, by analyzing the second normal similarity of the optimally transformed data set, a more stringent distribution characteristic is used for evaluation, which increases the accuracy of the analysis and reduces the inclusiveness. This makes the calculation of the second normal similarity more robust, and the optimal batch statistical method adaptively obtained based on the second normal similarity can significantly improve the batch statistical quality of skewed normal precision data, providing a more reliable basis for decision-making. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a flowchart of a method for batch statistical analysis of precision data based on data processing, provided in Embodiment 1 of the present invention.
[0029] Figure 2 This is a schematic diagram of a normal distribution curve divided into two sub-curves according to an embodiment of the present invention. Detailed Implementation
[0030] Embodiments of this disclosure are described in detail below, with examples of these embodiments illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting it.
[0031] It should be noted that the terms "first," "second," etc., used in this disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure.
[0032] To illustrate the technical solution of the present invention, specific embodiments are described below.
[0033] See Figure 1 This is a flowchart of a method for batch statistical analysis of precision data based on data processing, as provided in Embodiment 1 of the present invention. Figure 1 As shown, the method may include:
[0034] Step S101: Obtain the set of precision data to be statistically analyzed in batches.
[0035] There are many scenarios for high-precision data, and the acquisition method needs to be selected according to the specific data source and data type. See the following example for details:
[0036] 1. Data Collection
[0037] Sensor data acquisition; applicable scenarios include industrial manufacturing, environmental monitoring, and other situations requiring real-time acquisition of physical parameters.
[0038] Log file collection: Applicable scenarios include system monitoring, troubleshooting, and other scenarios that require analysis of device or system operation logs.
[0039] API interface acquisition: Applicable scenarios are those that require obtaining data from third-party systems or services, such as calling the weather forecast API to obtain real-time weather data.
[0040] 2. Data Preprocessing
[0041] Data cleaning: This involves removing missing, duplicate, or outlier values from precision data to initially improve its accuracy; it also converts the data into a format suitable for analysis and eliminates the influence of units, such as data normalization and data discretization.
[0042] 3. Data Integration
[0043] Merging data from multiple data sources resolves naming conflicts and redundancy issues. For example, entity identification: standardizing the naming of the same entity across different data sources (e.g., "Customer ID" and "User Number"). Redundancy detection: identifying and removing redundant attributes through correlation analysis (e.g., Pearson coefficient). Numerical conflict resolution: unifying the scaling standards or encoding methods across different data sources (e.g., date format, currency unit).
[0044] Using the above method, at least one set of precision data to be statistically analyzed in batches can be obtained.
[0045] Step S102: Construct a normal distribution curve for the precision dataset based on the mean and standard deviation of the dataset. Divide the normal distribution curve into two sub-curves based on the axis of symmetry. Perform a coarse analysis of the normal distribution characteristics of the normal distribution curve based on the changing trends of the two sub-curves to obtain the first normal similarity.
[0046] The fluctuations in many precision data (such as dimensional tolerances, form and position errors, and surface roughness) are inherently random and independent, and they usually conform to a normal distribution. However, when precision data is affected by multiple independent random factors (such as machining errors, material fluctuations, and environmental interference), regardless of the distribution of these factors themselves, their sum will approach a normal distribution (Central Limit Theorem) when the sample size is large enough. For example, in machining, factors such as tool wear, machine tool vibration, and temperature changes act independently, and the theoretical dimensional error of the final product follows a normal distribution. However, in actual processing, the size of each batch of precision data sets to be statistically analyzed, as well as the influence of outliers or extreme values (such as abnormal dimensions caused by sudden tool breakage during machining, or systematically larger dimensions due to long-term machine tool wear), may cause the batch of precision data sets to not all exhibit characteristics similar to a normal distribution. If traditional batch statistical methods are used for this type of data, it will have an adverse effect on the results and cause serious errors.
[0047] Therefore, based on the above feature analysis, this embodiment of the invention constructs a normal distribution similarity scoring model to analyze and filter the precision data set to be batch statistically analyzed, in order to obtain precision data sets that do not conform to the characteristics of a normal distribution, and then obtains the optimal batch statistical method through data transformation for batch statistical analysis. Specifically, for analyzing and filtering the precision data set to be batch statistically analyzed by constructing a normal distribution similarity scoring model, the mean of the precision data set is first obtained. and data standard deviation Substituting this into the probability density function of the normal distribution, we construct the normal distribution curve of the precision dataset. The formula for calculating the probability density function is:
[0048]
[0049] It should be noted that the construction of the normal distribution curve is a prior art, and its probability density function is also a prior art, so it will not be elaborated here.
[0050] Then, using the axis of symmetry (i.e., the center point of the horizontal axis of the normal distribution curve) as the dividing line, the normal distribution curve is divided into two sub-curves, denoted as F1 and F2 respectively. Finally, based on the trend characteristics of the two sub-curves, a normal distribution similarity scoring model is constructed to obtain the first normal similarity of the precision dataset. The specific method for obtaining the first normal similarity is as follows:
[0051] Since the two sub-curves of the normal distribution curve should theoretically be axially symmetric, meaning the slopes of the lines connecting them should be opposites, we connect the maximum and minimum values of each sub-curve to obtain the slope of the connecting line for each sub-curve, denoted as [the slopes of the lines connecting the two sub-curves]. , ,like Figure 2 As shown. Based on the slope of each sub-curve, the sum of the slopes is obtained. This sum is then inversely normalized using an exponential function with the natural constant as the base, yielding a mirror characteristic value. If the two sub-curves are mirror-symmetric, meaning they conform to a normal distribution, the sum of their slopes is 0. Conversely, the larger the sum of their slopes, the less mirror-symmetric the two sub-curves are, and the smaller their corresponding mirror characteristic value, indicating a less conformity to a normal distribution.
[0052] Since the fluctuation amplitudes of the two sub-curves should be consistent when the normal distribution curve conforms to the characteristics of a normal distribution, the difference between the maximum and minimum values on each sub-curve is obtained. The absolute value of the difference is then used to inversely normalize the absolute value of the difference using an exponential function with the natural constant as the base, thus obtaining a similarity value for the curve fluctuation amplitudes. The smaller the absolute value of the difference, the more similar the curve fluctuation amplitudes are between the two sub-curves, and the more they conform to the characteristics of a normal distribution. The larger the similarity value of the curve fluctuation amplitudes is.
[0053] The first normal similarity is obtained by weighted summation of the mirror feature values and the curve fluctuation amplitude similarity values. The normal distribution similarity scoring model is as follows:
[0054]
[0055] in, This represents the first normal similarity. Indicates the first weight. This represents an exponential function with the natural constant as its base. This represents the slope of the straight line representing the first sub-curve. The slope of the straight line representing the second sub-curve. Indicates the second weight. This represents the difference between the maximum and minimum values on the first subcurve. This represents the difference between the maximum and minimum values on the second sub-curve, where | represents the absolute value sign.
[0056] It should be noted that, in this embodiment of the invention, both the amplitude of curve fluctuations and the slope of the straight line are considered important for the analysis of normal distribution characteristics; therefore, the following settings are made: .
[0057] After obtaining the first normal similarity score, a coarse similarity threshold of 0.7 is set. This threshold is not fixed and can be adjusted appropriately based on the rigor required in the scenario. If the first normal similarity score is greater than or equal to 0.7, it indicates that the precision data set conforms to a normal distribution and can be processed in batches using existing techniques. That is, the mean and standard deviation can be used to perform batch statistics on the precision data set without the need for clustering or complex processing. Conversely, if the first normal similarity score is less than 0.7, it indicates that the precision data set does not conform to a normal distribution and requires subsequent data transformation to improve the precision data that does not conform to the normal distribution, making it closer to a normal distribution or more compatible with existing batch statistical methods.
[0058] It should be noted that by analyzing the first normal similarity, redundant calculations for other similarity judgments can be avoided, and classification efficiency can be improved while ensuring accuracy, so as to quickly filter out the precise data set that does not conform to the characteristics of normal distribution.
[0059] Step S103: If the first normal similarity is less than the preset coarse similarity threshold, then at least two data transformation methods are used to transform the precision data set to obtain multiple transformed data sets. The skewness of each transformed data set is calculated. Based on all the skewnesses, the optimal transformed data set is obtained from all the transformed data sets.
[0060] After obtaining the precision data set that does not conform to the normal distribution characteristics through step S102, that is, the first normal similarity of the precision data set is less than the preset coarse similarity threshold, the precision data set is transformed using at least two data transformation methods to obtain multiple transformed data sets in order to obtain the optimal transformed data set.
[0061] This invention includes, but is not limited to, four data transformation methods: logarithmic transformation, square root transformation, Yeo-Johnson transformation, and exponential transformation. All four methods are existing technologies. Logarithmic transformation is suitable for right-skewed data (values > 0), and the transformation formula is as follows: Square root transformation, suitable for mildly right-skewed data (value > 0), transformation formula: Yeo-Johnson transformation, suitable for data containing 0 or negative values, conversion method: expand Box-cox to the negative range; exponential transformation, suitable for left-skewed data, conversion formula: .
[0062] A data transformation method is provided to obtain a corresponding transformed data set. In order to ensure that the precision data set belongs to the optimal data transformation, the skewness of each transformed data set is obtained. The calculation of skewness is a prior art. The data transformation method that makes the skewness closest to 0 is selected to obtain the optimal transformed data set. Specifically, the absolute value of the difference between each skewness and the constant 0 is calculated, and the transformed data set with the smallest absolute value of the difference is taken as the optimal transformed data set.
[0063] This completes the correction of the precision data set, resulting in an optimal transformed data set that is closer to a normal distribution.
[0064] Step S104: Construct the histogram curve of the optimal transformed data set and the target normal distribution curve. The horizontal axis of the histogram curve represents the transformed data, and the vertical axis represents the frequency of the transformed data. Use the histogram curve to perform a fine analysis of the normal distribution characteristics of the target normal distribution curve to obtain the second normal similarity. Based on the second normal similarity, adaptively obtain the optimal batch statistical method and use the optimal batch statistical method to perform batch statistics on the optimal transformed data set.
[0065] After obtaining the optimal transformed data set, a quadratic normal distribution similarity scoring model is constructed again for fine analysis of the normal distribution of the multi-transformed data, so as to obtain the optimal batch statistical method based on the analysis results. In this embodiment of the invention, firstly, a histogram curve of the optimal transformed data set is constructed, with the horizontal axis of the histogram curve representing the transformed data and the vertical axis representing the frequency of the transformed data; simultaneously, based on the mean and standard deviation of the data in the optimal transformed data set, a normal distribution curve of the optimal transformed data set is constructed according to the construction method of the normal distribution curve in step S102, denoted as the target normal distribution curve. Then, the histogram curve is superimposed using the target normal distribution curve, that is, the horizontal axes of the two are aligned, and the difference in the vertical coordinate between the i-th data point in the histogram curve and the i-th data point in the target normal distribution curve is calculated. Based on the difference in the vertical coordinate of each data point in the histogram curve, the corresponding mean square error is obtained. The mean square error is inversely normalized using an exponential function with the natural constant as the base, and the curve overlap is obtained.
[0066] Furthermore, the mean value corresponding to the target normal distribution curve is obtained. and standard deviation ,statistics The number of data points is calculated, and the proportion of the number of data points in the optimal transformed data set is obtained. The absolute value of the difference between the proportion and the theoretical proportion under the standard normal distribution is obtained. The absolute value of the difference is inversely normalized using an exponential function with the natural constant as the base, and the standard normal distribution similarity is obtained. Finally, the curve overlap degree and the standard normal distribution similarity are weighted and summed to obtain the second normal similarity.
[0067] In one embodiment, the quadratic normal distribution similarity scoring model is as follows:
[0068]
[0069] in, This represents the second normal similarity. Indicates the third weight. This represents an exponential function with the natural constant as the base, where n represents the number of data points on the histogram curve. This represents the ordinate value of the i-th data point on the histogram curve. This represents the ordinate value of the i-th data point on the target normal distribution curve. Indicates the fourth weight. Represents the target normal distribution curve The amount of data within, This represents the total number of data points in the optimal transformed dataset, and 0.68 indicates a downward-biased standard normal distribution. The theoretical proportion of the data within, where | represents the absolute value symbol.
[0070] It should be noted that, The smaller the value, the higher the overlap between the histogram curve and the target normal distribution curve. A result of 0 indicates complete overlap. Under inverse proportional normalization, the closer the result is to 1, the higher the overlap between the two curves. The smaller the value, the greater the overlap of the curves, and the more it conforms to the characteristics of a normal distribution; Used to characterize the standard normal distribution, it analyzes the proportion of data under the standard normal distribution in the total data volume to compare the difference between the target normal distribution curve and the standard normal distribution. The smaller the value, the more the optimal transformed data set conforms to a normal distribution, and the greater the corresponding second normal similarity. Since the two reference values and influence degrees of the two items used to calculate the second normal similarity are consistent, we set... No restrictions are imposed here.
[0071] The first-constructed normal distribution similarity scoring model targets the initial partitioning of the precision dataset to be batch-counted. At this stage, the amount of data to be calculated is enormous, and the calculation method is more intuitive, less complex, and more inclusive. It does not require each precision dataset to perfectly conform to a normal distribution; only the basic curve features need to conform to a normal distribution, thus improving overall operational efficiency. The second-order normal distribution similarity scoring model, constructed with optimally transformed data, targets the precision dataset after the initial partitioning, which does not conform to a normal distribution and has undergone optimal data transformation. The amount of data is significantly reduced. Since the data after transformation and the selection of the optimal transformation method should be subject to stricter distribution characteristics, this increases the accuracy of the calculation model and reduces its inclusiveness, making the overall calculation process more robust. The reduced data volume increases the complexity of the calculation model to a certain extent, but does not affect the overall computational efficiency and time complexity. Therefore, after obtaining the second normal similarity of the optimal transformed data set, the optimal transformed data set can be classified according to the second normal similarity to adaptively obtain the optimal batch statistical method: since the secondary scoring is more stringent, the precision similarity threshold is set to 0.6, and adjusted appropriately according to the rigor of the scenario. If the second normal similarity is greater than or equal to 0.6, the optimal transformed data set is considered to conform to the normal distribution characteristics, which can be achieved using existing batch statistical methods. The corresponding optimal batch statistical methods include mean and standard deviation. At the same time, batch statistics can achieve speed and simplicity in the detection and processing of precision data anomalies, for example, through "mean". "(like (Principle) can quickly identify outliers; if the second normal similarity is less than 0.6, it is considered that the optimal transformed data set does not conform to the normal distribution characteristics. For precision data sets that do not conform to the normal distribution, statistics that are not sensitive to skewness are directly used to replace them: (1) Replace the central tendency indicator: replace the mean with the median (the median is robust to extreme values) or use the trimmed mean method (such as the mean after removing the highest / lowest 5%); (2) Replace the dispersion indicator: replace the standard deviation with the interquartile range (IQR): IQR=Q3-Q1, and replace it with the absolute deviation of the median (MAD), which is the optimal batch statistical method including the median, trimmed mean and interquartile range. It should be noted that the above replacement methods are more suitable for precision data scenarios and have lower computational complexity, while ensuring effectiveness. There are also other replacement methods, such as binning and smoothing methods, hybrid model modeling, etc., which will not be elaborated here.
[0072] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for batch statistical analysis of precision data based on data processing, characterized in that, The method includes: Obtain the set of precision data to be analyzed in batches; A normal distribution curve for the precision data set is constructed based on the mean and standard deviation of the data set. The normal distribution curve is divided into two sub-curves based on the axis of symmetry. Based on the changing trends of the two sub-curves, a coarse analysis of the normal distribution characteristics of the normal distribution curve is performed to obtain the first normal similarity. If the first normal similarity is less than the preset coarse similarity threshold, then at least two data transformation methods are used, including: logarithmic transformation, square root transformation, Yeo-Johnson transformation and exponential transformation, to transform the precision data set respectively, to obtain multiple transformed data sets, to calculate the skewness of each of the transformed data sets, and to obtain the optimal transformed data set among all the transformed data sets based on all the skewnesses. Construct a histogram curve of the optimal transformed data set and a target normal distribution curve. The horizontal axis of the histogram curve represents the transformed data, and the vertical axis represents the frequency of the transformed data. Use the histogram curve to perform a fine analysis of the normal distribution characteristics of the target normal distribution curve to obtain a second normal similarity. Based on the second normal similarity, adaptively obtain the optimal batch statistical method and use the optimal batch statistical method to perform batch statistics on the optimal transformed data set. The step of performing a coarse analysis of the normal distribution characteristics of the normal distribution curve based on the changing trends of the two sub-curves to obtain the first normal similarity includes: The slope of each sub-curve is obtained, and the sum of the slopes is obtained. The sum of the slopes is then normalized inversely using an exponential function with the natural constant as the base, to obtain the mirror feature value. The difference between the maximum and minimum values on each sub-curve is obtained, and the absolute value of the difference is obtained. The absolute value of the difference is then inversely normalized using an exponential function with the natural constant as the base, to obtain the similarity value of the curve fluctuation amplitude. The first normal similarity is obtained by weighted summation of the mirror feature value and the curve fluctuation amplitude similarity value.
2. The method for batch statistical analysis of precision data based on data processing according to claim 1, characterized in that, After obtaining the first normal similarity, the following is also included: If the first normal similarity is greater than or equal to the preset coarse similarity threshold, then the mean and standard deviation are used to perform batch statistics on the precision data set.
3. The method for batch statistical analysis of precision data based on data processing according to claim 1, characterized in that, The step of obtaining the optimal transformed data set from all transformed data sets based on all skewnesses includes: Calculate the absolute value of the difference between each skewness and the constant 0, and combine the transformed datasets of the skewness corresponding to the smallest absolute value of the difference into the optimal transformed dataset.
4. The method for batch statistical analysis of precision data based on data processing according to claim 1, characterized in that, The step of performing a fine analysis of the normal distribution characteristics of the target normal distribution curve using the histogram curve to obtain the second normal similarity includes: Calculate the difference in the ordinate between the i-th data point in the histogram curve and the i-th data point in the target normal distribution curve. Based on the difference in the ordinate of each data point in the histogram curve, obtain the corresponding mean square error. Use an exponential function with the natural constant as the base to inversely normalize the mean square error to obtain the curve overlap. Obtain the mean value corresponding to the target normal distribution curve. and standard deviation ,statistics The number of data within the set is calculated as a percentage of the total number of data in the optimal transformed data set. The absolute value of the difference between the percentage and the theoretical percentage under the standard normal distribution is obtained. The absolute value of the difference is then inversely normalized using an exponential function with the natural constant as the base, to obtain the standard normal distribution similarity. The second normal similarity is obtained by weighted summation of the curve overlap and the standard normal distribution similarity.
5. The method for batch statistical analysis of precision data based on data processing according to claim 1, characterized in that, The step of adaptively obtaining the optimal batch statistics method based on the second normal similarity includes: If the second normal similarity is greater than or equal to the preset precision similarity threshold, the optimal batch statistical method includes the mean and standard deviation; if the second normal similarity is less than the preset precision similarity threshold, the optimal batch statistical method includes the median, trimmed mean, and interquartile range.
Citation Information
Patent Citations
Normal distribution-based internet big data mining method and system
CN105279257A
Enterprise purchase management method and system based on big data
CN117290799A