Precision data batch statistics method based on data processing
By constructing a normal distribution similarity scoring model and data conversion, the problem of low batch statistical accuracy of non-normal distribution precision data is solved, and more efficient and reliable batch statistical results are achieved.
Patent Information
- Application Number
- CN202511194650.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-08-26
AI Technical Summary
When existing technologies process precision data with non-normal distribution, batch statistical results have low accuracy and are easily affected by extreme values, resulting in distorted results.
By constructing a normal distribution similarity scoring model, we can screen out precision data sets that do not conform to the normal distribution, and use data transformation methods to make them close to the normal distribution. We can then construct a histogram curve for precision analysis and adaptively obtain the optimal batch statistics method.
It improves the accuracy and robustness of batch statistics, reduces sensitivity to extreme values, and provides a more reliable basis for decision-making.
Smart Images

Figure CN120763455A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a precision data batch statistics method based on data processing. Background Art
[0002] Precision data is a data set used to quantify the degree of closeness between measurement results and true values in processes such as systems and equipment. Its core goal is to evaluate and control the error range and ensure the reliability of data results. Batch statistics and analysis of precision data can reveal the hidden laws in the data and provide a scientific basis for quality control, system optimization and decision-making. Batch statistics processing methods usually need to be highly efficient, and can satisfy the need to summarize the overall situation through a single statistic. For example, the mean and standard deviation can fully describe the distribution of the current data set without the need for grouping or complex processing. At the same time, batch statistics can achieve rapid and concise processing of anomaly detection of precision data. For example, through "mean "(like Principle) can quickly identify outliers, thereby achieving effective batch statistics of precision data.
[0003] For the existing technology, precision data that obeys the normal distribution is theoretically most suitable for the above-mentioned batch statistical processing. However, precision data that does not obey the normal distribution may cause distorted results due to violation of statistical assumptions or concealment of internal structures when undergoing the above-mentioned batch statistical processing. That is, the calculated statistical results will be affected by extreme values and become abnormal, making the batch statistical effect unable to be guaranteed. Summary of the Invention
[0004] In view of this, an embodiment of the present invention provides a precision data batch statistics method based on data processing to solve the problems of low accuracy and abnormal results of precision data in batch statistical processing.
[0005] An embodiment of the present invention provides a method for batch counting of precision data based on data processing, the method comprising the following steps: Get the precision data set to be counted in batches; Constructing a normal distribution curve of the precision data set according to the data mean and data standard deviation of the precision data set, dividing the normal distribution curve into two sub-curves based on a symmetry axis, and performing a rough analysis of normal distribution characteristics on the normal distribution curve according to the change trends of the two sub-curves to obtain a first normal similarity; If the first normal similarity is less than a preset coarse similarity threshold, performing data conversion on the precision data set respectively using at least two data conversion methods to obtain multiple transformed data sets, calculating the skewness of each of the transformed data sets, and obtaining an optimal transformed data set from all the transformed data sets based on all the skewnesses; A histogram curve and a target normal distribution curve of the optimal conversion data set are constructed, where the horizontal axis of the histogram curve is the conversion data and the vertical axis is the frequency of the conversion data. The histogram curve is used to perform a precise analysis of the normal distribution characteristics of the target normal distribution curve to obtain a second normal similarity. Based on the second normal similarity, an optimal batch statistics method is adaptively obtained, and batch statistics are performed on the optimal conversion data set using the optimal batch statistics method.
[0006] Preferably, performing a rough analysis of normal distribution characteristics on the normal distribution curve according to the change trends of the two sub-curves to obtain a first normal similarity includes: Obtaining the straight line slope of each sub-curve respectively to obtain a straight line slope sum value, and performing inverse proportional normalization on the straight line slope sum value using an exponential function with a natural constant as a base to obtain a mirror eigenvalue; respectively obtaining the difference between the maximum value and the minimum value on each of the sub-curves to obtain the absolute value of the difference between the two differences, and inversely normalizing the absolute value of the difference using an exponential function with a natural constant as the base to obtain a curve fluctuation amplitude similarity value; A weighted sum is performed on the mirror image characteristic value and the curve fluctuation amplitude similarity value to obtain a first normal similarity.
[0007] Preferably, after obtaining the first normal similarity, the method further includes: If the first normal similarity is greater than or equal to a preset rough similarity threshold, batch statistics are performed on the precision data set using the mean and standard deviation.
[0008] Preferably, the data conversion methods include: logarithmic conversion, square root conversion, Yeo-Johnson conversion and exponential conversion.
[0009] Preferably, obtaining the optimal transformed data set from all transformed data sets according to all skewnesses includes: The absolute value of the difference between each skewness and the constant 0 is calculated, and the transformed data set of the skewness corresponding to the minimum absolute value of the difference is taken as the optimal transformed data set.
[0010] Preferably, the step of performing a precise analysis of normal distribution characteristics on the target normal distribution curve using the histogram curve to obtain a second normal similarity comprises: Calculating the ordinate difference between the i-th data point in the histogram curve and the i-th data point in the target normal distribution curve, obtaining the corresponding mean square error based on the ordinate difference corresponding to each data point in the histogram curve, and inversely normalizing the mean square error using an exponential function with a natural constant as the base to obtain the curve coincidence degree; Get the mean corresponding to the target normal distribution curve and standard deviation ,statistics , calculating the proportion of the data number in the total data number in the optimal transformation data set, obtaining the absolute value of the difference between the proportion and the theoretical proportion under the standard normal distribution, and using an exponential function with a natural constant as the base to inversely normalize the absolute value of the difference to obtain the standard normal distribution similarity; A weighted sum is performed on the curve coincidence and the standard normal distribution similarity to obtain a second normal similarity.
[0011] Preferably, the adaptively obtaining the optimal batch statistics method according to the second normal similarity includes: If the second normal similarity is greater than or equal to the preset precise similarity threshold, the optimal batch statistical method includes mean and standard deviation; if the second normal similarity is less than the preset precise similarity threshold, the optimal batch statistical method includes median, trimmed mean and interquartile range.
[0012] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: The present invention obtains a precision data set to be batch counted; constructs a normal distribution curve of the precision data set based on the data mean and data standard deviation of the precision data set, divides the normal distribution curve into two sub-curves based on the symmetry axis, and performs a rough analysis of the normal distribution characteristics of the normal distribution curve according to the change trends of the two sub-curves to obtain a first normal similarity; if the first normal similarity is less than a preset rough similarity threshold, uses at least two data conversion methods to perform data conversion on the precision data set respectively to obtain multiple converted data sets, calculates the skewness of each of the converted data sets, and obtains an optimal converted data set from all the converted data sets based on all skewnesses; constructs a histogram curve and a target normal distribution curve of the optimal converted data set, the horizontal axis of the histogram curve is the converted data, and the vertical axis is the frequency of the converted data, uses the histogram curve to perform a fine analysis of the normal distribution characteristics of the target normal distribution curve to obtain a second normal similarity, and according to the second normal similarity, adaptively obtains an optimal batch statistics method, and uses the optimal batch statistics method to perform batch statistics on the optimal converted data set. Among them, the first normal similarity is obtained to screen the precision data set that does not conform to the normal distribution characteristics, and then, by performing optimal data conversion on the precision data set that does not conform to the normal distribution characteristics, it is made closer to the normal distribution or more consistent with the batch statistics method under the existing technology. Therefore, by analyzing the second normal similarity of the optimal conversion data set, it is evaluated with more stringent distribution characteristics to increase the analysis accuracy and reduce inclusiveness, so that the calculation robustness of the second normal similarity is higher, and the optimal batch statistics method adaptively obtained according to the second normal similarity can significantly improve the batch statistical quality of skewed normal precision data and provide a more reliable basis for decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0014] Figure 1 This is a method flow chart of a precision data batch statistics method based on data processing provided by the first embodiment of the present invention; Figure 2 This is a schematic diagram of a normal distribution curve divided into two sub-curves provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0015] Embodiments of the present disclosure are described in detail below with reference to the attached drawing figures, wherein the embodiments given herein are by way of illustration only and are not intended to be limiting of the present disclosure. For the purpose of clarity, not all of the illustrative embodiments are described. In this description, reference is made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration embodiments for practicing the present disclosure. The drawings employed herein are intended merely to further illustrate, and not to limit, the embodiments of the present disclosure.
[0016] It should be noted that the terms "first", "second", and the like in the description of the present disclosure and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all the embodiments consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure.
[0017] In order to illustrate the technical solutions of the present application, the following will be described by specific embodiments.
[0018] Referring to Figure 1 , a method flow chart of a precision data batch statistical method based on data processing provided by Embodiment One of the present application is shown, which can include: Figure 1 Step S101, obtaining a precision data set to be batched.
[0019] The precision data has many scenarios, and its collection method needs to be selected according to the data source and data type. For specific reference, the following examples can be referred to: 1. Data collection Sensor collection; applicable scenarios are industrial manufacturing, environmental monitoring, and other scenarios that require real-time acquisition of physical parameters.
[0020] Log file collection: applicable scenarios are system monitoring, troubleshooting, and other scenarios that require analysis of device or system operation logs.
[0021] API interface collection: applicable scenarios are scenarios that require data from third-party systems or services, such as calling a weather forecast API to obtain real-time weather data.
[0022] 2. Data preprocessing Data cleaning: handle missing values, duplicate values, or outliers in precision data, and preliminarily improve the accuracy of the data; at the same time, convert the data into a format suitable for analysis, eliminate the dimension effect, for example: data normalization, data discretization.
[0023] 3. Data integration Merge data from multiple data sources to resolve naming conflicts and redundancies. For example, entity recognition: unify the naming of the same entity across different data sources (e.g., "Customer ID" and "User Number"). Redundancy detection: identify and remove redundant attributes through correlation analysis (e.g., Pearson coefficient). Numeric conflict resolution: unify scaling standards or encoding methods (e.g., date format, currency unit) across different data sources.
[0024] By using the above method, at least one precision data set to be batch counted can be obtained.
[0025] Step S102: construct a normal distribution curve of the precision data set based on the data mean and data standard deviation of the precision data set, divide the normal distribution curve into two sub-curves based on the symmetry axis, and perform a rough analysis of the normal distribution characteristics of the normal distribution curve based on the changing trends of the two sub-curves to obtain the first normal similarity.
[0026] The fluctuations of many precision data (such as dimensional tolerances, form and position errors, and surface roughness) are inherently random and independent, and generally conform to normal distribution characteristics. However, when precision data is influenced by the superposition of multiple independent random factors (such as machining errors, material fluctuations, and environmental interference), regardless of the distribution of these factors themselves, their sum will approach a normal distribution (central limit theorem) when the sample size is large enough. For example, in machining, factors such as tool wear, machine tool vibration, and temperature changes act independently, and the theoretical distribution of the final product dimensional error is normally distributed. However, in actual processing, the size of each precision data set to be batch counted, as well as the influence of outliers or extreme values (such as abnormal dimensions caused by sudden tool breakage during machining, systematically larger dimensions caused by long-term machine wear, etc.), will cause the precision data sets to be batch counted to not all exhibit characteristics similar to a normal distribution. Using traditional batch count methods for such data will adversely affect the results and produce serious errors.
[0027] Therefore, based on the above feature analysis, the embodiment of the present invention constructs a normal distribution similarity scoring model to analyze and screen the precision data set to be batch counted, so as to obtain the precision data set that does not conform to the normal distribution characteristics, and obtains the optimal batch statistics method for batch statistics through data conversion. Then, for analyzing and screening the precision data set to be batch counted by constructing a normal distribution similarity scoring model, first obtain the data mean of the precision data set and the standard deviation of the data , substitute it into the probability density function of the normal distribution to construct the normal distribution curve of the precision data set, where the calculation formula of the probability density function is:
[0028] It should be noted that the construction of the normal distribution curve belongs to the existing technology, and its probability density function also belongs to the existing technology, which will not be described in detail here.
[0029] Then, using the axis of symmetry (that is, the center point of the horizontal axis of the normal distribution curve) as the dividing line, the normal distribution curve is divided into two sub-curves, denoted as F1 and F2. Finally, based on the trend characteristics of the two sub-curves, a normal distribution similarity scoring model is constructed to obtain the first normal similarity of the precision data set. The specific method for obtaining the first normal similarity is: Since the two sub-curves of the normal distribution curve should be axisymmetric in theory, that is, the slopes of the straight lines between them should be opposite, the maximum and minimum values of the two sub-curves are connected respectively to obtain the slopes of the connecting lines of each sub-curve, which are recorded as 、 ,like Figure 2 As shown. Based on the straight line slope of each sub-curve, the sum of the straight line slopes is obtained, and the sum of the straight line slopes is inversely normalized using an exponential function with a natural constant as the base to obtain a mirror eigenvalue. If the two sub-curves are mirror-symmetrical, that is, they conform to the normal distribution characteristics, the sum of the straight line slopes is 0. The larger the sum of the straight line slopes, the less the two sub-curves conform to the mirror symmetry, and the smaller the corresponding mirror eigenvalue, which is less consistent with the normal distribution characteristics.
[0030] Since the fluctuation amplitudes of the two sub-curves should be consistent when the normal distribution curve conforms to the normal distribution characteristics, the difference between the maximum value and the minimum value on each sub-curve is obtained respectively, and the absolute value of the difference between the two differences is obtained. The absolute value of the difference is inversely normalized using an exponential function with a natural constant as the base to obtain the curve fluctuation amplitude similarity value. If the absolute value of the difference is smaller, it means that the curve fluctuation amplitudes between the two sub-curves are more similar, the corresponding characteristics are more consistent with the normal distribution, and the curve fluctuation amplitude similarity value is greater.
[0031] The mirror image feature value and the curve fluctuation amplitude similarity value are weighted and summed to obtain the first normal similarity. The normal distribution similarity scoring model is:
[0032] in, represents the first normal similarity, represents the first weight, represents an exponential function with a natural constant as base, represents the slope of the first sub-curve, represents the slope of the second sub-curve, represents the second weight, Represents the difference between the maximum and minimum values on the first sub-curve, represents the difference between the maximum value and the minimum value on the second sub-curve, and || represents the absolute value symbol.
[0033] It should be noted that in the embodiments of the present application, both the curve fluctuation amplitude and the straight line slope are considered important for the analysis of the normal distribution characteristics, and therefore, the first normal similarity is set to be greater than 0 and less than 1. .
[0034] After obtaining the first normal similarity, a coarse similarity threshold is set to 0.7, which is not limited here and can be appropriately increased or decreased according to the degree of rigor in the scene. If the first normal similarity is greater than or equal to 0.7, it means that the precision data set conforms to the normal distribution characteristics, and the existing technology can be used for statistical batch processing, that is, the mean and standard deviation can be used to statistically process the precision data set in batches without grouping or complex processing. On the contrary, if the first normal similarity is less than 0.7, it means that the precision data set does not conform to the normal distribution characteristics, and subsequent data conversion is needed to improve the precision data that does not conform to the normal distribution characteristics, so that it is closer to the normal distribution or more matched with the batch statistical method under the existing technology.
[0035] It should be noted that through the analysis of the first normal similarity, redundant calculations of other similarity judgments can be avoided, and the classification efficiency can be improved under the premise of ensuring accuracy to quickly screen out the precision data set that does not conform to the normal distribution characteristics.
[0036] In step S103, if the first normal similarity is less than the preset coarse similarity threshold, at least two data conversion methods are used to convert the precision data set, a plurality of converted data sets are obtained, the skewness of each converted data set is calculated, and the optimal converted data set is obtained from all converted data sets according to all skewness.
[0037] After obtaining the precision data set that does not conform to the normal distribution characteristics through step S102, that is, the first normal similarity of the precision data set is less than the preset coarse similarity threshold, at least two data conversion methods are used to convert the precision data set, a plurality of converted data sets are obtained, and the optimal converted data set is obtained.
[0038] In the embodiments of the present application, the data conversion methods include but are not limited to four kinds, which are logarithmic conversion, square root conversion, Yeo-Johnson conversion and exponential conversion. The four kinds of data conversion methods are existing technologies, wherein the logarithmic conversion is suitable for right-skewed data (value>0), and the conversion formula is: ; the square root conversion is suitable for mildly right-skewed data (value>0), and the conversion formula is: Yeo-Johnson transformation, suitable for data containing 0 or negative values, transformation method: extend Box-cox to negative value domain; exponential transformation, suitable for left-skewed data, transformation formula: .
[0039] A data conversion mode is obtained, and a corresponding conversion data set is obtained, in order to ensure the accuracy of the data set, the skewness of each conversion data set is obtained, the calculation of the skewness belongs to the prior art, the data conversion mode is selected to make the skewness closest to 0, to obtain the optimal conversion data set, specifically: calculate the absolute value of the difference between each skewness and the constant 0, and take the conversion data set of the skewness corresponding to the minimum absolute value as the optimal conversion data set.
[0040] At this point, the accuracy of the data set is corrected, and the optimal conversion data set closer to the normal distribution is obtained.
[0041] Step S104, construct a histogram curve of the optimal conversion data set and a target normal distribution curve, the horizontal axis of the histogram curve is the conversion data, and the vertical axis is the frequency of the conversion data, and the histogram curve is used to perform a normal distribution feature analysis on the target normal distribution curve to obtain a second normal similarity, and the optimal batch statistical mode is adaptively obtained according to the second normal similarity, and the optimal batch statistical mode is used to perform batch statistics on the optimal conversion data set.
[0042] After obtaining the optimal conversion data set, a secondary normal distribution similarity scoring model is constructed again, which is used for a normal distribution analysis of the data after multiple conversions, so as to obtain an optimal batch statistical mode according to the analysis result. In the embodiment of the application, first, a histogram curve of the optimal conversion data set is constructed, the horizontal axis of the histogram curve is the conversion data, and the vertical axis is the frequency of the conversion data; at the same time, according to the construction mode of the normal distribution curve in step S102, a normal distribution curve of the optimal conversion data set is constructed, which is recorded as a target normal distribution curve, according to the data mean and data standard deviation of the optimal conversion data set. Then, the target normal distribution curve is used for superimposed processing of the histogram curve, that is, the horizontal axes of the two are aligned, the vertical coordinate difference between the i-th data point in the histogram curve and the i-th data point in the target normal distribution curve is calculated, and the corresponding mean square error is obtained according to the vertical coordinate difference corresponding to each data point in the histogram curve, and the mean square error is inversely proportional to the normalization by using the exponential function with the natural constant as the base, to obtain the curve coincidence degree. Further, the mean and the standard deviation corresponding to the target normal distribution curve are obtained, and the , calculate the proportion of the data number in the total data number in the optimal transformation data set, obtain the absolute value of the difference between the proportion and the theoretical proportion under the standard normal distribution, use an exponential function with a natural constant as the base to inversely normalize the absolute value of the difference, and obtain the standard normal distribution similarity; finally, perform weighted summation on the curve coincidence and the standard normal distribution similarity to obtain the second normal similarity.
[0043] In one embodiment, the quadratic normal distribution similarity score model is:
[0044] in, represents the second normal similarity, represents the third weight, represents an exponential function with a natural constant as the base, n represents the number of data points on the histogram curve, Represents the vertical coordinate value of the i-th data point on the histogram curve, Represents the ordinate value of the i-th data point on the target normal distribution curve, represents the fourth weight, Represents the target normal distribution curve The amount of data in represents the total number of data in the optimal transformation data set, and 0.68 represents the standard normal distribution downward The theoretical proportion of the number of data within, | | represents the absolute value symbol.
[0045] It should be noted that The smaller the value, the higher the degree of overlap between the histogram curve and the target normal distribution curve. When the result is 0, it means complete overlap. Under inverse proportional normalization, the closer the result is to 1, the higher the degree of overlap between the two curves. The smaller the value, the greater the curve overlap, and the more consistent it is with the normal distribution characteristics; It is used to characterize the standard normal distribution. By analyzing the proportion of the number of data under the standard normal distribution in the total amount of data, the difference between the target normal distribution curve and the normal distribution standard can be compared. The smaller the value of , the more the optimal transformation data set conforms to the normal distribution, and the larger the corresponding second normal similarity. Since the two reference values and influence levels for calculating the second normal similarity are consistent, we set , there is no restriction here.
[0046] Among them, for the normal distribution similarity scoring model constructed for the first time, it is aimed at the first division of the precision data set to be batch counted. At this time, the number of data to be calculated is huge, and the calculation method is more intuitive, less complex, and more inclusive. It does not require each precision data set to fully conform to the normal distribution. Only the basic curve characteristics must conform to the normal distribution effect, which improves the overall operation efficiency; and for the optimal conversion data, a secondary normal distribution similarity scoring model is constructed. It is aimed at the precision data set that does not conform to the normal distribution after the first division and undergoes the optimal data conversion processing. At this time, the number of data faced is greatly reduced. Since the data after data conversion and selection of the optimal data conversion method should be subject to more stringent distribution characteristics, the accuracy of the calculation model can be increased and the inclusiveness can be reduced, making the entire calculation processing more robust. Increasing the complexity of the calculation model to a certain extent based on the reduction in the number of data will not affect the overall calculation efficiency and time complexity. Therefore, after obtaining the second normal similarity of the optimal conversion data set, the optimal conversion data set can be classified according to the second normal similarity to adaptively obtain the optimal batch statistical method: since the secondary scoring is more stringent, the precision similarity threshold is set to 0.6, and it is appropriately increased or decreased according to the rigor of the scene. If the second normal similarity is greater than or equal to 0.6, it is considered that the optimal conversion data set meets the normal distribution characteristics, which can be achieved by the batch statistical method in the prior art. The corresponding optimal batch statistical method includes mean and standard deviation. At the same time, batch statistics can achieve rapidity and simplicity in the detection of precision data anomalies. For example, by "mean "(like Principle) can quickly identify outliers; if the second normal similarity is less than 0.6, it is considered that the optimal transformation data set does not conform to the normal distribution characteristics. For the precision data set that does not conform to the normal distribution, a statistic that is insensitive to skewness is directly used to replace it: (1) Replace the central tendency index: use the median instead of the mean (the median is robust to extreme values) or use the trimmed mean method (such as the mean after removing the highest / lowest 5%); (2) Replace the dispersion index: use the interquartile range (IQR) instead of the standard deviation: IQR=Q3-Q1, and use the median absolute deviation (MAD) instead, that is, the optimal batch statistics method includes the median, trimmed mean and interquartile range. It should be noted that the above-mentioned alternative methods are more suitable for precision data scenarios and have lower computational complexity, while ensuring effectiveness. There are also other alternative methods, such as binning and smoothing methods, mixed model modeling, etc., which will not be described in detail here.
[0047] The above examples are only used to illustrate the technical solutions of the present application, but not limit the present application; although the present application has been described in detail with reference to the foregoing examples, those ordinarily skilled in the art should understand: the technical solutions recorded in the foregoing examples can still be modified, or some technical features can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A precision data batch statistics method based on data processing, characterized in that: The method comprises: Get the precision data set to be counted in batches; Constructing a normal distribution curve of the precision data set according to the data mean and data standard deviation of the precision data set, dividing the normal distribution curve into two sub-curves based on a symmetry axis, and performing a rough analysis of normal distribution characteristics on the normal distribution curve according to the change trends of the two sub-curves to obtain a first normal similarity; If the first normal similarity is less than a preset coarse similarity threshold, performing data conversion on the precision data set respectively using at least two data conversion methods to obtain multiple transformed data sets, calculating the skewness of each of the transformed data sets, and obtaining an optimal transformed data set from all the transformed data sets based on all the skewnesses; A histogram curve and a target normal distribution curve of the optimal conversion data set are constructed, where the horizontal axis of the histogram curve is the conversion data and the vertical axis is the frequency of the conversion data. The histogram curve is used to perform a precise analysis of the normal distribution characteristics of the target normal distribution curve to obtain a second normal similarity. Based on the second normal similarity, an optimal batch statistics method is adaptively obtained, and batch statistics are performed on the optimal conversion data set using the optimal batch statistics method.
2. The method for batch statistics of precision data based on data processing according to claim 1, characterized in that: The step of performing a rough analysis of normal distribution characteristics on the normal distribution curve according to the change trends of the two sub-curves to obtain a first normal similarity includes: Obtaining the straight line slope of each sub-curve respectively to obtain a straight line slope sum value, and performing inverse proportional normalization on the straight line slope sum value using an exponential function with a natural constant as a base to obtain a mirror eigenvalue; respectively obtaining the difference between the maximum value and the minimum value on each of the sub-curves to obtain the absolute value of the difference between the two differences, and inversely normalizing the absolute value of the difference using an exponential function with a natural constant as the base to obtain a curve fluctuation amplitude similarity value; A weighted sum is performed on the mirror image characteristic value and the curve fluctuation amplitude similarity value to obtain a first normal similarity.
3. The method for batch statistics of precision data based on data processing according to claim 1, characterized in that: After obtaining the first normal similarity, it also includes: If the first normal similarity is greater than or equal to a preset rough similarity threshold, batch statistics are performed on the precision data set using the mean and standard deviation.
4. The method for batch statistics of precision data based on data processing according to claim 1, characterized in that: The data conversion methods include: logarithmic conversion, square root conversion, Yeo-Johnson conversion and exponential conversion.
5. The method for batch statistics of precision data based on data processing according to claim 1, characterized in that: The step of obtaining the optimal transformed data set from all transformed data sets according to all skewnesses includes: The absolute value of the difference between each skewness and the constant 0 is calculated, and the transformed data set of the skewness corresponding to the minimum absolute value of the difference is taken as the optimal transformed data set.
6. The method for batch statistics of precision data based on data processing according to claim 1, characterized in that: The step of performing a precise analysis of normal distribution characteristics on the target normal distribution curve using the histogram curve to obtain a second normal similarity includes: Calculating the ordinate difference between the i-th data point in the histogram curve and the i-th data point in the target normal distribution curve, obtaining the corresponding mean square error based on the ordinate difference corresponding to each data point in the histogram curve, and inversely normalizing the mean square error using an exponential function with a natural constant as the base to obtain the curve coincidence degree; Get the mean corresponding to the target normal distribution curve and standard deviation ,statistics , calculating the proportion of the data number in the total data number in the optimal transformation data set, obtaining the absolute value of the difference between the proportion and the theoretical proportion under the standard normal distribution, and using an exponential function with a natural constant as the base to inversely normalize the absolute value of the difference to obtain the standard normal distribution similarity; A weighted sum is performed on the curve coincidence and the standard normal distribution similarity to obtain a second normal similarity.
7. The method for batch statistics of precision data based on data processing according to claim 1, characterized in that: Adaptively obtaining an optimal batch statistics method according to the second normal similarity includes: If the second normal similarity is greater than or equal to the preset precise similarity threshold, the optimal batch statistical method includes mean and standard deviation; if the second normal similarity is less than the preset precise similarity threshold, the optimal batch statistical method includes median, trimmed mean and interquartile range.
Citation Information
Patent Citations
Normal distribution-based internet big data mining method and system
CN105279257A
Enterprise purchase management method and system based on big data
CN117290799A
Real-time detection method for operation fault of mining transport vehicle
CN120524179A
Domain Generalization via Batch Normalization Statistics
US20230122207A1