An Internet data quality assessment method and system
By defining the accuracy of Internet data sets and deducing their credibility using sampling methods, the problem that the existing technology cannot accurately evaluate Internet data quality is solved, and efficient and rapid data quality evaluation and dynamic adjustment capabilities are achieved.
Patent Information
- Application Number
- CN202110615173.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-02
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-06-02
AI Technical Summary
The prior art cannot accurately evaluate the quality of Internet data, cannot scientifically and effectively give confidence intervals for data quality evaluation values, and cannot cope with flexible and changeable business demand scenarios.
By defining the accuracy p of the Internet data set, and using sampling method to extract data samples n, deduce the relationship between the accuracy p of the Internet data set, the credibility relationship between the data sample accuracy p and the accuracy of the data sample, quantify the differences between p, and realize data quality evaluation.
It realizes efficient and rapid Internet data quality assessment, can dynamically adjust according to different data and business needs, and reduces the cost of assessment work and personnel costs.
Smart Images

Figure CN113256135B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet big data, and particularly to an Internet data quality evaluation method and system. Background Art
[0002] With the popularization of the Internet, enterprises are massively transforming to Internet-based operations. A large amount of enterprise operation information is released through the Internet, such as enterprise acquisition information, investment information, civil engineering information, housing transaction information, equity transfer information, and major project information. For tax authorities, enterprises are the main bodies involved in taxation. By analyzing and mining enterprise tax-related data in the Internet, more valuable information can be brought to tax source management.
[0003] In the face of a vast amount of Internet data, manual inspection has a high workload and high personnel costs. At the same time, manual inspection is inevitably prone to misjudgment. Existing evaluation methods cannot accurately evaluate the quality of Internet data, cannot scientifically and effectively give the confidence interval of the data quality evaluation value, and cannot cope with flexible business requirement scenarios. Summary of the Invention
[0004] In view of this, the present invention proposes an Internet data quality evaluation method and system to facilitate Internet data quality control personnel to efficiently and quickly evaluate data quality.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] The method includes:
[0007] S1 Define the accuracy rate of the Internet data set as p;
[0008] S2 Use sampling to extract a data sample n from the Internet data set, and the accuracy rate of the data sample is
[0009] S3 Data modeling, and deduce the credibility relationship among the accuracy rate p of the Internet data set, the data sample n, and the accuracy rate of the data sample ;
[0010] S4 Quantify the difference between p and , that is, the accuracy rate of the data sample can accurately represent the accuracy rate p of the Internet data set;
[0011] S5 Experimentally verify the effectiveness of data modeling in Internet data quality evaluation problems.
[0012] Further, when the credibility of the accuracy rate p of the Internet data set is 90% - 100%, select the sampling method with the smallest sample size, and the sample size is an integer greater than 1.
[0013] Further, the sample size can be selected as a fixed quantity or set as a fixed ratio according to the actual data set.
[0014] Further, if the correctness of the Internet data set follows a Bernoulli distribution, then the expectation of the accuracy rate p of the Internet data set is E(x) = p, the variance is D(x) = p(1 - p), and the standard deviation is
[0015] Further, according to De Moivre's central limit theorem, under the same sampling method, the accuracy rates of the data samples calculated by multiple samplings follow a normal distribution, and the average value μ = p. Further, the accuracy rates of the data samples calculated by multiple samplings The formula for the standard deviation is
[0016] Further, the accuracy rate of the Internet data set sample is converted into a standard normal distribution through a transformation function.
[0017] Further, define the accuracy rate difference of the Internet data set as Δp, and its value range is 0 < Δp < 1. That is, the acceptable inspection accuracy rate is the closed interval from p - Δp to p + Δp. Define η to represent the credible probability that the sampling inspection result falls within the acceptable inspection accuracy rate interval. Through the following formula
[0018] The probability distribution function of the normal distribution is
[0019] The normalization processing function of the normal distribution
[0020] can be changed into the probability distribution function of the standard normal distribution N(0, 1)
[0021] The function represents the cumulative function of the standard normal distribution from negative infinity to x. Then, through the transformation function the accuracy rate distribution of the Internet data sampling inspection can be converted into a standard normal distribution, where: g(x) represents the cumulative function of the standard normal distribution from negative infinity to x, p is the accuracy rate of the Internet data set, Δp is the accuracy rate difference of the Internet data set, and n represents the number of sampled data samples. The formula one is derived as
[0022] Further, through premise assumptions and modeling derivations, it can be obtained that the number of sampled data samples n, the accuracy rate p of the Internet data set, and the accuracy rate difference Δp ultimately determine the credible probability.
[0023] A method and system for evaluating the quality of Internet data according to the present invention, the system includes:
[0024] S1 data import synchronization, where files such as CSV, Excel, and database SQL can be imported to import the original data to be inspected. At the same time, system integration can be performed to provide data import synchronization functions, and the data set can be synchronized once or incrementally.
[0025] S2 Automatic sampling, according to Formula 1: It can be seen that under the same Δp and η conditions, the closer the accuracy rate p of the data set is to 0.5, the larger the sampling quantity n required; in the first inspection of the Internet data set, the accuracy rate p of the data set is default set to 0.5, and the system gives default values for other parameters. The same Δp takes a value of 0.03, and η takes a value greater than or equal to 0.9. According to the above Formula 1, the system automatically calculates the minimum sample size n required for sampling and performs random sampling to form the sampling data to be inspected; if it is not the first inspection of the Internet data set, the accuracy rate p of the previous Internet data set is used as the default value for calculation.
[0026] S3 Data inspection, the system provides a convenient interface operation for comparing data inspection with the original data, simplifies the work of comparing with the original data in Internet data inspection, and improves the efficiency of Internet data quality inspection.
[0027] S4 Data quality assessment Based on the data inspection results, the Internet data quality assessment results are given, including the accuracy rate p of the Internet data set, the accuracy rate difference Δp of the Internet data set, and the evaluation confidence probability η.
[0028] An Internet data quality assessment method and system proposed by the present invention have the following advantages and beneficial effects:
[0029] Combining the professional knowledge of probability theory and mathematical statistics, using scientific statistical inference methods, by designing reasonable simulation data to compare with real data, an Internet data quality assessment method applicable to large-scale data is given, which can be dynamically adjusted according to different data and different business requirements, realizing a perfect sampling inspection assessment system, facilitating Internet data quality control personnel to efficiently and quickly conduct data quality assessment. The quality assessment system is easy to operate and integrate, improving the efficiency of Internet data quality assessment from an engineering perspective and further reducing the cost of Internet data quality assessment work.
[0030] According to the following detailed description of the exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The drawings included in the specification and constituting a part of the specification, together with the specification, illustrate the exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present invention.
[0032] Figure 1 This is the flowchart of a method for evaluating the quality of Internet data according to the present invention;
[0033] Figure 2 This is the block diagram of the data quality evaluation system;
[0034] Figure 3 This is Experiment 1, an experiment with a raw dataset of 10,000, a two-sided verification graph of the simulated and approximate power levels and the sample size;
[0035] Figure 4 This is Experiment 1, an experiment with a raw dataset of 100,000, a two-sided verification graph of the simulated and approximate power levels and the sample size;
[0036] Figure 5 This is Experiment 2, an experiment with a raw dataset of 10,000, a two-sided verification graph of the simulated and approximate power levels and the sample size;
[0037] Figure 6 This is Experiment 2, an experiment with a raw dataset of 100,000, a two-sided verification graph of the simulated and approximate power levels and the sample size; Detailed implementation manners
[0038] In order to make the technical problems, technical solutions and beneficial effects solved by the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0039] As Figure 1 shown, a method for evaluating the quality of Internet data includes:
[0040] S1 Define the accuracy rate of the Internet dataset as p;
[0041] S2 Use the sampling method to extract a data sample n from the Internet dataset, and the accuracy rate of the data sample is
[0042] S3 Data modeling, and deduce the credibility relationship among the accuracy rate p of the Internet dataset, the data sample n and the accuracy rate of the data sample ;
[0043] S4 Quantify the difference between p and , that is, the accuracy rate of the data sample can accurately represent the accuracy rate p of the Internet dataset;
[0044] S5 Experimentally verify the effectiveness of data modeling in the problem of Internet data quality evaluation.
[0045] Preferably, under the confidence level where the accuracy p of the Internet dataset is between 90% and 100%, the sampling method with the smallest sample size is selected, and the sample size is an integer greater than 1.
[0046] Preferably, the sample size can be selected as a fixed quantity or set as a fixed ratio according to the actual dataset.
[0047] Preferably, if the correctness of the Internet dataset follows a Bernoulli distribution, the expectation of the accuracy p of the Internet dataset is E(x) = p, the variance is D(x) = p(1 - p), and the standard deviation is
[0048] Preferably, according to De Moivre's central limit theorem, under the same sampling method, the accuracy of the data samples calculated by multiple samplings follows a normal distribution, and the mean μ = p. Further, the accuracy of the data samples calculated by multiple samplings The formula for the standard deviation is
[0049] Further, through a transformation function, the distribution of the accuracy of the data samples of the Internet dataset is converted into a standard normal distribution.
[0050] Further, define the accuracy difference of the Internet dataset as Δp, and the value range is 0 < Δp < 1. That is, the acceptable inspection accuracy is the closed interval from p - Δp to p + Δp. Define η to represent the credible probability that the sampling inspection result falls within the acceptable inspection accuracy interval. Through the following formula
[0051] The probability distribution function of the normal distribution is
[0052] The standardization processing function of the normal distribution
[0053] can be changed into the probability distribution function of the standard normal distribution N(0, 1)
[0054] The function represents the cumulative function of the standard normal distribution from negative infinity to x. Then, through the transformation function the accuracy distribution of the Internet data sampling inspection can be converted into a standard normal distribution, where: g(x) represents the cumulative function of the standard normal distribution from negative infinity to x, p is the accuracy of the Internet dataset, Δp is the accuracy difference of the Internet dataset, and n represents the number of sampled data samples. The formula one is derived as
[0055] Preferably, through premise assumptions and modeling derivations, it can be obtained that the extracted data sample n, the accuracy rate p of the Internet dataset, and the accuracy rate difference Δp ultimately determine the credible probability.
[0056] As Figure 2 shown, an Internet data quality assessment system includes:
[0057] S1 Data import synchronization, where files such as CSV, Excel, and database SQL can be imported, the original data to be inspected is imported, and at the same time, system integration can provide data import synchronization functions, and the dataset can be synchronized once or incrementally.
[0058] S2 Automatic sampling, according to Formula 1: It can be known that in the case of the same Δp and η, the closer the accuracy rate p of the dataset is to 0.5, the larger the sampling quantity n required; in the first inspection of the Internet dataset, the accuracy rate p of the dataset is default set to 0.5, and the system gives default values for other parameters. The same Δp takes a value of 0.03, and η takes a value greater than or equal to 0.9. According to the said Formula 1, the system automatically calculates the minimum sample size n required for sampling and conducts random sampling to form the sampling data to be inspected; if it is not the first inspection of the Internet dataset, it is calculated based on the accuracy rate p of the previous Internet dataset as the default value.
[0059] S3 Data inspection, the system provides a convenient interface operation for data inspection and comparison with the original data, simplifies the work of comparing with the original data in Internet data inspection, and improves the efficiency of Internet data quality inspection.
[0060] S4 Data quality assessment Based on the data inspection results, the Internet data quality assessment results are given, including the accuracy rate p of the Internet dataset, the accuracy rate difference Δp of the Internet dataset, and the evaluation credible probability η.
[0061] Experimental data proves the effectiveness of the Internet data quality assessment method for Internet data quality assessment. For different datasets, different accuracy rates of the original Internet dataset, as well as the allowable accuracy error range and credible probability value, the required minimum sampling sample size can be calculated; this method can minimize the sample sampling under the basic requirements, thereby reducing the workload and cost of large-scale Internet data quality assessment.
[0062] To verify the accuracy of the assessment method, the following generates test data through simulation experiments. In this experiment, data with a total sample size of 10,000 and 100,000 are generated respectively. According to the standard deviation of the Bernoulli distribution being It can be seen that the accuracy rate p also has a greater impact on the entire evaluation method. In this paper, three accuracy rates of 0.1, 0.3, and 0.5 are selected as the experimental criteria, and the maximum allowable accuracy error Δp is taken as 0.03 and 0.05.
[0063] The calculated credible probability is denoted as the theoretical credible probability, and the credible probability obtained through experiments is denoted as the experimental credible probability. Based on the experimental data generated above, simple random sampling is performed 1000 times according to different sampling sample sizes n to obtain experimental values, and the effectiveness of the model is verified by comparing the differences between the experimental values and the theoretical values.
[0064] Experiment 1: Δp = 0.03
[0065] (1) Experiment with an original dataset of 10,000
[0066]
[0067] (2) Experiment with an original dataset of 100,000
[0068]
[0069] Experiment 2: Δp = 0.05
[0070] (1) Experiment with an original dataset of 10,000
[0071]
[0072] (2) Experiment with an original dataset of 100,000
[0073]
[0074] It can be seen from observing the experimental data that:
[0075] 1. The difference between the theoretical credible probability obtained by the model and the actual experimental credible probability is small.
[0076] 2. The size of the original data set in the experiment has no influence on the experimental results.
[0077] In several embodiments provided by the present application, it should be understood that the disclosed systems and methods can also be implemented in other ways. The above-described embodiments are merely illustrative. For example, the method flowcharts and system block diagrams in the drawings show the data quality inspection calculation architectures, functions, and operations that can be implemented by the systems, methods, and computer program products according to the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, depending on the actual data situation. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0078] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention. It should be noted that similar reference numerals and letters in the present disclosure denote similar items. Therefore, once an item is defined in the present disclosure, it does not need to be further defined and explained in the subsequent content.
Claims
1. An Internet data quality assessment method, characterized in that, the method includes: S1 Define the accuracy rate of the Internet dataset as ; S2 uses sampling to extract a data sample n from the Internet dataset and statistically calculates the accuracy of the data sample as ; S3 Data Modeling, Deducing the Accuracy Rate of Internet Datasets and the relationship between data sample n and the confidence level of the data sample accuracy rate ; If the correctness of the S31 Internet dataset follows a Bernoulli distribution, then the expected value of the accuracy of the Internet dataset is , the variance is , and the standard deviation is ; According to De Moivre's central limit theorem, under the same sampling method, the accuracy of data samples calculated by multiple samplings follows a normal distribution, with an average value , and further, the accuracy of data samples calculated by multiple samplings The formula for the standard deviation is ; The probability distribution function of the S33 normal distribution is , and the normalization processing function of the normal distribution can be transformed into a standard normal distribution probability distribution function , , and then through the transformation function the accuracy rate of the data sample in sampling inspection can be distributed and converted into a standard normal distribution; S34 defines the accuracy difference of the Internet dataset as , with a value range of . That is, the acceptable inspection accuracy is to in the closed interval. Define as the confidence probability that the sampling inspection result falls within the acceptable inspection accuracy interval, and the following formula is derived: ; where: represents the cumulative function of the standard normal distribution from negative infinity to , is the accuracy of the Internet dataset, is the accuracy difference of the Internet dataset, and n represents the data samples drawn from the Internet dataset; S4 Quantitative Evaluation and the difference between, i.e., the accuracy of the data sample can accurately represent the accuracy of the Internet dataset ; S5 Experimentally verify the effectiveness of data modeling in Internet data quality assessment problems.
2. The method according to claim 1, characterized in that, Accuracy of Internet dataset Under the confidence level of 90% - 100%, select the sampling method with the smallest sample size, and the sample size is an integer greater than 1.
3. The method according to claim 2, characterized in that, the data sample size is set at a fixed ratio according to the actual data set.
4. The method according to claim 1, characterized in that, It can be derived through premise assumptions and modeling that the extracted data sample n and the accuracy rate of the Internet dataset and the difference in the accuracy rate of the Internet dataset ultimately determine the credible probability.
5. An Internet data quality assessment system, characterized in that, the system includes: S1 Data import synchronization, where CSV, Excel, and database SQL files are imported, the original data to be inspected is imported, and at the same time, system integration provides a data import synchronization function, and the data set realizes one-time synchronization or incremental synchronization; S2 automatic sampling, according to Formula 1: It can be known that under the same and circumstances, the closer the accuracy rate of the data set is to 0.5, the larger the data sample n to be extracted; in the first Internet data set check, the accuracy rate of the data set is defaulted to 0.5, and the system gives default values for other parameters. The same value is 0.03, the value is greater than or equal to 0.
9. According to the said Formula 1, the system automatically calculates the minimum data sample n to be sampled and conducts random sampling to form the sampling data to be checked; if it is not the first time to check the Internet data set, it is calculated according to the accuracy rate of the previous Internet data set as the default value; S3 Data inspection, the system provides a convenient interface operation for data inspection and comparison with the original data, simplifies the work of comparing with the original data in Internet data inspection, and improves the efficiency of Internet data quality inspection; The S4 data quality assessment gives the Internet data quality assessment results based on the data inspection results, including the accuracy rate of the Internet data set , the difference in the accuracy rate of the Internet data set , and the evaluation credibility probability .
Citation Information
Patent Citations
Data quality evaluation method and apparatus, computer readable storage medium, and terminal
CN107633257A
Single detection credibility evaluation method of binary soft classifier based on sample contribution rate
CN111898697A