A data quality analysis method and system based on big data
Through the big data-based data quality analysis method, the problem of inefficiency in traditional methods when processing large-scale data is solved, a comprehensive and in-depth evaluation of data quality is achieved, analysis efficiency and accuracy are improved, and data-driven decision-making is provided.
Patent Information
- Application Number
- CN202510169791.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-17
AI Technical Summary
Traditional data quality analysis methods are inefficient when processing large-scale data, and cannot conduct comprehensive and in-depth quality assessments of massive data in a timely and effective manner, and it is difficult to comprehensively consider the multi-dimensional characteristics of the data.
The data quality analysis method based on big data is adopted, and the uniqueness and integrity of the data are checked by acquiring and preprocessing monitoring data and status data, extracting the access, timeliness and standardized status characteristics of the data, and building a data quality analysis model, combining technical means such as statistics, feature extraction, index construction and abnormal screening to achieve automated analysis of data quality.
It improves the efficiency and accuracy of data quality analysis, can more comprehensively evaluate the multi-dimensional quality characteristics of data, ensure the reliability and effectiveness of data in various fields, and provide a solid data foundation for scientific decision-making and precise business operations.
Smart Images

Figure CN119646472B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of quality analysis, and in particular to a data quality analysis method and system based on big data. Background Art
[0002] With the rapid development of information technology and the advent of the big data era, data has become an important asset for enterprises and society. In the big data environment, factors such as the surge in data volume, diverse data sources and complex data formats have put forward higher requirements on data quality. High-quality data quality is of great significance for decision-making, business optimization, trend forecasting, etc.
[0003] Traditional data quality analysis methods mostly rely on manual review and rule judgment. This method is inefficient when processing large-scale data and cannot conduct comprehensive and in-depth quality assessment of massive data in a timely and effective manner. At the same time, the multi-dimensional feature analysis of data is not perfect, and it is difficult to comprehensively consider the quality factors such as data integrity, accessibility, timeliness, standardization and consistency. Automated processing and analysis of data through big data can help improve the efficiency and accuracy of data quality analysis. Therefore, a data quality analysis method and system based on big data technology is designed to overcome the shortcomings of existing data quality analysis models, so as to ensure the reliability and effectiveness of big data applications in various fields and provide a solid data foundation for scientific decision-making and precise business operations based on big data. Summary of the invention
[0004] The purpose of the present invention is to provide a data quality analysis method and system based on big data.
[0005] To achieve the above object, the present invention is implemented according to the following technical solutions:
[0006] The present invention comprises the following steps:
[0007] Acquire monitoring data and status data, and preprocess the monitoring data and the status data;
[0008] Check the uniqueness of the monitoring data, count the integrity of the monitoring data to determine the integrity index, and perform feature extraction on the status data to obtain access status data, timeliness status data, and specification status data;
[0009] Input the access status data and the timeliness status data into a status function to obtain an accessibility index and a timeliness index, and determine a normative index according to the normative status data and the standard data;
[0010] Performing abnormal screening on the monitoring data, and determining a consistency index according to the abnormal screening results;
[0011] A data quality analysis model is constructed, and the monitoring data and status data to be analyzed are input into the data quality analysis model to obtain quality analysis results.
[0012] Furthermore, the method for determining the integrity index includes:
[0013] Check the uniqueness constraints in the monitoring data to ensure the uniqueness of the data entity, and eliminate duplicate data according to the uniqueness constraints; the uniqueness constraints include the number of the monitoring subject, monitoring location and monitoring category;
[0014] Statistical methods are used to count the missing data of monitoring data, the K-nearest neighbor algorithm is used to distinguish the missing nature of monitoring data, the cause of the missing is determined and labeled, the filling method is used to process the missing segment data, and the completeness index is calculated based on the missing proportion and missing nature; the missing nature includes required fields and non-required fields.
[0015] Furthermore, the method for obtaining the accessibility index and the timeliness index includes:
[0016] The access status data is input into the first status function to obtain the accessibility index, which is expressed as:
[0017] ,
[0018] in is the accessibility index, is the visit weight, is the access time weight, is the total number of data accesses, is the effective access volume of data, is the number of visitors, is the total access time of the data, is the access time of the data corresponding to the effective access volume, is the data transmission channel utilization, is the network utilization, is the bit error rate of data transmission, is the probability that the data can be accessed under various conditions;
[0019] The aging state data is input into the second state function to obtain the aging index, which is expressed as:
[0020] ,
[0021] in is the timeliness index, is the time delay weight, is the effective period weight, To update the delay, To handle delays, To query delay, is the transmission delay, is the delay coefficient, The remaining validity period of the data. is the validity period of the data standard, is the data update frequency, The data flow speed.
[0022] Furthermore, the method for determining the normative index includes:
[0023] Define the standard state data feature vector of monitoring data as , For the i A canonical state data feature vector, To monitor the characteristic indicators of the normative status data, d is the characteristic indicator dimension, and the standard data group is set according to the standard data format. The characteristic vector of the standard state data corresponding to the standard data is defined as , is the characteristic index of the standard normative status data, calculates the similarity of the characteristic vectors of two normative status data and determines the normative index, the expression is:
[0024] ,
[0025] ,
[0026] in To monitor the similarity between the feature vector of the specification status data and the feature vector of the standard specification status data, is the normative index, , is the comprehensive similarity weight, and , To monitor the state of the data set, is a set of standard specification status data, To monitor the characteristic vector modulus length of the standard status data, is the standard norm state data feature vector modulus, is the resolution coefficient, , are the mean and variance of the monitoring specification status data, , are the mean and variance of the standard specification state data respectively.
[0027] Furthermore, the method for determining the consistency index includes:
[0028] Construct an abnormal data screening model, input the monitoring data into the abnormal data screening model to obtain abnormal screening results, and determine the consistency index based on the abnormal screening results;
[0029] The steps of constructing an abnormal data screening model include: dividing the monitoring data into safe samples and interference samples, optimizing the interference samples, forming a new data set with the optimized interference samples and safe samples and labeling them, establishing an SVM classifier, training the model with a training set, and evaluating the model performance with a test set; the SVM classifier is used to verify the label category of the new data set;
[0030] The specific steps of dividing the safe samples and the interference samples are as follows: using the hypersphere covering algorithm to determine the distance probability and noise probability of the hypersphere samples, using the fast nearest neighbor search algorithm to determine the distance probability of the nearest neighbor search samples, performing weighted fusion on the distance probability of the hypersphere samples and the distance probability of the nearest neighbor search samples to obtain the boundary probability, determining the boundary samples according to the boundary probability, determining the noise samples according to the noise probability, and performing a union operation on the boundary samples and the noise samples to obtain the interference sample set;
[0031] The specific steps for optimizing interference samples are as follows: initialization, mutation, crossover and selection operations are performed on the interference sample set based on the JADE algorithm. The expressions of mutation and crossover operations are:
[0032] ,
[0033] ,
[0034] in For the i The value after the interference sample mutation, For the i The value after the interference samples cross, For the i The fitness of interference samples, For the dimension, is the original interference sample, For the i The scaling factor for the interfering samples, is the interference sample value corresponding to the optimal solution in the population, , For two different interference samples in The value of the dimension, is a constant less than 1, is the fitness control factor, is the average fitness of the population, is the crossover probability in each generation;
[0035] Update the scaling factor and crossover probability, the expression is:
[0036] ,
[0037] ,
[0038] in For the The scaling factor for the iteration, c is the fitness scaling factor, For the t The scaling factor for the iteration, n is the number of interference samples, For the The crossover probability at the iteration, For the t The crossover probability at the iteration;
[0039] The specific steps of establishing the SVM classifier are: constructing the Lagrangian function to determine the optimization target, dualizing the original problem, and outputting the classification decision function. The expressions of the optimization target and the classification decision function are:
[0040] ,
[0041] ,
[0042] in To optimize the goal, is the decision function, , is the Lagrange multiplier, , is the weight of the Lagrange multiplier, is the number of Lagrange multipliers, is the kernel function, , For the input dataset The sample data, m is the number of input data, are data categories, representing interference samples and normal samples respectively, is the regularization parameter, As constraints, is the penalty function, The optimal solution A sample of is the Gaussian kernel function, z is the center point of the kernel function, is the width parameter;
[0043] The abnormal data are obtained based on the classification results of the abnormal data screening model, and the consistency index is determined according to the abnormal data proportion and abnormal amplitude.
[0044] Further, the method for obtaining the quality analysis result includes:
[0045] The completeness index, accessibility index, timeliness index, standardization index and consistency index are combined into an index set, and the index set is divided into a training set and a test set in a ratio of 7:3 using a random forest algorithm.
[0046] Constructing a data quality analysis model, wherein the data quality analysis model uses a BP neural network to learn the data relationship between the indicator sets and outputs a data quality score, sets an indicator threshold in the output layer to mark the input indicator, and outputs the data quality score and the mark; the neural network hidden layer uses a ReLU activation function, and the output layer uses a linear activation function;
[0047] The mean square error loss function is used to measure the difference between the predicted value and the actual value, the SGD optimizer is used to adjust the hyperparameters of the model, and the test set is used to evaluate the model performance;
[0048] The monitoring data and status data to be analyzed are input into the data quality analysis model to obtain data quality scores and labels.
[0049] In a second aspect, a data quality analysis system based on big data includes:
[0050] Data acquisition module: used to obtain monitoring data and status data, and pre-process the monitoring data and status data;
[0051] Data processing module: used to determine the integrity index by counting the integrity of the monitoring data, to obtain the accessibility index and the timeliness index, to determine the normative index according to the normative status data and the standard data, and to determine the consistency index by performing abnormal screening on the monitoring data;
[0052] Screening model module: used to build an abnormal data screening model, input monitoring data into the abnormal data screening model to obtain abnormal screening results, and determine the consistency index based on the abnormal screening results;
[0053] Scoring model module: constructing a data quality analysis model, inputting the monitoring data and status data to be analyzed into the data quality analysis model to obtain quality analysis results;
[0054] Intelligent supervision module: used to store, view and manage the monitoring data, the status data and the quality analysis results, and display the quality analysis results to the user.
[0055] The beneficial effects of the present invention are:
[0056] The present invention is a data quality analysis method and system based on big data. Compared with the prior art, the present invention has the following technical effects:
[0057] The present invention can improve the data preprocessing capabilities and enhance the model adaptability in data quality analysis through data statistics, feature extraction, index construction, similarity calculation, anomaly screening and model building steps, thereby improving the efficiency and accuracy of data quality analysis. The data quality analysis technology can be optimized, which can greatly save resources and improve work efficiency. It can realize the analysis of data quality and provide support for data-driven decision-making. It is of great significance to improving the management and utilization efficiency of data. It can adapt to different data quality analysis systems based on big data and the terminal analysis needs of data quality analysis based on big data of different users, and has a certain universality. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 The present invention is a flowchart of the steps of a data quality analysis method based on big data. DETAILED DESCRIPTION
[0059] The present invention is further described below by means of specific embodiments. The illustrative embodiments and descriptions of the present invention are used to explain the present invention but are not intended to limit the present invention.
[0060] The present invention provides a data quality analysis method and system based on big data, comprising the following steps:
[0061] like Figure 1 As shown, in this embodiment, the following steps are included:
[0062] Acquire monitoring data and status data, and pre-process the monitoring data and the status data;
[0063] Check the uniqueness of the monitoring data, count the integrity of the monitoring data to determine the integrity index, and perform feature extraction on the status data to obtain access status data, timeliness status data, and specification status data;
[0064] Input the access status data and the timeliness status data into a status function to obtain an accessibility index and a timeliness index, and determine a normative index according to the normative status data and the standard data;
[0065] Performing abnormal screening on the monitoring data, and determining a consistency index according to the abnormal screening results;
[0066] A data quality analysis model is constructed, and the monitoring data and status data to be analyzed are input into the data quality analysis model to obtain quality analysis results.
[0067] In this embodiment, the method for determining the integrity index includes:
[0068] Check the uniqueness constraints in the monitoring data to ensure the uniqueness of the data entity, and eliminate duplicate data according to the uniqueness constraints; the uniqueness constraints include the number of the monitoring subject, monitoring location and monitoring category;
[0069] The missing data of the monitoring data are counted by the statistical method, the missing nature of the monitoring data is distinguished by the K-nearest neighbor algorithm, the missing causes are determined and labeled, the missing segment data is processed by the filling method, and the completeness index is calculated according to the missing ratio and missing nature; the missing nature includes required fields and non-required fields;
[0070] In the actual evaluation, taking the stress monitoring data of a bridge bottom as an example, three groups of monitoring data were randomly selected for integrity statistics: 1. Data identification 001-A1-C1, stress value missing 1.2% / required field / sensor failure, ambient temperature missing 0.5% / non-required field / data transmission error; 2. Data identification 001-A5-C2, traffic flow missing 0.6% / non-required field / accidental omission; 3. Data identification 003-A1-C3, weather conditions missing 0.5% / non-required field / data transmission error;
[0071] The completeness index was calculated based on the missingness ratio and the nature of the missingness. The weight of the required field was 2 and the weight of the non-required field was 1. The completeness indexes of the three sets of data were 97.1%, 99.4%, and 99.5%, respectively.
[0072] In this embodiment, the method for obtaining the accessibility index and the timeliness index includes:
[0073] The access status data is input into the first status function to obtain the accessibility index, which is expressed as:
[0074] ,
[0075] in is the accessibility index, is the visit weight, is the access time weight, is the total number of data accesses, is the effective access volume of data, is the number of visitors, is the total access time of the data, is the access time of the data corresponding to the effective access volume, is the data transmission channel utilization, is the network utilization, is the bit error rate of data transmission, is the probability that the data can be accessed under various conditions;
[0076] The aging state data is input into the second state function to obtain the aging index, which is expressed as:
[0077] ,
[0078] in is the timeliness index, is the time delay weight, is the effective period weight, To update the delay, To handle delays, To query delay, is the transmission delay, is the delay coefficient, The remaining validity period of the data. is the validity period of the data standard, is the data update frequency, is the data flow speed;
[0079] In the actual evaluation, the access status data corresponding to the three groups of monitoring data are (data identification, total visits, effective visits, number of visitors, total access time, effective access time, transmission channel utilization, network utilization, transmission bit error rate, probability of access): (001-A1-C1, 100, 90, 10, 1000s, 900s, 0.8, 0.9, 0.01, 0.95), (001-A5-C2, 120, 110, 15, 1200s, 1100s, 0.85, 0.95, 0.02, 0.90), (003-A1-C3, 90, 80, 8, 900s, 800s, 0.75, 0.85, 0.015, 0.98);
[0080] Get the visit weight =0.4, access time weight =0.6, the access status data is input into the first status function to obtain accessibility indexes of 1.5468, 1.6827, and 1.3703;
[0081] The time-effectiveness status data corresponding to the three groups of monitoring data are (data identification, update delay, processing delay, query delay, transmission delay, delay coefficient, remaining validity period, standard validity period, update frequency, circulation speed): (001-A1-C1, 30s, 20s, 10s, 5s, 0.1, 5min, 10min, 1Hz, 100 units / minute), (001-A5-C2, 25s, 15s, 8s, 4s, 0.15, 4min, 10min, 1Hz, 120 units / minute), (003-A1-C3, 28s, 18s, 9s, 3s, 0.12, 6min, 10min, 1Hz, 110 units / minute);
[0082] Take the time delay weight =0.7, effective period weight =0.3, and the aging state data is input into the second state function to obtain the aging index of 2.8516, 2.7712, and 2.5543.
[0083] In this embodiment, the method for determining the normative index includes:
[0084] Define the standard state data feature vector of monitoring data as , For the i A canonical state data feature vector, To monitor the characteristic indicators of the normative status data, d is the characteristic indicator dimension, and the standard data group is set according to the standard data format. The characteristic vector of the standard state data corresponding to the standard data is defined as , is the characteristic index of the standard normative status data, calculates the similarity of the characteristic vectors of two normative status data and determines the normative index, the expression is:
[0085] ,
[0086] ,
[0087] in To monitor the similarity between the feature vector of the specification status data and the feature vector of the standard specification status data, is the normative index, , is the comprehensive similarity weight, and , To monitor the state of the data set, is a set of standard specification status data, To monitor the characteristic vector modulus length of the standard status data, is the standard norm state data feature vector modulus, is the resolution coefficient, , are the mean and variance of the monitoring specification status data, , are the mean and variance of the standard specification state data, respectively;
[0088] In the actual evaluation, the standard status data corresponding to the three groups of monitoring data are as follows (data identification: integer ratio, floating point ratio, string ratio, required data length, non-required data length, text ratio, symbol ratio, timestamp ratio, tag ratio, data integrity, data size): (001-A1-C1: 60%, 30%, 10%, 100 characters, 50 characters, 70%, 5%, 10%, 15%, 0.971, 1MB), (001-A5-C2: 65%, 25%, 10%, 120 characters, 60 characters, 75%, 3%, 10%, 12%, 0.994, 1.15MB), (001-A5-C2: 70%, 20%, 10%, 150 characters, 70 characters, 80%, 2%, 10%, 18%, 0.995, 1.09MB), the standard data is (50%, 40%, 10%, 130 characters, 70 characters, 85%, 1%, 5%, 9%, 1, 1.1MB);
[0089] Take the comprehensive similarity weight ,The similarities of the normative state data feature vectors of the three sets of monitoring data and the standard data are calculated to be 0.9445, 0.8333, and 0.7781, and the normative indexes of the three sets of data are 2.1175, 2.2927, and 2.2178.
[0090] In this embodiment, the method for determining the consistency index includes:
[0091] Construct an abnormal data screening model, input the monitoring data into the abnormal data screening model to obtain abnormal screening results, and determine the consistency index based on the abnormal screening results;
[0092] The steps of constructing an abnormal data screening model include: dividing the monitoring data into safe samples and interference samples, optimizing the interference samples, forming a new data set with the optimized interference samples and safe samples and labeling them, establishing an SVM classifier, training the model with a training set, and evaluating the model performance with a test set; the SVM classifier is used to verify the label category of the new data set;
[0093] The specific steps of dividing the safe samples and the interference samples are as follows: using the hypersphere covering algorithm to determine the sample point distance matrix, the radius of different types of hyperspheres, calculating the number of times the sample points are covered by different types of hyperspheres, determining the hypersphere sample distance probability and the noise probability according to the calculation results of the hypersphere covering algorithm, using the fast nearest neighbor search algorithm to calculate the distance between each sample point and the nearest sample point of the same category and determine the nearest neighbor search sample distance probability, performing weighted fusion on the hypersphere sample distance probability and the nearest neighbor search sample distance probability to obtain the boundary probability, determining the boundary sample according to the boundary probability, determining the noise sample according to the noise probability, and performing a union operation on the boundary sample and the noise sample to obtain the interference sample set;
[0094] The specific steps for optimizing interference samples are as follows: initialization, mutation, crossover and selection operations are performed on the interference sample set based on the JADE algorithm. The expressions of mutation and crossover operations are:
[0095] ,
[0096] ,
[0097] in For the i The value after the interference sample mutation, For the i The value after the interference samples cross, For the i The fitness of interference samples, For the dimension, is the original interference sample, For the i The scaling factor for the interfering samples, is the interference sample value corresponding to the optimal solution in the population, , For two different interference samples in The value of the dimension, is a constant less than 1, is the fitness control factor, is the average fitness of the population, is the crossover probability in each generation;
[0098] Update the scaling factor and crossover probability, the expression is:
[0099] ,
[0100] ,
[0101] in For the The scaling factor for the iteration, c is the fitness scaling factor, For the t The scaling factor for the iteration, n is the number of interference samples, For the The crossover probability at the iteration, For the t The crossover probability at the iteration;
[0102] The specific steps of establishing the SVM classifier are: constructing the Lagrangian function to determine the optimization target, dualizing the original problem, and outputting the classification decision function. The expressions of the optimization target and the classification decision function are:
[0103] ,
[0104] ,
[0105] in To optimize the goal, is the decision function, , is the Lagrange multiplier, , is the weight of the Lagrange multiplier, is the number of Lagrange multipliers, is the kernel function, , For the input dataset The sample data, m is the number of input data, are data categories, representing interference samples and normal samples respectively, is the regularization parameter, As constraints, is the penalty function, The optimal solution A sample of is the Gaussian kernel function, z is the center point of the kernel function, is the width parameter;
[0106] Abnormal data are obtained based on the classification results of the abnormal data screening model, and the consistency index is determined according to the abnormal data proportion and abnormal amplitude;
[0107] In the actual evaluation, the three groups of monitoring data were input into the abnormal data screening model to obtain the abnormal screening results (data identification: average abnormal amplitude, abnormal proportion, consistency index): (001-A1-C1: 1, 5%, 0.9500), (001-A5-C2: 1.5, 3%, 0.9550), (001-A5-C2: 0.8, 7%, 0.9440).
[0108] In this embodiment, the method for obtaining the quality analysis result includes:
[0109] The completeness index, accessibility index, timeliness index, standardization index and consistency index are combined into an index set, and the index set is divided into a training set and a test set in a ratio of 7:3 using a random forest algorithm.
[0110] Constructing a data quality analysis model, wherein the data quality analysis model uses a BP neural network to learn the data relationship between the indicator sets and outputs a data quality score, sets an indicator threshold in the output layer to mark the input indicator, and outputs the data quality score and the mark; the neural network hidden layer uses a ReLU activation function, and the output layer uses a linear activation function;
[0111] The mean square error loss function is used to measure the difference between the predicted value and the actual value, the SGD optimizer is used to adjust the hyperparameters of the model, and the test set is used to evaluate the model performance;
[0112] Input the monitoring data and status data to be analyzed into the data quality analysis model to obtain data quality scores and labels;
[0113] In the actual evaluation, the monitoring data and status data to be analyzed are input into the data quality analysis model to obtain data quality scores and markings: (001-A1-C1: 4.9141 points, some required data are missing), (001-A5-C2: 4.7220 points), (001-A5-C2: 4.0198 points, poor data standardization).
[0114] In a second aspect, a data quality analysis system based on big data includes:
[0115] Data acquisition module: used to obtain monitoring data and status data, and pre-process the monitoring data and status data;
[0116] Data processing module: used to determine the integrity index by counting the integrity of the monitoring data, to obtain the accessibility index and the timeliness index, to determine the normative index according to the normative status data and the standard data, and to determine the consistency index by performing abnormal screening on the monitoring data;
[0117] Screening model module: used to build an abnormal data screening model, input monitoring data into the abnormal data screening model to obtain abnormal screening results, and determine the consistency index based on the abnormal screening results;
[0118] Scoring model module: constructing a data quality analysis model, inputting the monitoring data and status data to be analyzed into the data quality analysis model to obtain quality analysis results;
[0119] Intelligent supervision module: used to store, view and manage the monitoring data, the status data and the quality analysis results, and display the quality analysis results to the user.
[0120] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A data quality analysis method based on big data, characterized in that: The following steps are involved: S1. Acquire monitoring data and status data, and pre-process the monitoring data and the status data; S2, checking the uniqueness of the monitoring data, counting the integrity of the monitoring data to determine the integrity index, and performing feature extraction on the status data to obtain access status data, timeliness status data and specification status data; S3, inputting the access status data and the timeliness status data into a status function to obtain an accessibility index and a timeliness index, and determining a normative index according to the normative status data and the standard data; S4, performing abnormal screening on the monitoring data, and determining a consistency index according to the abnormal screening result; S5. Construct a data quality analysis model, and input the monitoring data and status data to be analyzed into the data quality analysis model to obtain quality analysis results; The method for obtaining the accessibility index and the timeliness index includes: The access status data is input into the first status function to obtain the accessibility index, which is expressed as: , in is the accessibility index, is the visit weight, is the access time weight, is the total number of data accesses, is the effective access volume of data, is the number of visitors, is the total access time of the data, is the access time of the data corresponding to the effective access volume, is the data transmission channel utilization, is the network utilization, is the bit error rate of data transmission, is the probability that the data can be accessed under various conditions; The aging state data is input into the second state function to obtain the aging index, which is expressed as: , in is the timeliness index, is the time delay weight, is the effective period weight, To update the delay, To handle delays, To query delay, is the transmission delay, is the delay coefficient, The remaining validity period of the data. is the validity period of the data standard, is the data update frequency, The data flow speed.
2. According to the data quality analysis method based on big data as described in claim 1, it is characterized in that: The method for determining the integrity index comprises: Checking the uniqueness constraints in the monitoring data and eliminating duplicate data according to the uniqueness constraints; the uniqueness constraints include the number of the monitoring subject, monitoring location and monitoring category; Statistical methods are used to count the missing data of monitoring data, the K-nearest neighbor algorithm is used to distinguish the missing nature of monitoring data, the cause of the missing is determined and labeled, the filling method is used to process the missing segment data, and the completeness index is calculated based on the missing proportion and missing nature; the missing nature includes required fields and non-required fields.
3. According to the data quality analysis method based on big data as claimed in claim 1, it is characterized in that: The method for determining the normative index comprises: Define the standard state data feature vector of monitoring data as , For the i A canonical state data feature vector, To monitor the characteristic indicators of the normative status data, d is the characteristic indicator dimension, and the standard data group is set according to the standard data format. The characteristic vector of the standard state data corresponding to the standard data is defined as , is the characteristic index of the standard normative status data, calculates the similarity of the characteristic vectors of two normative status data and determines the normative index, the expression is: , , in To monitor the similarity between the feature vector of the specification status data and the feature vector of the standard specification status data, is the normative index, , is the comprehensive similarity weight, and , To monitor the state of the data set, is a set of standard specification status data, To monitor the characteristic vector modulus length of the standard status data, is the standard norm state data feature vector modulus, is the resolution coefficient, , are the mean and variance of the monitoring specification status data, , are the mean and variance of the standard specification state data respectively.
4. According to the data quality analysis method based on big data as claimed in claim 1, it is characterized in that: The method for determining the consistency index comprises: Construct an abnormal data screening model, input the monitoring data into the abnormal data screening model to obtain abnormal screening results, and determine the consistency index based on the abnormal screening results; The steps of constructing an abnormal data screening model include: dividing the monitoring data into safe samples and interference samples, optimizing the interference samples, forming a new data set with the optimized interference samples and safe samples and labeling them, establishing an SVM classifier, training the model with a training set, and evaluating the model performance with a test set; the SVM classifier is used to verify the label category of the new data set; The specific steps of dividing the safe samples and the interference samples are as follows: using the hypersphere covering algorithm to determine the distance probability and noise probability of the hypersphere samples, using the fast nearest neighbor search algorithm to determine the distance probability of the nearest neighbor search samples, performing weighted fusion on the distance probability of the hypersphere samples and the distance probability of the nearest neighbor search samples to obtain the boundary probability, determining the boundary samples according to the boundary probability, determining the noise samples according to the noise probability, and performing a union operation on the boundary samples and the noise samples to obtain the interference sample set; The specific steps for optimizing interference samples are as follows: initialization, mutation, crossover and selection operations are performed on the interference sample set based on the JADE algorithm. The expressions of mutation and crossover operations are: , , in For the i The value after the interference sample mutation, For the i The value after the interference samples cross, For the i The fitness of interference samples, For the dimension, is the original interference sample, For the i The scaling factor for the interfering samples, is the interference sample value corresponding to the optimal solution in the population, , For two different interference samples in The value of the dimension, is a constant less than 1, is the fitness control factor, is the average fitness of the population, is the crossover probability in each generation; Update the scaling factor and crossover probability, the expression is: , , in For the The scaling factor for the iteration, c is the fitness scaling factor, For the t The scaling factor for the iteration, n is the number of interference samples, For the The crossover probability at the iteration, For the t The crossover probability at the iteration; The specific steps of establishing the SVM classifier are: constructing the Lagrangian function to determine the optimization target, dualizing the original problem, and outputting the classification decision function. The expressions of the optimization target and the classification decision function are: , , in To optimize the goal, is the decision function, , is the Lagrange multiplier, , is the weight of the Lagrange multiplier, is the number of Lagrange multipliers, is the kernel function, , For the input dataset The sample data, m is the number of input data, are data categories, representing interference samples and normal samples respectively, is the regularization parameter, As constraints, is the penalty function, The optimal solution A sample of is the Gaussian kernel function, z is the center point of the kernel function, is the width parameter; The abnormal data are obtained based on the classification results of the abnormal data screening model, and the consistency index is determined according to the abnormal data proportion and abnormal amplitude.
5. According to the data quality analysis method based on big data as claimed in claim 1, it is characterized in that: The method for obtaining the quality analysis result comprises: The completeness index, accessibility index, timeliness index, standardization index and consistency index are combined into an index set, and the index set is divided into a training set and a test set in a ratio of 7:3 using a random forest algorithm. Constructing a data quality analysis model, wherein the data quality analysis model uses a BP neural network to learn the data relationship between the indicator sets and outputs a data quality score, sets an indicator threshold in the output layer to mark the input indicator, and outputs the data quality score and the mark; the neural network hidden layer uses a ReLU activation function, and the output layer uses a linear activation function; The mean square error loss function is used to measure the difference between the predicted value and the actual value, the SGD optimizer is used to adjust the hyperparameters of the model, and the test set is used to evaluate the model performance; The monitoring data and status data to be analyzed are input into the data quality analysis model to obtain data quality scores and labels.
6. A data quality analysis system based on big data, used to execute the method according to any one of claims 1 to 5, characterized in that: include: Data acquisition module: used to obtain monitoring data and status data, and pre-process the monitoring data and status data; Data processing module: used to determine the integrity index by counting the integrity of the monitoring data, to obtain the accessibility index and the timeliness index, to determine the normative index according to the normative status data and the standard data, and to determine the consistency index by performing abnormal screening on the monitoring data; Screening model module: used to build an abnormal data screening model, input monitoring data into the abnormal data screening model to obtain abnormal screening results, and determine the consistency index based on the abnormal screening results; Scoring model module: constructing a data quality analysis model, inputting the monitoring data and status data to be analyzed into the data quality analysis model to obtain quality analysis results; Intelligent supervision module: used to store, view and manage the monitoring data, the status data and the quality analysis results, and display the quality analysis results to the user.
Citation Information
Patent Citations
Device state monitoring data quality evaluation system
CN107491381A
Data quality evaluation method and system based on big data analysis
CN116467292A