Table data quality automatic discrimination method and system based on multi-dimensional indexes
By adopting automatic discrimination method based on multi-dimensional indicators in data quality detection, and using technical means such as random forest and Gaussian hybrid models, the problems of limited accuracy and subjectivity of detection results in the existing technology are solved, and more efficient and reliable data quality detection and governance are achieved.
Patent Information
- Application Number
- CN202510321892.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art ignores the overall structure and semantic information of the data in data quality detection, resulting in limited accuracy of the detection results, lack of clear basis and standards, and has great subjectivity and cannot adapt to changeable business scenarios.
The automatic discrimination method of table data quality based on multi-dimensional indicators is adopted. By obtaining the metadata table information of the business system, a data table recognition and filtering model of the random forest classification algorithm is constructed, and the field types are identified and classified. The Gaussian mixed model's expected maximization clustering algorithm is used to build a detection rule matching model to realize automatic detection and verification of data quality.
It improves the accuracy and reliability of data quality detection, reduces manual intervention and manual operations, improves data quality governance efficiency, and reduces labor and time costs.
Smart Images

Figure CN120144576A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data quality detection, and mainly relates to a method and system for automatically discriminating the quality of table data based on multi-dimensional indicators. Background Art
[0002] With the rapid development of information technology, the generation of data has shown an explosive growth. The wide application of Internet of Things devices, the popularity of social media, the prosperity of e-commerce, etc. have led to an exponential growth in the amount of data, forming a vast amount of data resources. At the same time, the sources of data have become increasingly diversified, including structured data, semi-structured data, and unstructured data. The complexity of data is not only reflected in the diversity of data types, but also in the relevance and dependence between data.
[0003] Data quality problems are common difficulties faced by enterprises and organizations in the process of data management. Inaccurate, incomplete, inconsistent, and outdated data may lead to serious consequences. In terms of business decisions, incorrect data may cause enterprises to make wrong market forecasts, investment decisions, and strategic plans, resulting in huge economic losses. In the field of scientific research, data quality problems may lead to deviations and errors in research results, affecting the accuracy and reliability of scientific discoveries. In government management, inaccurate population data, economic data, etc. may affect the formulation and implementation effects of policies, having an adverse impact on social stability and development. At present, the level of intelligence and automation in the data quality detection industry is still low, and the popularization rate is small. Most enterprises still rely on manual rules and reviews to supplement the deficiencies of automated detection.
[0004] The Chinese invention patent with the publication number "CN117762914A" discloses a "Data Quality Detection Method and System", which specifically discloses "obtaining at least one piece of data to be detected in a data source and disassembling the data to be detected into the states of data characters; constructing screening rules and grouping data characters with the same type of character information together; constructing a data analysis model with different systems installed inside. Compared with traditional data quality detection methods, the present invention can detect data quality problems more precisely by disassembling the data to be detected into the states of data characters. Secondly, by constructing screening rules and a data analysis model, the information within the data characters can be extracted more accurately, and multiple data character items can be obtained. Thirdly, by comparing historical data character metrics with the data character items, a more accurate item difference score can be obtained. Finally, by constructing discrimination rules and setting discrimination thresholds, it can be more accurately determined whether the data quality meets the standard", but this method only disassembles the data to be detected into the states of data characters, ignoring the overall structure and semantic information of the data, resulting in limited accuracy of the detection results; in addition, the process of constructing discrimination rules and setting discrimination thresholds in this method lacks clear basis and standards, with relatively large subjectivity; at the same time, this method does not fully consider the diversity and complexity of business systems and cannot adapt to changing business scenarios. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present application provides a method and system for automatically discriminating the quality of table data based on multi-dimensional indicators.
[0006] The technical solution of the present application is as follows:
[0007] On the one hand, the present invention proposes a method for automatically discriminating the quality of table data based on multi-dimensional indicators, and the method includes:
[0008] Obtain the meta-table information of the business system, including table description, field description, and field type, and construct a meta-table information data set; use the random forest classification algorithm to construct a data table identification and screening model, and the data table identification and screening model performs data cleaning on the meta-table information to obtain a data table to be detected;
[0009] Construct a field type identification model including a field major type identification model and a field minor type identification model; use the field major type identification model to identify and classify the fields of the data table to be detected to obtain the field major types; use the field minor type identification model to perform feature extraction and type annotation on the major field types to obtain the field minor types;
[0010] A normal data quality detection library is created based on the professional knowledge of business personnel and the information quality rule verification rules of the historical metadata table, and a detection rule matching model is constructed using an expectation maximization clustering algorithm based on a Gaussian mixture model; the field small types are matched based on the normal data quality detection library and the detection rule matching model to obtain a detection matching result;
[0011] A closed-loop quality problem verification model is constructed to perform quality inspection and verification on the detection and matching results, and obtain the quality inspection results of the data table to be inspected.
[0012] Preferably, a random forest classification algorithm is used to construct a data table identification and screening model, specifically:
[0013] The metadata table information data set is subjected to data preprocessing, including type conversion, data normalization and feature extraction; the type conversion specifically includes converting the text type fields in the metadata table information data set into binary types; the data normalization specifically includes normalizing the data value type fields in the metadata table information data set; the feature extraction specifically includes calculating the correlation coefficient between each field and irrelevant data, selecting the fields whose correlation coefficient is within a preset interval, and generating a feature set; the irrelevant data includes an empty table, a low-heat table, an intermediate table, a log table, a low-heat field and an empty data field;
[0014] Initialize the random forest model and preset initial hyperparameters, including the number of decision trees, the maximum depth of the decision tree, the number of features randomly selected for each node, and the minimum number of samples;
[0015] A self-service sampling method is used to randomly extract samples from the metadata table information data set with replacement to construct a decision tree with a preset number of trees; at each node of the decision tree, a feature with a preset number of features is randomly selected from the feature set, and the information gain of the randomly selected feature for the current node is calculated; the feature with the largest information gain is selected as the optimal feature for node division;
[0016] Recursively construct decision trees until the maximum depth of the decision tree reaches a preset value or the number of samples of the node is less than the preset minimum number of samples, and then all decision trees are combined into a random forest to obtain a data table recognition and screening model;
[0017] The metadata table information data set is input into the data table identification and screening model to obtain a data table prediction result; irrelevant data in the metadata table information data set is eliminated according to the data table prediction result to obtain a data table to be detected.
[0018] Preferably, the field category recognition model is constructed based on automatic machine learning technology and random forest classification algorithm, specifically:
[0019] Select and configure an automated machine learning platform, preset the search algorithm, evaluation metrics, and maximum running time, and define the hyperparameter space for the random forest;
[0020] Use the data table to be detected as the input to start the search process. The automated machine learning platform searches for the optimal solution of the random forest hyperparameters in the defined random forest hyperparameter space according to the preset search algorithm, and calculates the evaluation metrics of the current optimal solution;
[0021] Iterate the search process until the preset maximum running time is reached. The automated machine learning platform selects the optimal solution corresponding to the maximum value of the evaluation metrics as the hyperparameter combination of the random forest model to obtain the field major type recognition model;
[0022] Input the data table to be detected into the field major type recognition model to obtain the prediction results of the major types to which each field in the data table to be detected belongs, and obtain the field major types; the field major types include coding type, text type, numerical type, date type, and time series type.
[0023] Preferably, use the field minor type recognition model to perform feature extraction and type annotation on the major field types, specifically:
[0024] Calculate the mean, variance, and maximum and minimum values as the features of the numerical major field;
[0025] Calculate the data difference between adjacent fields as the feature of the time series major field;
[0026] Perform word segmentation on the text major field, analyze the part-of-speech of each word segment, and use the keyword extraction algorithm to extract the key information in the text major field as the feature of the text major field;
[0027] Extract the date information in the date major field, including year, month, day, week, and quarter, and calculate the difference between adjacent dates. Use the date information and the difference as the features of the date major field;
[0028] Analyze the hierarchical structure and coding rules of the coding major field, and extract the upper-level coding and the major field of the deep hierarchy as the features of the coding major field;
[0029] Input the extracted features of each type into the field minor type recognition model to obtain the field minor types.
[0030] Preferably, use the expectation-maximization algorithm based on the Gaussian mixture model to construct a detection rule matching model, specifically:
[0031] Initialize the parameters of the Gaussian mixture model, including the mean, covariance matrix, and weight of each Gaussian distribution, and preset the number of clusters;
[0032] The optimal Gaussian mixture model parameters are obtained by using the expectation maximization algorithm, and the expectation maximization clustering algorithm includes an expectation step and a maximization step;
[0033] In the expectation step, the posterior probability of the Gaussian distribution is calculated, which is expressed by the formula:
[0034]
[0035] In the formula, γ ik represents the posterior distribution that the i-th small field belongs to the k-th Gaussian distribution; x i represents the i-th field small type; N(x i |μ k , λ k ) represents the probability density function of the i-th field small type under the k-th Gaussian distribution; ω k represents the weight of the k-th Gaussian distribution; μ k represents the mean of the k-th Gaussian distribution; λ k represents the covariance matrix of the k-th Gaussian distribution; k represents the index value of the k-th Gaussian distribution; K represents the number of clusters; i represents the index value of the i-th field small type;
[0036] In the maximization step, the Gaussian mixture model parameters are updated according to the posterior probability, which is expressed by the formula:
[0037]
[0038]
[0039] In the formula, M represents the number of field small types; T represents the transpose operation;
[0040] The expectation step and the maximization step are iteratively executed until the maximum number of iterations is reached, and the optimal Gaussian mixture model parameters are obtained to obtain a detection rule matching model.
[0041] Preferably, the field small types are matched based on the normal data quality detection library and the detection rule matching model, specifically:
[0042] The normal data quality detection library includes typical feature templates for various field small types; the detection types in the detection rule matching model include outlier detection, time series difference detection, enumeration detection, name detection, size comparison detection, and missing value detection;
[0043] The input field small types are compared with the corresponding feature templates and detection types to determine whether the input field small types conform to the typical features and detection rules, and a detection matching result is obtained;
[0044] If it meets the criteria, it is determined that the current field sub-type is a valid match and marked as a normal field; if it does not meet any of them, it is determined that the current field sub-type is an invalid match and marked as an abnormal field;
[0045] For the abnormal fields, start the abnormal handling process, and further confirm the authenticity and accuracy of the data by means of manual review, or correct the abnormal fields according to business rules.
[0046] Preferably, a closed-loop quality problem verification model is constructed to perform quality detection and verification on the detection and matching results. Specifically:
[0047] The closed-loop quality problem verification model includes a data sampling module, a quality verification module, a result evaluation module, and a feedback correction module;
[0048] The data sampling module presets a sampling method and selects verification samples from the detection and matching results; the quality verification module uses built-in business knowledge, industry standards, and detection experience to verify the verification samples; the result evaluation module calculates the evaluation indicators of the verification samples to obtain an evaluation result; the feedback correction model feeds back the evaluation result to the discrimination system, and the discrimination system analyzes the problems existing in the detection and matching results and formulates corresponding correction opinions.
[0049] On the other hand, the present invention also proposes an automatic table data quality discrimination system based on multi-dimensional indicators. The system includes a data acquisition module, a field identification module, a quality detection module, and a result output module, where:
[0050] The data acquisition module is used to obtain the meta-data table information of the business system, including table descriptions, field descriptions, and field types; use the random forest classification algorithm to construct a data table identification and screening model, and the data table identification and screening model performs data cleaning on the meta-data table information to obtain a data table to be detected; transmit the data table to be detected to the field identification module;
[0051] The field identification module is used to construct a field type identification model including a field major type identification model and a field sub-type identification model; use the field major type identification model to identify and classify the fields of the data table to be detected to obtain the field major type; use the field sub-type identification model to perform feature extraction and type annotation on the large field type to obtain the field sub-type;
[0052] The quality detection module is used to create a normal data quality detection library according to the professional knowledge of business personnel and the quality rule verification rules of historical meta-data table information, and use the expectation maximization clustering algorithm based on the Gaussian mixture model to construct a detection rule matching model; perform matching on the field sub-types based on the normal data quality detection library and the detection rule matching model to obtain a detection and matching result;
[0053] Build a closed-loop quality problem verification model to perform quality inspection and verification on the detection and matching results, and obtain the quality inspection results of the data table to be detected;
[0054] The result output module is used to display the quality inspection results of the data table to be detected.
[0055] On the other hand, the present invention also proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a method for automatically discriminating the quality of table data based on multi-dimensional indicators as described in any embodiment of the present invention.
[0056] On the other hand, the present invention also proposes a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements a method for automatically discriminating the quality of table data based on multi-dimensional indicators as described in any embodiment of the present invention.
[0057] Compared with the prior art, the beneficial effects of the present invention are:
[0058] 1) The present invention provides a method and system for automatically discriminating the quality of table data based on multi-dimensional indicators. By using machine learning algorithms and natural language processing technologies, it automatically identifies and classifies data tables and field types, ensuring that quality inspection rules can accurately match data features, avoiding errors and mismatches in traditional manual rule-making, and improving the accuracy and reliability of data quality inspection;
[0059] 2) The present invention provides a method and system for automatically discriminating the quality of table data based on multi-dimensional indicators. By constructing an intelligent detection model, it realizes the automation of the data quality discrimination process, reduces manual intervention and manual operations, and improves the overall data quality governance efficiency;
[0060] 3) The present invention provides a method and system for automatically discriminating the quality of table data based on multi-dimensional indicators. Through automated data quality detection and repair, it reduces the labor cost and time cost of enterprises in data quality management, can achieve rapid data quality assessment and problem repair, and reduces the high labor cost required for traditional data quality management. Brief Description of the Drawings
[0061] Figure 1 It is a flowchart of the method according to an embodiment of the present invention. Detailed Embodiments
[0062] The specific embodiments of the present invention will be described below to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0063] The present invention provides the following technical solution: A method and system for automatically discriminating the quality of table data based on multi-dimensional indicators.
[0064] Embodiment 1
[0065] Specifically refer to Figure 1 , this embodiment provides a method for automatically discriminating the quality of table data based on multi-dimensional indicators. The specific steps include:
[0066] S1. Obtain the meta-table information of the business system, including table description, field description, and field type, and construct a meta-table information data set;
[0067] S2. Use the random forest classification algorithm to construct a data table recognition and screening model, and perform data preprocessing on the meta-table information data set, including type conversion, data normalization, and feature extraction; the type conversion specifically converts the text-type fields in the meta-table information data set into binary types; the data normalization specifically normalizes the data value-type fields in the meta-table information data set; the feature extraction specifically calculates the correlation coefficients between each field and the irrelevant data, selects the fields with the correlation coefficients within the preset interval, and generates a feature set; the irrelevant data includes empty tables, low-popularity tables, intermediate tables, log tables, low-popularity fields, and empty data fields;
[0068] Initialize the random forest model, and preset the initial hyperparameters, including the number of decision trees, the maximum depth of the decision tree, the number of randomly selected features for each node, and the minimum number of samples;
[0069] Use the bootstrap sampling method to randomly draw samples with replacement from the meta-table information data set to construct decision trees with a preset number of trees; at each node of each decision tree, randomly select a preset number of features from the feature set, and calculate the information gain of the randomly selected features for the current node; select the feature with the maximum information gain as the optimal feature for node division;
[0070] Recursively construct decision trees until the maximum depth of the decision tree reaches the preset value or the number of samples at the node is less than the preset minimum number of samples, and obtain all decision trees combined into a random forest to obtain a data table recognition and screening model;
[0071] S3. The data table recognition and screening model performs data cleaning on the meta-data table information. The meta-data table information dataset is input into the data table recognition and screening model to obtain a data table prediction result. According to the data table prediction result, irrelevant data in the meta-data table information dataset is removed to obtain a data table to be detected.
[0072] S4. Build a field type recognition model including a major field type recognition model and a minor field type recognition model.
[0073] S41. Use the major field type recognition model to identify and classify the fields of the data table to be detected to obtain major field types.
[0074] The major field type recognition model is built based on automated machine learning technology and a random forest classification algorithm. An automated machine learning platform is selected and configured, a search algorithm, evaluation metrics, and a maximum running time are preset, and a random forest hyperparameter space is defined.
[0075] Taking the data table to be detected as the input, start the search process. The automated machine learning platform searches for the optimal solution of the random forest hyperparameters in the defined random forest hyperparameter space according to the preset search algorithm, and calculates the evaluation metrics of the current optimal solution.
[0076] Iterate the search process until the preset maximum running time is reached. The automated machine learning platform selects the optimal solution corresponding to the maximum value of the evaluation metrics as the hyperparameter combination of the random forest model to obtain the major field type recognition model.
[0077] Input the data table to be detected into the major field type recognition model to obtain the prediction result of the major category to which each field of the data table to be detected belongs, and obtain major field types. The major field types include coding type, text type, numerical type, date type, and time series type.
[0078] S42. Use the minor field type recognition model to perform feature extraction and type annotation on the major field types, specifically:
[0079] Calculate the mean, variance, and maximum and minimum values as the features of the numerical major field.
[0080] Calculate the data difference between adjacent fields as the feature of the time series major field.
[0081] Perform word segmentation on the text major field, analyze the part-of-speech of each word segment, and use a keyword extraction algorithm to extract the key information in the text major field as the feature of the text major field.
[0082] Extract the date information in the date major field, including year, month, day, week, and quarter, and calculate the difference between adjacent dates. Use the date information and the difference as the features of the date major field.
[0083] Analyze the hierarchical structure and coding rules of coded large fields, and extract the upper-level codes and large fields at the deep level as the features of coded large fields;
[0084] Input the extracted various features into the field sub-type recognition model to obtain the field sub-types;
[0085] S5. Create a normal data quality detection library according to the professional knowledge of business personnel and the information quality rule verification rules of the historical metadata table. The normal data quality detection library includes typical feature templates of various field sub-types;
[0086] S6. And construct a detection rule matching model using the expectation-maximization clustering algorithm based on the Gaussian mixture model;
[0087] Initialize the Gaussian mixture model parameters, including the mean, covariance matrix, and weight of each Gaussian distribution, and preset the number of clusters;
[0088] Use the expectation-maximization algorithm to obtain the optimal Gaussian mixture model parameters. The expectation-maximization clustering algorithm includes an expectation step and a maximization step;
[0089] The expectation step calculates the posterior probability of the Gaussian distribution, which is expressed by the formula:
[0090]
[0091] In the formula, γ ik represents the posterior distribution of the i-th small field belonging to the k-th Gaussian distribution; x i represents the i-th field sub-type; N(x i |μ k ,λ k ) represents the probability density function of the i-th field sub-type under the k-th Gaussian distribution; ω k represents the weight of the k-th Gaussian distribution; μ k represents the mean of the k-th Gaussian distribution; λ k represents the covariance matrix of the k-th Gaussian distribution; k represents the index value of the k-th Gaussian distribution; K represents the number of clusters; i represents the index value of the i-th field sub-type;
[0092] The maximization step updates the Gaussian mixture model parameters according to the posterior probability, which is expressed by the formula:
[0093]
[0094] In the formula, M represents the number of field sub-types; T represents the transpose operation;
[0095] Iteratively execute the expectation step and the maximization step until the maximum number of iterations is reached, obtain the optimal Gaussian mixture model parameters, and obtain the detection rule matching model;
[0096] S7. Match the small field types based on the normal data quality detection library and the detection rule matching model to obtain detection matching results;
[0097] The detection types in the detection rule matching model include outlier detection, time series difference detection, enumeration detection, name detection, size comparison detection, and missing value detection;
[0098] Compare the input small field types with the corresponding feature templates and detection types to determine whether the input small field types all conform to the typical features and detection rules, and obtain detection matching results;
[0099] If they conform, determine that the current small field type is a valid match and mark it as a normal field; if any one does not conform, determine that the current small field type is an invalid match and mark it as an abnormal field;
[0100] For the abnormal fields, start the abnormal handling process, and further confirm the authenticity and accuracy of the data by means of manual review, or correct the abnormal fields according to business rules;
[0101] S8. Construct a closed-loop quality problem verification model to perform quality detection and verification on the detection matching results, and obtain the quality detection results of the data table to be detected;
[0102] The closed-loop quality problem verification model includes a data sampling module, a quality verification module, a result evaluation module, and a feedback correction module;
[0103] The data sampling module presets a sampling method and selects verification samples from the detection matching results; the quality verification module uses the built-in business knowledge, industry standards, and detection experience to verify the verification samples; the result evaluation module calculates the evaluation indicators of the verification samples to obtain an evaluation result; the feedback correction model feeds back the evaluation result to the discrimination system, and the discrimination system analyzes the problems existing in the detection matching results and formulates corresponding correction opinions.
[0104] Embodiment 2
[0105] This embodiment provides an automatic table data quality discrimination system based on multi-dimensional indicators. The system includes a data acquisition module, a field identification module, a quality detection module, and a result output module, where:
[0106] The data acquisition module is used to obtain the metadata table information of the business system, including table descriptions, field descriptions, and field types; construct a data table identification and screening model using the random forest classification algorithm, and the data table identification and screening model performs data cleaning on the metadata table information to obtain the data table to be detected; transmit the data table to be detected to the field identification module;
[0107] The field identification module is used to construct a field type identification model including a major field type identification model and a minor field type identification model; use the major field type identification model to identify and classify the fields of the data table to be detected to obtain the major field types; use the minor field type identification model to perform feature extraction and type annotation on the major field types to obtain the minor field types;
[0108] The quality detection module is used to create a normal data quality detection library according to the professional knowledge of business personnel and the quality rule verification rules of historical metadata table information, and construct a detection rule matching model using the expectation maximization clustering algorithm based on the Gaussian mixture model; perform matching on the minor field types based on the normal data quality detection library and the detection rule matching model to obtain a detection matching result;
[0109] Construct a closed-loop quality problem verification model using the statistical sampling algorithm to perform quality detection and verification on the detection matching result, and obtain the quality detection result of the data table to be detected;
[0110] The result output module is used to display the quality detection result of the data table to be detected.
[0111] Embodiment 3
[0112] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a method for automatically discriminating the quality of table data based on multi-dimensional indicators as described in any embodiment of the present invention.
[0113] Embodiment 4
[0114] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements a method for automatically discriminating the quality of table data based on multi-dimensional indicators as described in any embodiment of the present invention.
[0115] It should be noted that the systems, electronic devices, and computer-readable storage media described in the present invention are all based on the same principle as the method described in Embodiment 1, and will not be elaborated here.
[0116] The above are only embodiments of the present invention, and thus do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of the present invention.
Claims
1. A method for automatically determining table data quality based on multidimensional indicators, characterized in that: The method comprises: Obtain metadata table information of the business system, including table description, field description and field type, and construct a metadata table information data set; use the random forest classification algorithm to construct a data table identification and screening model, and the data table identification and screening model performs data cleaning on the metadata table information to obtain a data table to be detected; Constructing a field type recognition model includes a field major category recognition model and a field minor category recognition model; using the field major category recognition model to recognize and classify the fields of the data table to be detected to obtain the field major type; using the field minor category recognition model to extract features and type label the major field type to obtain the field minor type; A normal data quality detection library is created based on the professional knowledge of business personnel and the information quality rule verification rules of the historical metadata table, and a detection rule matching model is constructed using an expectation maximization clustering algorithm based on a Gaussian mixture model; the field small types are matched based on the normal data quality detection library and the detection rule matching model to obtain a detection matching result; A closed-loop quality problem verification model is constructed to perform quality inspection and verification on the detection and matching results, and obtain the quality inspection results of the data table to be inspected.
2. The method for automatically determining table data quality based on multidimensional indicators according to claim 1, characterized in that: The random forest classification algorithm is used to build a data table identification and screening model, specifically: The metadata table information data set is subjected to data preprocessing, including type conversion, data normalization and feature extraction; the type conversion specifically includes converting the text type fields in the metadata table information data set into binary types; the data normalization specifically includes normalizing the data value type fields in the metadata table information data set; the feature extraction specifically includes calculating the correlation coefficient between each field and irrelevant data, selecting the fields whose correlation coefficient is within a preset interval, and generating a feature set; the irrelevant data includes an empty table, a low-heat table, an intermediate table, a log table, a low-heat field and an empty data field; Initialize the random forest model and preset initial hyperparameters, including the number of decision trees, the maximum depth of the decision tree, the number of features randomly selected for each node, and the minimum number of samples; A self-service sampling method is used to randomly extract samples from the metadata table information data set with replacement to construct a decision tree with a preset number of trees; at each node of the decision tree, a feature with a preset number of features is randomly selected from the feature set, and the information gain of the randomly selected feature for the current node is calculated; the feature with the largest information gain is selected as the optimal feature for node division; Recursively construct decision trees until the maximum depth of the decision tree reaches a preset value or the number of samples of the node is less than the preset minimum number of samples, and then all decision trees are combined into a random forest to obtain a data table recognition and screening model; The metadata table information data set is input into the data table identification and screening model to obtain a data table prediction result; irrelevant data in the metadata table information data set is eliminated according to the data table prediction result to obtain a data table to be detected.
3. The method for automatically determining table data quality based on multidimensional indicators according to claim 1, characterized in that: The field category recognition model is built based on automatic machine learning technology and random forest classification algorithm, specifically: Select and configure the automatic machine learning platform, preset the search algorithm, evaluation index and maximum running time, and define the random forest hyperparameter space; The data table to be detected is used as input to start the search process, and the automatic machine learning platform searches for the optimal solution of the random forest hyperparameters in the defined random forest hyperparameter space according to the preset search algorithm, and calculates the evaluation index of the current optimal solution; The search process is iterated until the preset maximum running time is reached. The automatic machine learning platform selects the optimal solution corresponding to the maximum value of the evaluation index as the hyperparameter combination of the random forest model to obtain the field category recognition model. The data table to be tested is input into the field category recognition model to obtain the prediction result of the category to which each field of the data table to be tested belongs, and the field category is obtained; the field category includes coding category, text category, numerical category, date category and time series category.
4. The method for automatically determining table data quality based on multidimensional indicators according to claim 1, characterized in that: The field sub-class recognition model is used to extract features and label the large field type, specifically: Calculate the mean, variance and maximum value as the characteristics of the large numerical field; Calculating data differences between adjacent fields as features of the time series large field; Perform word segmentation on large text fields, analyze the part of speech of each word, and use keyword extraction algorithms to extract key information from large text fields as features of large text fields. Extracting date information from a large date field, including year, month, day, week and quarter, and calculating the difference between adjacent dates, and using the date information and the difference as features of the large date field; Analyze the hierarchical structure and coding rules of large coding fields, and extract large fields of upper coding and deep level as the characteristics of large coding fields; The various extracted features are input into the field sub-category recognition model to obtain the field sub-type.
5. The method for automatically determining table data quality based on multidimensional indicators according to claim 1, characterized in that: The detection rule matching model is constructed using the expectation maximization algorithm based on the Gaussian mixture model, specifically: Initialize the Gaussian mixture model parameters, including the mean, covariance matrix and weight of each Gaussian distribution, and preset the number of clusters; Utilizing the expectation maximization algorithm to obtain optimal Gaussian mixture model parameters, the expectation maximization clustering algorithm includes an expectation step and a maximization step; The expectation step calculates the posterior probability of the Gaussian distribution, which is expressed as: In the formula, γ ik Indicates that the posterior distribution of the i-th small field belongs to the k-th Gaussian distribution; x i Indicates the small type of the i-th field; N(x i |μ k ,λ k ) represents the probability density function of the i-th field type under the k-th Gaussian distribution; ω k represents the weight of the kth Gaussian distribution; μ k represents the mean of the kth Gaussian distribution; λ k represents the covariance matrix of the kth Gaussian distribution; k represents the index value of the kth Gaussian distribution; K represents the number of clusters; i represents the index value of the i-th field small type; The maximization step updates the Gaussian mixture model parameters according to the posterior probability, which is expressed as: Where M represents the number of small field types; T represents the transposition operation; The expected step and the maximization step are iterated until the maximum number of iterations is reached, the optimal Gaussian mixture model parameters are obtained, and the detection rule matching model is obtained.
6. The method for automatically determining table data quality based on multidimensional indicators according to claim 1, characterized in that: The field subtype is matched based on the normal data quality detection library and the detection rule matching model, specifically: The normal data quality detection library includes typical feature templates of various field types; the detection types in the detection rule matching model include outlier detection, temporal difference detection, enumeration detection, name detection, size comparison detection and missing value detection; Compare the input field subtype with the corresponding feature template and the detection type to determine whether the input field subtype meets the typical features and detection rules, and obtain a detection matching result; If they match, the current field type is determined to be a valid match and marked as a normal field; if they do not match either of the two, the current field type is determined to be an invalid match and marked as an abnormal field; For the abnormal fields, the abnormal handling process is initiated, and the authenticity and accuracy of the data are further confirmed by manual review, or the abnormal fields are corrected according to business rules.
7. The method for automatically determining table data quality based on multidimensional indicators according to claim 1, characterized in that: A closed-loop quality problem verification model is constructed to perform quality inspection and verification on the detection and matching results, specifically: The closed-loop quality problem verification model includes a data sampling module, a quality verification module, a result evaluation module and a feedback correction module; The data sampling module presets a sampling method and selects verification samples from the detection and matching results; the quality verification module verifies the verification samples using built-in business knowledge, industry standards and detection experience; the result evaluation module calculates the evaluation index of the verification samples and obtains the evaluation results; The feedback correction model feeds back the evaluation result to the discrimination system, and the discrimination system analyzes and detects problems in the matching result and formulates corresponding correction suggestions.
8. A table data quality automatic judgment system based on multi-dimensional indicators, characterized in that: The system includes a data acquisition module, a field identification module, a quality detection module and a result output module, wherein: The data acquisition module is used to obtain metadata table information of the business system, including table description, field description and field type; use the random forest classification algorithm to build a data table identification and screening model, the data table identification and screening model performs data cleaning on the metadata table information to obtain a data table to be detected; the data table to be detected is transmitted to the field identification module; The field identification module is used to construct a field type identification model including a field major category identification model and a field minor category identification model; the field major category identification model is used to identify and classify the fields of the data table to be detected to obtain the major field type; the field minor category identification model is used to extract features and mark the types of the major field types to obtain the minor field types; The quality detection module is used to create a normal data quality detection library according to the professional knowledge of the business personnel and the historical metadata table information quality rule verification rules, and use the expectation maximization clustering algorithm based on the Gaussian mixture model to build a detection rule matching model; based on the normal data quality detection library and the detection rule matching model, the field small type is matched to obtain a detection matching result; Constructing a closed-loop quality problem verification model to perform quality inspection and verification on the detection and matching results, and obtaining quality inspection results of the data table to be inspected; The result output module is used to display the quality inspection result of the data table to be inspected.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for automatically determining the quality of table data based on multidimensional indicators as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements a method for automatically determining the quality of table data based on multidimensional indicators as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Data quality detection method and system
CN117762914A