Construction method of data quality evaluation model
By constructing a quality assessment model based on evaluation index vectors and feature matrices, the problem that existing technologies cannot fully reflect the data quality characteristics under complex business scenarios is solved, and dynamic weighted high-precision assessment is achieved, improving the intelligence and stability of the model.
Patent Information
- Application Number
- CN202511901264.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-05-15
AI Technical Summary
Existing data quality assessment methods cannot fully reflect the multidimensional quality characteristics under complex business scenarios, and lack dynamic optimization mechanisms, resulting in assessment results that lack specificity and accuracy.
By determining the evaluation index vector of the evaluation dataset to be evaluated, the training quality feature matrix and the test quality feature matrix are determined, a quality evaluation model is constructed and optimized to achieve high-precision evaluation with dynamic weighting.
It enhances the model's ability to capture data quality characteristics, improves the intelligence and stability of the quality assessment model, and increases its adaptability and accuracy in complex data scenarios.
Smart Images

Figure CN122045173A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of quality assessment technology, and in particular to a method for constructing a data quality assessment model. Background Technology
[0002] Data quality assessment models are primarily used to evaluate the quality level of datasets under specific business scenarios and application requirements, and are commonly applied in data mining, machine learning, and big data analytics. Currently, most data quality assessment methods rely on empirical rules or single statistical methods, such as manually defined quality features based on the proportion of missing values or consistency checks. However, these methods typically fail to comprehensively reflect the multidimensional quality characteristics of data in complex business scenarios. Furthermore, existing models often employ fixed weights for assessment metrics, making it difficult to adapt to the varying importance of different data features to the assessment results, leading to a lack of specificity and accuracy. For complex, multi-source datasets, traditional static assessment models struggle to capture diverse data quality issues and lack dynamic optimization mechanisms for model performance.
[0003] Therefore, the present invention provides a method for constructing a data quality assessment model. Summary of the Invention
[0004] This invention provides a method for constructing a data quality assessment model. It involves determining the assessment index vector of the dataset to be assessed, identifying the training quality feature matrix and the test quality feature matrix, constructing a quality assessment model based on the training set and the training quality feature matrix, and evaluating and optimizing the quality assessment performance based on the assessment index vector, the test set, and the test quality feature matrix. This method enhances the model's ability to capture data quality characteristics, achieves high-precision assessment with dynamic weighting, improves the intelligence and stability of the quality assessment model, and enhances its adaptability in complex data quality management and analysis scenarios.
[0005] This invention provides a method for constructing a data quality assessment model, comprising:
[0006] 101: Collect the dataset to be evaluated and obtain the business requirements and application scenarios of the dataset;
[0007] 102: Based on the business requirements and application scenarios of the dataset to be evaluated, determine the evaluation index vector of the dataset to be evaluated;
[0008] 103: Based on the dataset to be evaluated and the evaluation metric vector, determine the training quality feature matrix and the test quality feature matrix;
[0009] 104: Construct a quality assessment model based on the training set and training quality feature matrix, evaluate the quality assessment performance based on the assessment index vector, test set and test quality feature matrix, and optimize it.
[0010] According to the method for constructing a data quality assessment model provided by the present invention, the dataset to be assessed includes multiple sets of sub-datasets to be assessed.
[0011] According to a method for constructing a data quality assessment model provided by the present invention, based on the business requirements and application scenarios of the dataset to be assessed, an assessment index vector for the dataset to be assessed is determined, including:
[0012] Analyze the business requirements of the dataset to be evaluated and determine multiple evaluation characteristics of the dataset based on business objectives;
[0013] Based on the importance of each evaluation feature to the business objective, multiple evaluation features based on business objectives are ranked to determine the first sub-evaluation indicator vector based on business needs.
[0014] Analyze the application scenarios of the dataset to be evaluated and determine multiple evaluation features of the dataset based on the business scenarios.
[0015] Based on the importance of each evaluation feature to the business scenario, multiple evaluation features based on the business scenario are ranked to determine the second sub-evaluation index vector based on the business scenario.
[0016] The evaluation feature vector is determined based on the first sub-evaluation index vector and the second sub-evaluation index vector.
[0017] Determine whether the first sub-evaluation index vector and the second sub-evaluation index vector have the same evaluation features. If so, record the number of the same evaluation features, mark each of the same evaluation features in the same way, and adjust the position of each of the same evaluation features in the evaluation feature vector.
[0018] The dataset to be evaluated is obtained based on industry standards and evaluation indicator templates in the field of data quality. A general evaluation vector for the dataset to be evaluated is determined based on the evaluation indicator templates.
[0019] The adjusted evaluation feature vector and the general evaluation vector are compared item by item to determine the evaluation index vector of the dataset to be evaluated, and the evaluation features that are consistent with the comparison results are uniformly labeled.
[0020] According to a method for constructing a data quality assessment model provided by the present invention, based on the dataset to be assessed and the assessment index vector, the method determines the quality feature matrix of the dataset to be assessed, including:
[0021] Feature extraction is performed on each sub-dataset to be evaluated in the dataset to be evaluated, and the data features of each sub-dataset to be evaluated are determined based on each evaluation feature in the evaluation index vector.
[0022] Based on all data features of each set of sub-datasets to be evaluated and the position of the evaluation feature corresponding to each data feature in the evaluation index vector, the quality feature vector of each set of sub-datasets to be evaluated is determined.
[0023] The dataset to be evaluated is divided into a training set and a test set. The training set includes multiple sets of sub-datasets to be evaluated, and the test set includes multiple sets of sub-datasets to be evaluated.
[0024] Based on the quality feature vectors of all the subsets to be evaluated in the training set, determine the training quality feature matrix;
[0025] The test quality feature matrix is determined based on the quality feature vectors of all the subsets to be evaluated in the test set.
[0026] According to a method for constructing a data quality assessment model provided by the present invention, the assessment model is constructed based on a training set and a training quality feature matrix, including:
[0027] Each set of datasets to be evaluated in the training set and the quality feature vectors of each set of datasets to be evaluated in the training quality feature matrix are input into the model for supervised learning training to build a quality evaluation model.
[0028] According to the present invention, a method for constructing a data quality assessment model is provided, which evaluates and optimizes the quality assessment performance based on an assessment index vector, a test set, and a test quality feature matrix, including:
[0029] Each sub-dataset to be evaluated in the test set is input into the quality assessment model to determine the predicted quality vector for each sub-dataset to be evaluated in the test set.
[0030] Based on the evaluation index vector, the test quality feature matrix, and the predicted quality vectors of all datasets to be evaluated in the test set, the prediction accuracy of the quality evaluation model is determined.
[0031] If the prediction accuracy is lower than the preset accuracy threshold, the quality assessment model is optimized based on the prediction accuracy, the prediction quality vectors of all sub-datasets to be evaluated in the test set, and the test quality feature matrix.
[0032] According to a method for constructing a data quality assessment model provided by the present invention, the predicted accuracy value of the quality assessment model is determined based on the test quality feature matrix and the predicted quality vectors of all datasets to be evaluated in the test set, including:
[0033] Based on each evaluation feature in the evaluation index vector, determine the comprehensive weight of each data feature in the quality feature vector of the subset to be evaluated;
[0034]
[0035] ; ; ;
[0036] in, This represents the comprehensive weight of the j-th data feature. This indicates that the j-th evaluation feature with the same label has the same weight. This indicates that the a-th evaluation feature with the same label has the same weight. Let b represent the uniform weight of the b-th evaluation feature with uniform labeling. This represents the uniform weight of the j-th evaluation feature with uniform labeling. Let represent the feature weight of the c-th evaluation feature that does not have the same label or a uniform label. Let represent the feature weight of the j-th evaluation feature that has no identical labels and no uniform labels. Indicates the first indicator function, This indicates the second indicator function. This indicates the third indicator function. The label representing the j-th evaluation feature. Indicates the same marker, Indicates a uniform markup;
[0037] The prediction accuracy of the quality assessment model is calculated based on the test quality feature matrix, the predicted quality vectors of all datasets to be evaluated, and the combined weights of all data features in each quality feature vector of the test quality feature matrix.
[0038] According to a method for constructing a data quality assessment model provided by the present invention, calculating the prediction accuracy of the quality assessment model includes:
[0039] ;
[0040] in, N1 represents the predicted accuracy of the quality assessment model, N2 represents the number of subsets to be evaluated in the test set, and N2 represents the number of evaluation features in the evaluation metric vector. This represents the j-th data feature of the quality feature vector of the i-th subset of the test set to be evaluated. This represents the j-th predicted feature of the predicted quality vector of the i-th subset of the test set to be evaluated. This represents the weight adjustment factor for evaluation characteristics that do not have the same label or a uniform label. This indicates the impact factors based on industry standards and data quality. This indicates the number of evaluation features with the same label. This indicates the number of evaluation features with uniform labels.
[0041] Compared with the prior art, the beneficial effects of this application are as follows:
[0042] By determining the evaluation metric vectors for the dataset to be evaluated, and identifying the training and testing quality feature matrices, a quality evaluation model is constructed based on the training set and training quality feature matrices. The performance of the quality evaluation is then assessed and optimized based on the evaluation metric vectors, the test set, and the testing quality feature matrix. This enhances the model's ability to capture data quality characteristics, achieves high-precision evaluation with dynamic weighting, improves the intelligence and stability of the quality evaluation model, and enhances its adaptability in complex data quality management and analysis scenarios. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0044] Figure 1 This is a flowchart illustrating a method for constructing a data quality assessment model according to an embodiment of the present invention. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0046] Example 1:
[0047] This invention provides a method for constructing a data quality assessment model, such as... Figure 1 As shown, it includes:
[0048] 101: Collect the dataset to be evaluated and obtain the business requirements and application scenarios of the dataset;
[0049] 102: Based on the business requirements and application scenarios of the dataset to be evaluated, determine the evaluation index vector of the dataset to be evaluated;
[0050] 103: Based on the dataset to be evaluated and the evaluation metric vector, determine the training quality feature matrix and the test quality feature matrix;
[0051] 104: Construct a quality assessment model based on the training set and training quality feature matrix, evaluate the quality assessment performance based on the assessment index vector, test set and test quality feature matrix, and optimize it.
[0052] In this embodiment, a dataset to be evaluated is collected, and the business requirements and application scenarios of the dataset are defined.
[0053] In this embodiment, the dataset to be evaluated represents the set of data whose quality needs to be evaluated, which may include structured data (such as tables), unstructured data (such as text or images), etc.
[0054] In this embodiment, business requirements represent specific goals or expectations closely related to the application of the dataset, such as the requirements of tasks like data analysis and predictive models.
[0055] In this embodiment, the application scenario refers to the specific environment in which the data is used, such as data management in financial risk control, medical diagnosis, or industrial manufacturing.
[0056] In this embodiment, the evaluation index vector represents a set of standard indicators used to evaluate data quality, such as completeness (whether there are missing values), accuracy (whether the data is authentic), and consistency (whether the data is conflicting).
[0057] In this embodiment, the quality feature matrix for training and testing is constructed using the dataset to be evaluated and the evaluation index vector.
[0058] In this embodiment, a quality assessment model is constructed using a training set and a training quality feature matrix, and the model performance is optimized by combining a test set.
[0059] The beneficial effects of the above technical solution are as follows: By determining the evaluation index vector of the dataset to be evaluated, the training quality feature matrix and the test quality feature matrix are determined. A quality evaluation model is constructed based on the training set and the training quality feature matrix. The quality evaluation performance is evaluated and optimized based on the evaluation index vector, the test set, and the test quality feature matrix. This enhances the model's ability to capture data quality characteristics, achieves high-precision evaluation with dynamic weighting, improves the intelligence and stability of the quality evaluation model, and enhances the model's adaptability in complex data quality management and analysis scenarios.
[0060] Example 2:
[0061] This invention provides a method for constructing a data quality assessment model, wherein the dataset to be assessed includes multiple sub-datasets to be assessed.
[0062] In this embodiment, the subset to be evaluated is a subset of the dataset to be evaluated, and each subset can correspond to a different data source.
[0063] The beneficial effects of the above technical solution are: determining the dataset to be evaluated can improve data quality and provide a data basis for determining the training quality feature matrix and the testing quality feature matrix.
[0064] Example 3:
[0065] This invention provides a method for constructing a data quality assessment model, which determines an assessment index vector for the dataset to be assessed based on the business requirements and application scenarios of the dataset to be assessed, including:
[0066] Analyze the business requirements of the dataset to be evaluated and determine multiple evaluation characteristics of the dataset based on business objectives;
[0067] Based on the importance of each evaluation feature to the business objective, multiple evaluation features based on business objectives are ranked to determine the first sub-evaluation indicator vector based on business needs.
[0068] Analyze the application scenarios of the dataset to be evaluated and determine multiple evaluation features of the dataset based on the business scenarios.
[0069] Based on the importance of each evaluation feature to the business scenario, multiple evaluation features based on the business scenario are ranked to determine the second sub-evaluation index vector based on the business scenario.
[0070] The evaluation feature vector is determined based on the first sub-evaluation index vector and the second sub-evaluation index vector.
[0071] Determine whether the first sub-evaluation index vector and the second sub-evaluation index vector have the same evaluation features. If so, record the number of the same evaluation features, mark each of the same evaluation features in the same way, and adjust the position of each of the same evaluation features in the evaluation feature vector.
[0072] The dataset to be evaluated is obtained based on industry standards and evaluation indicator templates in the field of data quality. A general evaluation vector for the dataset to be evaluated is determined based on the evaluation indicator templates.
[0073] The adjusted evaluation feature vector and the general evaluation vector are compared item by item to determine the evaluation index vector of the dataset to be evaluated, and the evaluation features that are consistent with the comparison results are uniformly labeled.
[0074] In this embodiment, the adjusted evaluation index vector and the general evaluation vector are compared item by item. For example, the "timeliness" in the adjusted evaluation index vector is directly mapped to the "data delay" in the general evaluation vector.
[0075] In this embodiment, the business requirements of the dataset to be evaluated are analyzed, multiple evaluation features based on business objectives are extracted, and the importance of business objectives is ranked according to these features to generate the first sub-evaluation index vector.
[0076] In this embodiment, the first sub-evaluation index vector represents an index vector composed of evaluation features selected and ranked according to business objectives.
[0077] In this embodiment, the application scenarios of the dataset to be evaluated are analyzed, the quality evaluation features of the data in specific scenarios are extracted, their importance is ranked, and a second sub-evaluation index vector is generated.
[0078] In this embodiment, the second sub-evaluation index vector represents the index vector after the evaluation features associated with a specific scenario are sorted by importance.
[0079] In this embodiment, the first sub-evaluation index vector and the second sub-evaluation index vector are combined to generate a comprehensive evaluation feature vector. If the two have the same evaluation features, the quantity is recorded, they are uniformly marked, and the positions of these features in the vector are adjusted.
[0080] In this embodiment, a general evaluation index template based on industry standards and data quality is obtained, a general evaluation vector of the dataset to be evaluated is extracted, the adjusted evaluation feature vector is compared with the general evaluation vector item by item to form the final evaluation index vector, and the evaluation features that match the comparison results are uniformly marked.
[0081] The beneficial effects of the above technical solution are as follows: Based on the business requirements and application scenarios of the dataset to be evaluated, the evaluation index vector of the dataset to be evaluated can be determined, which can achieve the uniformity and compatibility of the evaluation index vector, and ensure that the evaluation model can not only adapt to the customized needs of specific scenarios, but also have industry universality.
[0082] Example 4:
[0083] This invention provides a method for constructing a data quality assessment model, which determines the quality feature matrix of the dataset to be assessed based on the dataset to be assessed and the assessment index vector, including:
[0084] Feature extraction is performed on each sub-dataset to be evaluated in the dataset to be evaluated, and the data features of each sub-dataset to be evaluated are determined based on each evaluation feature in the evaluation index vector.
[0085] Based on all data features of each set of sub-datasets to be evaluated and the position of the evaluation feature corresponding to each data feature in the evaluation index vector, the quality feature vector of each set of sub-datasets to be evaluated is determined.
[0086] The dataset to be evaluated is divided into a training set and a test set. The training set includes multiple sets of sub-datasets to be evaluated, and the test set includes multiple sets of sub-datasets to be evaluated.
[0087] Based on the quality feature vectors of all the subsets to be evaluated in the training set, determine the training quality feature matrix;
[0088] The test quality feature matrix is determined based on the quality feature vectors of all the subsets to be evaluated in the test set.
[0089] In this embodiment, the quality feature vector is in the form of (accuracy value, timeliness value, consistency value).
[0090] In this embodiment, each sub-dataset in the dataset to be evaluated is analyzed, and the corresponding data features are extracted based on each evaluation feature defined in the evaluation index vector.
[0091] In this embodiment, data features represent characteristics extracted from a subset of data sets that are associated with the evaluation metrics.
[0092] In this embodiment, all data features are organized in the order of the evaluation index vector to form the quality feature vector of each sub-dataset.
[0093] In this embodiment, the quality feature vector quantifies the quality attributes of the corresponding subset of data in numerical form.
[0094] In this embodiment, the dataset to be evaluated is divided into a training set and a test set, and each set contains multiple sub-datasets to be evaluated.
[0095] The beneficial effects of the above technical solution are as follows: Based on the dataset to be evaluated and the evaluation index vector, the quality feature matrix of the dataset to be evaluated can be determined, which can improve the automation level and accuracy of the evaluation, and is particularly suitable for quality management and analysis scenarios of complex data.
[0096] Example 5:
[0097] This invention provides a method for constructing a data quality assessment model, which builds the assessment model based on a training set and a training quality feature matrix, including:
[0098] Each set of datasets to be evaluated in the training set and the quality feature vectors of each set of datasets to be evaluated in the training quality feature matrix are input into the model for supervised learning training to build a quality evaluation model.
[0099] In this embodiment, each set of sub-datasets to be evaluated and its corresponding quality feature vectors are input into the model for supervised learning training.
[0100] The beneficial effects of the above technical solution are: constructing an evaluation model based on the training set and training quality feature matrix can reduce the limitations of excessive reliance on human rules and experience, and improve the evaluation accuracy, adaptability and efficiency.
[0101] Example 6:
[0102] This invention provides a method for constructing a data quality assessment model, which evaluates and optimizes quality assessment performance based on an assessment index vector, a test set, and a test quality feature matrix, including:
[0103] Each sub-dataset to be evaluated in the test set is input into the quality assessment model to determine the predicted quality vector for each sub-dataset to be evaluated in the test set.
[0104] Based on the evaluation index vector, the test quality feature matrix, and the predicted quality vectors of all datasets to be evaluated in the test set, the prediction accuracy of the quality evaluation model is determined.
[0105] If the prediction accuracy is lower than the preset accuracy threshold, the quality assessment model is optimized based on the prediction accuracy, the prediction quality vectors of all sub-datasets to be evaluated in the test set, and the test quality feature matrix.
[0106] In this embodiment, each subset of data to be evaluated in the test set is input into a trained quality assessment model, and the model generates corresponding prediction results based on the input quality feature vector.
[0107] In this embodiment, the subset of data to be evaluated in the test set represents the subset of data that did not participate in model training and is used to evaluate model performance.
[0108] In this embodiment, the predicted quality vector represents the output generated by the quality assessment model based on the quality feature vector of the input subset, and represents the predicted quality level.
[0109] In this embodiment, the model's prediction accuracy is calculated and the model performance is evaluated based on the evaluation index vector, the test quality feature matrix (true quality characteristics), and the predicted quality vector output by the model.
[0110] In this embodiment, if the prediction accuracy is lower than the preset accuracy threshold, it indicates that the model performance is substandard and needs to be optimized.
[0111] In this embodiment, the accurate threshold is used as the minimum acceptable standard for evaluating model performance and serves as the basis for determining optimization triggers.
[0112] In this embodiment, based on existing prediction results and test data, the model parameters or algorithm structure are adjusted to enhance the model's ability to learn data characteristics and improve prediction accuracy.
[0113] The beneficial effects of the above technical solution are as follows: by evaluating and optimizing the quality assessment performance based on the evaluation index vector, test set, and test quality feature matrix, the intelligence and stability of the quality assessment model can be improved, and the adaptability to the dynamic quality assessment needs of complex and multi-source data can be enhanced.
[0114] Example 7:
[0115] This invention provides a method for constructing a data quality assessment model, which determines the prediction accuracy of the quality assessment model based on a test quality feature matrix and the predicted quality vectors of all datasets to be evaluated in the test set, including:
[0116] Based on each evaluation feature in the evaluation index vector, determine the comprehensive weight of each data feature in the quality feature vector of the subset to be evaluated;
[0117]
[0118] ; ; ;
[0119] in, This represents the comprehensive weight of the j-th data feature. This indicates that the j-th evaluation feature with the same label has the same weight. This indicates that the a-th evaluation feature with the same label has the same weight. Let b represent the uniform weight of the b-th evaluation feature with uniform labeling. This represents the uniform weight of the j-th evaluation feature with uniform labeling. Let represent the feature weight of the c-th evaluation feature that does not have the same label or a uniform label. Let represent the feature weight of the j-th evaluation feature that has no identical labels and no uniform labels. Indicates the first indicator function, This indicates the second indicator function. This indicates the third indicator function. The label representing the j-th evaluation feature. Indicates the same marker, Indicates a uniform markup;
[0120] The prediction accuracy of the quality assessment model is calculated based on the test quality feature matrix, the predicted quality vectors of all datasets to be evaluated, and the combined weights of all data features in each quality feature vector of the test quality feature matrix.
[0121] In this embodiment, the first indicator function This represents an indicator function that determines whether the j-th evaluation feature has the same label.
[0122] In this embodiment, the second indicator function This represents an indicator function for determining whether the j-th evaluation feature has a uniform label.
[0123] In this embodiment, the second indicator function This represents an indicator function that determines whether the j-th evaluation feature does not simultaneously have the same label and a uniform label.
[0124] In this embodiment, for each evaluation feature in the evaluation index vector, its comprehensive weight in the different data features of the subset to be evaluated is determined.
[0125] In this embodiment, the comprehensive weight represents the importance weight of the evaluation feature to different data features, reflecting the degree of contribution of each data feature to the overall quality assessment.
[0126] In this embodiment, the model's prediction accuracy is calculated using the test quality feature matrix (true value), the predicted quality vectors of all datasets to be evaluated (model output), and the combined weights.
[0127] In this embodiment, the test quality feature matrix contains the true quality feature vectors in the test set, which is the standard answer for evaluating the model's prediction accuracy.
[0128] In this embodiment, the predicted quality vector represents the predicted output generated by the quality assessment model based on the input, which is used to compare with the true value.
[0129] In this embodiment, the prediction accuracy is determined by the deviation between the true value and the predicted value, as well as the comprehensive weights, to obtain the overall accuracy index of the model on different data features.
[0130] The beneficial effects of the above technical solution are as follows: Based on the test quality feature matrix and the predicted quality vectors of all datasets to be evaluated in the test set, the prediction accuracy of the quality assessment model can be determined, which can improve the sensitivity to complex data characteristics, realize dynamic weighted high-precision evaluation, and refine the evaluation of model performance.
[0131] Example 8:
[0132] This invention provides a method for constructing a data quality assessment model, and calculating the prediction accuracy of the quality assessment model, including:
[0133] ;
[0134] in, N1 represents the predicted accuracy of the quality assessment model, N2 represents the number of subsets to be evaluated in the test set, and N2 represents the number of evaluation features in the evaluation metric vector. This represents the j-th data feature of the quality feature vector of the i-th subset of the test set to be evaluated. This represents the j-th predicted feature of the predicted quality vector of the i-th subset of the test set to be evaluated. This represents the weight adjustment factor for evaluation characteristics that do not have the same label or a uniform label. This indicates the impact factors based on industry standards and data quality. This indicates the number of evaluation features with the same label. This indicates the number of evaluation features with uniform labels.
[0135] In this embodiment, This indicates the number of evaluation features that do not have the same label or a uniform label.
[0136] The beneficial effects of the above technical solution are: calculating the predicted accuracy of the quality assessment model can achieve high-precision assessment with dynamic weighting, thereby refining the evaluation of model performance.
[0137] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0138] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a data quality assessment model, characterized in that, include: 101: Collect the dataset to be evaluated and obtain the business requirements and application scenarios of the dataset; 102: Based on the business requirements and application scenarios of the dataset to be evaluated, determine the evaluation index vector of the dataset to be evaluated; 103: Based on the dataset to be evaluated and the evaluation metric vector, determine the training quality feature matrix and the test quality feature matrix; 104: Construct a quality assessment model based on the training set and training quality feature matrix, evaluate the quality assessment performance based on the assessment index vector, test set and test quality feature matrix, and optimize it.
2. The method for constructing a data quality assessment model according to claim 1, characterized in that, The dataset to be evaluated includes multiple sub-datasets to be evaluated.
3. The method for constructing a data quality assessment model according to claim 1, characterized in that, Based on the business requirements and application scenarios of the dataset to be evaluated, the evaluation metric vector for the dataset to be evaluated is determined, including: Analyze the business requirements of the dataset to be evaluated and determine multiple evaluation characteristics of the dataset based on business objectives; Based on the importance of each evaluation feature to the business objective, multiple evaluation features based on business objectives are ranked to determine the first sub-evaluation indicator vector based on business needs. Analyze the application scenarios of the dataset to be evaluated and determine multiple evaluation features of the dataset based on the business scenarios. Based on the importance of each evaluation feature to the business scenario, multiple evaluation features based on the business scenario are ranked to determine the second sub-evaluation index vector based on the business scenario. The evaluation feature vector is determined based on the first sub-evaluation index vector and the second sub-evaluation index vector. Determine whether the first sub-evaluation index vector and the second sub-evaluation index vector have the same evaluation features. If so, record the number of the same evaluation features, mark each of the same evaluation features in the same way, and adjust the position of each of the same evaluation features in the evaluation feature vector. The dataset to be evaluated is obtained based on industry standards and evaluation indicator templates in the field of data quality. A general evaluation vector for the dataset to be evaluated is determined based on the evaluation indicator templates. The adjusted evaluation feature vector and the general evaluation vector are compared item by item to determine the evaluation index vector of the dataset to be evaluated, and the evaluation features that are consistent with the comparison results are uniformly labeled.
4. The method for constructing a data quality assessment model according to claim 2, characterized in that, Based on the dataset to be evaluated and the evaluation metric vectors, the quality feature matrix of the dataset to be evaluated is determined, including: Feature extraction is performed on each sub-dataset to be evaluated in the dataset to be evaluated, and the data features of each sub-dataset to be evaluated are determined based on each evaluation feature in the evaluation index vector. Based on all data features of each set of sub-datasets to be evaluated and the position of the evaluation feature corresponding to each data feature in the evaluation index vector, the quality feature vector of each set of sub-datasets to be evaluated is determined. The dataset to be evaluated is divided into a training set and a test set. The training set includes multiple sets of sub-datasets to be evaluated, and the test set includes multiple sets of sub-datasets to be evaluated. Based on the quality feature vectors of all the subsets to be evaluated in the training set, determine the training quality feature matrix; The test quality feature matrix is determined based on the quality feature vectors of all the subsets to be evaluated in the test set.
5. The method for constructing a data quality assessment model according to claim 1, characterized in that, An evaluation model is constructed based on the training set and the training quality feature matrix, including: Each set of datasets to be evaluated in the training set and the quality feature vectors of each set of datasets to be evaluated in the training quality feature matrix are input into the model for supervised learning training to build a quality evaluation model.
6. The method for constructing a data quality assessment model according to claim 1, characterized in that, The quality assessment performance is evaluated and optimized based on the evaluation metric vector, test set, and test quality feature matrix, including: Each sub-dataset to be evaluated in the test set is input into the quality assessment model to determine the predicted quality vector for each sub-dataset to be evaluated in the test set. Based on the evaluation index vector, the test quality feature matrix, and the predicted quality vectors of all datasets to be evaluated in the test set, the prediction accuracy of the quality evaluation model is determined. If the prediction accuracy is lower than the preset accuracy threshold, the quality assessment model is optimized based on the prediction accuracy, the prediction quality vectors of all sub-datasets to be evaluated in the test set, and the test quality feature matrix.
7. The method for constructing a data quality assessment model according to claim 6, characterized in that, Based on the test quality feature matrix and the predicted quality vectors of all datasets to be evaluated in the test set, the predicted accuracy of the quality evaluation model is determined, including: Based on each evaluation feature in the evaluation index vector, determine the comprehensive weight of each data feature in the quality feature vector of the subset to be evaluated; ; ; ; ; in, This represents the comprehensive weight of the j-th data feature. This indicates that the j-th evaluation feature with the same label has the same weight. This indicates that the a-th evaluation feature with the same label has the same weight. Let b represent the uniform weight of the b-th evaluation feature with uniform labeling. This represents the uniform weight of the j-th evaluation feature with uniform labeling. Let represent the feature weight of the c-th evaluation feature that does not have the same label or a uniform label. Let represent the feature weight of the j-th evaluation feature that has no identical labels and no uniform labels. Indicates the first indicator function, This indicates the second indicator function. This indicates the third indicator function. The label representing the j-th evaluation feature. Indicates the same marker, Indicates a uniform markup; The prediction accuracy of the quality assessment model is calculated based on the test quality feature matrix, the predicted quality vectors of all datasets to be evaluated, and the combined weights of all data features in each quality feature vector of the test quality feature matrix.
8. The method for constructing a data quality assessment model according to claim 7, characterized in that, Calculate the predicted accuracy of the quality assessment model, including: ; in, N1 represents the predicted accuracy of the quality assessment model, N2 represents the number of subsets to be evaluated in the test set, and N2 represents the number of evaluation features in the evaluation metric vector. This represents the j-th data feature of the quality feature vector of the i-th subset of the test set to be evaluated. This represents the j-th predicted feature of the predicted quality vector of the i-th subset of the test set to be evaluated. This represents the weight adjustment factor for evaluation characteristics that do not have the same label or a uniform label. This indicates the impact factors based on industry standards and data quality. This indicates the number of evaluation features with the same label. This indicates the number of evaluation features with uniform labels.