Data set quality evaluation method and device, equipment and storage medium
By establishing a multi-dimensional data set quality evaluation system, including compliance, content, scale and value attribute evaluation, the problem of inaccurate evaluation of data set quality in the existing technology is solved, and the comprehensive and accurate data set quality evaluation is achieved.
Patent Information
- Application Number
- CN202510591796.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-08
AI Technical Summary
The existing technology cannot accurately evaluate the quality of data sets, and mainly focuses on basic quality indicators such as data integrity, accuracy and consistency, and cannot comprehensively measure the quality of data sets.
A multi-dimensional data set quality evaluation system is adopted, including a compliance attribute evaluation module, a content attribute evaluation module, a scale attribute evaluation module and a value attribute evaluation module. By obtaining the data set to be evaluated, the quality score is determined based on the preset weight, and the quality level is determined based on the preset score threshold.
A multi-dimensional comprehensive assessment of the quality of the data set is achieved, which improves the accuracy and effectiveness of the evaluation, and ensures the scientificity and practicality of the quality evaluation of the data set.
Smart Images

Figure CN120448737A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data evaluation technology, and in particular to a data set quality evaluation method, apparatus, device and storage medium. Background Art
[0002] In the era of booming big data and machine learning technologies, the quality of data sets has become a key factor in determining model performance and reliability.
[0003] However, current dataset evaluation methods still have significant limitations. Most of these traditional methods focus only on basic quality indicators such as data completeness, accuracy, and consistency, and are unable to accurately assess the quality of datasets. Summary of the Invention
[0004] The present invention provides a data set quality assessment method, apparatus, device and storage medium to solve the problem of being unable to accurately assess the quality of a data set.
[0005] According to one aspect of the present invention, a method for assessing data set quality is provided, the method comprising:
[0006] Obtaining a dataset to be evaluated, and evaluating the dataset based on a pre-established dataset quality evaluation system to obtain at least one evaluation indicator; wherein the dataset quality evaluation system includes a compliance attribute evaluation module, a content attribute evaluation module, a scale attribute evaluation module, and a value attribute evaluation module;
[0007] Determining a quality score corresponding to the data set to be evaluated based on at least one evaluation indicator and a preset weight corresponding to the evaluation indicator;
[0008] A quality level corresponding to the data set to be evaluated is determined based on the quality score and a preset score threshold.
[0009] According to another aspect of the present invention, a data set quality assessment device is provided, the device comprising:
[0010] An evaluation module, configured to obtain a dataset to be evaluated and evaluate the dataset based on a pre-established dataset quality evaluation system to obtain at least one evaluation indicator; wherein the dataset quality evaluation system includes a compliance attribute evaluation module, a content attribute evaluation module, a scale attribute evaluation module, and a value attribute evaluation module;
[0011] a score determination module, configured to determine a quality score corresponding to the dataset to be evaluated based on at least one evaluation indicator and a preset weight corresponding to the evaluation indicator;
[0012] The grade determination module is configured to determine a quality grade corresponding to the data set to be evaluated based on the quality score and a preset score threshold.
[0013] According to another aspect of the present invention, an electronic device is provided, comprising:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the dataset quality assessment method according to any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the dataset quality assessment method according to any embodiment of the present invention when executed.
[0018] The technical solution of an embodiment of the present invention obtains a dataset to be evaluated and evaluates it based on a pre-established dataset quality assessment system to obtain at least one evaluation indicator. The dataset quality assessment system includes a compliance attribute assessment module, a content attribute assessment module, a scale attribute assessment module, and a value attribute assessment module. The dataset to be evaluated is comprehensively evaluated in multiple dimensions. Then, a quality score corresponding to the dataset to be evaluated is determined based on at least one evaluation indicator and a preset weight corresponding to the evaluation indicator. The quality score of the dataset to be evaluated is accurately determined, and finally, the quality level corresponding to the dataset to be evaluated is determined based on the quality score and a preset score threshold. This solves the problem of being unable to accurately assess the quality of a dataset and achieves the beneficial effect of improving the accuracy and effectiveness of dataset quality assessment.
[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 This is a flowchart of a data set quality assessment method provided in accordance with the first embodiment of the present invention;
[0022] Figure 2 This is a flowchart of a data set quality assessment method provided in accordance with the second embodiment of the present invention;
[0023] Figure 3 2 is a schematic diagram of the structure of a data set quality assessment device provided according to the third embodiment of the present invention;
[0024] Figure 4 3 is a schematic diagram of the structure of an electronic device for implementing the data set quality assessment method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0026] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0027] Example 1
[0028] Figure 1 A flowchart of a data set quality assessment method is provided for the first embodiment of the present invention. This embodiment is applicable to data set quality assessment situations. The method can be performed by a data set quality assessment device. The data set quality assessment device can be implemented in the form of hardware and / or software. The assessment of the rescue force deployment operation can be configured in an electronic device. Figure 1 As shown, the method includes:
[0029] S110. Obtain a data set to be evaluated, and evaluate the data set to be evaluated based on a pre-established data set quality evaluation system to obtain at least one evaluation indicator; wherein the data set quality evaluation system includes a compliance attribute evaluation module, a content attribute evaluation module, a scale attribute evaluation module, and a value attribute evaluation module.
[0030] The dataset to be evaluated can be understood as the dataset waiting for quality assessment. The evaluation indicators can be understood as a series of quantitative or non-quantitative standards used to measure the performance of the evaluated entity during the evaluation process.
[0031] Specifically, a data set to be evaluated is obtained, and a comprehensive and holistic evaluation is performed on the data set to be evaluated based on the compliance attribute evaluation module, content attribute evaluation module, scale attribute evaluation module, and value attribute evaluation module of a pre-established data set quality evaluation system to obtain at least one evaluation indicator.
[0032] Optionally, the evaluating the dataset to be evaluated based on a pre-established dataset quality evaluation system to obtain at least one evaluation indicator includes: determining a target evaluation method corresponding to the dataset to be evaluated based on the pre-established dataset quality evaluation system; wherein the target evaluation method includes at least one of a 5-point scoring method, an expert scoring method, a 01 scoring method and a benchmark scoring method; and evaluating the dataset to be evaluated based on the target evaluation method to obtain at least one evaluation indicator.
[0033] The 0-1 scoring method assigns a score of 0 or 1 based on the presence or absence of the corresponding indicator. For example, if a field in a dataset is correctly assigned a value, it is scored 1, otherwise it is scored 0. The 5-point method assigns a score from 1 to 5 based on the correspondence between qualitative indicators and multiple specific evaluation indicators. For example, the frequency of dataset updates is scored from 1, 3, or 5 based on real-time updates, daily updates, or weekly updates. The benchmarking method assigns a score from 0 to 100 based on the quantitative data collected from quantitative indicators using the formula "quantitative data / total dataset * 100." For example, the data integrity measurement of null value rate is scored as null value rate = (number of missing or empty records / total number of records) × 100%. The expert scoring method cites the opinions, works, or experiences of authoritative experts as the basis for argumentation, decision-making, or evaluation, and assigns a score from 0 to 100. For example, content attributes account for 35%, compliance attributes for 15%, and scale and value attributes each for 25%. As can be seen, content attributes are the core of data quality assessment and account for the largest proportion.
[0034] Specifically, the target evaluation method corresponding to the data set to be evaluated is determined through the data set quality evaluation system, and a comprehensive evaluation is performed on each item to be evaluated of the data set to be evaluated (for example, security system, security control and data authorization, etc.) based on the target evaluation method to obtain the evaluation indicators corresponding to each item to be evaluated.
[0035] Optionally, the data set to be evaluated is evaluated based on a pre-established data set quality evaluation system to obtain at least one evaluation indicator, including at least one of the following: performing compliance attribute evaluation on the data set to be evaluated through the compliance attribute evaluation module to obtain a compliance attribute indicator; performing content attribute evaluation on the data set to be evaluated through the content attribute evaluation module to obtain a content attribute indicator; performing scale attribute evaluation on the data set to be evaluated through the scale attribute evaluation module to obtain a scale attribute indicator; performing value attribute evaluation on the data set to be evaluated through the value attribute evaluation module to obtain a value attribute indicator.
[0036] Data compliance can be understood as ensuring that businesses or organizations comply with relevant laws, regulations, industry guidelines, and standards when processing and using data, as well as requirements for protecting user privacy and data security. This includes compliance requirements for data collection, storage, processing, transmission, and sharing. Compliance with relevant laws and regulations is the foundation for businesses to collect, process, and apply data legally and in compliance with regulations. Compliance requires businesses or individuals to implement necessary security measures to protect the security and integrity of data. These primarily include the use of security controls such as encryption, access control, and data backup and recovery, as well as the establishment of privacy protection, risk assessment, security controls, and emergency response measures to prevent data leakage, damage, or misuse.
[0037] Optionally, the compliance attribute evaluation module includes a data standardization evaluation unit and a data security evaluation unit; the content attribute evaluation module includes a data accuracy evaluation unit, a data consistency evaluation unit, a data timeliness evaluation unit and a data integrity evaluation unit; the scale attribute evaluation module includes a quantity evaluation unit, a type evaluation unit and a coverage area evaluation unit; the value attribute module includes a utility evaluation unit, a benefit evaluation unit and a scenario evaluation unit.
[0038] Specifically, the four dimensions are content attributes, scale attributes, value attributes, and compliance attributes. Content attributes are divided into four indicators: accuracy, consistency, timeliness, and completeness; scale attributes are divided into three indicators: growth, diversity, and comprehensiveness; value attributes are divided into two indicators: utility and efficiency; and compliance attributes are divided into two indicators: standardization and security. Each of these indicators is divided into three levels, totaling 32 specific evaluation indicators. The four first-level indicators are designed to comprehensively categorize big data's basic attributes (natural attributes and social attributes) and the entire data lifecycle. For example, based on the national information technology-data quality evaluation indicators, indicators can be refined for each stage of data sources, data processes, and data applications.
[0039] Optionally, the compliance attribute module consists of a data standardization assessment unit and a data security assessment unit. The data standardization assessment unit includes laws and regulations, data authorization, and data ethics, while the data security assessment unit includes security policies and security controls. Exemplarily, the compliance attribute module accounts for 15% of the high-quality dataset evaluation indicator system.
[0040] Optionally, the compliance attribute evaluation module performs a compliance attribute evaluation on the data set to be evaluated to obtain a compliance attribute indicator, including: performing a normative evaluation on the data set to be evaluated by the data normative evaluation unit and the 01 scoring method to obtain at least one normative evaluation indicator; performing a security evaluation on the data set to be evaluated by the data security evaluation unit and the expert scoring method to obtain at least one security evaluation indicator; and determining the compliance attribute indicator based on at least one normative evaluation indicator and at least one security evaluation indicator.
[0041] For example, the evaluation indicators of the compliance attribute dimension are shown in the following table:
[0042]
[0043] Table 1
[0044] For example, laws and regulations - whether data collection is legal, and whether the data set complies with relevant laws and regulations on national security, personal privacy, commercial secrets, etc., 1 point if there is a legal basis, otherwise 0 point. Data authorization - data application authorization, whether authorization and permission are obtained for the use of the data set, 1 point if authorization and permission are obtained, otherwise 0 point. Data ethics - whether the data set contains discriminatory or biased information, 1 point if it does not contain discriminatory or biased information, otherwise 0 point. Security system - data compliance with relevant systems such as privacy protection, risk assessment, security control, emergency response and security review. Experts score whether it complies with data security systems based on their experience. Security control - the adoption of security control technologies such as data encryption, data backup, and access control. Experts score whether security controls are adopted based on their experience.
[0045] Optionally, after performing a normative evaluation on the dataset to be evaluated by using the data normative evaluation unit and the 01 scoring method to obtain at least one normative evaluation indicator, the method further includes:
[0046] If at least one of the normative evaluation indicators is 0, the dataset to be evaluated is determined to be an unqualified dataset and the evaluation is terminated.
[0047] It's worth noting that compliance attributes are a crucial element in protecting data security and possess a "veto power." That is, if any of the compliance attributes—laws and regulations, data authorization, and data ethics—are scored 0 using the 0-1 method, the dataset will no longer participate in subsequent indicator evaluations and will be deemed unqualified (low-quality). If all three are scored 1, the dataset will continue to be evaluated using the scoring tool.
[0048] Optionally, the content attribute module is composed of a data accuracy assessment unit, a data consistency assessment unit, a data timeliness assessment unit, and a data integrity assessment unit. The data accuracy assessment unit includes data content, data format, and data source; the data timeliness assessment unit includes format consistency, annotation consistency, and semantic consistency; and the data integrity assessment unit includes data latency, data update frequency, data validity period, data null value rate, data duplication rate, field integrity, and record integrity. Exemplarily, the content attribute module accounts for 35% of the high-quality dataset evaluation indicator system.
[0049] Optionally, the content attributes of a dataset are central to data quality assessment (for governments, etc.). The quality of data values directly impacts the outcomes of data use. Data's "three multiplies" (generated by multiple systems, linked to multiple departments, and stored in multiple formats) result in lengthy data collection and sharing processes, making it difficult to ensure data timeliness, accuracy, and completeness, and resulting in varying quality. Improving data quality requires establishing a comprehensive quality control system encompassing data collection, governance, sharing, and open access. Data content primarily refers to elements related to the data values themselves and is the most crucial dimension of a dataset. For example, whether data values are objectively accurate encompasses four aspects: accuracy, consistency, timeliness, and completeness. Accuracy is currently recognized by the international statistical community as a fundamental component of data quality and a key component and basis for data quality assessment by government statistical agencies. Consistency refers to the degree of data synchronization. If the same data differs during data processing, it indicates that data quality issues arose during the processing process. Timeliness: For government data, the time it takes from collection to the creation of a formal dataset is a significant factor. The shorter this timeframe, the closer the data update is to the day of use, and the newer the dataset, the higher its timeliness. Integrity is the most basic guarantee of data quality. The lower the integrity of the data, the less it can fully reflect the real situation, and the lower the quality of the data.
[0050] For example, the evaluation indicators of the content attribute dimension are shown in the following table:
[0051]
[0052]
[0053] Table 2
[0054] Specifically, accuracy: Data content—whether the data content meets expectations. At the content dimension, deep learning, natural language processing, and other technologies can be used to vectorize data descriptions. Combined with similarity calculations, the data content quality level can be assessed based on a similarity threshold. Benchmarking is used to calculate the similarity threshold.
[0055] Specifically, consistency: Data format - whether the data format complies with regulatory requirements. 1 point is awarded for compliance with custom specifications; otherwise, 0 points are awarded. Does the data format comply with the provisions of 5.5.3 of GB / T19488.1? Data source - the authority of the data source. 1 point is awarded for data directly obtained (first-hand) or official data; otherwise, 0 points are awarded. Format consistency - whether the dataset format remains consistent during storage and transmission. 1 point is awarded for consistency during the data formatting process; otherwise, 0 points are awarded. Annotation consistency - whether data is consistently annotated by multiple annotators. For consistency between annotators, Fleiss's Kappa coefficient is calculated using benchmarking, calculating the percentage of consistently annotated samples out of all annotated samples. The value ranges from -1 to 1. A higher value indicates better consistency between annotators. Semantic consistency - whether the definition or connotation of the data remains consistent across different stages and systems. 1 point is awarded for providing unified data definitions and data encoding; 0 points are awarded for ambiguous or incorrect definitions.
[0056] Specifically, timeliness: Data latency - the delay between data refreshes or updates. Data latency = data query time - last data update time. Data update frequency - the actual number of data updates within a period, or the difference in days between the last update date of a data resource and the data collection date. Data validity period - whether data remains valid after generation. Based on business needs and data characteristics, set a data validity threshold to check whether the data is within the validity period.
[0057] Specifically, completeness: Data null rate - whether the dataset contains null values. Data null rate = (number of missing or empty records / total number of records) × 100%. Data duplication rate - whether the dataset contains duplicate data. Data duplication rate = (number of duplicate values / total data volume) × 100%. Field completeness - whether each field in the dataset is correctly assigned a value, without omissions. If all fields in the dataset are correctly assigned a value, 1 point is awarded; if not, 0 points are awarded. Record completeness - whether each record in the dataset contains all required fields. If all required fields are included, 1 point is awarded; if not, 0 points are awarded.
[0058] Optionally, the scale attribute module consists of a quantity assessment unit, a type assessment unit, and a coverage assessment unit. The quantity assessment unit includes the total amount of data and data growth trends; the type assessment unit includes data types and data distribution; the type assessment unit is used to assess the diversity of the dataset to be evaluated; and the coverage assessment unit includes data domains and data themes. The coverage assessment unit includes an assessment of the comprehensiveness of the dataset to be evaluated. Exemplarily, the scale attribute module accounts for 25% of the high-quality dataset evaluation indicator system.
[0059] Specifically, data scale is a core characteristic for measuring data quality, primarily reflected in the amount of data, reflecting its richness and breadth. It is primarily measured through two dimensions: magnitude and comprehensiveness. Magnitude assesses the overall scale of the data. On the one hand, it measures the overall data volume, a direct indicator of data scale. On the other hand, it assesses data growth trends. Diversity refers to the diversity of data types and the breadth of data distribution, which helps provide a more comprehensive perspective. Comprehensiveness of data refers to the effort to cover all relevant data areas and topics during data processing and analysis to maximize the overall scale of the data. It is a key principle in data quality assessment and is crucial for decision support and business optimization.
[0060] For example, the evaluation indicators of the scale attribute dimension are shown in the following table:
[0061]
[0062]
[0063] Table 3
[0064] For example, order of magnitude: total amount of data - whether the dataset contains a large number of data points or data records, and the size of the hard disk space occupied by the dataset. Data growth trend - the speed of data generation and update, growth rate = the number of growth within a certain time interval / total number * 100%. Diversity: data type - whether the data is multimodal data, 1 point if the data types exceed the preset number of types, otherwise 0 point. Data distribution - whether the data is distributed on different values or categories, common measurement methods for distribution consistency include measurement methods based on sample weighting, measurement methods based on hypothesis testing, and measurement methods based on various measurement functions, etc. Among them, the measurement method based on the KL distance measurement function. Comprehensiveness: whether the dataset covers multiple application fields, 1 point if the application fields covered by the dataset exceed the preset number of fields, otherwise 0 point. Data topic - whether the dataset provides data content for multiple business topics, 1 point if the business topics provided by the dataset exceed the preset number of business topics, otherwise 0 point.
[0065] Optionally, the value attribute module consists of a utility evaluation unit, a benefit evaluation unit, and a scenario evaluation unit. The utility evaluation unit includes usability, understandability, and accessibility; the benefit evaluation unit includes market benefits and adaptation costs; and the scenario evaluation unit includes scenario diversity, business adaptability, and application feasibility. Exemplarily, the value attribute module accounts for 25% of the high-quality dataset evaluation indicator system.
[0066] Specifically, the value of data can only be fully tapped in specific application scenarios. Different industries and fields have different demands for data, so the application scenarios of data elements are very broad, including industry, finance, medical care, agriculture, transportation, electricity, smart cities, etc. The value attributes of data are mainly measured in three aspects: utility, efficiency, and scenario. Utility is mainly reflected in the support and assistance that can be provided for various activities and decisions, and is specifically broken down into ease of use, understandability, and accessibility. Efficiency is mainly an evaluation of the value of data being applied, mainly including market benefits and adaptation costs. Scenario quality characteristics describe data quality from the perspective of data usage scenarios. It refers to the degree to which data quality needs to be guaranteed in the system when data is used under specified conditions, such as timeliness and accessibility. It needs to be integrated with business processes and is an extended dimension category of the measurement framework.
[0067] For example, the evaluation indicators of the scale attribute dimension are shown in the following table:
[0068]
[0069]
[0070] Table 4
[0071] For example, utility: Usability measures the degree to which a dataset is usable by users. A dataset's subject classification is scientifically and rationally reasonable, with no overlap, and fuzzy classification is scored as 1; visualization is scored as 1, and non-visualization is scored as 0. Datasets with searchable tags are scored as 1, and 0 if not. Understandability measures the dataset's availability of a brief introduction or explanation, with 1 point awarded, and 0 if not. Accessibility measures the availability of data for application, i.e., whether the data is easily accessible when needed. Effectiveness measures Market returns, which measures the benefits generated by the dataset in the application scenario, and whether the dataset incurs losses and costs during collection, management, and application. Adaptation costs, which measures the costs incurred in the application scenario, and whether the dataset can generate benefits during collection, management, and application. Scenario-specific diversity measures whether the dataset is applicable to a single scenario or multiple scenarios, and the possibility of cross-scenario application. Business adaptability measures the dataset's ability to adapt to different business scenarios and requirements, and the degree to which the data is adaptable to scenario applications. Application feasibility measures the effectiveness of the data in actual business scenarios.
[0072] S120: Determine a quality score corresponding to the dataset to be evaluated based on at least one evaluation indicator and a preset weight corresponding to the evaluation indicator.
[0073] The preset weights may be pre-set based on experience, and are not limited in this embodiment. The quality score may be understood as a quantitative score of the comprehensive quality of the data set.
[0074] Specifically, data quality scores are primarily calculated using a weighted approach. For example, weights are assigned to the 32 items in the data quality assessment, and the final score is obtained through a weighted summation. When grading data quality overall, different weights should be assigned to the 32 items based on the actual application of the data. The accuracy, completeness, and effective use of data are fundamental to data application. Therefore, data accuracy and completeness scores should be assigned higher weights. If data timeliness is a high requirement in certain scenarios, the data timeliness score should be assigned a higher weight.
[0075] S130: Determine a quality level corresponding to the dataset to be evaluated based on the quality score and a preset score threshold.
[0076] The preset score threshold can be set based on experience and is not limited in this embodiment. The quality level can be divided into five levels: A, B, C, D, and E. It can also be divided into qualified or unqualified, or divided into high quality, medium-high quality, ordinary quality, and low quality, etc., which are not limited in this embodiment.
[0077] For example, the quality level of the dataset to be evaluated is divided into the following table:
[0078]
[0079] Table 5
[0080] The technical solution of an embodiment of the present invention obtains a dataset to be evaluated and evaluates it based on a pre-established dataset quality assessment system to obtain at least one evaluation indicator. The dataset quality assessment system includes a compliance attribute assessment module, a content attribute assessment module, a scale attribute assessment module, and a value attribute assessment module. The dataset to be evaluated is comprehensively evaluated in multiple dimensions. Then, a quality score corresponding to the dataset to be evaluated is determined based on at least one evaluation indicator and a preset weight corresponding to the evaluation indicator. The quality score of the dataset to be evaluated is accurately determined, and finally, the quality level corresponding to the dataset to be evaluated is determined based on the quality score and a preset score threshold. This solves the problem of being unable to accurately assess the quality of a dataset and achieves the beneficial effect of improving the accuracy and effectiveness of dataset quality assessment.
[0081] Example 2
[0082] Figure 2This is a flowchart of a dataset quality assessment method provided in Example 2 of the present invention. This embodiment further refines how the above-described embodiment determines a quality score corresponding to the dataset to be assessed based on at least one assessment metric and a preset weight corresponding to the assessment metric. Optionally, determining the quality score corresponding to the dataset to be assessed based on at least one assessment metric and a preset weight corresponding to the assessment metric includes performing a weighted sum of the at least one assessment metric and the preset weight corresponding to the assessment metric to determine the quality score corresponding to the dataset to be assessed.
[0083] like Figure 2 As shown, the method includes:
[0084] S210. Obtain a data set to be evaluated, and evaluate the data set to be evaluated based on a pre-established data set quality evaluation system to obtain at least one evaluation indicator; wherein the data set quality evaluation system includes a compliance attribute evaluation module, a content attribute evaluation module, a scale attribute evaluation module, and a value attribute evaluation module.
[0085] Specifically, the compliance attribute module is the foundation of high-quality datasets. This module ensures the legality, standardization, and accuracy of datasets through rigorous review and control of data compliance. It adheres to relevant laws, regulations, and industry standards, comprehensively managing data sources, collection methods, storage formats, and access permissions, effectively preventing risks such as data leakage and misuse. This directly impacts the quality and reliability of datasets, which in turn impacts subsequent data analysis and application. When constructing datasets, the development of the compliance attribute module is crucial to ensure high quality and credibility. Legal and regulatory assessments assess whether data sets comply with relevant laws and regulations, such as those concerning national security, personal privacy, and commercial confidentiality.
[0086] Specifically, data authorization assesses whether authorization and permission are obtained for the use of a dataset. Data ethics assesses whether a dataset contains discriminatory or biased information. Security systems assess whether the data complies with relevant systems such as privacy protection, risk assessment, security control, emergency response, and security review. Security controls assess the use of security control technologies such as data encryption, data backup, and access control.
[0087] Optionally, the content attribute module consists of data content, data format, data source, format consistency, labeling consistency, semantic consistency, data delay, data update frequency, data validity period, data null rate, data duplication rate, field integrity, and record integrity. The content attribute module accounts for 35% of the high-quality dataset evaluation index system.
[0088] The optional data content attribute module, as a core component of a dataset, directly and profoundly reflects the overall quality of the dataset, serving as a direct and critical basis for dataset quality assessment. This module not only covers key elements such as the data's specific content, format specifications, and provenance, but also addresses multiple dimensions such as data integrity, consistency, and timeliness, comprehensively and deeply reflecting the dataset's quality characteristics. Through rigorous data screening, organization, and verification, this module ensures the dataset's accuracy, reliability, and practicality, providing a solid foundation for subsequent data application and analysis. During the dataset's construction and management, we prioritize the development and optimization of the content attribute module to ensure it meets high-quality standards in all aspects, fully demonstrating the dataset's value and meeting the demand for high-quality data. Data content assesses whether the data content meets the expectations of high-quality dataset standards. Data format assesses whether the data format meets regulatory requirements.
[0089] Specifically, data provenance assesses the authority of the data source. Format consistency assesses whether the format of the dataset remains consistent during storage and transmission. Annotation consistency assesses whether the data is consistently annotated by multiple annotators.
[0090] Semantic consistency evaluates whether the definition or connotation of data remains consistent across different stages and systems. Data latency evaluates the delay in refreshing or updating data. Data update frequency evaluates the actual number of data updates within a period of time based on the data set collection rules. Data validity period evaluates whether the data remains valid within the time range after generation. Data duplication rate evaluates whether there is duplicate data in the data set. Field completeness evaluates whether each field in the data set is correctly assigned a value without omission. Record completeness evaluates whether each record in the data set contains all necessary fields. The scale attribute module consists of the total amount of data, data growth trend, data type, data distribution, data domain, and data theme. The scale attribute module accounts for 25% of the high-quality data set evaluation indicator system.
[0091] Specifically, the scale attribute module is an essential component for measuring and building high-quality datasets. It plays a crucial role in ensuring the comprehensiveness and accuracy of datasets. It provides multi-dimensional information such as dataset size, growth rate, and type distribution, revealing the inherent structure and characteristics of the data. This ensures that datasets have sufficient sample size to support complex data analysis and model training, ensuring that datasets can continue to meet business development requirements. By providing comprehensive and accurate dataset feature information, the accuracy of data analysis and decision-making is improved, providing strong support for business development and innovation.
[0092] Specifically, the total data volume evaluates whether the dataset contains a large number of data points or data records. The data growth trend assesses the speed of data generation and update. The data type evaluates whether the data is multimodal. The data distribution evaluates whether the data is distributed across different values or categories. The data domain evaluates whether the dataset covers various application areas. The value attribute module consists of ease of use, understandability, accessibility, market benefits, adaptation cost, scenario diversity, business adaptability, and application feasibility. The value attribute module accounts for 25% of the high-quality dataset evaluation indicator system.
[0093] The optional Value Attributes module provides a systematic and structured approach to assessing, understanding, and optimizing value. By defining metrics such as ease of use, understandability, accessibility, market benefits, adaptation costs, scenario diversity, business adaptability, and application feasibility, the Value Attributes module provides a comprehensive understanding of data value, accurately grasps high-quality datasets, adds value to dataset users, and provides a data element basis for promoting the development of the digital economy.
[0094] Specifically, usability evaluates the degree to which a dataset is available to users. Comprehensibility evaluates the degree to which the dataset is understood by users, and whether the dataset has a relevant introduction or explanation. Accessibility evaluates the availability of data when it is needed, and whether the data is easy to obtain when needed. Market benefits evaluate the benefit measurement generated by the dataset in the application scenario, and whether the dataset incurs losses and costs during collection, governance, and application. Adaptation costs evaluate the cost measurement generated by the dataset in the application scenario, and whether the dataset can bring benefits during collection, governance, and application. Scenario diversity evaluates whether the dataset is applied to a single scenario or multiple scenarios, and the possibility of cross-scenario application of the dataset. Business adaptability evaluates the ability of the dataset to be applicable to different business scenarios and needs, and the degree of adaptability of the data in scenario applications. Application feasibility evaluates the application effect of the data in actual business scenarios.
[0095] Optionally, the data collection method includes qualitative analysis indicators and quantitative analysis indicators.
[0096] For example, for indicators requiring qualitative analysis, the evaluation factors are mainly derived from descriptive statistical analysis and text analysis of relevant laws and regulations, policies, annual plans and work programs, standards and specifications, news reports, and other materials. The search methods mainly include using search engines to search for keywords in relevant laws and regulations and policy texts, standards and specifications, annual work plans, news reports, information from data open management agencies, and expert opinions in the relevant industries of data elements. Data is collected and judged through manual observation, keyword search, and inviting leading industry expert groups to conduct research and judgment on various corporate public websites and various data-related platforms.
[0097] For example, for indicators requiring quantitative analysis, evaluation factors are primarily derived through automatic capture of publicly available authorized information from various enterprises and data platforms, combined with manual observation to collect relevant authorized information. This information is then subjected to descriptive statistical analysis, cross-sectional analysis, text analysis, and spatial analysis. For indicators requiring expert evaluation and questionnaire surveys, stratified sampling surveys and field surveys of representative local enterprises are used for collection and analysis.
[0098] The acquisition, storage, use, and processing of data in the technical solution of this application are in compliance with the laser tube regulations of national laws and regulations.
[0099] S220: Perform a weighted summation on at least one of the evaluation indicators and a preset weight corresponding to the evaluation indicator to determine a quality score corresponding to the data set to be evaluated.
[0100] Exemplarily, the first-level indicator score corresponding to the dataset to be evaluated is determined by the following formula:
[0101]
[0102] Among them, Q n Y is the specific score of the third-level indicators of each attribute. n W is the specific indicator weight of each attribute’s third-level indicator. n is the weight of the first-level indicator. n It is the score of the first-level indicator.
[0103] It is worth noting that when assessing data quality, if the data compliance dimension is judged to be 0 using the 0-1 method, no further scoring is performed on other dimensions, and the entire dataset is judged to be of low quality. If the data compliance attribute is judged to be 1, calculations can continue according to the scoring formula.
[0104] Exemplarily, the quality score corresponding to the dataset to be evaluated is determined by the following formula:
[0105]
[0106] Where X1 is the compliance attribute indicator. X2 is the content attribute indicator. X3 is the scale attribute indicator. X4 is the value attribute indicator. W1 is the preset weight of the compliance attribute indicator. W2 is the preset weight of the content attribute indicator. W3 is the preset weight of the scale attribute indicator. W4 is the preset weight of the value attribute indicator.
[0107] S230: Determine a quality level corresponding to the dataset to be evaluated based on the quality score and a preset score threshold.
[0108] The technical solution of the embodiment of the present invention determines a quality score corresponding to the dataset to be evaluated by performing a weighted summation of at least one evaluation indicator and a preset weight corresponding to the evaluation indicator, thereby improving the scientificity and practicality of data quality assessment.
[0109] Example 3
[0110] Figure 3 This is a schematic diagram of the structure of a data set quality assessment device provided in the third embodiment of the present invention. Figure 3 As shown, the apparatus includes: an evaluation module 310 , a score determination module 320 and a grade determination module 330 .
[0111] Among them, the evaluation module 310 is used to obtain the data set to be evaluated, and evaluate the data set to be evaluated based on a pre-established data set quality evaluation system to obtain at least one evaluation indicator; wherein, the data set quality evaluation system includes a compliance attribute evaluation module, a content attribute evaluation module, a scale attribute evaluation module and a value attribute evaluation module; the score determination module 320 is used to determine the quality score corresponding to the data set to be evaluated based on at least one evaluation indicator and a preset weight corresponding to the evaluation indicator; the grade determination module 330 is used to determine the quality grade corresponding to the data set to be evaluated based on the quality score and a preset score threshold.
[0112] The technical solution of an embodiment of the present invention obtains a dataset to be evaluated and evaluates it based on a pre-established dataset quality assessment system to obtain at least one evaluation indicator. The dataset quality assessment system includes a compliance attribute assessment module, a content attribute assessment module, a scale attribute assessment module, and a value attribute assessment module. The dataset to be evaluated is comprehensively evaluated in multiple dimensions. Then, a quality score corresponding to the dataset to be evaluated is determined based on at least one evaluation indicator and a preset weight corresponding to the evaluation indicator. The quality score of the dataset to be evaluated is accurately determined, and finally, the quality level corresponding to the dataset to be evaluated is determined based on the quality score and a preset score threshold. This solves the problem of being unable to accurately assess the quality of a dataset and achieves the beneficial effect of improving the accuracy and effectiveness of dataset quality assessment.
[0113] Optionally, the evaluation module includes:
[0114] An evaluation method determination unit, configured to determine a target evaluation method corresponding to the dataset to be evaluated based on a pre-established dataset quality evaluation system; wherein the target evaluation method includes at least one of a 5-point scoring method, an expert scoring method, a 0.1 scoring method, and a benchmark scoring method;
[0115] An evaluation indicator determination unit is used to evaluate the data set to be evaluated based on the target evaluation method to obtain at least one evaluation indicator.
[0116] Optionally, the evaluation module includes at least one of the following units:
[0117] a compliance assessment unit, configured to perform compliance attribute assessment on the dataset to be assessed using the compliance attribute assessment module to obtain a compliance attribute indicator;
[0118] a content evaluation unit, configured to perform content attribute evaluation on the to-be-evaluated data set through the content attribute evaluation module to obtain a content attribute index;
[0119] a scale assessment unit, configured to perform scale attribute assessment on the dataset to be assessed using the scale attribute assessment module to obtain a scale attribute index;
[0120] The value evaluation unit is used to perform value attribute evaluation on the data set to be evaluated through the value attribute evaluation module to obtain a value attribute index.
[0121] Optionally, the compliance attribute evaluation module includes a data standardization evaluation unit and a data security evaluation unit; the content attribute evaluation module includes a data accuracy evaluation unit, a data consistency evaluation unit, a data timeliness evaluation unit and a data integrity evaluation unit; the scale attribute evaluation module includes a quantity evaluation unit, a type evaluation unit and a coverage area evaluation unit; the value attribute module includes a utility evaluation unit, a benefit evaluation unit and a scenario evaluation unit.
[0122] Optionally, the compliance assessment unit includes:
[0123] a normative evaluation subunit, configured to perform a normative evaluation on the dataset to be evaluated by using the data normative evaluation unit and the 01 scoring method, so as to obtain at least one normative evaluation indicator;
[0124] A security assessment subunit, configured to perform a security assessment on the data set to be assessed using the data security assessment unit and an expert scoring method to obtain at least one security assessment indicator;
[0125] The compliance indicator determination subunit is used to determine the compliance attribute indicator based on at least one of the normative evaluation indicators and at least one of the security evaluation indicators.
[0126] Optionally, the device also includes an unqualified data set determination module, which is used to perform a normative evaluation on the data set to be evaluated through the data normative evaluation unit and the 01 scoring method to obtain at least one normative evaluation indicator. If at least one of the normative evaluation indicators is 0, determine that the data set to be evaluated is an unqualified data set and end the evaluation.
[0127] Optionally, the score determination module is specifically configured to:
[0128] A weighted sum is performed on at least one of the evaluation indicators and a preset weight corresponding to the evaluation indicator to determine a quality score corresponding to the data set to be evaluated.
[0129] The data set quality assessment device provided in the embodiment of the present invention can execute the data set quality assessment method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0130] Example 4
[0131] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0132] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0133] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0134] The processor 11 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for dataset quality assessment.
[0135] In some embodiments, the method for data set quality assessment can be implemented as a computer program that is tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method for data set quality assessment described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the method for data set quality assessment in any other suitable manner (e.g., by means of firmware).
[0136] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0137] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0138] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0139] To provide interaction with a service acquirer, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the service acquirer; and a keyboard and pointing device (e.g., a mouse or trackball), through which the service acquirer can provide input to the electronic device. Other types of devices can also be used to provide interaction with the service acquirer; for example, the feedback provided to the service acquirer can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the service acquirer can be received in any form (including acoustic input, voice input, or tactile input).
[0140] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a service acquirer computer having a graphical service acquirer interface or a web browser through which a service acquirer can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0141] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0142] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0143] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A dataset quality assessment method, characterized in that: include: Obtaining a dataset to be evaluated, and evaluating the dataset based on a pre-established dataset quality evaluation system to obtain at least one evaluation indicator; wherein the dataset quality evaluation system includes a compliance attribute evaluation module, a content attribute evaluation module, a scale attribute evaluation module, and a value attribute evaluation module; Determining a quality score corresponding to the data set to be evaluated based on at least one evaluation indicator and a preset weight corresponding to the evaluation indicator; A quality level corresponding to the data set to be evaluated is determined based on the quality score and a preset score threshold.
2. The method according to claim 1, characterized in that The data set to be evaluated is evaluated based on a pre-established data set quality evaluation system to obtain at least one evaluation indicator, including: Determining a target evaluation method corresponding to the dataset to be evaluated based on a pre-established dataset quality evaluation system; wherein the target evaluation method includes at least one of a 5-point scoring method, an expert scoring method, a 01 scoring method, and a benchmark scoring method; The to-be-evaluated data set is evaluated based on the target evaluation method to obtain at least one evaluation indicator.
3. The method according to claim 1, characterized in that The dataset to be evaluated is evaluated based on a pre-established dataset quality evaluation system to obtain at least one evaluation indicator, including at least one of the following: Performing compliance attribute evaluation on the dataset to be evaluated by the compliance attribute evaluation module to obtain a compliance attribute index; Performing content attribute evaluation on the data set to be evaluated by the content attribute evaluation module to obtain content attribute indicators; Performing scale attribute evaluation on the dataset to be evaluated by the scale attribute evaluation module to obtain a scale attribute index; The value attribute evaluation module performs value attribute evaluation on the data set to be evaluated to obtain a value attribute index.
4. The method according to claim 1, wherein The compliance attribute evaluation module includes a data standardization evaluation unit and a data security evaluation unit; the content attribute evaluation module includes a data accuracy evaluation unit, a data consistency evaluation unit, a data timeliness evaluation unit and a data integrity evaluation unit; the scale attribute evaluation module includes a quantity evaluation unit, a type evaluation unit and a coverage area evaluation unit; the value attribute module includes a utility evaluation unit, a benefit evaluation unit and a scenario evaluation unit.
5. The method according to claim 4, characterized in that The step of performing compliance attribute evaluation on the dataset to be evaluated by the compliance attribute evaluation module to obtain compliance attribute indicators includes: Performing a normative evaluation on the dataset to be evaluated by using the data normative evaluation unit and the 01 scoring method to obtain at least one normative evaluation indicator; Performing a security assessment on the data set to be assessed using the data security assessment unit and the expert scoring method to obtain at least one security assessment indicator; The compliance attribute indicator is determined based on at least one of the normative evaluation indicators and at least one of the security evaluation indicators.
6. The method according to claim 5, characterized in that After performing a normative evaluation on the dataset to be evaluated by the data normative evaluation unit and the 01 scoring method to obtain at least one normative evaluation indicator, the method further includes: If at least one of the normative evaluation indicators is 0, the dataset to be evaluated is determined to be an unqualified dataset and the evaluation is terminated.
7. The method according to claim 1, characterized in that The determining of a quality score corresponding to the dataset to be evaluated based on at least one evaluation indicator and a preset weight corresponding to the evaluation indicator includes: A weighted sum is performed on at least one of the evaluation indicators and a preset weight corresponding to the evaluation indicator to determine a quality score corresponding to the data set to be evaluated.
8. A data set quality assessment device, characterized in that: include: An evaluation module, configured to obtain a dataset to be evaluated and evaluate the dataset based on a pre-established dataset quality evaluation system to obtain at least one evaluation indicator; wherein the dataset quality evaluation system includes a compliance attribute evaluation module, a content attribute evaluation module, a scale attribute evaluation module, and a value attribute evaluation module; a score determination module, configured to determine a quality score corresponding to the dataset to be evaluated based on at least one evaluation indicator and a preset weight corresponding to the evaluation indicator; The grade determination module is configured to determine a quality grade corresponding to the data set to be evaluated based on the quality score and a preset score threshold.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the dataset quality assessment method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the dataset quality assessment method according to any one of claims 1 to 7 when executed.
Citation Information
Cited By
Unstructured data value evaluation method and related equipment thereof
CN121412324A