Knowledge graph-based data set quality evaluation method and device, and electronic equipment
By proposing a knowledge graph-based dataset quality assessment method, which utilizes graph attention networks to generate embedded representations of assessment rules, the problem of inappropriate selection of assessment rules in existing technologies is solved, and efficient and accurate dataset quality assessment is achieved.
Patent Information
- Application Number
- CN202511501705.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-10-21
AI Technical Summary
In existing technologies, dataset quality assessment methods struggle to accurately select the most suitable assessment rules, impacting assessment efficiency and accuracy.
A dataset quality assessment method based on knowledge graphs is constructed. By acquiring knowledge graphs, target assessment rules are determined based on the feature information of the dataset to be assessed. Graph attention networks are used to generate embedded representations of the assessment rules. Finally, the most suitable assessment rule is determined through similarity calculation and matching degree to assess the dataset quality.
It improves the efficiency and accuracy of dataset quality assessment, enables precise evaluation of various quality indicators of the dataset, reduces manual searching and workload, and enhances the interpretability of the assessment results.
Smart Images

Figure CN120974130B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data evaluation technology, such as a knowledge graph-based dataset quality evaluation method and apparatus, and electronic equipment. Background Technology
[0002] Currently, in the fields of artificial intelligence and machine learning, datasets are fundamental for model training and validation. High-quality datasets can significantly improve model performance, including accuracy, robustness, and generalization ability. However, datasets may contain issues such as labeling errors, missing data, data inconsistencies, and imbalanced data distribution, all of which directly affect the training effect of the model. Therefore, dataset quality assessment is necessary.
[0003] In related technologies, a dataset quality assessment method is disclosed, which evaluates the dataset based on multiple preset evaluation dimensions, determines the evaluation result of the dataset in each preset evaluation dimension, and then evaluates the quality of the dataset based on preset quality assessment rules and the evaluation results.
[0004] In the process of implementing the embodiments of this disclosure, at least the following problems were found in the related art:
[0005] In related technologies, the evaluation method is based on preset evaluation dimensions and preset quality evaluation rules. For different datasets, it is difficult to accurately select the most suitable evaluation rule, which affects the evaluation efficiency and accuracy.
[0006] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0007] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.
[0008] This disclosure provides a knowledge graph-based dataset quality assessment method, apparatus, and electronic device to determine suitable target assessment rules based on the dataset to be assessed, thereby improving assessment efficiency and accuracy.
[0009] In some embodiments, a knowledge graph-based dataset quality assessment method includes: acquiring a knowledge graph for dataset quality assessment; determining target assessment rules from the knowledge graph based on the feature information of the dataset to be assessed; and performing quality assessment on the dataset to be assessed according to the target assessment rules to obtain assessment results.
[0010] Optionally, the knowledge graph is constructed as follows: collect evaluation rules and results of historical datasets from data sources; encode the collected data to obtain encoded data; extract entities, relationships, and attributes from the encoded data to form structured data; entities include dataset type, evaluation metrics, evaluation methods, data sources, and dataset versions; relationships include the types of relationships between entities; attributes include detailed information about the entities; and a knowledge graph is constructed based on the structured data.
[0011] Optionally, based on the feature information of the dataset to be evaluated, the target evaluation rule is determined from the knowledge graph, including: extracting the feature information of the dataset to be evaluated to obtain a feature vector; the feature information includes the data type, data volume, data distribution, and content features of the dataset to be evaluated; determining candidate evaluation rules based on the similarity between the feature vector and the embedding representation of each evaluation rule in the knowledge graph; and determining the target evaluation rule based on the matching degree between the candidate evaluation rules and the dataset to be evaluated.
[0012] Optionally, the embedding representation of each evaluation rule in the knowledge graph is generated as follows: In the knowledge graph, entity features and relation features related to the evaluation rule are extracted to obtain an initial feature vector for each evaluation rule; based on a graph attention network, each initial feature vector is aggregated with the feature vectors of adjacent entities to generate the embedding representation of each evaluation rule; wherein, the embedding representation of the evaluation rule is:
[0013]
[0014] B is the embedding representation of the evaluation rule, σ is the non-linear activation function, K is the number of attention heads in the graph attention network, and N is the number of attention heads. i Let α be the set of adjacent entities. ij (k) W represents the attention weight of the k-th attention head. (k) Let h be the weight matrix of the k-th attention head, L be the number of layers in the graph attention network, and h be the weight matrix of the k-th attention head. i (L) Let i be the embedding representation of the Lth layer, i be the index of the evaluation rule, and j be the index of the vector entity.
[0015] Optionally, the similarity between the feature vector and the embedding representation of the evaluation rule is calculated according to the following formula:
[0016]
[0017] Where D is the similarity, β is the dynamic adjustment factor, M is the weight vector representing the importance weight of each feature, ⨀ represents element-wise multiplication, A is the feature vector, and B is the embedding representation.
[0018] Optionally, a target evaluation rule is determined based on the matching degree between the candidate evaluation rules and the dataset to be evaluated, including one or more of the following: determining the target evaluation rule based on the matching degree between the evaluation metrics of the dataset to be evaluated and the evaluation metrics covered by the candidate evaluation rules; determining the target evaluation rule based on the matching degree between the dataset type of the dataset to be evaluated and the dataset type to which the candidate evaluation rules apply; determining the user's commonly used evaluation rules for the dataset type based on the user's historical usage records and the dataset type of the dataset to be evaluated, and determining the target evaluation rule based on the matching degree between the commonly used evaluation rules and the candidate evaluation rules.
[0019] Optionally, the dataset quality assessment method further includes: after obtaining the target assessment rule, adjusting the target assessment rule according to the data to be assessed; specifically including: obtaining the characteristic information of the dataset to be assessed; the characteristic information includes data distribution, noise level, and data volume; and adjusting the assessment index of the target assessment rule according to the characteristic information of the dataset to be assessed.
[0020] Optionally, the dataset quality assessment method further includes: comparing the assessment results of the dataset to be assessed with the assessment results of historical datasets using the same objective assessment rule; and displaying the knowledge graph, objective assessment rule, assessment results of the dataset to be assessed, and the comparative analysis results of the dataset to be assessed and historical datasets in a visual interface.
[0021] In some embodiments, a knowledge graph-based dataset quality assessment apparatus includes a processor and a memory storing program instructions, the processor being configured to execute the knowledge graph-based dataset quality assessment method described above when the program instructions are executed.
[0022] In some embodiments, the electronic device includes: an electronic device body; and a dataset quality assessment device based on a knowledge graph, as described above, installed on the electronic device body.
[0023] The knowledge graph-based dataset quality assessment method, apparatus, and electronic device provided in this disclosure can achieve the following technical effects:
[0024] In this embodiment, the knowledge graph used for dataset quality assessment can structurally manage dispersed dataset quality assessment knowledge, making the quality assessment knowledge more systematic and easier to store, query, and update. Based on the feature information of the dataset to be assessed, suitable target assessment rules can be obtained from the knowledge graph, reducing the time and workload required for pre-setting assessment rules. Finally, the dataset to be assessed is evaluated according to the target assessment rules, enabling the most appropriate evaluation of each quality indicator of the dataset, thereby improving assessment efficiency and accuracy.
[0025] The above general description and the description below are exemplary and illustrative only and are not intended to limit this application. Attached Figure Description
[0026] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein:
[0027] Figure 1 This is a schematic diagram of a dataset quality assessment method based on knowledge graphs provided in an embodiment of this disclosure;
[0028] Figure 2 This is a schematic diagram of a method for constructing a knowledge graph provided in an embodiment of this disclosure;
[0029] Figure 3 This is a schematic diagram of another knowledge graph-based dataset quality assessment method provided in this disclosure embodiment;
[0030] Figure 4 This is a schematic diagram of another knowledge graph-based dataset quality assessment method provided in this disclosure embodiment;
[0031] Figure 5 This is a schematic diagram of a dataset quality assessment device based on a knowledge graph provided in an embodiment of this disclosure;
[0032] Figure 6 This is a schematic diagram of another knowledge graph-based dataset quality assessment device provided in this embodiment. Detailed Implementation
[0033] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.
[0034] The terms "first," "second," etc., used in the technical solutions described in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.
[0035] Unless otherwise stated, the term "multiple" means two or more.
[0036] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.
[0037] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0038] The term "correspondence" can refer to an association or binding relationship. The correspondence between A and B means that there is an association or binding relationship between A and B.
[0039] Combination Figure 1 As shown, this disclosure provides a dataset quality assessment method based on knowledge graphs. The execution entity of this dataset quality assessment method can be a processor, and the dataset quality assessment method includes:
[0040] S101, the processor acquires the knowledge graph used for dataset quality assessment.
[0041] S102, the processor determines the target evaluation rules from the knowledge graph based on the feature information of the dataset to be evaluated.
[0042] S103, the processor performs a quality assessment on the dataset to be evaluated according to the target evaluation rules and obtains the evaluation results.
[0043] In this embodiment, the knowledge graph used for dataset quality assessment can structurally manage dispersed dataset quality assessment knowledge, making the quality assessment knowledge more systematic and easier to store, query, and update. Based on the feature information of the dataset to be assessed, suitable target assessment rules can be obtained from the knowledge graph, reducing the time and workload required for pre-setting assessment rules. Finally, the dataset to be assessed is evaluated according to the target assessment rules, enabling the most appropriate evaluation of each quality indicator of the dataset, thereby improving assessment efficiency and accuracy.
[0044] Optionally, the knowledge graph can be constructed as follows: collect evaluation rules and evaluation results of historical datasets from data sources; encode the collected data to obtain encoded data; extract entities, relationships and attributes from the encoded data to form structured data; and construct the knowledge graph based on the structured data.
[0045] Combination Figure 2 As shown, this disclosure provides a method for constructing a knowledge graph. The execution entity of this method may be a processor, and the method includes:
[0046] S201, the processor collects evaluation rules and evaluation results from historical datasets from the data source.
[0047] S202, the processor encodes the collected data to obtain encoded data.
[0048] S203, the processor extracts entities, relationships and attributes from the encoded data to form structured data.
[0049] The S204 processor builds knowledge graphs based on structured data.
[0050] In this embodiment, evaluation rules and results from historical datasets are collected from data sources and encoded to form structured data, thus achieving the systematization and structuring of knowledge. This centralizes evaluation rules and results scattered across different data sources, creating a unified knowledge system that facilitates management and utilization. Structured data makes knowledge easier to query and retrieve, allowing evaluators to quickly find the required evaluation rules and historical results. Through encoding and structuring, knowledge becomes more standardized, facilitating reuse across different projects and scenarios and reducing repetitive work.
[0051] Optionally, the evaluation rules include evaluation rules for image datasets and evaluation rules for text datasets. Evaluation rules are specific standards and methods for measuring dataset quality. They define how to evaluate various quality metrics of the dataset, such as accuracy, completeness, consistency, diversity, and distribution balance.
[0052] Optionally, the image dataset evaluation rules include: accuracy evaluation rules, integrity evaluation rules, consistency evaluation rules, and diversity evaluation rules.
[0053] In one specific embodiment, the accuracy evaluation rules for the image dataset include: checking the accuracy of image annotations using annotation tools, such as using the YOLO (You Only Look Once) object detection algorithm to detect objects in the image, comparing the detection results with the manually annotated results, and calculating the accuracy rate. If the accuracy rate is below 90%, the annotations need to be reviewed again.
[0054] In one specific embodiment, the integrity assessment rule for the image dataset includes checking whether there are missing image files or corresponding annotation files in the image dataset. This can be achieved by writing a Python script to traverse the image folder and annotation folder, comparing the number of files and their names to see if they match.
[0055] In one specific embodiment, the consistency evaluation rules for the image dataset include ensuring that all images in the image dataset have the same size and format. For example, the OpenCV library is used to read image files, and the width, height, and format of the images are checked; if they are inconsistent, they are converted or removed.
[0056] In one specific embodiment, the diversity evaluation rule for the image dataset includes: evaluating whether the image dataset contains images under different conditions such as angle, lighting, and background. Data augmentation techniques (such as rotation, flipping, and brightness adjustment) can be used to generate more diverse images, or clustering algorithms (such as K-means) can be used to cluster image features and analyze the diversity of the clustering results.
[0057] Optionally, the text dataset evaluation rules include: accuracy evaluation rules, completeness evaluation rules, consistency evaluation rules, and semantic consistency evaluation rules.
[0058] In one specific embodiment, the accuracy evaluation rules for the text dataset include: checking the accuracy of text annotation, for example, for a sentiment analysis dataset, manually annotating the sentiment tendency (positive, negative, neutral) of each text, then using a sentiment analysis model (such as a BERT-based model) to make predictions, comparing the prediction results with the manually annotated results, and calculating the accuracy.
[0059] In one specific embodiment, the integrity assessment rules for the text dataset include checking whether there are missing text content or corresponding annotation labels in the text dataset. The integrity and correspondence between text files and annotation files can be checked by writing a Python script.
[0060] In one specific embodiment, the consistency evaluation rules for the text dataset include ensuring that the encoding format, delimiters, etc., of the text in the text dataset are consistent. For example, the UTF-8 encoding format is used uniformly, and specific delimiters (such as commas, tabs, etc.) are used uniformly to separate different text contents.
[0061] In one specific embodiment, the semantic consistency evaluation rule for the text dataset includes checking whether there are semantic contradictions or inconsistencies in the text dataset. For example, semantic similarity calculation methods (such as cosine similarity to calculate the similarity between text vectors) are used to discover similar texts but with different annotations.
[0062] Optionally, the evaluation results include annotation accuracy, completeness score, consistency score, diversity score, and semantic consistency score. The evaluation results are specific numerical values and conclusions obtained after evaluating the dataset according to the evaluation rules, reflecting the dataset's performance on each quality indicator.
[0063] Optionally, the collected data is encoded to obtain encoded data, including: parsing and processing the evaluation rules and evaluation results in text form, extracting key information from the text; and organizing the extracted key information into a structured encoding format to obtain encoded data.
[0064] In this embodiment, text data from different sources and in different formats are converted into a unified structured format to facilitate subsequent processing and analysis. Based on the evaluation rules and key information in the evaluation results, ambiguity and redundancy in the text can be eliminated, ensuring the accuracy and consistency of the data.
[0065] Optionally, the evaluation rules and results in text form are parsed and processed to extract key information from the text, including: cleaning the text data to remove irrelevant characters and stop words; and using natural language processing techniques to extract key information from the text.
[0066] In this embodiment, natural language processing technology is used to extract key information from text through text classification algorithms and entity recognition algorithms, such as TF-IDF and LDA (Latent Dirichlet Allocation) models. This information includes dataset names, evaluation metric names, and evaluation rule names. Finally, the extracted key information is organized into a structured encoding format, such as JSON or XML. For example, the name of the evaluation rule, its scope of application, and the evaluation method can be stored as different fields.
[0067] Optionally, entities, relationships, and attributes can be extracted from the encoded data to form structured data, including: using entities as nodes, relationships as edges, and attributes as attributes of nodes or edges to construct a graph structure and form structured data.
[0068] Optionally, entities include: dataset type (such as image, text, time series, etc.), evaluation metrics (such as accuracy, completeness, consistency, etc.), evaluation methods (such as statistical analysis, machine learning models, etc.), data source (such as database, file system, etc.), dataset version, etc.
[0069] Optionally, the relationship includes: the type of relationship between entities, such as "applies to" (a certain evaluation metric applies to a certain dataset type), "contains" (a certain dataset contains a specific evaluation metric), "generates" (a certain evaluation method generates a specific evaluation result), etc.
[0070] Optionally, the attributes include: detailed information about the entity, such as the size of the dataset, resolution, annotation accuracy value, and specific values of the evaluation metrics.
[0071] Optionally, the dataset quality assessment method also includes: storing the constructed knowledge graph in a graph database; periodically checking for changes in the data source, and updating the knowledge graph based on the new data when new assessment rules or assessment results appear in the data source.
[0072] In this embodiment, storing the knowledge graph in a graph database ensures its integrity and consistency. The graph database also provides efficient graph traversal and query capabilities, enabling rapid retrieval of specific entities, relationships, and attributes, thus improving the query efficiency of the knowledge graph. Regularly checking for changes in the data source and updating the knowledge graph ensures that the data in the knowledge graph is always up-to-date, reflecting the latest evaluation rules and results.
[0073] Optionally, use Python's Scrapy or Beautiful Soup libraries for data crawling, combined with scheduled task tools such as Apache Airflow or Cron; based on data comparison algorithms, identify newly added or modified data according to the crawled data; based on incremental update algorithms (such as Neo4j's APOC process), integrate the updated data into the existing knowledge graph to ensure data consistency and integrity.
[0074] Optionally, based on the feature information of the dataset to be evaluated, the target evaluation rule is determined from the knowledge graph, including: extracting the feature information of the dataset to be evaluated to obtain a feature vector; the feature information includes the data type, data volume, data distribution, and content features of the dataset to be evaluated; determining candidate evaluation rules based on the similarity between the feature vector and the embedding representation of each evaluation rule in the knowledge graph; and determining the target evaluation rule based on the matching degree between the candidate evaluation rules and the dataset to be evaluated.
[0075] Combination Figure 3 As shown, this disclosure provides another method for evaluating the quality of datasets based on knowledge graphs, including:
[0076] S301, the processor acquires a knowledge graph for dataset quality assessment.
[0077] S302, the processor extracts the feature information of the dataset to be evaluated and obtains the feature vector.
[0078] S303, the processor determines candidate evaluation rules based on the similarity between the feature vector and the embedded representation of each evaluation rule in the knowledge graph.
[0079] S304, the processor determines the target evaluation rule based on the matching degree between the candidate evaluation rules and the dataset to be evaluated.
[0080] S305, the processor performs a quality assessment on the dataset to be evaluated according to the target evaluation rules, and obtains the evaluation results.
[0081] In this embodiment, the feature vector is obtained by extracting the feature information of the dataset to be evaluated. Then, the similarity between the feature vector and the embedding representation of each evaluation rule in the knowledge graph is calculated by combining the rich rules in the knowledge graph. This can accurately determine the few candidate evaluation rules with the highest similarity. Then, the most matching target evaluation rule is determined from the candidate evaluation rules, reducing the time and workload of manual search.
[0082] Optionally, the dataset's features include: data type, data volume, data distribution, content features, etc.
[0083] Optionally, feature information of the dataset to be evaluated can be extracted, including: determining feature information by analyzing the basic information, content and distribution of the data to be evaluated.
[0084] Optionally, the basic information of the dataset to be evaluated includes data type (image, text, etc.), data size, and data dimensions. For image datasets, information such as image resolution and color space can also be obtained; for text datasets, information such as average text length and vocabulary size can be obtained.
[0085] Optionally, data preprocessing tools and techniques can be used to analyze the content of the dataset to be evaluated. For example, OpenCV can be used to perform histogram analysis on image datasets to understand statistical characteristics such as brightness and contrast; NLP libraries such as NLTK or spaCy can be used to perform word frequency statistics, keyword extraction, and other analyses on text datasets.
[0086] Optionally, the distribution of the dataset to be evaluated can be analyzed, including: analyzing whether the class distribution of the dataset is balanced. For image classification datasets, the number of images in each class can be counted to calculate the class distribution balance; for text classification datasets, the number and proportion of text in each class can be counted.
[0087] Optionally, the embedding representation of each evaluation rule in the knowledge graph is generated as follows: in the knowledge graph, entity features and relation features related to the evaluation rule are extracted to obtain the initial feature vector of each evaluation rule; based on the graph attention network, each initial feature vector is aggregated with the feature vectors of adjacent entities to generate the embedding representation of each evaluation rule.
[0088] The embedding of the evaluation rule is represented as follows:
[0089]
[0090] B is the embedding representation of the evaluation rule, σ is the non-linear activation function, K is the number of attention heads in the graph attention network, and N is the number of attention heads. i Let α be the set of adjacent entities.ij (k) W represents the attention weight of the k-th attention head. (k) Let h be the weight matrix of the k-th attention head, L be the number of layers in the graph attention network, and h be the weight matrix of the k-th attention head. i (L) Let i be the embedding representation of the Lth layer, i be the index of the evaluation rule, and j be the index of the vector entity.
[0091] In this embodiment, the graph attention network can automatically learn the importance weights of neighboring entities. Through an attention mechanism, it assigns different weights to the neighboring nodes of each entity node, thereby aggregating information more accurately. Assume that each node corresponding to an evaluation rule in the knowledge graph has an initial feature vector h. i (0) Graph attention networks use an attention mechanism to calculate the attention weights between nodes. For each node i and its neighboring node j, the attention weight α is calculated using the following formula. ij :
[0092]
[0093] Where a is a learnable attention vector, W is a weight matrix used to map node features to a new space, and h i (l) and h j (l) Let i and j be the feature vectors of i and j in the l-th layer, respectively. i (l) ||Wh j (l) ] indicates a vector concatenation operation.
[0094] In each layer, the feature vector h of node i i (l+1) It can be updated by aggregating the feature vectors of its neighboring nodes:
[0095]
[0096] To capture richer node relationships and feature information, graph attention networks typically employ multi-head attention mechanisms. Assuming a multi-head attention mechanism, the output of each head can be represented as:
[0097]
[0098] After passing through a multi-layer graph attention network, the embedding representation B for each evaluation can be obtained.
[0099] Optionally, the similarity between the feature vector and the embedding representation of the evaluation rule is calculated according to the following formula:
[0100]
[0101] Where D is the similarity, β is the dynamic adjustment factor, M is the weight vector representing the importance weight of each feature, ⨀ represents element-wise multiplication, A is the feature vector, and B is the embedding representation.
[0102] In this embodiment, a weighting mechanism is introduced during the similarity calculation process to consider the importance of different features. Element-wise multiplication is performed using a weight vector M to assign different weights to different features. To further improve the adaptability of the similarity calculation, a dynamic adjustment factor β is introduced. This factor can dynamically adjust the similarity calculation results based on the feature distribution of the dataset. For example, if the feature distribution of the dataset is relatively uniform, β can be set to a smaller value; if the importance of some features is significantly higher than other features, β can be set to a larger value.
[0103] Optionally, the dynamic adjustment factor β can be determined as follows: Calculate the variance or standard deviation of the dataset features. If the feature distribution is relatively uniform, set β to a smaller value (e.g., 0.2); if the variance of some features is significantly larger, set β to a larger value (e.g., 0.8). Adjust the value of β dynamically based on user feedback on the recommendation results. If users are more satisfied with the weighted similarity results, the value of β can be appropriately increased; conversely, the value of β can be decreased. Use cross-validation to determine the optimal β value through experiments. Divide the dataset into training and validation sets, try different β values, and select the β value that yields the most accurate similarity calculation results on the validation set.
[0104] Optionally, a target evaluation rule is determined based on the matching degree between the candidate evaluation rules and the dataset to be evaluated, including one or more of the following: determining the target evaluation rule based on the matching degree between the evaluation metrics of the dataset to be evaluated and the evaluation metrics covered by the candidate evaluation rules; determining the target evaluation rule based on the matching degree between the dataset type of the dataset to be evaluated and the dataset type to which the candidate evaluation rules apply; determining the user's commonly used evaluation rules for the dataset type based on the user's historical usage records and the dataset type of the dataset to be evaluated, and determining the target evaluation rule based on the matching degree between the commonly used evaluation rules and the candidate evaluation rules.
[0105] In this embodiment, multiple matching strategies can accurately recommend applicable evaluation rules based on the specific characteristics of the dataset to be evaluated and the user's historical behavior, reducing the time and workload of manual searching. By comprehensively considering evaluation metrics, dataset type, and user preferences, the recommended evaluation rules ensure comprehensive coverage of all quality metrics of the dataset. Specifically, the matching strategy based on evaluation metric matching degree requires calculating the matching degree between the metrics to be evaluated in the dataset and the metrics covered by the candidate evaluation rules. For example, if the dataset to be evaluated needs to assess accuracy and completeness, a candidate evaluation rule has a high matching degree if it covers both metrics, and a low matching degree if it covers only one metric. In addition to considering the matching degree of evaluation metrics, the similarity between the dataset type and features to which the candidate evaluation rules are applicable and the dataset to be evaluated is also considered. Using a dataset feature similarity calculation method, combined with the matching degree of evaluation metrics, the most suitable evaluation rule is comprehensively determined. Furthermore, the user's historical usage records can be analyzed to understand the types of evaluation rules that the user frequently uses and prefers. For image datasets, if the user has frequently used image preprocessing rules and annotation accuracy evaluation rules and has given them high ratings, these rules and similar rules are given priority in the recommendation process.
[0106] Optionally, the dataset quality assessment method further includes: after obtaining the target assessment rule, adjusting the target assessment rule according to the data to be assessed; specifically including: obtaining the characteristic information of the dataset to be assessed; the characteristic information includes data distribution, noise level, and data volume; and adjusting the assessment index of the target assessment rule according to the characteristic information of the dataset to be assessed.
[0107] In this embodiment, different datasets may have different characteristics, such as data distribution, noise level, and data volume. Even if the feature vectors of the target evaluation rule and the dataset to be evaluated are highly similar, adjustments may still be necessary based on the specific characteristics of the dataset. Adjustments to the evaluation metrics of the target evaluation rule include adjusting the weights of the evaluation metrics, adjusting the thresholds, and adding or deleting evaluation metrics.
[0108] Optionally, the weighting of evaluation metrics can be adjusted as follows: if the dataset has a class imbalance problem, the weight of recall can be increased; if the dataset has a high noise level, the weight of robustness evaluation metrics can be increased.
[0109] Optionally, the threshold adjustment of the evaluation metrics includes: if the accuracy requirement of the dataset is high, the accuracy threshold can be increased; if the diversity of the dataset is low, the threshold of the diversity evaluation metric can be adjusted.
[0110] Optionally, the addition or deletion of evaluation metrics may include: for image datasets, adding evaluation metrics such as image resolution and color space; for text datasets, adding evaluation metrics such as text length and lexical richness; deleting evaluation metrics that are not applicable to the current dataset or affect the accuracy of the evaluation results; and deleting evaluation metrics that users consider unimportant based on user feedback.
[0111] Optionally, the dataset quality assessment method further includes: comparing the assessment results of the dataset to be assessed with the assessment results of historical datasets using the same objective assessment rule; and displaying the knowledge graph, objective assessment rule, assessment results of the dataset to be assessed, and the comparative analysis results of the dataset to be assessed and historical datasets in a visual interface.
[0112] Combination Figure 4 As shown, this disclosure provides another method for evaluating the quality of datasets based on knowledge graphs, including:
[0113] S401, the processor acquires a knowledge graph for dataset quality assessment.
[0114] S402, the processor determines the target evaluation rules from the knowledge graph based on the feature information of the dataset to be evaluated.
[0115] S403, the processor performs a quality assessment on the dataset to be evaluated according to the target evaluation rules and obtains the evaluation results.
[0116] S404, the processor compares and analyzes the evaluation results of the dataset to be evaluated with the evaluation results of historical datasets using the same objective evaluation rule.
[0117] The S405 processor displays the knowledge graph, target evaluation rules, evaluation results of the dataset to be evaluated, and comparative analysis results between the dataset to be evaluated and historical datasets in a visual interface.
[0118] In this embodiment, comparing the evaluation results of the dataset to be evaluated with those of historical datasets allows for the identification of differences and trends in the evaluation results of different datasets under the same evaluation rules. Through a visual interface, users can intuitively see the comparison between various quality indicators of the dataset to be evaluated and historical datasets, enhancing the interpretability of the evaluation results. Displaying the evaluation result trends of historical datasets helps users identify long-term changes in dataset quality, providing a basis for improvement measures.
[0119] Optionally, comparative analysis methods include: numerical comparison of indicators, distribution comparison, and trend comparison. Numerical comparison of indicators refers to directly comparing the values of the dataset to be evaluated with historical datasets on various quality assessment indicators (such as accuracy and completeness). For example, if the annotation accuracy rate of the dataset to be evaluated is 92%, and the average annotation accuracy rate of historical datasets is 85%, then the dataset to be evaluated can be considered superior to the historical dataset in terms of annotation accuracy. Distribution comparison refers to analyzing the distribution of the dataset to be evaluated and historical datasets, such as whether the category distribution is similar. Statistical tests (such as the chi-square test) can be used to test the significance of the distribution differences. Trend analysis involves arranging the evaluation results of the dataset to be evaluated and the evaluation results of historical datasets in chronological order and analyzing the changing trends of the evaluation indicators. For example, observing whether the annotation accuracy rate of the image dataset gradually increases over time.
[0120] Optionally, the comparative analysis may include the following comparison items: quality assessment metric values, dataset distribution characteristics, and the effectiveness of the evaluation rules. Quality assessment metric values include the values of specific evaluation indicators such as annotation accuracy, completeness score, and consistency score. Dataset distribution characteristics include category distribution ratios and data volume distributions. The effectiveness of the evaluation rules involves comparing the evaluation results of the current dataset with historical datasets under the same or similar evaluation rules. For example, using the same image preprocessing rules, compare the image quality assessment metrics of the current dataset and historical datasets after preprocessing.
[0121] In a specific embodiment, to implement the knowledge graph-based dataset quality assessment method provided in this disclosure, a knowledge graph needs to be constructed first. 1000 dataset quality assessment rules and results accumulated by an enterprise in multiple projects over the past year are collected. The collected knowledge is encoded to generate a standardized knowledge-encoded dataset. Entities, relationships, and attributes are extracted from the encoded data to construct a knowledge graph containing 500 entity nodes and 800 relationship edges. The constructed knowledge graph is stored in the Neo4j graph database, configured with a periodic update task, updating every two weeks. When a user uploads a new image dataset for evaluation, the system analyzes the dataset's characteristics and retrieves candidate evaluation rules related to the image dataset from the knowledge graph. Based on the search results, five most matching target evaluation rules are recommended to the user, including image data preprocessing rules and annotation accuracy evaluation rules. The system automatically performs quality assessment on the image dataset according to the target evaluation rules and displays the knowledge graph, target evaluation rules, evaluation results of the dataset to be evaluated, and comparative analysis results between the dataset to be evaluated and historical datasets in a visual interface. Users can intuitively view different entities and their relationships in the knowledge graph, the dataset to which the target evaluation rules apply, the evaluation metrics, and the reasons for the recommendations, while also quickly understanding the quality of the dataset.
[0122] Combination Figure 5 As shown in the illustration, this disclosure provides a dataset quality assessment device 500 based on a knowledge graph, including a knowledge acquisition module 501, a rule recommendation module 502, and a quality assessment module 503. The knowledge acquisition module 501 is configured to acquire a knowledge graph for dataset quality assessment. The rule recommendation module 502 is configured to determine target assessment rules from the knowledge graph based on the feature information of the dataset to be assessed. The quality assessment module 503 is configured to perform quality assessment on the dataset to be assessed according to the target assessment rules and obtain assessment results.
[0123] The knowledge graph-based dataset quality assessment device provided in this disclosure can determine suitable target assessment rules based on the dataset to be assessed, and perform the most appropriate assessment of various quality indicators of the dataset to be assessed, thereby improving assessment efficiency and accuracy.
[0124] Optionally, the knowledge graph-based dataset quality assessment device 500 further includes a knowledge construction module 504, a storage management module 505, and a visualization module 506. The knowledge construction module 504 is configured to collect evaluation rules and results from historical datasets from data sources; encode the collected data to obtain encoded data; extract entities, relationships, and attributes from the encoded data to form structured data; and construct a knowledge graph based on the structured data. The storage management module 505 is configured to store and manage the knowledge graph constructed by the knowledge construction module 504, and the knowledge acquisition module 501 can acquire the knowledge graph from the storage management module 505. The visualization module 506 is configured to display the knowledge graph acquired by the knowledge acquisition module 501 and the evaluation results output by the quality assessment module 503.
[0125] Combination Figure 6 As shown, this disclosure provides another knowledge graph-based dataset quality assessment device 600, including a processor 601 and a memory 602. Optionally, the device may further include a communication interface 603 and a bus 604. The processor 601, communication interface 603, and memory 602 can communicate with each other via the bus 604. The communication interface 603 can be used for information transmission. The processor 601 can call logical instructions in the memory 602 to execute the knowledge graph-based dataset quality assessment method of the above embodiments.
[0126] Furthermore, the logic instructions in the aforementioned memory 602 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.
[0127] The memory 602, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 601 executes functional applications and data processing by running the program instructions / modules stored in the memory 602, thereby implementing the knowledge graph-based dataset quality assessment method in the above embodiments.
[0128] The memory 602 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 602 may include high-speed random access memory and may also include non-volatile memory.
[0129] This disclosure provides an electronic device, including: an electronic device body, and the aforementioned knowledge graph-based dataset quality assessment device. The knowledge graph-based dataset quality assessment device is installed in the electronic device body. The installation relationship described herein is not limited to placement inside the electronic device, but also includes installation connections with other components of the electronic device, including but not limited to physical connections, electrical connections, or signal transmission connections. Those skilled in the art will understand that the knowledge graph-based dataset quality assessment device can be adapted to feasible electronic device bodies to achieve other feasible embodiments.
[0130] This disclosure provides a computer-readable storage medium storing computer-executable instructions configured to perform the aforementioned knowledge graph-based dataset quality assessment method.
[0131] The technical solutions of this disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this disclosure. The aforementioned storage medium can be a non-transitory storage medium, including: a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, and other media capable of storing program code.
[0132] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the technical solutions described herein. As used in the technical solutions described herein, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used herein refers to any and all possible combinations of one or more of the associated listed elements. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.
[0133] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0134] The methods and products disclosed in the embodiments herein (including but not limited to devices and equipment) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to implement this embodiment according to actual needs. In addition, the functional units in the embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
Claims
1. A dataset quality assessment method based on knowledge graphs, characterized in that, include: Obtain a knowledge graph for dataset quality assessment; construct the knowledge graph as follows: collect assessment rules and results of historical datasets from data sources; The collected data is encoded to obtain encoded data; Entities, relationships, and attributes are extracted from coded data to form structured data; entities include dataset type, evaluation metrics, evaluation methods, data source, and dataset version; relationships include the types of relationships between entities; The attributes include detailed information about the entity; Building knowledge graphs based on structured data; Based on the feature information of the dataset to be evaluated, the target evaluation rule is determined from the knowledge graph. Specifically, this includes: extracting the feature information of the dataset to be evaluated to obtain feature vectors; the feature information includes the data type, data volume, data distribution, and content features of the dataset to be evaluated; determining candidate evaluation rules based on the similarity between the feature vectors and the embedding representation of each evaluation rule in the knowledge graph; and determining the target evaluation rule based on the matching degree between the candidate evaluation rules and the dataset to be evaluated. The quality of the dataset to be evaluated is assessed according to the target evaluation rules, and the evaluation results are obtained.
2. The dataset quality assessment method according to claim 1, characterized in that, The embedding representation of each evaluation rule in the knowledge graph is generated as follows: In the knowledge graph, entity features and relation features related to the evaluation rules are extracted to obtain the initial feature vector for each evaluation rule; Based on graph attention networks, each initial feature vector is aggregated with the feature vectors of neighboring entities to generate an embedded representation of each evaluation rule; The embedding of the evaluation rule is represented as follows: B is the embedding representation of the evaluation rule, σ is the non-linear activation function, K is the number of attention heads in the graph attention network, and N is the number of attention heads. i Let α be the set of adjacent entities. ij (k) W represents the attention weight of the k-th attention head. (k) Let h be the weight matrix of the k-th attention head, L be the number of layers in the graph attention network, and h be the weight matrix of the k-th attention head. i (L) Let i be the embedding representation of the Lth layer, i be the index of the evaluation rule, and j be the index of the vector entity.
3. The dataset quality assessment method according to claim 1, characterized in that, The similarity between the feature vector and the embedding representation of the evaluation rule is calculated using the following formula: Where D represents the similarity, β is the dynamic adjustment factor, and M is the weight vector, representing the importance weight of each feature. Let A represent element-wise multiplication, where A is the feature vector and B is the embedding representation.
4. The dataset quality assessment method according to claim 1, characterized in that, Based on the matching degree between the candidate evaluation rules and the dataset to be evaluated, the target evaluation rule is determined, including one or more of the following: The target evaluation rule is determined based on the degree of matching between the evaluation metrics of the dataset to be evaluated and the evaluation metrics covered by the candidate evaluation rules. The target evaluation rule is determined based on the degree of matching between the dataset type of the dataset to be evaluated and the dataset type to which the candidate evaluation rule applies. Based on the user's historical usage records and the dataset type of the dataset to be evaluated, determine the user's commonly used evaluation rules for that dataset type, and determine the target evaluation rule based on the matching degree between the commonly used evaluation rules and the candidate evaluation rules.
5. The dataset quality assessment method according to any one of claims 1 to 4, characterized in that, Also includes: After obtaining the target evaluation rules, adjust the target evaluation rules according to the data to be evaluated; Specifically, it includes: Obtain characteristic information of the dataset to be evaluated; characteristic information includes data distribution, noise level, and data volume; Based on the characteristics of the dataset to be evaluated, adjust the evaluation metrics of the target evaluation rules.
6. The dataset quality assessment method according to any one of claims 1 to 4, characterized in that, Also includes: The evaluation results of the dataset to be evaluated are compared and analyzed with the evaluation results of historical datasets using the same objective evaluation rule; The knowledge graph, target evaluation rules, evaluation results of the dataset to be evaluated, and comparative analysis results between the dataset to be evaluated and historical datasets are displayed in a visual interface.
7. A dataset quality assessment device based on knowledge graphs, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute the knowledge graph-based dataset quality assessment method as described in any one of claims 1 to 6 when running the program instructions.
8. An electronic device, characterized in that, include: The electronic device itself; The knowledge graph-based dataset quality assessment device as described in claim 7 is installed on the electronic device body.
Citation Information
Patent Citations
Service data quality evaluation method for data exchange
CN114971140A
Credit assessment method and system based on knowledge graph and deep learning
CN115358754A