Data value density evaluation method and device
By using multi-dimensional evaluation metrics and Gaussian kernel function to calculate metric weights, combined with data distillation technology, the problem of inaccurate dataset value density assessment in existing technologies is solved, achieving a comprehensive assessment of dataset value density and improving task applicability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INST OF AUTOMATION CHINESE ACAD OF SCI
- Filing Date
- 2025-06-10
- Publication Date
- 2026-05-05
AI Technical Summary
Existing methods for evaluating the value density of datasets cannot comprehensively and accurately assess the value density of datasets, especially when considering the adaptability of the dataset itself to specific task scenarios.
Multiple evaluation metrics (such as completeness, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain bias, leakage ratio, and prediction accuracy) are used for evaluation. The weights of the metrics are calculated based on the Gaussian kernel function, and the value density of the dataset in the target task scenario is determined by combining data distillation techniques.
It enables a comprehensive assessment of the value density of datasets, improving accuracy and applicability in real-world task scenarios and ensuring that the assessment results better meet actual needs.
Smart Images

Figure CN120910495B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for evaluating data value density. Background Technology
[0002] Dataset value density is used to quantitatively measure the amount of effective knowledge carried per unit volume of data in a dataset and its adaptability to a specific task. Currently, methods for evaluating dataset value density mainly include: static data quality checks, single performance metrics, and random or simple policy compression. Static data quality checks primarily focus on assessing data completeness and repeatability, failing to comprehensively evaluate dataset value density. Single performance metrics assess dataset value density by training and testing the dataset on a specific (mature, classic) deep learning model and then evaluating the test results using a single metric (e.g., recall). However, a single metric is insufficient to characterize the multifaceted features of a dataset, easily leading to underestimation or overestimation of the data's potential information. Random or simple policy compression only evaluates a subset of the dataset, also failing to comprehensively assess dataset value density. Furthermore, data value density depends not only on the static quality of the data itself but also on its performance in a specific task scenario. Existing evaluation methods only consider the dataset itself and cannot accurately assess its value density.
[0003] In summary, existing methods for evaluating the value density of datasets only address the dataset itself and cannot comprehensively and accurately assess the value density of a dataset. Summary of the Invention
[0004] This invention provides a data value density assessment method and apparatus to solve the problem that existing technologies cannot comprehensively and accurately assess the value density of datasets.
[0005] This invention provides a method for evaluating data value density, comprising the following steps:
[0006] Determine the values of multiple evaluation metrics for the dataset to be evaluated;
[0007] Based on the values of each evaluation indicator, determine the weight of each evaluation indicator relative to the target task scenario.
[0008] Based on the index values and corresponding index weights of each evaluation index, the value density of the dataset to be evaluated relative to the target task scenario is assessed.
[0009] According to a data value density assessment method provided by the present invention, based on the index values of each assessment index, the index weight of each assessment index relative to the target task scenario is determined, including:
[0010] Based on the index values of each evaluation indicator and the default values of each evaluation indicator for the target task scenario, the task fit of each evaluation indicator is determined.
[0011] Based on the task fit of each evaluation indicator, the indicator weight of each evaluation indicator relative to the target task scenario is determined.
[0012] According to a data value density assessment method provided by the present invention, the task fit of each assessment indicator is determined based on the indicator values of each assessment indicator and the default values of each assessment indicator in the target task scenario. This includes: calculating the task fit of each assessment indicator using the following Gaussian kernel function. T Di :
[0013] ;
[0014] in, α i Indicates the first i The values of each evaluation indicator, T i Indicates the target task scenario for the first i The default values for each evaluation indicator This represents the bandwidth parameter of the Gaussian kernel. This represents an exponential function with the natural constant e as the base.
[0015] According to a data value density assessment method provided by the present invention, the index values of multiple assessment indicators for the dataset to be assessed are determined, including:
[0016] Data distillation is performed on the dataset to be evaluated to obtain a density ratio index between the distilled data subset and the dataset to be evaluated. The density ratio index represents the ratio of the average data volume of the distilled data subset divided by category to the average data volume of the dataset to be evaluated divided by category.
[0017] Calculate the completeness rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy of the dataset to be evaluated.
[0018] The multiple evaluation metrics include at least two of the following: density ratio, integrity rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy.
[0019] According to a data value density assessment method provided by the present invention, the intraclass diversity index is calculated according to the following formula:
[0020] ;
[0021] in, x i Var( represents the sample feature of the i-th sample of any class in the dataset to be evaluated.) x i () indicates the calculation of the variance of sample features for any of the categories. n This indicates the number of samples in any of the categories.
[0022] The inter-class diversity index is calculated using the following formula:
[0023] ;
[0024] in, c j Indicates the first element in the dataset to be evaluated. j The category center features of each category, Var( c j ) represents the calculation of the inner variance of the category center feature, wherein the first... j The category center feature of the first category is based on the first category. j The sample characteristics of all samples in each category are determined. K Indicates the number of categories.
[0025] According to a data value density assessment method provided by the present invention, the ambiguity index is calculated according to the following formula:
[0026] ;
[0027] in, x i Indicates the first element in the dataset to be evaluated. i Sample characteristics of each sample This indicates the preset test model pair. x i Predicted as the first k The probability of a class K Indicates the number of categories.
[0028] According to a data value density assessment method provided by the present invention, the leakage ratio index is calculated according to the following formula:
[0029] ;
[0030] in, m This represents the total number of samples in the test set within the dataset to be evaluated. n This represents the total number of samples in the training set of the dataset to be evaluated. x i Indicates the first test seti Sample characteristics of each sample y j Indicates the first training set j The sample features of each sample, cos( i ( x i , y j )) represents sample features x i and sample features y j cosine similarity, e Indicates the similarity threshold. i ( x i , y j ) represents the angle between two sample features, and II() represents the indicator function.
[0031] The present invention also provides a data value density assessment device, comprising the following modules:
[0032] The indicator value determination module is used to determine the indicator values of multiple evaluation indicators for the dataset to be evaluated.
[0033] The indicator weight determination module is used to determine the indicator weight of each evaluation indicator relative to the target task scenario based on the indicator value of each evaluation indicator.
[0034] The value density assessment module is used to assess the value density of the dataset to be assessed relative to the target task scenario based on the index values and corresponding index weights of each assessment index.
[0035] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the program to implement the data value density assessment method as described above.
[0036] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data value density assessment method as described above.
[0037] The data value density assessment method and apparatus provided by this invention evaluates the value density of the dataset to be evaluated using the values of multiple evaluation indicators. Based on the values of each evaluation indicator, the weight of each evaluation indicator relative to the target task scenario is determined. Based on the values of each evaluation indicator and their corresponding weights, the value density of the dataset to be evaluated relative to the target task scenario is assessed. Because multiple dimensions of evaluation indicators are used, and corresponding weights are determined for each evaluation indicator for the target task scenario, a comprehensive assessment of the value density of the dataset to be evaluated is achieved, effectively improving the accuracy and applicability of the value density of the dataset to be evaluated in real task scenarios. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0039] Figure 1 This is a flowchart illustrating the data value density assessment method provided by the present invention.
[0040] Figure 2 This is a schematic diagram of the data value density assessment device provided by the present invention.
[0041] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0043] The data value density assessment method of this invention, such as... Figure 1 As shown, the procedure includes steps S110 to S130.
[0044] Step S110: Determine the values of multiple evaluation metrics for the dataset to be evaluated, so as to assess the value density of the dataset from multiple evaluation dimensions and to more comprehensively evaluate the value density of the dataset. Multiple evaluation metrics may include: completeness rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy.
[0045] Step S120: Based on the values of each evaluation metric, determine the weight of each metric relative to the target task scenario. Since different datasets to be evaluated contain different data, for example, dataset A for classification includes multiple images of animals such as cats and dogs, primarily used for classification training, while dataset B for object recognition, such as face recognition, includes multiple images of faces, primarily used for object recognition. These two datasets obviously have different value densities for different training tasks. Dataset A has a higher value density for classification tasks (i.e., for classification training), while dataset B has a higher value density for object recognition tasks (i.e., for object recognition training). Therefore, to more accurately evaluate the value density of the datasets and their applicability to the task scenario, it is necessary to determine the weight of each evaluation metric relative to the target task scenario.
[0046] Step S130: Based on the index values and corresponding index weights of each evaluation index, evaluate the value density of the dataset to be evaluated relative to the target task scenario. Because the index weights of the dataset to be evaluated relative to the target task scenario are considered during the evaluation process, the accuracy and applicability of the value density of the dataset to be evaluated in real task scenarios are effectively improved.
[0047] In the data value density assessment method of this embodiment, the value density of the dataset to be assessed is evaluated using the values of multiple assessment indicators. Based on the values of each assessment indicator, the indicator weight of each assessment indicator relative to the target task scenario is determined. Based on the values of each assessment indicator and the corresponding indicator weights, the value density of the dataset to be assessed relative to the target task scenario is evaluated. Because multiple dimensions of assessment indicators are used and corresponding indicator weights are determined for the assessment indicators for the target task scenario, a comprehensive assessment of the value density of the dataset to be assessed is achieved, and the accuracy and applicability of the value density of the dataset to be assessed in real task scenarios are effectively improved.
[0048] In some embodiments, step S120 specifically includes:
[0049] Based on the values of each evaluation metric and the default values of each evaluation metric in the target task scenario, the task fit of each evaluation metric is determined. Task fit aims to reflect the degree of matching between the data characteristics in the dataset to be evaluated and the target task scenario. Since each evaluation metric is related to the data characteristics in the dataset, the relationship between the dataset to be evaluated and the target task scenario can be quantified by the metric values, thereby determining the task fit of each evaluation metric. The default values of each evaluation metric in the target task scenario are pre-set.
[0050] Based on the task fit of each evaluation indicator, the indicator weight of each evaluation indicator relative to the target task scenario is determined. The higher the task fit, the greater the corresponding indicator weight.
[0051] Specifically, for each evaluation metric, the metric weight is obtained by continuously adjusting the initial weight based on task fit. This weight adjustment is a method of modifying the metric weights of each evaluation metric according to the requirements of the target task scenario, ensuring that the metric weights of each dimension of the evaluation dataset better meet the needs of the target task scenario. In this embodiment, the metric weight update mechanism is driven by task fit, with the initial weight... Subsequent approval Iterative updates are performed in rounds. For the first i The evaluation indicators in the first t The weights of the indicators after the round of evaluation For the first i The rate of change of the weights of each evaluation indicator, based on task fit. Segmented control, satisfying:
[0052] .
[0053] The weight update mechanism of this indicator supports multiple rounds of iterative evaluation, provided that the task fit remains unchanged. This will continue to grow. Therefore, this paper sets the maximum number of evaluation rounds to 5. A convergence criterion is also introduced: if the difference between the weights of two consecutive indicators is lower than a preset difference threshold (e.g., 0.01), the iteration terminates. Furthermore, different task fits... By varying the growth rates, the weights of indicators corresponding to different levels of task relevance are differentiated, thereby ensuring the stability and efficiency of indicator weight adjustments and avoiding extreme inflation or meaningless duplicate evaluations.
[0054] In this embodiment, the task fit of each evaluation indicator is determined by the indicator value of each evaluation indicator and the default value of each evaluation indicator in the target task scenario. Then, the indicator weights are adjusted in multiple rounds based on the task fit to obtain more accurate indicator weights.
[0055] In some embodiments, the task fit of each evaluation indicator is determined based on the indicator values of each evaluation indicator and the default values of each evaluation indicator for the target task scenario, including: calculating the task fit of each evaluation indicator according to the following Gaussian kernel function. T Di :
[0056] ;
[0057] in, α i Indicates the first i The values of each evaluation indicator, T i Indicates the target task scenario for the first i The default values for each evaluation indicator This represents the bandwidth parameter of the Gaussian kernel. This indicates an exponential function with the natural constant e as the base. (Default value) T i These are pre-defined values. The default values for different evaluation metrics vary for the same target task scenario. These values can be set according to the actual situation. For example, the default value for the inter-class diversity metric in the classification task scenario is greater than that in the object detection task scenario. In this way, for a dataset suitable for classification, the difference between the calculated inter-class diversity metric value and the default value of the inter-class diversity metric in the classification task scenario is small. Substituting this into the task fit calculation formula above, the inter-class diversity metric has a greater task fit with the classification task scenario. On the other hand, the difference between the calculated inter-class diversity metric value and the default value of the inter-class diversity metric in the object detection task scenario is larger, and the inter-class diversity metric has a smaller task fit with the object detection task scenario.
[0058] For example, the default values for each evaluation metric for different target task scenarios can be as follows:
[0059] Density ratio index: 0.5 can be used for both classification task scenarios and object detection task scenarios.
[0060] Data integrity rate and label consistency rate: 0.98 is suitable for classification tasks, and 0.96 is suitable for object detection tasks.
[0061] Noise detection index: 0.1 for classification tasks and 0.2 for target detection tasks.
[0062] Intra-class diversity index: 0.6 is suitable for classification tasks, and 0.75 is suitable for object detection tasks.
[0063] Inter-class diversity index: 0.9 is suitable for classification tasks, and 0.7 is suitable for object detection tasks.
[0064] Data complexity metrics: In classification tasks, the mean can be 0.6, the variance can be 0.65, and the entropy can be 0.7; in object detection tasks, the mean can be 0.85, the variance can be 0.8, and the entropy can be 0.85.
[0065] Ambiguity index: 0.3 is acceptable for classification tasks, and 0.45 is acceptable for object detection tasks.
[0066] Domain offset metric: 0.2 for classification tasks and 0.35 for object detection tasks.
[0067] Leakage ratio: 0.1 for classification tasks and 0.2 for object detection tasks.
[0068] Prediction accuracy metrics: 0.9 is suitable for classification tasks, and 0.8 is suitable for object detection tasks.
[0069] In some embodiments, in step S130, the value density Final Score of the dataset to be evaluated relative to the target task scenario is calculated according to the following formula.
[0070] .
[0071] Among them, dq[ i ] indicates the first i The indicator values of each evaluation indicator, aw[ i ] indicates the first i The weight of each evaluation metric relative to the target task scenario.
[0072] Furthermore, when the final score of the dataset to be evaluated is low relative to the value density of the target task scenario, the dataset can be optimized. For example, data augmentation, noise reduction, and label integrity checks can be performed to make the evaluation metrics more accurate, resulting in a higher evaluation score. For some standard and mature datasets (datasets that have undergone extensive model training and validation in target task scenarios and consistently yield good training results), a low final score indicates that the default values for some evaluation metrics in the target task scenario are unreasonable. Adjusting these default values will adjust the weights of some evaluation metrics, ultimately leading to a higher final score relative to the value density of the target task scenario.
[0073] In some embodiments, determining the values of multiple evaluation metrics for the dataset to be evaluated includes:
[0074] Data distillation is performed on the dataset to be evaluated to obtain a density ratio index between the distilled data subset and the dataset to be evaluated. The density ratio index represents the ratio of the average data volume of the distilled data subset divided by category to the average data volume of the dataset to be evaluated divided by category.
[0075] Calculate the following metrics for the dataset to be evaluated: completeness rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy.
[0076] The evaluation metrics can include at least two of the following: density ratio, completeness rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy. This allows for a more accurate assessment of the data value density of the dataset to be evaluated using multi-dimensional evaluation metrics.
[0077] In some embodiments, to effectively remove a large number of redundant and low-value samples from the dataset to be evaluated and improve the knowledge density of the core samples in the dataset, this embodiment introduces a data distillation algorithm for knowledge condensation and density quantization. Specifically, let the dataset to be evaluated be represented as... The sample distribution is initially analyzed using feature extraction and dimensionality reduction methods (such as Principal Component Analysis (PCA) or t-SNE). Subsequently, gradient matching, Difficulty Alignment Trajectory Matching (DATM) algorithm, or K-means clustering are used to select more representative core samples to form the distillation dataset. The specific formula for the distillation algorithm is as follows:
[0078] .
[0079] Where D′ is a feasible set of candidate subsets of the dataset D to be evaluated (D′ can be randomly selected from D, selected by average of categories, or selected from D′ with large entropy value), and arg min is a function that acts on a variable D′ to find the D′ with the minimum loss of the model on D, that is, the training effect of the model based on D′ and based on D is basically the same. To evaluate the expected value of samples on dataset D, This is the loss function. It is based on the data subset after distillation The trained model, on the input data x The predicted output, y Indicates input data x The corresponding tags.
[0080] To measure the effectiveness of distillation, the number of data points retained in each category within the distilled data subset is defined as IPC1 = |D distill | / K, where |D distill | represents the amount of data in the distilled subset, and K is the number of classes. IPC1 reflects the degree to which data distillation preserves data from different classes; a higher value indicates that more samples are retained on average for each class. IPC2 = |D| / K, where |D| represents the amount of data in the dataset to be evaluated, and the density ratio of the distilled subset to the dataset to be evaluated is defined as IPC1 / IPC2.
[0081] In some embodiments, the completeness of the dataset directly affects the training stability and prediction accuracy of the model. If a large number of missing or corrupted samples exist in the dataset, the model training process will be greatly affected, leading to problems such as overfitting or underfitting. Therefore, a completeness ratio is defined as the Complete Ratio, which measures the proportion of valid samples in the dataset. Specifically, it is defined as the proportion of valid samples to the total number of samples. The formula for calculating Complete Ratio is as follows:
[0082] .
[0083] Wherein, Valid Samples represents the number of valid samples in the dataset to be evaluated that are neither missing nor corrupted, and TotalSamples represents the total number of samples in the dataset to be evaluated.
[0084] In some embodiments, the accuracy and consistency of labels in a dataset directly affect the training effect and final performance of the model. Especially in supervised learning, datasets with low label consistency can lead to bias between the training and test sets, thus affecting the model's generalization ability. Therefore, a label consistency metric, Label Consistency, is defined, representing the proportion of samples with incorrect labels to the total number of samples. The formula for calculating Label Consistency is as follows:
[0085] .
[0086] in, Indicates the first i If a sample is incorrectly labeled, it is recorded as 1; if it is correctly labeled, it is recorded as 0. N This represents the total number of samples.
[0087] In some embodiments, noise refers to outlier samples in the dataset, which are typically generated by incorrect labeling, data input problems, or anomalies in the samples themselves. Noisy samples not only affect the model training process but may also cause the model to produce misleading results, thereby reducing the model's prediction accuracy. To effectively detect noisy samples, this embodiment combines dimensionality reduction techniques (e.g., t-distributed stochastic neighbor embedding, t-SNE) and clustering algorithms (e.g., density-based spatial clustering of applications with noise, DBSCAN) for noise sample identification.
[0088] Specifically, t-SNE is used to reduce high-dimensional data to two or three dimensions, visually demonstrating the distribution of data points. Noise samples are typically scattered and far from other sample categories. The DBSCAN clustering algorithm is used to identify low-density regions, which usually contain noise samples.
[0089] For example: The dataset to be evaluated has a total of N There are n samples, where the feature representation of each data point is as follows: After standardization and t-SNE dimensionality reduction, a low-dimensional feature representation is obtained. Then, the cluster centers and distances are calculated, and the mean vector of all the reduced-dimensional features is... , No. i The Euclidean distance from each sample to the cluster center is Generate a threshold and determine noise: Calculate the mean for all distances. and standard deviation Construct a threshold, and the mean vector of all dimensionality-reduced features is:
[0090] .
[0091] Finally, a threshold is selected. k (For example: when k=3) i k For each sample Calculate the Euclidean distance from the cluster center. If satisfied If a sample is deemed noise, it is considered a noise sample. This embodiment can effectively detect outliers in the data and distinguish between noise samples and valid samples. Finally, the ratio of the number of noise samples to the total number of samples is calculated as the noise detection index.
[0092] In some embodiments, intra-class diversity is used to measure the differences between samples of the same class. Moderate intra-class diversity helps improve the model's generalization ability, enabling the model to better learn the diverse features of the class. Therefore, the intra-class diversity metric Class Intra-Diversity is defined, specifically calculated using intra-class variance:
[0093] .
[0094] in, x i Var( represents the sample feature of the i-th sample of any class in the dataset to be evaluated.) x i () indicates the calculation of the variance of sample features for any of the categories. n This indicates the number of samples in any of the categories.
[0095] Inter-class diversity measures the discriminative power between different classes, reflecting the dataset's ability to differentiate between categories. Datasets with high inter-class diversity help models better distinguish between classes, improving their adaptability to complex classification tasks; conversely, low inter-class diversity can lead to blurred boundaries between classes, affecting the model's classification accuracy. Therefore, we define the inter-class diversity metric Class Inter-Diversity, specifically calculated using inter-class variance:
[0096] .
[0097] in, c j Indicates the first element in the dataset to be evaluated. j The category center features of each category, Var( c j ) represents the calculation of the inner variance of the category center feature, wherein the first... j The category center feature of the first category is based on the first category. j The sample characteristics of all samples in each category are determined. K Indicates the number of categories.
[0098] In this embodiment, by introducing intra-class variance and inter-class variance as core indicators, the internal differences between different categories in the dataset and the discriminative power between categories are effectively distinguished, thereby improving the granularity and discriminative power of the dataset feature representation.
[0099] In some embodiments, the data statistical complexity metric is used to measure the overall information content and diversity of a dataset. A high data statistical complexity metric value indicates that it contains more potential information and can provide richer features for model learning. In this embodiment, the mean value is typically used. m Standard deviation s The statistical complexity of a dataset is evaluated using entropy H(data). Data statistical complexity metrics include the mean, variance, and entropy of the dataset. The entropy value of a dataset reflects the diversity and information content of the data; the higher the entropy value, the more complex the data. Therefore, the mean is defined separately. m Standard deviation s Entropy H(data) is used as an indicator of data statistical complexity. The specific formulas for the mean and variance of the dataset to be evaluated are as follows:
[0100] .
[0101] x i This represents the sample feature of the i-th sample in the dataset to be evaluated.
[0102] The entropy H(data) of the dataset to be evaluated can be measured by the distribution uncertainty, as shown in the following formula:
[0103] .
[0104] in, g i This represents the probability of element i in the data. For example, when the data is an image, g i This represents the pixel value of the i-th pixel in the image. When the data is text data, g i This represents the i-th text content (word or character) in the text. For image data, in order to calculate... First, the image is converted to grayscale, and then a histogram of the grayscale values is plotted.
[0105] .
[0106] In this case, the values of i and j are both [0, 255], but they are different. i is the gray level of the probability being calculated, and j is the index variable when summing all gray levels. Pixel value g i The frequency is calculated using the total number of all pixels in the denominator. For the grayscale value range, [0, 255] is used for calculation.
[0107] In some embodiments, ambiguity refers to the degree of ambiguity of a sample at the class boundary. Higher ambiguity increases the risk of model misjudgment and misleads the decision boundary. Therefore, in this embodiment, the model is tested to predict the samples and output results. The entropy of the output results is then calculated to quantify the ambiguity of the samples at the class boundary; this entropy is the ambiguity index, and the specific calculation formula is as follows:
[0108] ;
[0109] in, x i Indicates the first element in the dataset to be evaluated. i Sample characteristics of each sample This indicates the preset test model pair. x i Predicted as the first k The probability of a class K Indicates the number of categories.
[0110] Preferably, multiple test models can be used to test the first... i Different prediction results are obtained by predicting the sample features of each sample, and multiple corresponding entropies are obtained. The average of the multiple entropies is taken to obtain the first entropy. i The entropy value of each sample, i.e., the ambiguity index. Entropy obtained from the test results of multiple test models better reflects the ambiguity of the samples; that is, the ambiguity index value is more accurate. It should be noted that each sample feature... x i Each of these corresponds to an entropy value. Therefore, the ambiguity index of the dataset to be evaluated is a vector composed of the entropy values of all sample features. The index value of the ambiguity index is defined as the sum of the values of each element in the vector.
[0111] In some embodiments, domain offset is used to measure the degree of difference in the data distribution of the training and test sets in the feature space. If the distributions of the training and test sets are inconsistent, the model's performance in practical applications will significantly decrease, especially in transfer learning or cross-domain tasks, where domain offset is a key factor leading to unsatisfactory results. Therefore, a domain offset metric is defined to measure the degree of difference in the data distribution of the training and test sets in the feature space of the dataset to be evaluated. Specifically, the Jensen-Shannon divergence metric JSD is used to characterize the domain offset metric.
[0112] .
[0113] in, P and Q Let represent the data distributions of the training and test sets in the feature space, respectively. The average distribution of the training and test sets. The Kullback-Leibler divergence (KL divergence) measures the difference between a distribution P and the mean distribution M, and is specifically defined as follows:
[0114] .
[0115] in, x This represents the index of each dimension of all feature vectors in the sample space after flattening. D KL ( Q || M )and D KL ( P || M The meaning of ) is similar, and will not be repeated here.
[0116] In some embodiments, a high degree of repetition or similarity between test samples and training set samples means that the test set contains a large number of samples that are identical or extremely similar to the training set samples, leading to a decrease in the objectivity and generalization of the model's test results. Therefore, a leakage ratio metric is defined to characterize the repetition or similarity between test samples and training set samples. Specifically, the formula for calculating the leakage ratio metric is as follows:
[0117] .
[0118] in, m This represents the total number of samples in the test set within the dataset to be evaluated. n This represents the total number of samples in the training set of the dataset to be evaluated. x i Indicates the first test set i Sample characteristics of each sample y j Indicates the first training set j The sample features of each sample, cos( i ( x i , y j )) represents sample features x i and sample features y j cosine similarity, e Indicates the similarity threshold. i ( x i , y j ) represents the angle between two sample features, and II() represents the indicator function.
[0119] In some embodiments, if a dataset achieves high prediction accuracy on several test models (mature classic models such as ResNet, VGG, etc.), it usually indicates that it has strong feature separability and high data quality, making it suitable for more complex tasks. Therefore, a prediction accuracy metric is defined to evaluate the feature representation ability of a dataset. The prediction accuracy metric represents the average prediction accuracy of the dataset under evaluation on multiple test models. Specifically, the formula for calculating prediction accuracy is as follows:
[0120] .
[0121] Correct Predictions represents the number of correct predictions, Total Predictions represents the total number of predictions, and the prediction accuracy metric is the average of the accuracy values across multiple test models.
[0122] The data value density assessment device provided by the present invention is described below. The data value density assessment device described below can be referred to in correspondence with the data value density assessment method described above.
[0123] The data value density assessment device of this invention, such as Figure 2 As shown, it includes:
[0124] The indicator value determination module 210 is used to determine the indicator values of multiple evaluation indicators for the dataset to be evaluated.
[0125] The indicator weight determination module 220 is used to determine the indicator weight of each evaluation indicator relative to the target task scenario based on the indicator value of each evaluation indicator.
[0126] The value density assessment module 230 is used to assess the value density of the dataset to be assessed relative to the target task scenario based on the index values and corresponding index weights of each assessment index.
[0127] The data value density assessment device provided by this invention evaluates the value density of the dataset to be evaluated using the values of multiple evaluation indicators. Based on the values of each evaluation indicator, it determines the indicator weight of each evaluation indicator relative to the target task scenario. Based on the indicator values and corresponding indicator weights of each evaluation indicator, it evaluates the value density of the dataset to be evaluated relative to the target task scenario. Because it uses multiple dimensions of evaluation indicators and determines the corresponding indicator weights for the evaluation indicators for the target task scenario, it achieves a comprehensive evaluation of the value density of the dataset to be evaluated and effectively improves the accuracy and applicability of the value density of the dataset to be evaluated in real task scenarios.
[0128] In some embodiments, the indicator weight determination module 220 is specifically used to determine the task fit of each evaluation indicator based on the indicator value of each evaluation indicator and the default value of each evaluation indicator for the target task scenario; and to determine the indicator weight of each evaluation indicator relative to the target task scenario based on the task fit of each evaluation indicator.
[0129] In some embodiments, the indicator weight determination module 220 is specifically used to calculate the task fit of each evaluation indicator according to the following Gaussian kernel function. T Di :
[0130] .
[0131] in, α i Indicates the first i The values of each evaluation indicator, T i Indicates the target task scenario for the first i The default values for each evaluation indicator This represents the bandwidth parameter of the Gaussian kernel. This represents an exponential function with the natural constant e as the base.
[0132] In some embodiments, the indicator value determination module 210 is specifically used for:
[0133] Data distillation is performed on the dataset to be evaluated to obtain a density ratio index between the distilled data subset and the dataset to be evaluated. The density ratio index represents the ratio of the average data volume of the distilled data subset divided by category to the average data volume of the dataset to be evaluated divided by category.
[0134] Calculate the completeness rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy of the dataset to be evaluated.
[0135] The multiple evaluation metrics include at least two of the following: density ratio, integrity rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy.
[0136] In some embodiments, the intra-class diversity index is calculated using the following formula:
[0137] .
[0138] in, x iVar( represents the sample feature of the i-th sample of any class in the dataset to be evaluated.) x i () indicates the calculation of the variance of sample features for any of the categories. n This indicates the number of samples in any of the categories.
[0139] The inter-class diversity index is calculated using the following formula:
[0140] .
[0141] in, c j Indicates the first element in the dataset to be evaluated. j The category center features of each category, Var( c j ) represents the calculation of the inner variance of the category center feature, wherein the first... j The category center feature of the first category is based on the first category. j The sample characteristics of all samples in each category are determined. K Indicates the number of categories.
[0142] In some embodiments, the ambiguity index is calculated using the following formula:
[0143] .
[0144] in, x i Indicates the first element in the dataset to be evaluated. i Sample characteristics of each sample This indicates the preset test model pair. x i Predicted as the first k The probability of a class K Indicates the number of categories.
[0145] In some embodiments, the leakage ratio is calculated using the following formula:
[0146] .
[0147] in, m This represents the total number of samples in the test set within the dataset to be evaluated. n This represents the total number of samples in the training set of the dataset to be evaluated. x i Indicates the first test set i Sample characteristics of each sample y j Indicates the first training set j The sample features of each sample, cos( i ( xi , y j )) represents sample features x i and sample features y j cosine similarity, e Indicates the similarity threshold. i ( x i , y j ) represents the angle between two sample features, and II() represents the indicator function.
[0148] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3 As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communication bus 340, wherein the processor 310, the communications interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 can invoke logical instructions in the memory 330 to execute a data value density assessment method, which includes:
[0149] Determine the values of multiple evaluation metrics for the dataset to be evaluated.
[0150] Based on the values of each evaluation indicator, the weight of each evaluation indicator relative to the target task scenario is determined.
[0151] Based on the index values and corresponding index weights of each evaluation index, the value density of the dataset to be evaluated relative to the target task scenario is assessed.
[0152] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0153] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is capable of executing the data value density assessment method provided by the above methods, the method comprising:
[0154] Determine the values of multiple evaluation metrics for the dataset to be evaluated.
[0155] Based on the values of each evaluation indicator, the weight of each evaluation indicator relative to the target task scenario is determined.
[0156] Based on the index values and corresponding index weights of each evaluation index, the value density of the dataset to be evaluated relative to the target task scenario is assessed.
[0157] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the data value density assessment method provided by the methods described above, the method comprising:
[0158] Determine the values of multiple evaluation metrics for the dataset to be evaluated.
[0159] Based on the values of each evaluation indicator, the weight of each evaluation indicator relative to the target task scenario is determined.
[0160] Based on the index values and corresponding index weights of each evaluation index, the value density of the dataset to be evaluated relative to the target task scenario is assessed.
[0161] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0162] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for assessing data value density, characterized in that, include: Determine the values of multiple evaluation metrics for the dataset to be evaluated. The multiple evaluation metrics include at least two of the following: density ratio, completeness rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy. The dataset to be evaluated includes image datasets and text datasets. Based on the values of each evaluation indicator, the weight of each evaluation indicator relative to the target task scenario is determined. The target task scenario includes: classification task scenario and target recognition task scenario. The dataset to be evaluated has different value densities for different target task scenarios. Based on the index values and corresponding index weights of each evaluation index, the value density of the dataset to be evaluated relative to the target task scenario is assessed. Based on the values of each evaluation indicator, determine the weight of each evaluation indicator relative to the target task scenario, including: Based on the index values of each evaluation indicator and the default values of each evaluation indicator for the target task scenario, the task fit of each evaluation indicator is determined. Based on the task fit of each evaluation indicator, the indicator weight of each evaluation indicator relative to the target task scenario is determined. Specifically, based on the index values of each evaluation indicator and the default values of each evaluation indicator for the target task scenario, the task fit of each evaluation indicator is determined, including: calculating the task fit of each evaluation indicator using the following Gaussian kernel function. T Di : ; in, α i Indicates the first i The values of each evaluation indicator, T i Indicates the target task scenario for the first i The default values for each evaluation indicator This represents the bandwidth parameter of the Gaussian kernel. This represents an exponential function with the natural constant e as the base.
2. The data value density assessment method according to claim 1, characterized in that, Determine the values of multiple evaluation metrics for the dataset to be evaluated, including: Data distillation is performed on the dataset to be evaluated to obtain a density ratio index between the distilled data subset and the dataset to be evaluated. The density ratio index represents the ratio of the average data volume of the distilled data subset divided by category to the average data volume of the dataset to be evaluated divided by category. Calculate the following metrics for the dataset to be evaluated: completeness rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy.
3. The data value density assessment method according to claim 2, characterized in that, The intra-class diversity index is calculated using the following formula: ; in, x i Var( represents the sample feature of the i-th sample of any class in the dataset to be evaluated.) x i () indicates the calculation of the variance of sample features for any of the categories. n This indicates the number of samples in any of the categories. The inter-class diversity index is calculated using the following formula: ; in, c j Indicates the first element in the dataset to be evaluated. j The category center features of each category, Var( c j ) represents the calculation of the inner variance of the category center feature, wherein the first... j The category center feature of the first category is based on the first category. j The sample characteristics of all samples in each category are determined. K Indicates the number of categories.
4. The data value density assessment method according to claim 2, characterized in that, The ambiguity index is calculated using the following formula: ; in, x i Indicates the first element in the dataset to be evaluated. i Sample characteristics of each sample This indicates the preset test model pair. x i Predicted as the first k The probability of a class K Indicates the number of categories.
5. The data value density assessment method according to claim 2, characterized in that, The leakage ratio is calculated using the following formula: ; in, m This represents the total number of samples in the test set within the dataset to be evaluated. n This represents the total number of samples in the training set of the dataset to be evaluated. x i Indicates the first test set i Sample characteristics of each sample y j Indicates the first training set j The sample features of each sample, cos( θ ( x i , y j )) represents sample features x i and sample features y j cosine similarity, ε Indicates the similarity threshold. θ ( x i , y j ) represents the angle between two sample features, and II() represents the indicator function.
6. A data value density assessment device, characterized in that, include: The indicator value determination module is used to determine the indicator values of multiple evaluation indicators for the dataset to be evaluated. The multiple evaluation indicators include at least two of the following: density ratio, completeness rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy. The dataset to be evaluated includes an image dataset and a text dataset. The indicator weight determination module is used to determine the indicator weight of each evaluation indicator relative to the target task scenario based on the indicator value of each evaluation indicator. The target task scenario includes a classification task scenario and an object recognition task scenario. The dataset to be evaluated has different value densities for different target task scenarios. The value density assessment module is used to assess the value density of the dataset to be assessed relative to the target task scenario based on the index values and corresponding index weights of each assessment index. The indicator weight determination module is specifically used to determine the task fit of each evaluation indicator based on the indicator value of each evaluation indicator and the default value of each evaluation indicator in the target task scenario; and to determine the indicator weight of each evaluation indicator relative to the target task scenario based on the task fit of each evaluation indicator. The indicator weight determination module is specifically used to calculate the task fit of each evaluation indicator according to the following Gaussian kernel function. T Di : ; in, α i Indicates the first i The values of each evaluation indicator, T i Indicates the target task scenario for the first i The default values for each evaluation indicator This represents the bandwidth parameter of the Gaussian kernel. This represents an exponential function with the natural constant e as the base.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the data value density assessment method as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the data value density assessment method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Data set multi-scale evaluation method, system and equipment and storage medium
CN115543975A
Contribution degree assessment method and apparatus, and communication device and storage medium
WO2025066801A1