Data value density evaluation method and device

By using multi-dimensional evaluation metrics and task fit adjustment methods, the problem of the inability to comprehensively evaluate the value density of datasets in existing technologies is solved, thereby improving accuracy and applicability in specific task scenarios.

CN120910495AActive Publication Date: 2025-11-07INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510771402.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-11-07
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

Existing methods for evaluating the value density of datasets cannot comprehensively and accurately assess the value density of datasets, nor can they take into account the performance of datasets in specific task scenarios.

Method used

The dataset's value density is evaluated using multiple assessment metrics (such as completeness, label consistency, noise detection, and intra-class diversity). The weights of each assessment metric are determined based on the target task scenario. The task fit and metric weights are calculated using a Gaussian kernel function, and the dataset is optimized using a data distillation algorithm.

Benefits of technology

It enables a comprehensive assessment of the value density of datasets, improving their accuracy and applicability in real-world task scenarios. By adjusting the weights of indicators through multi-dimensional evaluation and task fit, it enhances the applicability and value density of datasets in specific tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910495A_ABST
    Figure CN120910495A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and provides a data value density evaluation method and device, and the method comprises the steps: determining the index values of a plurality of evaluation indexes of a to-be-evaluated data set; based on the index value of each evaluation index, determining the index weight of each evaluation index relative to the target task scene; and based on the index value of each evaluation index and the corresponding index weight, evaluating the value density of the to-be-evaluated data set relative to the target task scene. The value density of the to-be-evaluated data set is evaluated by adopting the index values of the plurality of evaluation indexes, the index weight of each evaluation index relative to the target task scene is determined based on the index value of each evaluation index, and the value density of the to-be-evaluated data set is evaluated based on the index value of each evaluation index and the corresponding index weight. And evaluating the value density of the to-be-evaluated data set relative to the target task scene, thereby realizing comprehensive and accurate evaluation of the value density of the to-be-evaluated data set.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a data value density evaluation method and device. BACKGROUND

[0002] Dataset Value Density is used to quantitatively measure the effective knowledge amount carried by unit data quantity in a dataset and its adaptation ability in a specific task. At present, the evaluation methods of dataset value density mainly include static data quality check, single performance index measurement, and random or simple strategy compression. Static data quality check mainly focuses on the evaluation of data completeness and repeatability, and cannot comprehensively evaluate the value density of the dataset. Single performance index measurement is to train and test the dataset in a specific (mature classic) deep learning model, and then evaluate the test results by a certain index (for example: recall rate), so as to evaluate the value density of the dataset. A single index is difficult to depict the characteristics of the dataset in many aspects, and is easy to underestimate or overestimate the potential information of the dataset. The evaluation method of random or simple strategy compression only evaluates a part of the dataset, and also cannot comprehensively evaluate the value density of the dataset. Moreover, the value density of the data depends not only on the static quality of the data itself, but also on its performance in specific task scenarios. The existing evaluation methods only consider the dataset itself, and cannot accurately evaluate the value density of the dataset.

[0003] In summary, the existing evaluation methods of dataset value density are only for the dataset itself, and cannot comprehensively and accurately evaluate the value density of the dataset. SUMMARY

[0004] The present application provides a data value density evaluation method and device to solve the problem that the existing technology cannot comprehensively and accurately evaluate the value density of the dataset.

[0005] The present application provides a data value density evaluation method, comprising the following steps: determining the index values of a plurality of evaluation indexes of a to-be-evaluated dataset; based on the index values of the evaluation indexes, determining the index weights of the evaluation indexes relative to a target task scenario; based on the index values of the evaluation indexes and the corresponding index weights, evaluating the value density of the to-be-evaluated dataset relative to the target task scenario.

[0006] According to the data value density evaluation method provided by the present application, based on the index values of the evaluation indexes, the index weights of the evaluation indexes relative to a target task scenario are determined, which comprises: based on the index values of the evaluation indexes and the default values of the evaluation indexes for the target task scenario, determining the task fit degrees of the evaluation indexes; Based on the task fit of each evaluation indicator, the indicator weight of each evaluation indicator relative to the target task scenario is determined.

[0007] According to a data value density assessment method provided by the present invention, the task fit of each assessment indicator is determined based on the indicator values ​​of each assessment indicator and the default values ​​of each assessment indicator in the target task scenario. This includes: calculating the task fit of each assessment indicator using the following Gaussian kernel function. T Di : ; in, α i Indicates the first i The values ​​of each evaluation indicator, T i Indicates the target task scenario for the first i The default values ​​for each evaluation indicator This represents the bandwidth parameter of the Gaussian kernel. This represents an exponential function with the natural constant e as the base.

[0008] According to a data value density assessment method provided by the present invention, the index values ​​of multiple assessment indicators for the dataset to be assessed are determined, including: Data distillation is performed on the dataset to be evaluated to obtain a density ratio index between the distilled data subset and the dataset to be evaluated. The density ratio index represents the ratio of the average data volume of the distilled data subset divided by category to the average data volume of the dataset to be evaluated divided by category. Calculate the completeness rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy of the dataset to be evaluated. The multiple evaluation metrics include at least two of the following: density ratio, integrity rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy.

[0009] According to a data value density assessment method provided by the present invention, the intraclass diversity index is calculated according to the following formula: ; in, x i Var( represents the sample feature of the i-th sample of any class in the dataset to be evaluated.) x i () indicates the calculation of the variance of sample features for any of the categories.n This indicates the number of samples in any of the categories. The inter-class diversity index is calculated using the following formula: ; in, c j Indicates the first element in the dataset to be evaluated. j The category center features of each category, Var( c j ) represents the calculation of the inner variance of the category center feature, wherein the first... j The category center feature of the first category is based on the first category. j The sample characteristics of all samples in each category are determined. K Indicates the number of categories.

[0010] According to a data value density assessment method provided by the present invention, the ambiguity index is calculated according to the following formula: ; in, x i Indicates the first element in the dataset to be evaluated. i Sample characteristics of each sample This indicates the preset test model pair. x i Predicted as the first k The probability of a class K Indicates the number of categories.

[0011] According to a data value density assessment method provided by the present invention, the leakage ratio index is calculated according to the following formula: ; in, m This represents the total number of samples in the test set within the dataset to be evaluated. n This represents the total number of samples in the training set of the dataset to be evaluated. x i Indicates the first test set i Sample characteristics of each sample y j Indicates the first training set j The sample features of each sample, cos( θ ( x i , y j )) represents sample features x i and sample features y j cosine similarity, ε Indicates the similarity threshold. θx i y j represents the included angle between two sample features, and II() represents an indicator function.

[0012] The application further provides a data value density evaluation device, comprising the following modules: An index value determination module is configured to determine index values of a plurality of evaluation indexes of a data set to be evaluated; An index weight determination module is configured to determine index weights of the evaluation indexes relative to a target task scenario based on the index values of the evaluation indexes; A value density evaluation module is configured to evaluate a value density of the data set to be evaluated relative to the target task scenario based on the index values of the evaluation indexes and the corresponding index weights.

[0013] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the data value density evaluation method according to any one of the above when executing the program.

[0014] The application further provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the data value density evaluation method according to any one of the above.

[0015] The data value density evaluation method and device provided by the application evaluate the value density of the data set to be evaluated through the index values of the plurality of evaluation indexes, determine the index weights of the evaluation indexes relative to the target task scenario based on the index values of the evaluation indexes, and evaluate the value density of the data set to be evaluated relative to the target task scenario based on the index values of the evaluation indexes and the corresponding index weights. Since the plurality of evaluation indexes in multiple dimensions are adopted and the corresponding index weights of the evaluation indexes are determined for the target task scenario, the value density of the data set to be evaluated is comprehensively evaluated, and the accuracy and applicability of the value density of the data set to be evaluated in the real task scenario are effectively improved. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0017] Figure 1 is a flowchart of the data value density evaluation method provided by the application. ​​

[0018] Figure 2 This is a schematic diagram of the data value density assessment device provided by the present invention.

[0019] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0021] The data value density assessment method of this invention, such as... Figure 1 As shown, the procedure includes steps S110 to S130.

[0022] Step S110: Determine the values ​​of multiple evaluation metrics for the dataset to be evaluated, so as to assess the value density of the dataset from multiple evaluation dimensions and to more comprehensively evaluate the value density of the dataset. Multiple evaluation metrics may include: completeness rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy.

[0023] Step S120: Based on the values ​​of each evaluation metric, determine the weight of each metric relative to the target task scenario. Since different datasets contain different data, for example, dataset A for classification includes multiple images of animals such as cats and dogs, primarily used for classification training, while dataset B for object recognition, such as face recognition, includes multiple images of faces, primarily used for object recognition. These two datasets obviously have different value densities for different training tasks. Dataset A has a higher value density for classification tasks (i.e., for classification training), while dataset B has a higher value density for object recognition tasks (i.e., for object recognition training). Therefore, to more accurately evaluate the value density of the datasets and their applicability to the task scenario, it is necessary to determine the weight of each evaluation metric relative to the target task scenario.

[0024] Step S130: Based on the index values ​​and corresponding index weights of each evaluation index, evaluate the value density of the dataset to be evaluated relative to the target task scenario. Because the index weights of the dataset to be evaluated relative to the target task scenario are considered during the evaluation process, the accuracy and applicability of the value density of the dataset to be evaluated in real task scenarios are effectively improved.

[0025] In the data value density evaluation method of the embodiment, the value density of the to-be-evaluated data set is evaluated by using the index values of multiple evaluation indexes, and the index weight of each evaluation index relative to the target task scene is determined based on the index values of the evaluation indexes. The value density of the to-be-evaluated data set relative to the target task scene is evaluated based on the index values of the evaluation indexes and the corresponding index weights. Since multiple evaluation indexes in multiple dimensions are used and the corresponding index weights of the evaluation indexes are determined for the target task scene, comprehensive evaluation of the value density of the to-be-evaluated data set is realized, and the accuracy and applicability of the value density of the to-be-evaluated data set in the real task scene are effectively improved.

[0026] In some embodiments, step S120 specifically includes: Based on the index values of the evaluation indexes and the default values of the evaluation indexes for the target task scene, the task fit degree of each evaluation index is determined. The task fit degree is intended to reflect the matching degree of the data features in the to-be-evaluated data set and the target task scene. Since each evaluation index is related to the data features in the data set, the relationship between the to-be-evaluated data set and the target task scene can be quantified by the index values of the evaluation indexes, so as to determine the task fit degree of each evaluation index. The default values of the evaluation indexes for the target task scene are set in advance.

[0027] Based on the task fit degrees of the evaluation indexes, the index weight of each evaluation index relative to the target task scene is determined. The higher the task fit degree, the greater the corresponding index weight.

[0028] Specifically, for each evaluation index, the index weight is obtained by continuously adjusting the initial weight based on the task fit degree. The weight adjustment is a way to modify the index weight of each evaluation index according to the requirements of the target task scene, so as to ensure that the index weights of the evaluation indexes in each dimension in the to-be-evaluated data set can better meet the requirements of the target task scene. In this embodiment, the initial weight is adjusted by in each round of iteration. Wherein, is the index weight of the i th evaluation index after the t th round of evaluation, is the weight change rate of the i th evaluation index, which is controlled by the task fit degree and satisfies: .

[0029] The index weight updating mechanism supports multiple rounds of iterative evaluation. In the case where the task fit degree is unchanged, This will continue to grow. Therefore, this paper sets the maximum number of evaluation rounds to 5. A convergence criterion is also introduced: if the difference between the weights of two consecutive indicators is lower than a preset difference threshold (e.g., 0.01), the iteration terminates. Furthermore, different task fits... By varying the growth rates, the weights of indicators corresponding to different levels of task relevance are differentiated, thereby ensuring the stability and efficiency of indicator weight adjustments and avoiding extreme inflation or meaningless duplicate evaluations.

[0030] In this embodiment, the task fit of each evaluation indicator is determined by the indicator value of each evaluation indicator and the default value of each evaluation indicator in the target task scenario. Then, the indicator weights are adjusted in multiple rounds based on the task fit to obtain more accurate indicator weights.

[0031] In some embodiments, the task fit of each evaluation indicator is determined based on the indicator values ​​of each evaluation indicator and the default values ​​of each evaluation indicator for the target task scenario, including: calculating the task fit of each evaluation indicator according to the following Gaussian kernel function. T Di : ; in, α i Indicates the first i The values ​​of each evaluation indicator, T i Indicates the target task scenario for the first i The default values ​​for each evaluation indicator This represents the bandwidth parameter of the Gaussian kernel. This indicates an exponential function with the natural constant e as the base. (Default value) T i These are pre-defined values. The default values ​​for different evaluation metrics vary for the same target task scenario. These values ​​can be set according to the actual situation. For example, the default value for the inter-class diversity metric in the classification task scenario is greater than that in the object detection task scenario. In this way, for a dataset suitable for classification, the difference between the calculated inter-class diversity metric value and the default value of the inter-class diversity metric in the classification task scenario is small. Substituting this into the task fit calculation formula above, the inter-class diversity metric has a greater task fit with the classification task scenario. On the other hand, the difference between the calculated inter-class diversity metric value and the default value of the inter-class diversity metric in the object detection task scenario is larger, and the inter-class diversity metric has a smaller task fit with the object detection task scenario.

[0032] For example, the default values ​​for each evaluation metric for different target task scenarios can be as follows: Density ratio index: 0.5 can be used for both classification task scenarios and object detection task scenarios.

[0033] Data integrity index and label consistency index: 0.98 for classification task scenarios and 0.96 for target detection task scenarios.

[0034] Noise detection index: 0.1 for classification task scenarios and 0.2 for target detection task scenarios.

[0035] Intra-class diversity index: 0.6 for classification task scenarios and 0.75 for target detection task scenarios.

[0036] Inter-class diversity index: 0.9 for classification task scenarios and 0.7 for target detection task scenarios.

[0037] Data statistical complexity index: 0.6 for mean, 0.65 for variance, and 0.7 for entropy in classification task scenarios; 0.85 for mean, 0.8 for variance, and 0.85 for entropy in target detection task scenarios.

[0038] Ambiguity index: 0.3 for classification task scenarios and 0.45 for target detection task scenarios.

[0039] Domain shift index: 0.2 for classification task scenarios and 0.35 for target detection task scenarios.

[0040] Leakage ratio index: 0.1 for classification task scenarios and 0.2 for target detection task scenarios.

[0041] Prediction accuracy index: 0.9 for classification task scenarios and 0.8 for target detection task scenarios.

[0042] In some embodiments, in step S130, the value density Final Score of the data set to be evaluated relative to the target task scenario is calculated as follows.

[0043] .

[0044] wherein dq[ i ] represents the index value of the i i th evaluation index, and aw[ i ] represents the index weight of the i i th evaluation index relative to the target task scenario.

[0045] Further, when the value density Final Score of the to-be-evaluated data set relative to the target task scene is low, the to-be-evaluated data set can be optimized and fed back, for example, data enhancement processing is performed on the sample, data is denoised, whether the label is complete is checked, and the like, so that the index values of the evaluation indexes are more accurate, and a higher evaluation score is obtained. For some standard and mature data sets (data sets that can obtain good training results after a large number of model training and verification of the target task scene), if the score of the Final Score is low, it indicates that the default value of the target task scene for part of the evaluation indexes is unreasonable, and the default value of the target task scene for part of the evaluation indexes can be adjusted, so as to adjust the weight of part of the evaluation indexes, and finally make the score of the value density Final Score relative to the target task scene reach a higher value.

[0046] In some embodiments, determining the index values of the plurality of evaluation indexes of the to-be-evaluated data set comprises: performing data distillation on the to-be-evaluated data set to obtain a density ratio index of the distilled data subset and the to-be-evaluated data set, the density ratio index representing a ratio of an average value of data amount classified by categories in the distilled data subset to an average value of data amount classified by categories in the to-be-evaluated data set.

[0047] calculating the completeness index, the label consistency index, the noise detection index, the intra-class diversity index, the inter-class diversity index, the data statistical complexity index, the ambiguity index, the domain shift index, the leakage ratio index, and the prediction accuracy index of the to-be-evaluated data set.

[0048] The plurality of evaluation indexes can include at least two of the density ratio, the completeness index, the label consistency index, the noise detection index, the intra-class diversity index, the inter-class diversity index, the data statistical complexity index, the ambiguity index, the domain shift index, the leakage ratio index, and the prediction accuracy index, so that the data value density of the to-be-evaluated data set is more accurately evaluated in multiple dimensions.

[0049] In some embodiments, in order to effectively remove a large number of redundant and low-value samples in the to-be-evaluated data set and improve the knowledge density of core samples in the data set, a data distillation algorithm is introduced in this embodiment for knowledge condensation and density quantification. Specifically, the to-be-evaluated data set is represented as The sample distribution is preliminarily analyzed by a feature extraction and dimension reduction method (such as principal component analysis (PCA) or t-SNE), and then a gradient matching strategy, a difficulty alignment trajectory matching algorithm (DATM), or K-means clustering is used to select more representative core samples to constitute a distilled data set. The specific distillation algorithm formula is as follows: .

[0050] where D' is a candidate subset of the data set D to be evaluated (D' can be randomly selected from D, selected by class average or selected by entropy value), arg min is a function acting on a variable D' to find the D' that minimizes the loss of the model on D, that is, the training effect of the model based on D' and based on D is basically the same, is the expectation of the sample in the data set D to be evaluated, is the loss function, is the model trained according to the distilled data subset , the prediction output of the input data x , y represents the corresponding label of the input data x .

[0051] To measure the distillation effect, the number of each category retained in the distilled data subset is defined, that is, the average value of the data amount divided by category IPC1=|D distill | / K, where |D distill | represents the data amount in the distilled data subset, and K is the number of categories. IPC1 reflects the retention degree of different categories of data by data distillation, and the higher the value, the more samples are retained on average. IPC2=|D| / K, |D| represents the data amount in the data set to be evaluated, and the density ratio index of the distilled data subset and the data set to be evaluated is defined as IPC1 / IPC2.

[0052] In some embodiments, the integrity of the data set directly affects the training stability and prediction accuracy of the model. If there are a large number of missing or damaged samples in the data set, the training process of the model will be greatly affected, resulting in problems such as overfitting or underfitting. Therefore, the complete rate index Complete Ratio is defined, and the complete rate index of the data set to be evaluated is used to measure the proportion of effective samples in the data set, which is specifically defined as the proportion of the number of effective samples to the total number of samples. The calculation formula of CompleteRatio is as follows: .

[0053] where Valid Samples represents the number of valid samples without missing and damage in the data set to be evaluated, and TotalSamples represents the total number of samples in the data set to be evaluated.

[0054] In some embodiments, the accuracy and consistency of labels in a dataset directly affect the training effect and final performance of the model. Especially in supervised learning, datasets with low label consistency can lead to bias between the training and test sets, thus affecting the model's generalization ability. Therefore, a label consistency metric, Label Consistency, is defined, representing the proportion of samples with incorrect labels to the total number of samples. The formula for calculating Label Consistency is as follows: .

[0055] in, Indicates the first i If a sample is incorrectly labeled, it is recorded as 1; if it is correctly labeled, it is recorded as 0. N This represents the total number of samples.

[0056] In some embodiments, noise refers to outlier samples in the dataset, which are typically generated by incorrect labeling, data input problems, or anomalies in the samples themselves. Noisy samples not only affect the model training process but may also cause the model to produce misleading results, thereby reducing the model's prediction accuracy. To effectively detect noisy samples, this embodiment combines dimensionality reduction techniques (e.g., t-distributed stochastic neighbor embedding, t-SNE) and clustering algorithms (e.g., density-based spatial clustering of applications with noise, DBSCAN) for noise sample identification.

[0057] Specifically, t-SNE is used to reduce high-dimensional data to two or three dimensions, visually demonstrating the distribution of data points. Noise samples are typically scattered and far from other sample categories. The DBSCAN clustering algorithm is used to identify low-density regions, which usually contain noise samples.

[0058] For example: The dataset to be evaluated has a total of N There are n samples, where the feature representation of each data point is as follows: After standardization and t-SNE dimensionality reduction, a low-dimensional feature representation is obtained. Then, the cluster centers and distances are calculated, and the mean vector of all the reduced-dimensional features is... , No. i The Euclidean distance from each sample to the cluster center is Generate a threshold and determine noise: Calculate the mean for all distances. and standard deviation Construct a threshold, and the mean vector of all dimensionality-reduced features is: .

[0059] Finally, a threshold is selected k (e.g., take k = 3 θ k ), for each sample , calculate its Euclidean distance to the cluster center , if it satisfies , it is considered that the sample is a noise sample. The embodiment can effectively detect outliers in the data, and distinguish noise samples from effective samples. Finally, the ratio of the number of noise samples to the total number of samples is calculated as a noise detection index.

[0060] In some embodiments, the intra-class diversity index is used to measure the difference between samples of the same class. Moderate intra-class diversity helps improve the generalization ability of the model, so that the model can better learn the diverse features of the class. Therefore, the intra-class diversity index Class Intra-Diversity is defined, specifically, the Class Intra-Diversity is calculated by the intra-class variance: .

[0061] Wherein, x i represents the sample feature of the i-th sample of any class in the data set to be evaluated, Var( x i ) represents the calculation of the sample feature variance of the any class, n represents the number of samples of the any class.

[0062] The inter-class diversity index is used to measure the discrimination between different classes, reflecting the class discrimination ability of the data set. A data set with high inter-class diversity can help the model better distinguish different classes and improve the model's adaptability to complex classification tasks. If the inter-class diversity is low, it may lead to blurred boundaries between different classes, affecting the classification accuracy of the model. Therefore, the inter-class diversity index Class Inter-Diversity is defined, specifically, the Class Inter-Diversity is calculated by the inter-class variance: .

[0063] Wherein, c j represents the class center feature of the i-th class in the data set to be evaluated, Var( j j ) represents the calculation of the inter-class variance of the class center feature, the i-th class c j ​The category center feature of each category is determined based on the sample features of all samples in the category, j . K represents the number of categories.

[0064] In this embodiment, by introducing the intra-class variance and the inter-class variance as core indexes, the internal differences of different categories in the data set and the discrimination degree between categories are effectively distinguished, thereby improving the fineness and discrimination ability of the feature representation of the data set.

[0065] In some embodiments, the data statistical complexity index is used to measure the overall information amount and diversity of the data set. A high index value of the data statistical complexity index means that more potential information is contained, which can provide more rich features for model learning. In this embodiment, the statistical complexity of the data set is usually evaluated by the mean μ , standard deviation σ and entropy H(data), that is, the data statistical complexity index includes the mean, variance and entropy of the data set. The entropy value of the data set reflects the diversity and information amount of the data, and the higher the entropy value, the more complex the data. Therefore, the mean μ , standard deviation σ and entropy H(data) are defined as the data statistical complexity index respectively. The specific formulas of the mean and variance of the data set to be evaluated are as follows: .

[0066] x i represents the sample feature of the i-th sample in the data set to be evaluated.

[0067] The entropy value H(data) of the data set to be evaluated can be measured by the distribution uncertainty, and the specific formula is as follows: .

[0068] wherein, g i represents the probability of the i-th element in the data. For example, when the data is an image, g i represents the pixel value of the i-th pixel in the image, and when the data is text data, g i represents the i-th text content (word or character) in the text. For image data, in order to calculate , first, the image is grayed, and the histogram of the gray value is counted: .

[0069] wherein, the values of i and j are both [0, 255], but they are different, iis the gray level of the probability being calculated, and j is the index variable when summing all gray levels. Pixel value g i The frequency is calculated using the total number of all pixels in the denominator. For the grayscale value range, [0, 255] is used for calculation.

[0070] In some embodiments, ambiguity refers to the degree of ambiguity of a sample at the class boundary. Higher ambiguity increases the risk of model misjudgment and misleads the decision boundary. Therefore, in this embodiment, the model is tested to predict the samples and output results. The entropy of the output results is then calculated to quantify the ambiguity of the samples at the class boundary; this entropy is the ambiguity index, and the specific calculation formula is as follows: ; in, x i Indicates the first in the dataset to be evaluated i Sample characteristics of each sample This indicates the preset test model pair. x i Predicted as the first k The probability of a class K Indicates the number of categories.

[0071] Preferably, multiple test models can be used to test the first... i Different prediction results are obtained by predicting the sample features of each sample, and multiple corresponding entropies are obtained. The average of the multiple entropies is taken to obtain the first entropy. i The entropy value of each sample, i.e., the ambiguity index. Entropy obtained from the test results of multiple test models better reflects the ambiguity of the samples; that is, the ambiguity index value is more accurate. It should be noted that each sample feature... x i Each of these corresponds to an entropy value. Therefore, the ambiguity index of the dataset to be evaluated is a vector composed of the entropy values ​​of all sample features. The index value of the ambiguity index is defined as the sum of the values ​​of each element in the vector.

[0072] In some embodiments, domain offset is used to measure the degree of difference in the data distribution of the training and test sets in the feature space. If the distributions of the training and test sets are inconsistent, the model's performance in practical applications will significantly decrease, especially in transfer learning or cross-domain tasks, where domain offset is a key factor leading to unsatisfactory results. Therefore, a domain offset metric is defined to measure the degree of difference in the data distribution of the training and test sets in the feature space of the dataset to be evaluated. Specifically, the Jensen-Shannon divergence metric JSD is used to characterize the domain offset metric. .

[0073] in,P and Q respectively represent the data distribution of the training set and the test set in the feature space, is the average distribution of the training set and the test set distribution, is the Kullback-Leibler divergence (KL divergence) used to measure the difference between distribution P and average distribution M, which is specifically defined as: .

[0074] wherein, x represents the index of each dimension after flattening all feature vectors in the sample space, D KL ( Q || M ) and D KL ( P || M ) have similar meanings, which will not be repeated here.

[0075] In some embodiments, the degree of repetition or similarity between the test sample and the training set sample is high, which means that the test set contains a large number of samples that are identical or extremely similar to the training set samples, resulting in a decrease in the objectivity and generalization of the model test results. Therefore, the leakage ratio indicator is defined to represent the degree of repetition or similarity between the test sample and the training set sample, and specifically, the calculation formula of the leakage ratio indicator Leaking Ratio is as follows: .

[0076] wherein, m represents the total number of samples in the test set in the to-be-evaluated data set, n represents the total number of samples in the training set in the to-be-evaluated data set, x i represents the sample feature of the i-th sample in the test set, i j represents the sample feature of the i-th sample in the training set, and cos( y ( j i , θ j )) represents the cosine similarity between the sample feature x i and the sample feature y j , x i represents a similarity threshold, y j ( ε Di , θ i ​​represents the included angle between the features of two samples, and II() represents an indicator function.

[0077] In some embodiments, if a data set can achieve high prediction accuracy on several test models (mature classic models such as ResNet, VGG, etc.), it generally indicates that it has strong feature separability and high data quality, and is suitable for more complex tasks. Therefore, a prediction accuracy indicator is defined to evaluate the feature expression ability of the data set, and the prediction accuracy indicator represents the average value of the prediction accuracy of the data set in multiple test models. Specifically, the calculation formula of the prediction accuracy Accuracy is as follows: .

[0078] Correct Predictions is the number of correct predictions, Total Predictions is the total number of predictions, and the prediction accuracy indicator is the average value of the Accuracy corresponding to multiple test models.

[0079] The data value density evaluation device provided by the present application is described below, and the data value density evaluation device described below can be correspondingly referred to the data value density evaluation method described above.

[0080] The data value density evaluation device of the embodiment of the present application, as shown in Figure 2 , comprises: The index value determination module 210 is configured to determine the index values of the plurality of evaluation indexes of the data set to be evaluated.

[0081] The index weight determination module 220 is configured to determine the index weights of the evaluation indexes with respect to the target task scene based on the index values of the evaluation indexes.

[0082] The value density evaluation module 230 is configured to evaluate the value density of the data set to be evaluated with respect to the target task scene based on the index values of the evaluation indexes and the corresponding index weights.

[0083] The data value density evaluation device provided by the present application evaluates the value density of the data set to be evaluated through the index values of the plurality of evaluation indexes, determines the index weights of the evaluation indexes with respect to the target task scene based on the index values of the evaluation indexes, and evaluates the value density of the data set to be evaluated with respect to the target task scene based on the index values of the evaluation indexes and the corresponding index weights. Since the evaluation indexes of multiple dimensions are adopted and the evaluation indexes are determined with corresponding index weights with respect to the target task scene, the value density of the data set to be evaluated is comprehensively evaluated, and the accuracy and applicability of the value density of the data set to be evaluated in the real task scene are effectively improved.

[0084] In some embodiments, the indicator weight determination module 220 is specifically used to determine the task fit of each evaluation indicator based on the indicator value of each evaluation indicator and the default value of each evaluation indicator for the target task scenario; and to determine the indicator weight of each evaluation indicator relative to the target task scenario based on the task fit of each evaluation indicator.

[0085] In some embodiments, the indicator weight determination module 220 is specifically used to calculate the task fit of each evaluation indicator according to the following Gaussian kernel function. T Di : .

[0086] in, α i Indicates the first i The values ​​of each evaluation indicator, T i Indicates the target task scenario for the first i The default values ​​for each evaluation indicator This represents the bandwidth parameter of the Gaussian kernel. This represents an exponential function with the natural constant e as the base.

[0087] In some embodiments, the indicator value determination module 210 is specifically used for: Data distillation is performed on the dataset to be evaluated to obtain a density ratio index between the distilled data subset and the dataset to be evaluated. The density ratio index represents the ratio of the average data volume of the distilled data subset divided by category to the average data volume of the dataset to be evaluated divided by category. Calculate the completeness rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy of the dataset to be evaluated. The multiple evaluation metrics include at least two of the following: density ratio, integrity rate, label consistency, noise detection, intra-class diversity, inter-class diversity, data statistical complexity, ambiguity, domain offset, leakage ratio, and prediction accuracy.

[0088] In some embodiments, the intra-class diversity index is calculated using the following formula: .

[0089] in, x i Var( represents the sample feature of the i-th sample of any class in the dataset to be evaluated.) x i () indicates the calculation of the variance of sample features for any of the categories.n This indicates the number of samples in any of the categories. The inter-class diversity index is calculated using the following formula: .

[0090] in, c j Indicates the first in the dataset to be evaluated j The category center features of each category, Var( c j ) represents the calculation of the inner variance of the category center feature, wherein the first... j The category center feature of the first category is based on the first category. j The sample characteristics of all samples in each category are determined. K Indicates the number of categories.

[0091] In some embodiments, the ambiguity index is calculated using the following formula: .

[0092] in, x i Indicates the first element in the dataset to be evaluated. i Sample characteristics of each sample This indicates the preset test model pair. x i Predicted as the first k The probability of a class K Indicates the number of categories.

[0093] In some embodiments, the leakage ratio is calculated using the following formula: .

[0094] in, m This represents the total number of samples in the test set within the dataset to be evaluated. n This represents the total number of samples in the training set of the dataset to be evaluated. x i Indicates the first test set i Sample characteristics of each sample y j Indicates the first training set j The sample features of each sample, cos( θ ( x i , y j )) represents sample features x i and sample features y j cosine similarity, εdenotes a similarity threshold value, θ x i y j denotes an included angle between two sample features, and II() denotes an indicator function.

[0095] Figure 3 An example of an entity structure diagram of an electronic device is shown in Figure 3 The electronic device can include a processor 310, a communications interface 320, a memory 330, and a communications bus 340, wherein the processor 310, the communications interface 320, and the memory 330 can communicate with each other through the communications bus 340. The processor 310 can invoke a logical instruction in the memory 330 to execute a data value density evaluation method, which includes: Determining an index value of a plurality of evaluation indexes of a to-be-evaluated data set.

[0096] Based on the index value of each evaluation index, determining an index weight of each evaluation index with respect to a target task scenario.

[0097] Based on the index value of each evaluation index and the corresponding index weight, evaluating the value density of the to-be-evaluated data set with respect to the target task scenario.

[0098] In addition, the logical instruction in the memory 330 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0099] On the other hand, the present application also provides a computer program product, which includes a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the data value density evaluation method provided by the above-mentioned methods, which includes: ​​Determine an index value of each evaluation index of the to-be-evaluated data set.

[0100] Determine an index weight of each evaluation index relative to the target task scenario based on the index value of each evaluation index.

[0101] Evaluate the value density of the to-be-evaluated data set relative to the target task scenario based on the index value of each evaluation index and the corresponding index weight.

[0102] In another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the data value density evaluation method provided by the above method, the method comprising: Determine an index value of each evaluation index of the to-be-evaluated data set.

[0103] Determine an index weight of each evaluation index relative to the target task scenario based on the index value of each evaluation index.

[0104] Evaluate the value density of the to-be-evaluated data set relative to the target task scenario based on the index value of each evaluation index and the corresponding index weight.

[0105] The device embodiments described above are only illustrative, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0106] From the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software product, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, server, or network device, etc.) execute the method described in each embodiment or some parts of the embodiment.

[0107] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method of data value density assessment, characterized by, The method comprises the following steps: determining index values of a plurality of evaluation indexes of a data set to be evaluated; based on the index values of each evaluation index, determining the index weight of each evaluation index relative to the target task scene; based on the index values of each evaluation index and the corresponding index weight, evaluating the value density of the data set to be evaluated relative to the target task scene.

2. The data value density evaluation method of claim 1, wherein, Based on the index values of each evaluation index, the index weight of each evaluation index relative to the target task scene is determined, which comprises: based on the index values of each evaluation index and the default value of each evaluation index for the target task scene, determining the task fit degree of each evaluation index; based on the task fit degree of each evaluation index, determining the index weight of each evaluation index relative to the target task scene.

3. The data value density evaluation method of claim 2, wherein, Based on the index value of each evaluation index and the default value of each evaluation index for the target task scene, the task fitness of each evaluation index is determined, including calculating the task fitness of each evaluation index according to the following Gaussian kernel function T Di : ; wherein, α i an index value representing an evaluation index, i T i a default value representing a target task scenario for an evaluation index, i a bandwidth parameter representing a Gaussian kernel, an exponential function with a natural constant e as a base.​​ 4. The data value density evaluation method according to any one of claims 1 to 3, characterized by, The index values of a plurality of evaluation indexes of a data set to be evaluated are determined, which comprises: data distillation is performed on the data set to be evaluated to obtain a density ratio index of the distilled data subset and the data set to be evaluated, and the density ratio index represents the ratio of the average data amount classified by categories in the distilled data subset to the average data amount classified by categories in the data set to be evaluated; the completeness index, the label consistency index, the noise detection index, the intra-class diversity index, the inter-class diversity index, the data statistical complexity index, the ambiguity index, the domain shift index, the leakage ratio index and the prediction accuracy index of the data set to be evaluated are calculated; The plurality of evaluation indexes include at least two evaluation indexes in the density ratio, the completeness index, the label consistency index, the noise detection index, the intra-class diversity index, the inter-class diversity index, the data statistical complexity index, the ambiguity index, the domain shift index, the leakage ratio index and the prediction accuracy index.

5. The data value density evaluation method of claim 4, wherein, The intra-class diversity index is calculated according to the following formula: ; wherein, x i represents the sample feature of the i-th sample of any category in the data set to be evaluated, Var( x i represents the variance of the sample feature of any category, n represents the number of samples of any category; The inter-class diversity index is calculated according to the following formula: ; in, c j Indicates the first element in the dataset to be evaluated. j The category center features of each category, Var( c j ) represents the calculation of the inner variance of the category center feature, wherein the first... j The category center feature of the first category is based on the first category. j The sample characteristics of all samples in each category are determined. K Indicates the number of categories.

6. The data value density evaluation method of claim 4, wherein, The ambiguity index is calculated according to the following formula: ; wherein, x i denotes a sample feature of an i-th sample in the data set to be evaluated, i denotes a probability predicted by a preset test model for the i-th sample to be in a j-th class, x i a j-th class, k K denotes a number of classes.​​ 7. The data value density evaluation method of claim 4, wherein, The leakage ratio index is calculated according to the following formula: ; in, m This represents the total number of samples in the test set within the dataset to be evaluated. n This represents the total number of samples in the training set of the dataset to be evaluated. x i Indicates the first test set i Sample characteristics of each sample y j Indicates the first training set j The sample features of each sample, cos( θ ( x i , y j )) represents sample features x i and sample features y j cosine similarity, ε Indicates the similarity threshold. θ ( x i , y j ) represents the angle between two sample features, and II() represents the indicator function.

8. A data value density assessment apparatus, characterized by, The method comprises the following steps: an index value determination module for determining index values of a plurality of evaluation indexes of a data set to be evaluated; an index weight determination module for determining the index weight of each evaluation index relative to the target task scene based on the index values of each evaluation index; a value density evaluation module for evaluating the value density of the data set to be evaluated relative to the target task scene based on the index values of each evaluation index and the corresponding index weight.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The processor executes the computer program to realize the data value density evaluation method of any one of claims 1 to 7. 10.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the data value density evaluation method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data value evaluation system and method

    CN110659926A

  • Data set multi-scale evaluation method, system and equipment and storage medium

    CN115543975A

  • Contribution degree assessment method and apparatus, and communication device and storage medium

    WO2025066801A1