A method for value assessment and sampling of a dataset
By constructing an evaluation model for the individual value and redundancy of data and sampling high-value data sets, the problems of inaccurate data value evaluation and redundant utilization in existing technologies are solved, achieving efficient data sampling and target task effectiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-19
- Publication Date
- 2026-04-03
AI Technical Summary
Existing data valuation methods fail to accurately reflect the value of data sets, and the redundant value between individual data points is not effectively utilized, resulting in high data collection costs and high computational complexity.
Establish a data individual value assessment model and a function to describe the degree of value redundancy among data individuals, construct a data set value assessment model, and sample high-value data sets from the data sampling space using a greedy method or a global optimization method.
While ensuring the effectiveness of the target task, we should reasonably assess the value of the dataset, reduce data collection costs and computational complexity, and improve the quality of the dataset.
Smart Images

Figure CN115525869B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer science and technology, and in particular to a data set value assessment technique, specifically a data set value assessment and sampling method. Background Technology
[0002] In recent years, data-driven modeling methods have been widely applied in various fields such as computer vision, system fault diagnosis, and state prediction. The superior performance of data-driven models relies on large amounts of training data, but data acquisition in many engineering scenarios is often difficult, time-consuming, and costly. Previous research has found that the data that plays a crucial role in the performance of the target task is often only a portion of the dataset; noise, outliers, and other poor data can actually have a negative impact on the target task. Therefore, sampling high-value datasets can significantly reduce data acquisition costs and computational complexity while ensuring the effectiveness of the target task.
[0003] Patent CN114926204A discloses a data processing device and method based on data value. This method calculates the value of individual data points and selects data points with high individual values to form a set based on the calculation results. This method can provide the data set with the largest sum of individual data values. However, in practice, it has been found that redundant value exists among the data points in the data set. If one data point is already included in the data set, even if other data points in the neighborhood have high individual values, their contribution to the overall value of the data set is minimal. Therefore, the sum of individual data values cannot effectively reflect the value of the data set.
[0004] Therefore, this invention comprehensively considers the value of individual data points and the degree of value redundancy between data points during the construction of a data set value assessment model, and proposes a data set value assessment and sampling method. The method described in this invention can reasonably assess the value of a data set, guide data sampling in data-driven analysis, modeling, and decision-making tasks, thereby improving the quality of the data set and effectively reducing data collection costs while ensuring the effectiveness of the target task. Summary of the Invention
[0005] The purpose of this invention is to address the problem that existing data value assessments cannot accurately reflect the value of data. This invention provides a data set value assessment and sampling method that can reasonably evaluate the value of a data set and thus achieve the target task with a small-scale, high-value data set. This method can guide data sampling in data-driven analysis, modeling, and decision-making tasks, thereby improving the quality of the data set and effectively reducing data collection costs while ensuring the effectiveness of the target task.
[0006] The technical solution of this invention is:
[0007] A method for evaluating the value of a dataset is characterized by the following steps: First, an evaluation model is established to assess the value of individual data points, and a function is established to describe the degree of value redundancy among individual data points; then, a value evaluation model for the dataset is constructed by comprehensively considering the value of individual data points and the degree of value redundancy among individual data points.
[0008] The method for establishing the evaluation model for the value of individual data is one of the following:
[0009] The value of individual data points is assessed by calculating the gain of each data point on the target task, and then an assessment model for evaluating the value of individual data points is established.
[0010] The value of individual data is evaluated by calculating the gain of individual data in scenarios similar to the target task, and then an evaluation model for evaluating the value of individual data is established.
[0011] The value of individual data points is assessed based on domain knowledge of the data generation context, and an assessment model is then established to evaluate the value of individual data points.
[0012] The gain calculation method can be obtained by calculating the Shapley value of the data individual for the target task.
[0013] The domain knowledge assessment of the data generation scenario can evaluate the value of individual data based on a preliminary understanding of the target task or the data generation mechanism. For example, in a surface measurement task, coordinate points with smaller radii of curvature have greater individual data value for the target task.
[0014] The methods used to establish the evaluation model for assessing the value of individual data can be machine learning algorithms such as least squares, Gaussian process regression, and neural networks.
[0015] The function describing the degree of value redundancy among individual data items can be established using one of the following methods:
[0016] The degree of redundancy between individual data is inversely proportional to the distance between them. A set of data that are close to each other generates greater redundancy value. The distances include, but are not limited to, Euclidean distance and Mahalanobis distance.
[0017] The degree of redundancy between individual data is directly proportional to the correlation between them. A set of data with greater correlation generates greater redundancy value. The correlation can be represented by kernel functions, membership functions, etc.
[0018] The function describing the degree of value redundancy among individual data items can take one of the following forms:
[0019] Gaussian kernel function:
[0020] Laplace kernel function:
[0021] Inverse multiquadratic kernel function:
[0022] where \(x\) , , o , p , i , , , ,
[0030] , i , <00}) represents i data individuals x1, x2, ..., x i The data set {x1,x2,…,x} is composed of i The value of} is determined by sampling the data within the data sampling space. First, the first data individual x1 that maximizes v({x1}) is sampled using a greedy method or a global optimization method. Then, the second data individual x2 that maximizes v({x1,x2}) - v({x1}) is sampled using a greedy method or a global optimization method. If the set {x1,x2} does not meet the target task performance requirements, the third data individual x3 that maximizes v({x1,x2,x3}) - v({x1,x2}) is sampled using a greedy method or a global optimization method. This process is repeated iteratively until the target task performance requirements are met.
[0031] The beneficial effects of this invention are:
[0032] The data set value assessment method proposed in this invention can reasonably assess the value of the data set, and the sampling method for improving the value of the data set can sample high-value data sets for achieving the target task. It can guide data sampling in data-driven analysis, modeling and decision-making tasks, thereby improving the quality of the data set and effectively reducing data collection costs while ensuring the effectiveness of the target task. Attached Figure Description
[0033] Figure 1 This is a flowchart illustrating the implementation of the present invention.
[0034] Figure 2 This is a comparison chart of the results of the method described in a specific embodiment of the present invention and the random sampling method.
[0035] Figure 3 This is a comparison chart of the results of the method described in a specific embodiment of the present invention and the sampling method that maximizes the sum of individual data values (referred to as individual value sampling). Detailed Implementation
[0036] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0037] like Figure 1-3 As shown.
[0038] A method for evaluating and sampling the value of a dataset is provided, which can reasonably assess the value of the dataset and guide data sampling in data-driven analysis, modeling, and decision-making tasks, thereby improving the quality of the dataset and effectively reducing data acquisition costs while ensuring the effectiveness of the target task. The invention will be further described below with reference to the accompanying drawings and a modeling task for composite material curing processes.
[0039] Task Description: The objective is to model the curing process of composite materials. Specifically, this involves using finite element analysis (FEM) software to acquire curing state data under different curing processes and establishing a mapping model from single-insulation curing process parameters (a total of six parameters: heating rate, cooling rate, insulation temperature, insulation time, air-part convective heat transfer coefficient, and air-mold convective heat transfer coefficient) to thermal hysteresis. While FEM software offers high accuracy, its efficiency is low; therefore, sampling high-value datasets will reduce data collection costs.
[0040] Specifically, the following steps are included:
[0041] Step 1: Establish an evaluation model to assess the value of individual data points.
[0042] The method for establishing the evaluation model for the value of individual data is one of the following:
[0043] The value of individual data points is assessed by calculating the gain of each data point on the target task, and then an assessment model for evaluating the value of individual data points is established.
[0044] The value of individual data is evaluated by calculating the gain of individual data in scenarios similar to the target task, and then an evaluation model for evaluating the value of individual data is established.
[0045] The value of individual data points is assessed based on domain knowledge of the data generation context, and an assessment model is then established to evaluate the value of individual data points.
[0046] The gain calculation method can be obtained by calculating the Shapley value of the data individual for the target task.
[0047] The domain knowledge assessment of the data generation scenario can evaluate the value of individual data based on a preliminary understanding of the target task or the data generation mechanism. For example, in a surface measurement task, coordinate points with smaller radii of curvature have greater individual data value for the target task.
[0048] The second approach is used here to establish an evaluation model for assessing the value of individual data points. Data points in scenarios similar to the target task are generated using a low-precision but efficient finite difference method. Specifically, 600 potential sampleable data points are randomly generated within a given range of six single-insulation curing process parameters. The thermal hysteresis corresponding to each of the 600 curing process parameters is obtained using the finite difference method. Then, Gaussian process regression is used as the fitting model. The value of each data point is assessed by calculating the Shapley value of the 600 data points relative to the target task. Furthermore, machine learning algorithms such as least squares, Gaussian process regression, and neural networks can be used to establish an evaluation model for assessing the value of individual data points; here, a neural network is used. It is worth noting that the established evaluation model for assessing the value of individual data points can be generalized to the value assessment of all potential sampleable data points beyond the aforementioned 600 data points.
[0049] Step 2: Establish a function to describe the degree of value redundancy among individual data.
[0050] The function describing the degree of value redundancy among individual data items can be established using one of the following methods:
[0051] The degree of redundancy between individual data is inversely proportional to the distance between them. A set of data that are close to each other generates greater redundancy value. The distances include, but are not limited to, Euclidean distance and Mahalanobis distance.
[0052] The degree of redundancy between individual data is directly proportional to the correlation between them. A set of data with greater correlation generates greater redundancy value. The correlation can be represented by kernel functions, membership functions, etc.
[0053] In this embodiment, the second method is used. The function describing the degree of value redundancy between individual data can be a Gaussian kernel function, a Laplace kernel function, an inverse multivariate quadratic kernel function, etc. Here, a Gaussian kernel function is used, specifically:
[0054]
[0055] Where, x i Let x represent the i-th data individual. j Let represent the j-th data individual, and σ be a parameter controlling the size of the Gaussian kernel function, which determines the influence range of the data individual. σ can be predefined or determined through parameter selection strategies such as cross-validation. In this embodiment, σ is used. 2 =50.
[0056] Step 3: Construct a value assessment model for the data set by comprehensively considering the value of individual data points and the degree of value redundancy among them. The calculation method is as follows:
[0057]
[0058] x′(x,S) = v(x) max{k(x,x1), …, k(x,x m )}, x1, …, x m ∈S
[0059] where v(S) is the value evaluation model of the data set, n is the number of potential data points in the sample space, N is the data set composed of n potential data points in the sample space, S is a data subset of the data set N, v(x) is the evaluation model of the value of the data individual, and k(x,x i )(i = 1, 2, …, m) is the function describing the value redundancy degree between data individuals, and m (0 < m ≤ n) is the number of data in S.
[0060] Step 4: Sample high-value data sets from the data sampling space according to user needs, including the following methods:
[0061] Given the sampling quantity p, based on the value evaluation model of the data set, denote v(S p ) as the value of the set S p composed of p data individuals, and sample the data set that maximizes v(S p ) from the data sampling space by using the greedy method or the global optimization method;
[0062] Given the target task performance requirement, based on the value evaluation model of the data set, denote x i as the i-th data individual in the sampling process, and v({x1, x2, …, x i}) as the value of the data set {x1, x2, …, x i} composed of i data individuals x1, x2, …, x i . In the data sampling space, first sample the first data individual x1 that maximizes v({x1}) by using the greedy method or the global optimization method; then sample the second data individual x2 that maximizes v({x1, x2}) - v({x1}) by using the greedy method or the global optimization method; if the set {x1, x2} does not meet the target task performance requirement, continue to sample the third data individual x3 that maximizes v({x1, x2, x3}) - v({x1, x2}) by using the greedy method or the global optimization method, and iterate the sampling until the target task performance requirement is met.
[0063] This embodiment uses the first method, sampling with a given sampling quantity, which is 50 in this case. This embodiment employs a greedy algorithm as the maximization optimization method, sampling 50 high-value data sets from the 600 finite difference data sets in step 1. Note that the evaluation model for the value of individual data points can be generalized to all potential sampleable data points; therefore, the sampling of the high-value data set is not limited to the 600 sampleable data sets used to establish the evaluation model for assessing the value of individual data points, but can also be sampled from all potential sampleable data points.
[0064] Step 5: The 50 sets of sampled finite difference data are used to guide the sampling of the finite element analysis software, thereby realizing the modeling task of composite material curing process. For the 50 sets of curing process parameters in the sampled high-value dataset, the corresponding thermal hysteresis is simulated using the finite element analysis software, that is, 50 sets of high-value finite element data are obtained. Then, the mapping model between the curing process parameters and the thermal hysteresis obtained by the finite element analysis software simulation is fitted by the Gaussian process regression method to complete the target task.
[0065] Task Results: To illustrate the value assessment of the dataset and the effectiveness and stability of the sampling method, except for the 600 sets of finite difference data fixed in step 1, the above steps were repeated 5 times, using 500 sets of finite difference data as the test dataset. The mean absolute error between the predicted values of the established mapping model and the simulation results of the finite element analysis software was used as the evaluation index. The task results of the 5 repeated experiments are shown in the table below:
[0066] Experiment number 1 2 3 4 5 Mean absolute error / K 4.92 4.99 4.85 4.59 5.01
[0067] Furthermore, to further illustrate the advantages of the data set value assessment method and high-value data set sampling method described in this invention, in the task of modeling composite material curing processes, the method was compared with a random sampling method and a sampling method that maximizes the sum of individual data values (referred to as individual value sampling) on 181 sets of sampling experiments with the number of sampled data ranging from 20 to 200 at intervals of 1. The mean absolute error was used as the evaluation index. Each set of experiments was repeated 10 times, and the average of the 10 results was taken as the final result. The results of the method and the random sampling method are compared as follows: Figure 2 As shown, the method is compared with the results of individual value sampling, for example... Figure 3 As shown in the figure, some of the results are shown in the table below.
[0068] Number of sampled data 25 50 75 100 125 150 175 200 Random sampling method / K 11.29 7.84 6.58 6.37 5.63 5.29 4.81 4.72 Individual value sampling / K 17.25 8.59 5.67 5.00 4.77 4.48 4.13 3.92 The method / K 5.96 4.98 4.68 4.50 4.30 4.13 3.97 3.83
[0069] The parts not covered in this invention are the same as those in the prior art and are implemented using existing technologies.
Claims
1. A method for evaluating the value of a dataset, characterized in that, Includes the following steps: First, an evaluation model is established to assess the value of individual data points, and a function is established to describe the degree of value redundancy among individual data points. Then, a value assessment model for the data set is constructed by comprehensively considering the value of individual data points and the degree of value redundancy among individual data points. One of the following methods can be used to establish a model for evaluating the value of individual data: The value of individual data points is assessed by calculating the gain of each data point on the target task, and then an assessment model for evaluating the value of individual data points is established. The value of individual data is evaluated by calculating the gain of individual data in scenarios similar to the target task, and then an evaluation model for evaluating the value of individual data is established. Assess the value of individual data points based on domain knowledge of the data generation context, and then establish an assessment model for evaluating the value of individual data points. type; The objective is to model the curing process of composite materials and establish a mapping model from single-insulation curing process parameters to thermal hysteresis. The single-insulation curing process parameters include heating rate, cooling rate, insulation temperature, insulation time, air-part convective heat transfer coefficient, and air-mold convective heat transfer coefficient. The data points are potential sampleable data points within the range of single-insulation curing process parameters. The function describing the degree of value redundancy among individual data items is established using one of the following methods: The degree of redundancy between individual data is inversely proportional to the distance between them. A set of data that are close to each other generates greater redundancy value. The distance includes Euclidean distance and Mahalanobis distance. The degree of redundancy between individual data is directly proportional to the correlation between them. A set of data with greater correlation generates greater redundancy value. The correlation can be represented by kernel functions and membership functions.
2. The method according to claim 1, characterized in that, The gain calculation method is obtained by calculating the Shapley value of the data individual for the target task.
3. The method according to claim 1, characterized in that, The value assessment model for the aforementioned dataset is calculated using the following method: ;v′(x,S)=v(x)max{k(x,x1),…,k(x,x m )},x1,…,x m ∈S Where v(S) is the value assessment model of the dataset, n is the number of potential data points in the sample space, N is the dataset consisting of n potential data points in the sample space, S is a subset of the dataset N, v(x) is the value assessment model of individual data points, and k(x,x) is the value assessment model of individual data points. i Let be a function describing the degree of value redundancy among individual data points, i = 1, 2, ..., m, where m is the number of data points in S, 0 ≤ i ≤ 1, 2, ..., m, and m is the number of data points in S. <m≤n。 4. The method according to claim 1, characterized in that, The function describing the degree of value redundancy among individual data points can take one of the following forms: Gaussian kernel function: ; Laplace kernel function: Inverse quadratic kernel function: In the formula x i Let x represent the i-th data individual. j Let represent the j-th data individual, σ be the parameter controlling the size of the Gaussian kernel function, τ be the parameter controlling the size of the Laplace kernel function, and c be the parameter controlling the size of the inverse multivariate quadratic kernel function.
5. A sampling method for a high-value dataset, characterized in that, Based on the data set value assessment model described in claim 1, high-value data sets are sampled from the data sampling space according to user needs.
6. The method according to claim 5, characterized in that, The method of sampling high-value data sets from the data sampling space according to user needs includes the following: Given a sample size p, based on the value assessment model of the aforementioned dataset, let v(S) p Let S represent a set consisting of p data individuals. p The value is determined by sampling from the data sampling space using a greedy method or a global optimization method, which yields v(S) such that v(S) = ... p The largest dataset; Given the target task performance requirements, based on the value assessment model of the aforementioned dataset, let x be... i Let v({x1,x2,…,x...) represent the i-th data instance in the sampling process. i }) represents i data individuals x1, x2, ..., x i The data set {x1,x2,…,x} is composed of i The value of} is determined by sampling the data within the data sampling space. First, the first data individual x1 that maximizes v({x1}) is sampled using a greedy method or a global optimization method. Then, the second data individual x2 that maximizes v({x1,x2}) - v({x1}) is sampled using a greedy method or a global optimization method. If the set {x1,x2} does not meet the target task performance requirements, the third data individual x3 that maximizes v({x1,x2,x3}) - v({x1,x2}) is sampled using a greedy method or a global optimization method. This process is repeated iteratively until the target task performance requirements are met.