Training data evaluation method and device and electronic equipment
By adjusting the parameters of the initial evaluation model and combining the processing confidence of the sub-service model, the problem of time and calculation overhead in the training data value evaluation process in the prior art is solved, and a fast and accurate data value evaluation is achieved.
Patent Information
- Application Number
- CN202311595795.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-27
AI Technical Summary
When evaluating the value of training data, the prior art requires a large number of retraining models and calculation accuracy differences, resulting in excessive time and calculation overhead.
By obtaining the initial evaluation model to be trained and the second training data, the parameters of the initial evaluation model are adjusted to obtain the evaluation model for evaluating the value of the first training data. The method includes calculating the processing confidence of the sub-service model for the second training data and adjusting the model parameters based on the output evaluation results and confidence.
It realizes the rapid and accurate evaluation of the value of each first training data to the training service model, without the need for a large number of retraining the model and computing accuracy differences, saving time, computing and storage overhead.
Smart Images

Figure CN120045932A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method for evaluating training data, a method for training an evaluation model, a device, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the rapid development of deep learning, machine learning models are becoming more and more common in modern science and business. Since the amount of training data used for model training is very large, the training data used to train the model is usually provided by multiple data suppliers. Since the quality of data provided by different suppliers may vary greatly, in order to fairly allocate reasonable remuneration to different suppliers and to urge suppliers to provide better quality training data, the training data provided by the suppliers can be evaluated for value.
[0003] When evaluating the value of training data, related technologies can use combinations of different samples to retrain the model, remove the training data to be evaluated in the training data, use the remaining training data to retrain the model to obtain a retrained model, and calculate the value of the training data to be evaluated based on the difference in prediction accuracy between the retrained model and the model trained using all the training data.
[0004] However, in practice, since the training data includes a large amount of data, when using the above method to evaluate the value of each training data, it is necessary to retrain the model a large number of times. Each recalculation also requires calculating the accuracy difference of different models, resulting in a lot of time and computing overhead for evaluating the value of training data. Summary of the invention
[0005] The present application provides a training data evaluation method, an evaluation model training method, an apparatus, an electronic device, and a computer-readable storage medium, which can quickly evaluate the value of each training data to save the computational overhead in the process of evaluating the value of the training data. The specific method is as follows.
[0006] In a first aspect, an embodiment of the present application provides a method for training an evaluation model, wherein the evaluation model is used to evaluate the value of first training data for training a service model, and the method includes:
[0007] Obtaining an initial evaluation model to be trained and second training data, where the second training data has the same distribution as the first training data;
[0008] Inputting the second training data into the initial evaluation model to obtain an output evaluation result;
[0009] Calculating the processing confidence of the sub-service model on the second training data, the sub-service model is obtained by training the initial service model using the data in the sampling subset, the initial service model is the initial service model to be trained corresponding to the service model, and the sampling subset is a data set sampled from the first training data;
[0010] According to the output evaluation result and the processing confidence, the parameters of the initial evaluation model are adjusted to obtain an evaluation model for evaluating the value of the first training data for training the service model.
[0011] In a second aspect, an embodiment of the present application provides a method for evaluating training data, the method being used to evaluate the value of each first training data of a training service model, the method comprising:
[0012] Acquire first training data for training the service model;
[0013] Use a pre-trained evaluation model to evaluate the value of training the service model with each first training data to obtain an evaluation result, and the evaluation model is trained according to the training method described in any one of the first aspects.
[0014] In a third aspect, an embodiment of the present application further provides a training device for an evaluation model, wherein the evaluation model is used to evaluate the value of first training data for training a service model, and the device includes:
[0015] A first acquisition unit, used to acquire an initial evaluation model to be trained and second training data, where the second training data has the same distribution as the first training data;
[0016] An output unit, used for inputting the second training data into the initial evaluation model to obtain an output evaluation result;
[0017] a calculation unit, configured to calculate a processing confidence of a sub-service model on the second training data, wherein the sub-service model is obtained by training an initial service model using data in a sampling subset, the initial service model being an initial service model to be trained corresponding to the service model, and the sampling subset being a data set sampled from the first training data;
[0018] An adjustment unit is used to adjust the parameters of the initial evaluation model according to the output evaluation result and the processing confidence level to obtain an evaluation model for evaluating the value of the first training data for training the service model.
[0019] In a fourth aspect, an embodiment of the present application further provides an electronic device, including:
[0020] Processor; and
[0021] The memory is used to store a data processing program. After the electronic device is powered on and the program is run by the processor, the method described in any one of the first aspect or the second aspect is executed.
[0022] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium storing a data processing program, which is executed by a processor to perform the method described in any one of the first aspect or the second aspect.
[0023] Compared with the prior art, this application has the following advantages:
[0024] The training method of the evaluation model provided in the embodiment of the present application obtains an initial evaluation model to be trained and second training data. Since the trained evaluation model is used to evaluate the value of the first training data for training the service model, the second training data has the same distribution as the first training data, so that the trained evaluation model can evaluate the value of the first training data with the same distribution as the second training data. The second training data is input into the above-mentioned initial evaluation model to output the evaluation result, and then the processing confidence of the sub-service model on the second training data is calculated. Since the sub-service model is obtained by training the initial service model using the data in the sampling subset, the initial service model is the initial service model to be trained corresponding to the above-mentioned service model, and the sampling subset is the data obtained from the first training data. That is, the sub-service model has the same model structure and the same initialized model parameters as the above-mentioned service model, and the sub-service model is trained by data in a subset of the first training data. Therefore, there is a corresponding relationship between the processing confidence of the sub-service model on the second training data and the value of each second training data estimated by the estimation model that meets the estimation requirements. Therefore, based on the processing confidence of the sub-service model on the second training data and the evaluation result output by the initial evaluation model, the parameters of the initial evaluation model can be adjusted so that the evaluation result output by the adjusted evaluation model and the confidence output by the sub-service model satisfy the above-mentioned corresponding relationship, thereby obtaining an evaluation model for evaluating the value of the first training data for training the service model.
[0025] It can be seen that the training method provided in the present application can train an evaluation model for performing value evaluation on each first training data of the service model. The model trained using the method provided in the present application can quickly and accurately evaluate the value of each first training data to the training service model, without the need for a large number of retraining models or recalculating the accuracy differences between the different models trained, thereby saving the time, computational and storage costs required for training the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1A schematic diagram of a flow chart of a training method for an evaluation model provided in an embodiment of the present application;
[0027] Figure 2 A schematic diagram of an example of a training algorithm for training an evaluation model in an embodiment of the present application;
[0028] Figure 3 A schematic diagram of another example of a training algorithm for training an evaluation model in an embodiment of the present application;
[0029] Figure 4 A schematic diagram of a process for predicting the value of training data through an evaluation model in an embodiment of the present application;
[0030] Figure 5 A flowchart of a method for evaluating training data provided in an embodiment of the present application;
[0031] Figure 6 A structural block diagram of a training device for an evaluation model provided in an embodiment of the present application;
[0032] Figure 7 A structural block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.
[0034] With the rapid development of deep learning, machine learning models are becoming more and more common in modern science and business. Since the amount of training data used for model training is very large, the training data used to train the model is usually provided by multiple data suppliers. Since the quality of data provided by different suppliers may vary greatly, in order to fairly allocate reasonable remuneration to different suppliers and to urge suppliers to provide better quality training data, the training data provided by the suppliers can be evaluated for value.
[0035] For example, the Shapley value method can be used to evaluate the Shapley value of each training data. The Shapley value method refers to the combination of different samples to train the model, and then the marginal contribution is used to calculate the Shapley value of each training sample. Shapley can reflect the marginal contribution of each training data to the training model, thereby reflecting the contribution value of each training data.
[0036] For another example, the training data to be evaluated may be removed from the training data, and the model may be retrained using the remaining training data to obtain a retrained model. The value of the training data to be evaluated may be calculated based on the difference in prediction accuracy between the retrained model and the model trained using all the training data.
[0037] However, in practice, since the training data includes a large amount of data, when using the above method to evaluate the value of each training data, it is necessary to retrain the model a large number of times. Each recalculation also requires calculating the accuracy difference of different models, resulting in a lot of time and computing overhead for evaluating the value of training data.
[0038] In order to solve the above problems in the related art, the embodiment of the present application provides a training method for an evaluation model. The execution subject of the solution provided by the present application can be an electronic device, which can be a desktop computer, a laptop computer, a mobile device, a smart watch, a smart TV, a tablet computer, a server, etc., or other devices with data processing and display functions.
[0039] The evaluation model trained by the evaluation model training method provided in this application can be used to evaluate the value of the first training data for training the service model. Those skilled in the art can set the specific representation for value evaluation according to actual needs. The service model can be a classification model, an image recognition model, a speech repair model, a text extraction model, an intelligent dialogue model, etc., or other models for providing services, which are not specifically limited in this application.
[0040] like Figures 1 to 3 As shown, the training method of the evaluation model provided in the embodiment of the present application includes the following steps S110 to S140.
[0041] Step S110: Acquire an initial evaluation model to be trained and second training data, where the second training data has the same distribution as the first training data.
[0042] The above initial evaluation model is used to train an evaluation model so as to use the evaluation model to evaluate the value of the first training data for training the service model.
[0043] In the embodiment of the present application, the second training data and the first training data can be used to train the same model, that is, the second training data and the first training data can both be used to train the service model, that is, if the first training data can train the service model, then the second training data can also train the service model. The second training data and the first training data can be the same data or different data. The sub-service model subsequently trained by a subset of the first training data can achieve functions consistent with the service model. For example, if the service model can recognize faces, then the sub-service model can also recognize faces. The difference is that the sub-service model is a model trained with less data, so the recognition accuracy of the sub-service model is usually lower than that of the service model.
[0044] like Figure 2 , Figure 3 In the model training algorithm shown in FIG. 1 , the second training data is the data in the set D, and the first training data is the service model training data. Since the service model needs to be used to filter and obtain the effective second training data later, the algorithm input also includes the service model f D , Figure 2 , Figure 3 The model training algorithm shown in can output the evaluation model φ θ (x, y).
[0045] Optionally, the second training data has the same distribution as the first training data, specifically, the second training data and the first training data both satisfy Gaussian distribution, or both satisfy uniform distribution, or both satisfy other distributions. When the second training data has the same distribution as the first training data, the evaluation model obtained by subsequent training can be better matched with the service model, so that the evaluation model trained based on the sub-service model can more accurately evaluate the contribution value of the first training data to the service model training.
[0046] In the embodiment of the present application, since the evaluation model and the service model are used to implement different functions, the evaluation model and the service model are set to different model structures, that is, the model structures of the initial service model to be trained corresponding to the above-mentioned initial evaluation model and the above-mentioned service model are different. In this way, the functional requirements of different models can be better met.
[0047] The second training data may be provided by a data provider, or may be extracted by an electronic device from a user's behavior record, which is not specifically limited in this application.
[0048] Step S120: input the second training data into the initial evaluation model to obtain an output evaluation result.
[0049] The output evaluation result can be understood as the initial evaluation result, and the output evaluation result can be an output Shapley value, or an output contribution rate, etc. Since the Shapley value can effectively reflect the value of the training data, the above-mentioned output evaluation result can be an output Shapley value, so that the trained evaluation model can also output the Shapley value of the first training data to the trained service model. In step S120, each second training data can be input into the initial evaluation model to obtain each output evaluation result corresponding to each second training data.
[0050] like Figure 2 , Figure 3 As shown, the initial evaluation model is Initializeφ θ (x, y).
[0051] Step S130: Calculate the confidence level of the sub-service model in processing the second training data.
[0052] The sub-service model is obtained by training the initial service model using data in a sampling subset, wherein the initial service model is the initial service model to be trained corresponding to the service model, and the sampling subset is a data set sampled from the first training data.
[0053] When the above-mentioned evaluation model is used to output the Shapley value of the service model trained by the first training data, the sampling of the sampling subset can satisfy the Shapley Kernel distribution for calculating the Shapley value. Among them, the Shapley Kernel distribution is a sampling method for calculating the Shapley value, which determines the sampling probability of each training sample based on the kernel function. Specifically, the Shapley Kernel distribution calculates the sampling probability by applying the kernel function to the distance (or similarity) between training samples. The closer the distance or the higher the similarity, the greater the probability of being sampled.
[0054] In the embodiment of the present application, the data in each sampling subset of the first training data can be used to train the initial service model to obtain a sub-service model.
[0055] Optionally, before step S130, the following steps S10 to S30 of training the sub-service model may also be included.
[0056] Step S10: sampling from the first training data to obtain a sampling subset.
[0057] Step S20: input the data in the sampling subset into the initial service model to obtain an initial sub-service result.
[0058] Step S30: Based on the initial sub-service result and the service truth value marked by the data in the sampling subset, the parameters of the initial service model are adjusted to obtain a sub-service model.
[0059] When performing parameter adjustment in step S30, a preset number of training iterations may be obtained, and based on the initial sub-service result and the service truth value marked by the data in the sampling subset, the gradient descent method may be used to adjust the parameters of the initial service model for the number of iterations to obtain a sub-service model. The number of training iterations may be set according to actual needs. In order to reduce training consumption, the number of training iterations may be set relatively small, so that the sub-service model obtained through the previous rounds of training information can estimate the truth value of the original converged sub-service model.
[0060] This training method can be used to train the sub-service model.
[0061] For example, if the first training data includes n data, then the first training data includes 2 n The initial service model can be trained using each sampling subset, and a total of 2 n To reduce the training intensity, the sampling subsets can be distributed through the Shapley Kernel distribution to obtain sampling subsets that satisfy the Shapley Kernel distribution, and the initial model can be trained using the sampling subsets that satisfy the Shapley Kernel distribution. For example, if there are m sampling subsets that satisfy the Shapley Kernel distribution, then m sub-service models can be obtained through m sampling subsets.
[0062] In the process of executing steps S10 to S30, sampling and training of the sub-service model can be stopped when the subsequent step S140 meets the convergence condition. That is, after the parameters of the initial evaluation model are adjusted in step S140, sampling of the sampling subset and training of the sub-service model can be stopped. When step S140 does not adjust the parameters of the initial evaluation model, when a new round of gradient descent evaluation model training is performed, the sampling subset can be obtained again through Shapley Kernel distribution sampling, and a new sub-service model can be trained, and the processing confidence of the new sub-service model for the second training data is calculated, so as to adjust the parameters of the untrained evaluation model through the newly calculated processing confidence until the training process of the evaluation model converges to obtain the evaluation model.
[0063] In step S130, when calculating the processing confidence of the sub-service model for the second training data, the processing confidence of the sub-service model trained with the sampling subset satisfying the Shapley Kernel distribution for each second training data may be calculated.
[0064] Each sub-service model can be pre-trained based on a sampling subset, or can be trained in the process of training an evaluation model, which is not specifically limited in this application.
[0065] When calculating the processing confidence of the sub-service model on the second training data, since the second training data is marked with the corresponding service truth value, step S130 can specifically input the second training data into the sub-service model to obtain the sub-service result, and determine the processing confidence of the sub-service model on the second training data based on the obtained sub-service result and the service truth value marked by the second training data.
[0066] Specifically, the processing confidence of the sub-service model on the second training data can be determined based on the difference between the service result of the sub-service model on the second training data and the service truth value. For example, the predicted correct second training data that is the same as the service truth value can be selected from the service result of the sub-service model on the second training data, and the processing confidence of the sub-service model on the second training data can be determined based on the ratio between the predicted correct second training data and the second training data. The processing confidence of the sub-service model on the second training data can be quickly and accurately determined through the sub-service result and the service truth value output by the sub-service model.
[0067] Optionally, the above-mentioned calculation of the processing confidence of the sub-service model for the second training data can also be implemented in the following manner: determine the score of the sub-service model predicting the second training data corresponding to each prediction category (the score can be understood as the above-mentioned sub-service result), and based on the above-mentioned score and the service true value marked by the second training data, determine the normalized score of the correct prediction of the sub-service model for the second training data, and determine the normalized score as the processing confidence of the sub-service model for the second training data.
[0068] In one implementation, before step S120, the training method of the evaluation model may further include the following steps S110a to S110b.
[0069] Step S110a: Use the service model to process the second training data to obtain a service result.
[0070] Specifically, each second training data can be input into the above service model to obtain each service result corresponding to each second training data. Since the service model is a trained model, the service result output by the service model can be understood as the real service result corresponding to the second training data.
[0071] Step S110b: Select valid second training data from the second training data, the service result obtained by processing being the same as the marked service result.
[0072] Accordingly, the above step S120 can be implemented by the following steps: inputting the effective second training data in the second training data into the initial evaluation model. The above step S130 can be implemented by the following steps: calculating the processing confidence of the sub-service model on the effective second training data in the second training data.
[0073] Since the second training data is marked with corresponding service true values, when the service result of the second training data obtained through the service model is different from the service true value marked by the second data, it means that the service true value marked by the second training data may be inaccurate. Therefore, this part of the inaccurately marked second training data can be deleted, and the remaining valid second training data can be used to train the evaluation model to improve the efficiency and accuracy of model training.
[0074] In a specific embodiment, the processing confidence of the sub-service model on the second training data can also be understood as the utility of the sub-service model on the second training data.
[0075] Exemplarily, the confidence level of the sub-service model in processing the second training data may be calculated using the following formula (1).
[0076] v x,y (s)=link(f D (x),y)σ(f (D,s) (x) y ,
[0077]
[0078] In formula (1), v x,y (s) represents the confidence of the sub-service model in processing the second training data, link(f D (x), y) represents the above-mentioned effective second training data, f D (x) represents the above service model, and the score function σ(f (D,s) (x) y represents f (D,s) (x) softmax score on label y, f (D,s) (x) represents the service result of processing the second training sample x using the sub-service model trained with the sampling subset S, y represents the service truth value marked by the second training sample x, arg maxf D (x) represents the service result output by the service model for the second training data x.
[0079] Formula (1) represents the confidence of the sub-service model in processing the second training data through the softmax score. The Softmax function can convert the original score into a probability value, so that the probability of each category can reflect its relative importance in the whole.
[0080] Step S140: According to the output evaluation result and the processing confidence, the parameters of the initial evaluation model are adjusted to obtain an evaluation model for evaluating the value of the first training data for training the service model.
[0081] Specifically, step S140 can adjust the parameters of the initial evaluation model according to the difference between the output evaluation result corresponding to the second training data and the processing confidence corresponding to the second training data. Specifically, the parameters of the initial evaluation model can be adjusted based on the principle of reducing the difference between the output evaluation result corresponding to the second training data and the processing confidence corresponding to the second training data, until the difference between the output evaluation result corresponding to the second training data and the processing confidence corresponding to the second training data is less than the preset difference, and the evaluation model is obtained. Alternatively, the average output evaluation result of the output evaluation results corresponding to each second training data and the average processing confidence of the processing confidence corresponding to each second training data can be calculated, and the parameters of the initial evaluation model can be adjusted according to the average output evaluation result and the average processing confidence to obtain the above evaluation model.
[0082] Optionally, step S140 can be implemented according to the following steps: calculating the sum of the output evaluation results corresponding to each valid second training data, and the sum of the processing confidences corresponding to each valid second training data; adjusting the parameters of the initial evaluation model according to the sum of the output evaluation results and the sum of the processing confidences.
[0083] Specifically, the above output evaluation results can be corrected according to the output evaluation results corresponding to the second training data, the processing confidences corresponding to the sub-service models, and the processing confidences corresponding to the initial service model, and the parameters of the initial evaluation model can be adjusted based on the output evaluation results. Among them, the corrected output evaluation result is: initial output evaluation result + (processing confidence of the service model - processing confidence of the initial service model - sum of initial output evaluation results) ÷ number of first training data. The processing confidence corresponding to the initial service model is the processing confidence of the initial service model for the second training data. The calculation method of the processing confidence of the initial service model can refer to the processing confidence of the sub-service model, which will not be described in detail here. The processing confidence of the service model is the processing confidence of the service model for each second training data.
[0084] The processing confidence corresponding to the sub-service model is the processing confidence of the sub-service model for the second training data. For example, if there are m second training data, then the m processing confidences of the sub-service model for the m second training data can be determined. When adjusting the parameters of the initial evaluation model, the parameters of the initial evaluation model can be adjusted by the gradient descent method based on the difference between the sum of the output evaluation results and the sum of the processing confidences corresponding to the sub-service models, with the principle of reducing the difference, until the difference between the sum of the output evaluation results and the sum of the processing confidences corresponding to the sub-service models is less than the preset difference, thereby obtaining an evaluation model for evaluating the value of the first training data for training the service model.
[0085] Exemplarily, when the evaluation model is used to output the Shapley value of the first training data for the trained service model, the idea of KernelSHAP can be used to express the Shapley value of the second training data output by the initial evaluation model as a solution to a constrained weighted least squares optimization problem, thereby obtaining a loss function to train the initial evaluation model using the loss function. KernelSHAP is a method for explaining the prediction results of machine learning models, which is based on the SHAP (SHapley Additive exPlanations) framework and uses a kernel function to estimate the Shapley value.
[0086] Specifically, the solution of the constrained weighted least squares optimization problem corresponding to the Shapley value of the second training data can be expressed by the following formula (2).
[0087]
[0088] In formula (2), represents the solution of the constrained weighted least squares optimization problem corresponding to the Shapley value of the second training data, v x,y (s) represents the processing confidence of the sub-service model on the second training data, represents the confidence of the initial service model in processing the second training data, s T φ x,y represents the evaluation result of the initial evaluation model on the second training data output, It means finding the optimal φx, y to minimize the expected value of the following formula, E p(s) represents the expectation of the p(s) distribution, Right now represents the confidence of the service model in processing the second training data, n is the number of second training data, Indicates the weight selection method of the collection subset selected from the first training data, the denominator It represents the number of combinations of selecting collection subsets from n first training data multiplied by the size of the collection subset multiplied by n minus the size of the subset.
[0089] The above loss function can be expressed by the following formula (3).
[0090]
[0091] In formula (3), L(θ) represents the loss function, E x~X Indicates that the distribution of the second training data of the training evaluation model is the same as that of the first training data, y~U(Y) indicates that the service truth value marked by the second training data of the training evaluation model satisfies the average distribution, s~P(s) indicates that the training sample sampling of the sub-service model must satisfy the Shapley Kernel distribution P(s), Represents an evaluation model, the input of the evaluation model is the prediction sample (i.e., the second training data) and the prediction result of the service model for the prediction sample, and the output of the evaluation model is the evaluation result for the n first training data.
[0092] The above loss function L(θ) is Figure 2 , Figure 3 The loss L in , and finally get the adjusted parameter θ.
[0093] In a specific embodiment, when step S30 is based on a pre-set number of training iterations when performing parameter adjustment, the processing confidence of the sub-service model on the second training data can be calculated by the following formula (4), and the above loss function can be expressed by the following formula (5).
[0094]
[0095] In formula (4), v x,y,K (s) represents the confidence of the sub-service model trained by the iteration number K on the second training data. The score function express The softmax score on label y, represents the i-th sub-service model trained using the subset (D, s) of the second training data, D represents the entire training data set of the second training data, and s represents the subset selected from the training set D.
[0096]
[0097] In formula (5), L(θ,K) represents the loss function, v x,y,K (s) represents the processing confidence of the sub-service model trained based on the iteration number K on the second training data.
[0098] In one implementation, the above step S140 can be implemented according to the following steps S141 to S147.
[0099] Step S141: adjusting the parameters of the initial evaluation model according to the output evaluation result and the processing confidence, to obtain a first adjusted evaluation model.
[0100] The process of adjusting the evaluation model in step S141 can refer to the process of adjusting parameters in step S140, which will not be repeated here. The adjusted evaluation model obtained in step S141 can be understood as the model obtained after the initial evaluation model is trained once.
[0101] Step S142: set i=1;
[0102] Step S143: when the evaluation model after the i-th adjustment does not meet the convergence condition, resampling is performed from the first training data to obtain the i-th updated sampling subset.
[0103] The above-mentioned convergence conditions may be that the number of iterative training of the initial evaluation model reaches a preset number, the difference between the output evaluation result and the processing confidence is less than a preset difference, etc., but are not limited thereto.
[0104] The i-th updated sampling subset obtained by resampling is different from the i-1-th updated sampling subset. The 0-th updated sampling subset is the sampling subset in step S130. The subsets resampled each time can satisfy the same sampling distribution. For example, the i-th updated sampling subset and the i-1-th updated sampling subset are both sampled through the Shapley Kernel distribution.
[0105] Step S144: adjusting the parameters of the initial service model based on the data in the i-th updated sampling subset to obtain the i-th updated sub-service model.
[0106] The process of obtaining the ith updated sub-service model in step S144 is similar to the process of obtaining the sub-service model in steps S20 and S30, and will not be described in detail here.
[0107] Step S145: Calculate the confidence of the i-th update processing of the second training data by the i-th updated sub-service model.
[0108] The process of calculating the i-th update processing confidence in step S145 is similar to the process of calculating the processing confidence in step S130, and will not be repeated here.
[0109] Step S146: input the second training data into the evaluation model after the i-th adjustment to obtain the output evaluation result after the i-th adjustment.
[0110] Step S147: According to the i-th adjusted output evaluation result and the i-th updated processing confidence, adjust the parameters of the i-th adjusted evaluation model to obtain the i+1-th adjusted evaluation model, so that i=i+1, and execute steps S143 to S147 until the i-th adjusted evaluation model meets the convergence condition, and then determine the i-th adjusted evaluation model as the evaluation model for evaluating the value of the first training data for training the service model.
[0111] The process of adjusting the parameters of the evaluation model after the i-th adjustment in step S147 can refer to the relevant description in step S140, which will not be described in detail here.
[0112] In the process of adjusting the initial evaluation model through the iterative method, this embodiment simultaneously performs resampling of the sampling subset and synchronous updating of the sub-service model. The sub-service model can be flexibly trained based on whether the evaluation model converges, thereby training the sub-service model according to the training requirements of the evaluation model, thereby improving the training efficiency of the evaluation model. Synchronous training can also make the trained evaluation model more accurate.
[0113] In one implementation, when the evaluation model is used to output the Shapley value of the service model trained by the first training data, before step S140, the method may further include the following step S120a.
[0114] Step S130a: Correct the output evaluation results corresponding to each second training data to obtain a corrected output evaluation result, so that the corrected output evaluation result satisfies the validity of the Shapley value.
[0115] Specifically, the output evaluation result corresponding to the second training data can be corrected based on the principle that the corrected output evaluation result corresponding to the second training data satisfies the total utility value of the second training data, wherein the total utility value of the second training data refers to the processing confidence of the above service model on the second training data.
[0116] Correspondingly, step S140 may adjust the parameters of the initial evaluation model according to the corrected output evaluation result and the processing confidence.
[0117] This embodiment adjusts the effectiveness of the Shapley value of the output evaluation result so that the evaluation result output by the initial evaluation model meets the characteristic requirements of the Shapley value, thereby enabling the trained evaluation model to evaluate the Shapley value of the first training data.
[0118] Optionally, the output evaluation result corresponding to each second training data may be corrected by the following formula (6).
[0119]
[0120] In formula (6), φ before ← represents the evaluation result of the evaluation model. Indicates the processing confidence of the service model, Indicates the processing confidence of the initial serving model.
[0121] like Figure 2 , Figure 3 The following is the algorithm logic for evaluating model training. Figure 3 Relative to Figure 2 The difference is that Figure 3 The initial service model is iterated K times to adjust the parameters to reduce the time and system resources consumed in training the sub-service model.
[0122] The training method of the evaluation model provided in the embodiment of the present application obtains an initial evaluation model to be trained and second training data. Since the trained evaluation model is used to evaluate the value of the first training data for training the service model, the second training data has the same distribution as the first training data, so that the trained evaluation model can evaluate the value of the first training data with the same distribution as the second training data. The second training data is input into the above-mentioned initial evaluation model to output the evaluation result, and then the processing confidence of the sub-service model on the second training data is calculated. Since the sub-service model is obtained by training the initial service model using the data in the sampling subset, the initial service model is the initial service model to be trained corresponding to the above-mentioned service model, and the sampling subset is the data sampled from the first training data. That is, the sub-service model has the same model structure and the same initialized model parameters as the above-mentioned service model, and the sub-service model is trained by data in a subset of the first training data. Therefore, there is a certain corresponding relationship between the processing confidence of the sub-service model on the second training data and the value of each second training data estimated by the estimation model that meets the estimation requirements. Therefore, based on the processing confidence of the sub-service model on the second training data and the evaluation result output by the initial evaluation model, the parameters of the initial evaluation model can be adjusted so that the evaluation result output by the adjusted evaluation model and the confidence output by the sub-service model satisfy the above-mentioned corresponding relationship, thereby obtaining an evaluation model for evaluating the value of the first training data for training the service model.
[0123] It can be seen that the training method provided in the present application can train an evaluation model for performing value evaluation on each first training data of the service model. The model trained using the method provided in the present application can quickly and accurately evaluate the value of each first training data to the training service model, without the need for a large number of retraining models or recalculating the accuracy differences between the different models trained, thereby saving the time, computational and storage costs required for training the model.
[0124] The present application also provides a method for evaluating training data, the method being used to evaluate the value of each first training data of a training service model, such as Figure 4 , Figure 5 As shown, the method includes the following steps S210 to S220.
[0125] Step S210: Acquire first training data for training the service model.
[0126] Step S220: Use a pre-trained evaluation model to evaluate the value of training the service model with each first training data to obtain an evaluation result, wherein the evaluation model is trained according to the training method of the evaluation model described in any one of the above embodiments.
[0127] Exemplarily, each first training data for training the service model can be input into the evaluation model to obtain the evaluation results of each first training data. In order to improve the evaluation accuracy, the service data in the actual service process of the service model can also be input into the evaluation model to more accurately obtain the evaluation results of each first training data, and the service data includes the service request data input by the user and the service result data of the service model.
[0128] Optionally, step S210 can be implemented according to the following steps S211 to S212.
[0129] Using a pre-trained evaluation model to evaluate the Shapley value of each of the first training data for training the service model;
[0130] Obtaining a preset total value of first training data;
[0131] The values corresponding to each of the first training data are determined according to the total value and the Shapley values corresponding to each of the first training data.
[0132] In the present application embodiment, Figure 4As shown, when it is necessary to conduct a value assessment on the first training data for training a service model, if a customer requests the service model to output a service result, the service data corresponding to the user request (including the service request and the corresponding service result) can be input into the explanation model (i.e., the evaluation model of the present application), and the training set of the training service model is also input into the explanation model, so that the explanation model outputs the value of each first training data in the training set to the training service model.
[0133] Corresponding to the training method of the evaluation model provided in the embodiment of the present application, the embodiment of the present application also provides a training device for the evaluation model, wherein the evaluation model is used to evaluate the value of the first training data for training the service model, such as Figure 6 As shown, the device comprises:
[0134] A first acquisition unit 310 is used to acquire an initial evaluation model to be trained and second training data, where the second training data has the same distribution as the first training data;
[0135] An output unit 320, configured to input the second training data into the initial evaluation model to obtain an output evaluation result;
[0136] a calculation unit 330, configured to calculate a processing confidence of a sub-service model on the second training data, wherein the sub-service model is obtained by training an initial service model using data in a sampling subset, wherein the initial service model is an initial service model to be trained corresponding to the service model, and the sampling subset is a data set sampled from the first training data;
[0137] The adjustment unit 340 is used to adjust the parameters of the initial evaluation model according to the output evaluation result and the processing confidence level to obtain an evaluation model for evaluating the value of the first training data for training the service model.
[0138] Optionally, the computing unit is specifically used to input the second training data into the sub-service model to obtain a sub-service result; based on the sub-service result and the service truth value marked by the second training data, determine the processing confidence of the sub-service model on the data in the sampling subset.
[0139] Optionally, the device further comprises:
[0140] A sampling unit, used for sampling from the first training data to obtain a sampling subset;
[0141] The second adjustment unit is used to input the data in the sampling subset into the initial service model to obtain an initial sub-service result; based on the initial sub-service result and the service truth value marked by the data in the sampling subset, the initial service model is adjusted to obtain a sub-service model.
[0142] Optionally, the adjustment unit is specifically used to obtain a preset number of training iterations; based on the initial sub-service result and the service true value marked by the data in the sampling subset, the gradient descent method is used to adjust the parameters of the initial service model by the number of iterations to obtain a sub-service model.
[0143] Optionally, the device further comprises:
[0144] A selection unit is used to process the second training data using the service model to obtain a service result; and select valid second training data from the second training data, wherein the service result obtained by processing is the same as the marked service result;
[0145] The step of inputting the second training data into the initial evaluation model comprises:
[0146] inputting the valid second training data in the second training data into the initial evaluation model;
[0147] The calculating sub-service model's processing confidence of the second training data includes:
[0148] The confidence level of the sub-service model in processing the valid second training data in the second training data is calculated.
[0149] Optionally, the evaluation model is used to output a Shapley value of a service model trained by the first training data, and the sampling of the sampling subset satisfies a Shapley Kernel distribution for calculating the Shapley value.
[0150] Optionally, based on the output evaluation results and the processing confidences, the adjustment unit is specifically used to: calculate the sum of the output evaluation results corresponding to each of the valid second training data, and the sum of the processing confidences corresponding to each of the valid second training data; and adjust the parameters of the initial evaluation model based on the sum of the output evaluation results and the sum of the processing confidences.
[0151] Optionally, an embodiment of the present application further provides a training data evaluation device, the device is used to evaluate the value of each first training data of the training service model, the device includes:
[0152] A second acquisition unit, used to acquire first training data for training the service model;
[0153] An evaluation unit is used to evaluate the value of training the service model with each first training data using a pre-trained evaluation model to obtain an evaluation result, wherein the evaluation model is trained according to the training method according to any one of claims 1 to 7.
[0154] Corresponding to the training method of the evaluation model provided in the embodiment of the present application, the embodiment of the present application also provides an electronic device. Figure 7 As shown, the electronic device includes: a processor 401; and a memory 402, which is used to store a program of a training method for an evaluation model. After the electronic device is powered on and the program of the training method for the evaluation model is run by the processor, the following steps are performed:
[0155] Obtaining an initial evaluation model to be trained and second training data, where the second training data has the same distribution as the first training data;
[0156] Inputting the second training data into the initial evaluation model to obtain an output evaluation result;
[0157] Calculating the processing confidence of the sub-service model on the second training data, the sub-service model is obtained by training the initial service model using the data in the sampling subset, the initial service model is the initial service model to be trained corresponding to the service model, and the sampling subset is a data set sampled from the first training data;
[0158] According to the output evaluation result and the processing confidence, the parameters of the initial evaluation model are adjusted to obtain an evaluation model for evaluating the value of the first training data for training the service model.
[0159] Corresponding to the training method of the evaluation model provided in the embodiment of the present application, the embodiment of the present application provides a computer-readable storage medium storing a program of the training method of the evaluation model, which is executed by a processor to perform the following steps:
[0160] Obtaining an initial evaluation model to be trained and second training data, where the second training data has the same distribution as the first training data;
[0161] Inputting the second training data into the initial evaluation model to obtain an output evaluation result;
[0162] Calculating the processing confidence of the sub-service model on the second training data, the sub-service model is obtained by training the initial service model using the data in the sampling subset, the initial service model is the initial service model to be trained corresponding to the service model, and the sampling subset is a data set sampled from the first training data;
[0163] According to the output evaluation result and the processing confidence, the parameters of the initial evaluation model are adjusted to obtain an evaluation model for evaluating the value of the first training data for training the service model.
[0164] It should be noted that for the detailed description of the device, electronic device and computer-readable storage medium provided in the embodiments of the present application, reference can be made to the relevant description of the method in the first embodiment of the present application, which will not be repeated here.
[0165] Although the present application is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
[0166] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0167] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0168] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), random access memory (RAM) of other attributes, read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage media or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer-readable media does not include non-transitory media such as modulated data signals and carrier waves.
[0169] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0170] Although the present application is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
Claims
1. A training method for evaluating the model, It is characterized in that The evaluation model is used to evaluate the value of first training data for training a service model, and the method includes: Obtaining an initial evaluation model to be trained and second training data, where the second training data has the same distribution as the first training data; Inputting the second training data into the initial evaluation model to obtain an output evaluation result; Calculating the processing confidence of the sub-service model on the second training data, the sub-service model is obtained by training the initial service model using the data in the sampling subset, the initial service model is the initial service model to be trained corresponding to the service model, and the sampling subset is a data set sampled from the first training data; According to the output evaluation result and the processing confidence, the parameters of the initial evaluation model are adjusted to obtain an evaluation model for evaluating the value of the first training data for training the service model.
2. The method according to claim 1, It is characterized in that The calculating sub-service model's processing confidence of the second training data includes: Inputting the second training data into the sub-service model to obtain a sub-service result; Based on the sub-service result and the service truth value marked by the second training data, the processing confidence of the sub-service model on the second training data is determined.
3. The method according to claim 2, It is characterized in that Before inputting the second training data into the sub-service model, the method further includes: Sampling from the first training data to obtain a sampling subset; Inputting the data in the sampling subset into the initial service model to obtain an initial sub-service result; Based on the initial sub-service result and the service truth value marked by the data in the sampling subset, the parameters of the initial service model are adjusted to obtain a sub-service model.
4. The method according to claim 3, It is characterized in that The step of adjusting parameters of the initial service model based on the initial sub-service result and the service truth value marked by the data in the sampling subset to obtain the sub-service model includes: Get the preset number of training iterations; Based on the initial sub-service result and the service truth value marked by the data in the sampling subset, the initial service model is adjusted for the number of iterations using a gradient descent method to obtain a sub-service model.
5. The method according to claim 4, It is characterized in that The adjusting the parameters of the initial evaluation model according to the output evaluation result and the processing confidence to obtain an evaluation model for evaluating the value of the first training data for training the service model includes: According to the output evaluation result and the processing confidence, adjusting the parameters of the initial evaluation model to obtain a first adjusted evaluation model; Let i=1; When the evaluation model after the i-th adjustment does not meet the convergence condition, resampling is performed from the first training data to obtain the i-th updated sampling subset; Adjusting the parameters of the initial service model based on the data in the i-th updated sampling subset to obtain an i-th updated sub-service model; Calculating the confidence of the i-th update processing of the second training data by the i-th update sub-service model; Inputting the second training data into the evaluation model after the i-th adjustment to obtain the i-th adjusted output evaluation result; According to the i-th adjusted output evaluation result and the i-th updated processing confidence, the parameters of the i-th adjusted evaluation model are adjusted to obtain the i+1-th adjusted evaluation model, so that i=i+1, and the step of resampling from the first training data to obtain the i-th updated sampling subset when the i-th adjusted evaluation model does not meet the convergence condition is performed, until the i-th adjusted evaluation model meets the convergence condition, the i-th adjusted evaluation model is determined as the evaluation model for evaluating the value of the first training data for training the service model.
6. The method according to claim 1, It is characterized in that Before inputting the second training data into the initial evaluation model, the method further includes: Processing the second training data using the service model to obtain a service result; Selecting valid second training data from the second training data, wherein the service result obtained by processing is the same as the marked service result; The step of inputting the second training data into the initial evaluation model comprises: inputting the valid second training data in the second training data into the initial evaluation model; The calculating sub-service model's processing confidence of the second training data includes: The confidence level of the sub-service model in processing the valid second training data in the second training data is calculated.
7. The method according to claim 1, It is characterized in that The evaluation model is used to output the Shapley value of the service model trained by the first training data, and the sampling of the sampling subset satisfies the Shapley Kernel distribution for calculating the Shapley value.
8. The method according to claim 6, It is characterized in that The adjusting the parameters of the initial evaluation model according to the output evaluation result and the processing confidence level includes: Calculating the sum of the output evaluation results corresponding to the valid second training data, and the sum of the processing confidences corresponding to the valid second training data; The parameters of the initial evaluation model are adjusted according to the sum of the output evaluation results and the sum of the processing confidences.
9. A method for evaluating training data. It is characterized in that The method is used to evaluate the value of each first training data for training a service model, and the method includes: Acquire first training data for training the service model; Use a pre-trained evaluation model to evaluate the value of training the service model with each first training data to obtain an evaluation result, wherein the evaluation model is trained according to the training method according to any one of claims 1 to 8.
10. The method according to claim 9, It is characterized in that The evaluation model is trained by the training method according to claim 6; The using of the pre-trained evaluation model to evaluate the value of training the service model with each first training data to obtain an evaluation result includes: Using a pre-trained evaluation model to evaluate the Shapley value of each of the first training data for training the service model; Obtaining a preset total value of first training data; The values corresponding to each of the first training data are determined according to the total value and the Shapley values corresponding to each of the first training data.
11. A training device for evaluating a model, It is characterized in that The evaluation model is used to evaluate the value of the first training data for training the service model, and the device includes: A first acquisition unit, used to acquire an initial evaluation model to be trained and second training data, where the second training data has the same distribution as the first training data; An output unit, used for inputting the second training data into the initial evaluation model to obtain an output evaluation result; a calculation unit, configured to calculate a processing confidence of a sub-service model on the second training data, wherein the sub-service model is obtained by training an initial service model using data in a sampling subset, the initial service model being an initial service model to be trained corresponding to the service model, and the sampling subset being a data set sampled from the first training data; An adjustment unit is used to adjust the parameters of the initial evaluation model according to the output evaluation result and the processing confidence level to obtain an evaluation model for evaluating the value of the first training data for training the service model.
12. A device for evaluating training data, It is characterized in that The device is used to evaluate the value of each first training data for training the service model, and the device includes: A second acquisition unit, used to acquire first training data for training the service model; An evaluation unit is used to evaluate the value of training the service model with each first training data using a pre-trained evaluation model to obtain an evaluation result, wherein the evaluation model is trained according to the training method according to any one of claims 1 to 8.
13. An electronic device, It is characterized in that include: processor; as well as The memory is used to store a data processing program. After the electronic device is powered on and the program is run by the processor, the method according to any one of claims 1 to 10 is executed.
14. A computer-readable storage medium, It is characterized in that A data processing program is stored, and the program is run by a processor to execute the method according to any one of claims 1 to 10.