Data value evaluation method and device
Through the combination of gradient internal product and Shapley value theory, the rapid, accurate and low-cost problems of data value evaluation by multiple data providers in large language models are solved, and efficient data value evaluation and dynamic adjustment are achieved.
Patent Information
- Application Number
- CN202510387747.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to quickly, accurately and at low cost to evaluate the data value of multiple data providers, especially in large language model scenarios, where traditional methods are computationally expensive and unstable.
Using a gradient product-based method, by calculating the gradient data product on the target data set and support set of large language models, combined with Shapley value theory, a fast Shapley value calculation method is constructed to evaluate the contribution of data providers.
It realizes the accurate and efficient evaluation of data value distribution without training the model, reduces the computational cost, and is suitable for the dynamic adjustment needs of multiple data providers.
Smart Images

Figure CN120295876A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of machine learning technology, and in particular, to a method and device for evaluating data value, a computer-readable storage medium, and a computing device. Background Art
[0002] The development of modern machine learning models highly depends on large-scale and high-quality training data. Such data not only shapes the evolution path of the model's capabilities but also determines the upper limit of its performance. With the continuous development of cutting-edge technologies such as Large Language Models (LLMs), the demand for large-scale and high-quality training data has become even more urgent. However, the quality of data on the Internet is uneven, and the cost of collecting and validating high-quality data is high. At the same time, the phenomenon of data silos restricts the data circulation efficiency and becomes an important bottleneck restricting the implementation of artificial intelligence applications.
[0003] To address these challenges, it is crucial to build a framework that can accurately measure the value of data from different sources. As an emerging technology, data valuation aims to evaluate and quantify the value of certain training data in a specific context or domain scenario. Especially in the case of coexistence of multiple data providers, how to quickly and accurately evaluate the unique contributions of each data provider has become a key issue that urgently needs to be solved in the era of large models. Summary of the Invention
[0004] Embodiments of this specification describe a method and device for evaluating data value, which can quickly, accurately, and at low cost evaluate the value of data provided by a data provider, thereby meeting higher requirements in practical applications.
[0005] According to a first aspect, a method for evaluating data value is provided. The method includes: receiving an evaluation request for the value of a target data set in a number of domains; determining first gradient data of a pre-trained large language model on the target data set by parsing the evaluation request; for any target domain among the number of domains, obtaining second gradient data of the large language model on a support set in the target domain; and determining a target value of the target data set in the target domain based on the inner product between the first gradient data and the second gradient data, and incorporating it into the processing result of the evaluation request.
[0006] In one embodiment, obtaining the second gradient data of the large language model on a support set in the target domain includes: reading the pre-computed second gradient data from memory.
[0007] In one embodiment, the first gradient data includes a first gradient sum value among the gradients of each sample in the target data set; the second gradient data includes a second gradient sum value among the gradients of each sample in the support set; wherein, determining the target value of the target data set in the target domain includes: calculating the inner product between the first gradient sum value and the second gradient sum value as the target value.
[0008] Further, in a specific embodiment, calculating the inner product between the first gradient sum value and the second gradient sum value as the target value includes: respectively performing dimensionality reduction processing on the first gradient sum value and the second gradient sum value; performing an inner product operation based on the results of the dimensionality reduction processing to obtain the target value.
[0009] According to a second aspect, there is provided a method for evaluating data value, where multiple data parties respectively hold multiple data sets. This method is applied to a computing platform and includes the following steps: obtaining gradient data of a pre-trained large language model on each of the multiple data sets; obtaining gradient data of the large language model on a support set in a predetermined domain; taking each of the data sets as a target data set to be evaluated, and calculating its corresponding Shapley value as its target value in the predetermined domain; the calculation of the Shapley value includes: for each combination of data sets involved, determining the benefit of the combination based on the inner product between the gradient data corresponding to the combination and the gradient data on the support set.
[0010] In one embodiment, obtaining gradient data of a pre-trained large language model on each of the multiple data sets includes: receiving multiple gradient data from the multiple data parties correspondingly, where any one of the gradient data is determined by the corresponding data party based on the large language model deployed locally by it and the data set held by it.
[0011] Further, in a specific embodiment, the computing platform is a trusted platform independent of the multiple data parties; wherein, obtaining gradient data of a pre-trained large language model on each of the multiple data sets includes: receiving the multiple data sets from the multiple data parties correspondingly; locally calculating the gradient data of the large language model on each of the multiple data sets.
[0012] In a specific embodiment, obtaining gradient data of the large language model on a support set in a predetermined domain includes: locally calculating the gradient data of the large language model on the support set; or, reading the pre-calculated gradient data of the support set from memory.
[0013] In a specific embodiment, for each combination of involved data sets, based on the inner product between the gradient data corresponding to the combination and the gradient data on the support set, the benefit of the combination is determined, including: determining the third gradient sum value between the gradient data corresponding to the combination; based on the inner product between the third gradient sum value and the second gradient sum value corresponding to the gradient data on the support set, determining the benefit of the combination.
[0014] Furthermore, in an example, the gradient data on each data set includes the gradient sum value between the gradients corresponding to each sample in the data set; wherein, determining the third gradient sum value between the gradient data corresponding to the combination includes: performing a summation process on the gradient sum values corresponding to each data set in the combination to obtain the third gradient sum value.
[0015] In an example, the gradient data on each data set includes the gradients of each sample in the data set; wherein, determining the third gradient sum value between the gradient data corresponding to the combination includes: performing a deduplication process on the gradients of all samples corresponding to the combination; performing a summation process on the gradients remaining after the deduplication process to obtain the third gradient sum value.
[0016] In an example, the gradient data on the support set includes the gradient sum value between the gradients corresponding to each sample in the support set, and this gradient sum value is used as the second gradient sum value.
[0017] In an example, determining the benefit of the combination includes: performing a dimensionality reduction process on the third gradient sum value and the second gradient sum value; based on the result of the dimensionality reduction process, performing an inner product operation to obtain the benefit.
[0018] In an embodiment, each sample in the multiple data sets and the support set includes a task instruction and a corresponding expected output.
[0019] According to a third aspect, an evaluation device for data value is provided. The evaluation device includes:
[0020] An evaluation request receiving unit configured to receive an evaluation request for the value of the target data set in several fields; an evaluation request parsing unit configured to determine the first gradient data of the pre-trained large language model on the target data set by parsing the evaluation request; a gradient data determining unit configured to, for any target field in the several fields, obtain the second gradient data of the large language model on the support set of the target field; a target value determining unit configured to determine the target value of the target data set in the target field based on the inner product between the first gradient data and the second gradient data, and classify it into the processing result of the evaluation request.
[0021] According to a fourth aspect, there is provided an apparatus for evaluating data value, where multiple data parties respectively hold multiple data sets. The apparatus is integrated into a computing platform and includes:
[0022] A first gradient acquisition unit configured to acquire gradient data of a pre-trained large language model on each of the multiple data sets. A second gradient acquisition unit configured to acquire gradient data of the large language model on a support set in a predetermined domain. A target value calculation unit configured to use each of the data sets as a target data set to be evaluated, and calculate its corresponding Shapley value as its target value in the predetermined domain; the calculation of the Shapley value includes: for each data set combination involved, determining the gain of the combination based on the inner product between the gradient data corresponding to the combination and the gradient data on the support set.
[0023] According to a fifth aspect, there is provided a computer-readable storage medium having a computer program stored thereon, which when executed on a computer causes the computer to execute the method provided in the first aspect or the second aspect.
[0024] According to a sixth aspect, there is provided a computing device including a memory and a processor, where the memory stores executable code, and when the processor executes the executable code, the method provided in the first aspect or the second aspect is implemented.
[0025] In summary, by using the above methods and apparatuses disclosed in the embodiments of the present specification and based on gradient inner product tracking, it is possible to accurately and efficiently estimate the data value distribution in different domains without training a model. For scenarios where multiple data providers coexist, in order to further meet the dynamic adjustment requirements, the gradient tracking mechanism is combined with the Shapley value theory to construct a fast Shapley value calculation method, and the additivity of gradients is used to effectively evaluate the contribution of each data provider. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0027] Figure 1 Schematic diagram of the implementation architecture of the data value evaluation method disclosed in the embodiments of the present specification;
[0028] Figure 2 Schematic diagram of the flow steps of the data value evaluation method disclosed in the embodiments of the present specification, where the method involves a single data party;
[0029] Figure 3 This is the second schematic diagram of the process steps of the data value evaluation method disclosed in the embodiments of this specification. This method involves multiple data parties;
[0030] Figure 4 This is one of the schematic diagrams of the functional structure of the data value evaluation device disclosed in the embodiments of this specification;
[0031] Figure 5 This is the second schematic diagram of the functional structure of the data value evaluation device disclosed in the embodiments of this specification. Detailed implementation manners
[0032] The following describes the solution provided in this specification with reference to the accompanying drawings.
[0033] As mentioned above, data valuation (or data value evaluation) is an emerging technology aimed at quantifying the potential value of data in different fields and scenarios. It should be noted that the embodiments of this specification focus on the data value that can be generated by the data to be evaluated when it is assumed to be used for fine-tuning a large language model (or simply referred to as a large model) in specific downstream tasks in a specific field.
[0034] The embodiments of this specification disclose an improved data value evaluation scheme (hereinafter or simply referred to as the improved scheme), which can complete the data value evaluation of a data set in different downstream task scenarios at low cost, efficiently, and accurately.
[0035] To help understanding, the inventive concept of the improved scheme proposed by the applicant will be introduced first, and then the implementation framework and process steps of the improved scheme will be introduced.
[0036] I. Traditional data valuation schemes
[0037] Traditional data valuation schemes are mainly divided into schemes based on marginal contribution and schemes based on data impact.
[0038] The schemes based on marginal contribution evaluate the value by quantifying the impact of different subsets of data on the model performance. For example, the Leave-One-Out (LOO) strategy observes the change in model performance after removing a single data source to measure the importance of this data source. Another example is the scheme driven by the Shapley value theory, which evaluates the unique contribution of a certain data source by calculating the weighted average of the performance differences when a specific data source is included and not included in all possible subsets.
[0039] Although marginal contribution-based schemes provide relatively accurate valuations, they are computationally expensive, especially when dealing with large-scale datasets in the context of LLMs. Such schemes typically require a cumbersome retraining and testing process for each data source. For n data providers (or data suppliers, data parties, etc.), calculating the Shapley value for one provider requires considering (2 n - 1) different subset combinations, which makes the calculation extremely complex and resource-intensive.
[0040] Data impact-based schemes aim to reduce the valuation cost by estimating the impact of data on the model. Such schemes either rely on high-cost federated training across multiple data parties or result in inaccurate and unstable valuation results. Especially in the case of coexisting multiple data sources, there may be problems of metric conflicts. Therefore, it is difficult to apply such schemes to large language models in multi-domain and multi-data source scenarios.
[0041] In summary, traditional schemes have problems such as high computational costs, unstable and inaccurate valuations.
[0042] II. Theoretical Basis of Gradient Inner Product in the Improved Scheme
[0043] Based on the above observations and analyses of traditional schemes, the applicant constructed a gradient-based value measurement algorithm framework. For the instruction fine-tuning scenario of LLMs, assume the data to be valued is The k downstream domains that need to be valued are T = {t1, t2,..., t k}, and for each target domain t i ∈ T, there is a corresponding support set (or direction set) whose scale is usually small. When a large model M θ is given, it is hoped to measure the data value of D θ in the target domain t to be measured valu using M i as a benchmark.
[0044] During the training process, assume that a sample s a is used for a certain iterative training. After this iterative training, the model parameters change from θ a to θ a+1 . Therefore, the influence size of s a on the support sample is the decrease in the loss of this support sample d j :
[0045] L(d j , θ a ) - L(d j , θ a+1) (1)
[0046] For equation (1), perform a first-order Taylor expansion in the θ dimension:
[0047]
[0048] When the model M θ uses s a to update the model parameters using Stochastic Gradient Descent (SGD), we have:
[0049]
[0050] where η is the learning rate. Therefore, the influence of s a on d j is finally transformed into:
[0051]
[0052] For all d within the support set j and all s valu in the data D to be estimated a sum the meta-training processes respectively, and the contribution of the influence on the data D to be estimated valu for the target domain is:
[0053]
[0054] That is:
[0055]
[0056] Or, let Let We can obtain:
[0057]
[0058] Furthermore, we can omit η and the subscript a in formula (8), thus obtaining
[0059]
[0060] Therefore, the final value evaluation calculation can be for the data D to be estimated valu at the gradient of M θ and the support data D i sup at M θThe inner product sum of the gradients on. It should be noted that the above derivation process uses the SGD optimizer. In fact, the Adaptive Moment Estimation (Adam) optimizer can also be used for derivation, and the derivation result is still to calculate the inner product sum of the gradients to measure the data value, but the specific calculation of the gradients is different.
[0061] In the above introduction of the improvement scheme, based on the theoretical basis of the gradient inner product. Further, based on this theoretical basis, two key data evaluation scenarios are mainly considered. The first scenario involves a single data party, and a strategy is proposed to quickly estimate the static value distribution of the target data in one or more domains; the second scenario involves multiple data parties, and the interaction between multiple parties needs to be considered, so a dynamic adjustment mechanism based on the approximation of the Shapley value is proposed.
[0062] The implementation architectures in the above two key scenarios can all refer to Figure 1 , and the computing platform can pre-compute the support set D of a given LLM in different domains sup The gradient data on, and cache it. Then, according to the valuation request (or value query request) for the data D to be valued initiated by one or more data parties valu , calculate the gradient data of the given LLM on D valu , and calculate the inner product between this gradient data and the pre-cached gradient data ( Figure 1 Shown as Dot in), so as to determine the valuation result (or valuation result) of D valu .
[0063] The above introduces examples of the implementation architectures in two key scenarios. Next, the implementation steps in the above two key scenarios will be introduced in turn.
[0064] Figure 2 Shows the method flow for data value evaluation in the single data party scenario. The execution subject of this method can be any device, platform, server, or device cluster with computing and processing capabilities, such as a computing platform, or a data trading platform, etc. Figure 2 The method flow shown in includes the following steps:
[0065] Step S210, receive the evaluation request for the value of the evaluation target data set in several domains.
[0066] It should be noted that hereinafter, D valu is still used to represent the target data set. That is to say, D valu includes n samples, where n is generally a value not less than 2, and the scale can be large or small. These n samples can be used for fine-tuning of the LLM in actual applications. Any sample s aincluding task instruction x a and expected output y a or sample feature x a and sample label y a Exemplarily, the task types involved in n samples may include: text generation tasks, text classification tasks, question answering tasks, text filling tasks, translation tasks, text summarization tasks, and inference tasks, etc.
[0067] Several in the text refer to one or more. Exemplarily, the above several fields may be different industry fields, such as the medical field, the financial field, the education field, the customer service industry, the manufacturing industry, etc. It should be understood that the fields in the text can be referred to as application fields and can be divided according to actual needs, not limited to the industry.
[0068] For the content in the above evaluation request, in one embodiment, Figure 2 the execution entity of the method is a trusted computing platform (or simply referred to as a trusted platform), where trust can be achieved at the technical level, or alternatively, the computing platform can be endorsed by an authoritative institution. Based on this, the evaluation request may include the target data set of the target data party.
[0069] In another embodiment, the evaluation request includes the first gradient data of the given LLM on the target data set. It should be noted that the "first" in "the first gradient data" and the "second" and other similar terms elsewhere in the text are used to distinguish similar things and do not have other limiting effects such as sorting. In addition, the first gradient data can be calculated locally by the target data party based on the target data set and the given LLM. In this way, it is possible to avoid providing the target data set to the computing platform and prevent privacy leakage.
[0070] Further, in a specific embodiment, the above first gradient data may include the gradient of the LLM on any sample s a In another specific embodiment, the above first gradient data may include the sum of the gradients of the LLM on all samples to be valued, or D valu The sum of the gradients is referred to as the first gradient sum value in the text. In this way, compared with sending the gradients of each sample, the privacy protection intensity can be further enhanced and the communication overhead can be reduced. It should be understood that the gradient refers to the rate of change of the loss function with respect to the model parameters, and for its specific calculation, reference can be made to the prior art and will not be elaborated here.
[0071] From the above, a data value evaluation request sent by the target data party can be received.
[0072] Step S220, by parsing the evaluation request, determine the first gradient data of the pre-trained large language model on the target data set.
[0073] In one implementation, by parsing the evaluation request, the target data set included therein can be obtained, thereby calculating the first gradient data of the LLM thereon. In another implementation, the first gradient data included in the evaluation request can be directly parsed out.
[0074] From the above, the first gradient data can be obtained. It can be understood that by parsing the evaluation request, the above-mentioned several fields for which the valuation is made can also be obtained.
[0075] Step S230, for any target field among the several fields, obtain the second gradient data of the large language model on the support set in this target field.
[0076] It should be noted that hereinafter, is still used to represent the support set under the specific field i. For the sake of simplicity, the superscript i may be omitted. That is to say, D sup includes m samples, where m is generally a value not less than 2 and the scale is not large. These m samples can generally be used to test the prediction performance of the fine-tuned LLM under the specific field i. Any one of the samples d t includes the task instruction x t and the expected output y t . In addition, the task types involved in these m samples can include: text generation tasks, text classification tasks, question answering tasks, text filling tasks, translation tasks, text summarization tasks, and reasoning tasks, etc.
[0077] For the execution of this step, in one implementation, the pre-calculated second gradient data can be read from the memory (such as Figure 1 the cache shown). In another implementation, the second gradient data can be calculated locally in real time.
[0078] On the other hand, the second gradient data can include the gradient of the LLM on any sample d t , or it can also include the sum of the gradients of the LLM on all support samples, or rather the support set D sup . In the text, this sum of gradients may be referred to as the second gradient sum value. It should be noted that the calculation method of the gradient of the LLM on the support sample d t is generally configured to be the same as the calculation method of the gradient of the LLM on the sample s a to be valued.
[0079] From the above, the second gradient data can be obtained. It should be noted that the same LLM, or rather the same model parameters, are used for the calculation of this second gradient data and the above-mentioned first gradient data, and moreover, no update of the model parameters is required, thereby greatly reducing the valuation cost.
[0080] Step S240, based on the inner product between the first gradient data and the second gradient data, determining the target value of the target data set in the target domain, and including it in the processing result of the evaluation request.
[0081] In one implementation, an inner product operation may be performed on the first gradient sum value corresponding to the first gradient data and the second gradient sum value corresponding to the second gradient data, and the operation result may be used as the target value. For this, reference may be made to the above formula (8).
[0082] In another embodiment, each sample gradient in the first gradient data and each sample gradient in the second gradient data may be subjected to ergodic pairwise inner products, and then all inner product results may be summed up, so that the summed result may be used as the target value. For this, reference may be made to the above formula (6).
[0083] It should be noted that the inner product result obtained above can be used directly as the target value, or it can be used as the target value after scaling or other processing. The key is to satisfy: the target value is positively correlated with the inner product result between gradients.
[0084] On the other hand, considering that the model parameters of LLM are huge, and correspondingly, the dimensions of the first gradient data and the second gradient data are huge, in order to reduce the amount of calculation of the inner product operation, it is proposed to perform dimensionality reduction processing on the first gradient data and the second gradient data respectively, and then perform the inner product operation after reducing them to the same dimension. Exemplarily, the computing platform can perform dimensionality reduction processing on the first gradient sum value and the second gradient sum value respectively, and then calculate the inner product between the two gradient sum values after the dimensionality reduction processing as the target value. It should be understood that the dimensionality reduction processing can also be pre-placed as needed. For example, the target data party includes the dimensionality reduction result of the first gradient sum value in the evaluation request and sends it to the computing platform, which can reduce communication costs. For another example, the computing platform caches the second gradient sum value after dimensionality reduction processing, which can reduce storage overhead.
[0085] On the other hand, the above mentioned evaluation of data value under a given set of model parameters can actually be performed multiple times under multiple given sets of model parameters, and the average data value can be obtained. For example, for LLM, in the pre-training stage, the training set can be traversed multiple times, so that the model parameters after each traversal of the training set (or each epoch) can be used as the given model parameters.
[0086] From the above, we can get the target value of the target dataset in the target field, and by analogy, we can get the value distribution of the target dataset in the above fields. Furthermore, this value distribution can be used as the processing result of the evaluation request and fed back to the target data party who initiated the evaluation request.
[0087] The above describes the process steps of obtaining the value distribution of the data to be evaluated in several domains by sending an evaluation request to the computing platform in a single data provider scenario. In fact, if the data provider has access to the support sets of each domain in several domains, it can also complete the calculation of the value distribution locally.
[0088] Next, the data value evaluation process in the scenario of multiple data providers in the improvement scheme will be introduced. Specifically, the dynamic value correction strategy proposed in this scenario will be described first, including calculating the Shapley value of each data provider based on the gradient inner product, and then the specific process steps will be introduced.
[0089] For the dynamic valuation of multiple data providers, the main consideration is that in the actual data trading scenario, multiple data providers usually coexist, and their potential mutual influence may require dynamic adjustment of the value estimation. For example, when the data of two providers partially or completely overlaps, additional penalties need to be imposed on the data value. To achieve this dynamic correction, inspiration is drawn from the traditional Shapley value calculation.
[0090] Specifically, for data provider i, its Shapley value φ i is:
[0091]
[0092] where N is the complete set formed by multiple data providers; N\{i} represents the complement of {i}; S represents a subset of N\{i}, which can also be called a data provider coalition; |S| is the size of coalition S, that is, the number of data providers in coalition S; |N| represents the size of the complete set N, or the total number of multiple data providers;! represents the factorial operation, and v() can be called the characteristic function, revenue function, utility function or value function. To avoid confusion, v() is called the revenue function in this paper, and v(S) represents the revenue of coalition S; v(S∪{i}) - v(S) represents the marginal contribution of data provider i to coalition S.
[0093] It should be noted that in actual calculation, it is the dataset to be evaluated held by data provider i that participates in the calculation. In addition, in the traditional scheme, when calculating the Shapley value φ i , the revenue function v() used is usually the verification index of the model (such as prediction accuracy, etc.), which means that a large number of retraining and testing of the model are required. For example, assuming there are 3 data providers, the number of training iterations required is (2 3 -1 =) 7 times, and the number of tests is also 7 times, which results in a high calculation cost of the Shapley value.
[0094] Based on the above proposed gradient inner product theory, the above formula (8) can be used as the revenue function v(), that is:
[0095]
[0096] Since the model parameters will not be updated, for |N| data providers, only |N| gradient calculations are required, which reduces the computational cost of calculating the Shapley value φ i from the original (2 |N| -1) possible subsets to |N|, thus enabling dynamic adjustment with almost no additional cost.
[0097] The above introduces a dynamic value evaluation strategy for improving the calculation of the Shapley value based on the gradient inner product in the scenario of multiple data providers. Figure 3 Shows the method flow for data value evaluation in the scenario of multiple data providers. The execution entity of this method can be any device, platform, server, or device cluster with computing and processing capabilities, such as a computing platform or a data trading platform, etc. Figure 3 The method flow shown in
[0098] Step S310, obtain the gradient data of the pre-trained large language model on each of the multiple data sets. It should be noted that the multiple data sets are held by multiple data providers respectively.
[0099] In one implementation case, the computing platform is a trusted platform. At this time, the trusted platform can receive multiple data sets from multiple data providers respectively, so as to calculate the gradient data of the LLM on each of the data sets locally.
[0100] In another implementation case, the computing platform can receive multiple gradient data from multiple data providers respectively, and any one of the gradient data is determined by the corresponding data provider based on the LLM and the data set it holds.
[0101] From the above, the gradient data of the LLM on each data set to be valued can be obtained.
[0102] Step S320, obtain the gradient data of the large language model on the support set in a predetermined domain.
[0103] In one implementation manner, the gradient data of the pre-calculated support set can be read from the memory.
[0104] In another implementation manner, the gradient data of the LLM on the support set can be calculated locally.
[0105] From the above, the gradient data of the LLM on the support set can be obtained.
[0106] Step S330: Take each of the multiple data sets as the target data set to be evaluated, and calculate its corresponding Shapley value as its target value in the predetermined domain. The calculation of the Shapley value includes: for each data set combination involved, determine the gain of the combination based on the inner product between the gradient data corresponding to the combination and the gradient data on the support set.
[0107] The calculation method of the Shapley value has been introduced above. For better understanding, an example is given below.
[0108] Suppose there are three data parties, corresponding to holding data sets and Furthermore, suppose we want to calculate φ1, then the coalition At the data set level, For the calculation of the benefit function v(s), taking as an example, according to the above formula (10), we have:
[0109]
[0110] where is the gain of the combination of data sets and ; is the second gradient sum value corresponding to the gradient data obtained in the above step S320.
[0111] As for Considering the additivity of gradients, we can directly sum the gradient data of the LLM (model parameters θ) obtained in step S310 on and respectively to obtain the third gradient sum value. In addition, considering that and may have an intersection, that is:
[0112]
[0113] Correspondingly, we can first perform a deduplication process on , and then determine the sum of the gradients of the remaining samples as the third gradient sum value. In one implementation case, the gradient data received by the computing platform from each data party includes the gradients of each sample. At this time, we can directly perform a deduplication process on the gradients corresponding to all samples in and , and then sum the remaining gradients as the third gradient sum value. In another implementation case, the computing platform calculates the gradients after receiving the original data sets from each data party. At this time, the implementation of the deduplication operation is very flexible.
[0114] Combined with the above examples, an improved revenue function based on the gradient inner product is introduced, so that the Shapley value can be calculated at low cost and efficiently. Further, the Shapley values corresponding to multiple data providers can be normalized to obtain the proportion of the contribution of the multiple data sets to the prediction effect gain of the fine-tuned LLM in a predetermined field when assuming that the LLM is fine-tuned using the multiple data sets, that is, the target value.
[0115] According to an embodiment of another aspect, for the gradient data or gradient sum values mentioned above, dimensionality reduction processing can also be performed before the inner product calculation, and specifically when to perform dimensionality reduction can be flexibly set as needed.
[0116] According to an embodiment of yet another aspect, after determining the data set value of each data provider among multiple data providers, some data providers can be selected for fine-tuning the LLM, or alternatively, after fine-tuning the LLM using multiple data sets, a remuneration can be paid to the corresponding data providers according to the data set value.
[0117] From the above, it is possible to achieve low-cost and efficient data value evaluation in the scenario of multiple data providers.
[0118] In summary, by adopting the improved solution for data value evaluation disclosed in the embodiments of this specification, through gradient inner product tracking, it is possible to accurately and efficiently estimate the data value distribution in different fields without training. For the scenario where multiple data providers coexist, in order to further meet the dynamic adjustment requirements, the gradient tracking mechanism is combined with the Shapley value theory to construct a fast Shapley value calculation method, and the additivity of the gradient is used to effectively evaluate the contribution of each data provider. Such an evaluation framework is based on the basic theory of gradient inner product tracking, and further optionally, with optimizations such as gradient projection and fast Shapley value calculation, it can complete accurate data value estimation at low cost and high efficiency in different downstream task scenarios.
[0119] Corresponding to the above data value evaluation method, the embodiments of this specification also disclose an evaluation device, which can be integrated into a computing platform. Specifically as follows:
[0120] Figure 4 The illustrated evaluation device 400 includes the following functional units:
[0121] An evaluation request receiving unit 410 is configured to receive an evaluation request for the value of an evaluation target data set in a number of domains. An evaluation request parsing unit 420 is configured to determine first gradient data of the pre-trained large language model on the target data set by parsing the evaluation request. A gradient data determination unit 430 is configured to obtain second gradient data of the large language model on a support set in any target domain among the number of domains. A target value determination unit 440 is configured to determine a target value of the target data set in the target domain based on the inner product between the first gradient data and the second gradient data, and classify it into the processing result of the evaluation request.
[0122] In one embodiment, the gradient data determination unit 430 is specifically configured to: read the pre-computed second gradient data from the memory.
[0123] In one embodiment, the first gradient data includes a first gradient sum value between the gradients of each sample in the target data set; the second gradient data includes a second gradient sum value between the gradients of each sample in the support set; wherein the target value determination unit 440 is specifically configured to: calculate the inner product between the first gradient sum value and the second gradient sum value as the target value.
[0124] Furthermore, in a specific embodiment, the target value determination unit 440 is further configured to: perform dimensionality reduction processing on the first gradient sum value and the second gradient sum value respectively; perform an inner product operation based on the results of the dimensionality reduction processing to obtain the target value.
[0125] Figure 5 The illustrated evaluation device 500 is integrated into a computing platform and includes the following functional units:
[0126] A first gradient acquisition unit 510 is configured to acquire gradient data of the pre-trained large language model on each data set among the multiple data sets, and the multiple data sets respectively belong to multiple data parties. A second gradient acquisition unit 520 is configured to acquire gradient data of the large language model on a support set in a predetermined domain. A target value calculation unit 530 is configured to use each data set as a target data set to be evaluated, calculate its corresponding Shapley value as its target value in the predetermined domain; the calculation of the Shapley value includes: for each data set combination involved, determining the benefit of the combination based on the inner product between the gradient data corresponding to the combination and the gradient data on the support set.
[0127] In one embodiment, the first gradient acquisition unit 510 is specifically configured to: receive multiple gradient data from the multiple data parties, and any one of the gradient data is determined by the corresponding data party based on the large language model deployed locally by it and the data set it holds.
[0128] In one embodiment, the computing platform is a trusted platform independent of the multiple data parties; the first gradient acquisition unit 510 is configured to: receive the multiple data sets corresponding to the multiple data parties; calculate the gradient data of the large language model on each of the multiple data sets locally.
[0129] In one embodiment, the second gradient acquisition unit 520 is specifically configured to: calculate the gradient data of the large language model on the support set locally; or, read the gradient data of the support set calculated in advance from the memory.
[0130] In one embodiment, the target value calculation unit 530 is configured to determine the return of the combination, specifically including: determining the third gradient sum value between the gradient data corresponding to the combination; based on the inner product between the third gradient sum value and the second gradient sum value corresponding to the gradient data on the support set, determining the return of the combination.
[0131] Further, in a specific embodiment, the gradient data on each of the data sets includes the gradient sum value between the gradients corresponding to each sample in the data set; wherein, the target value calculation unit 530 is configured to determine the third gradient sum value, specifically including: performing a summation process on the gradient sum values corresponding to each data set in the combination to obtain the third gradient sum value.
[0132] In a specific embodiment, the gradient data on each of the data sets includes the gradients of each sample in the data set; the target value calculation unit 530 is configured to determine the third gradient sum value, specifically including: performing a deduplication process on the gradients of all samples corresponding to the combination; performing a summation process on the gradients remaining after the deduplication process to obtain the third gradient sum value.
[0133] In a specific embodiment, the gradient data on the support set includes the gradient sum value between the gradients corresponding to each sample in the support set, and this gradient sum value is used as the second gradient sum value.
[0134] In a specific embodiment, the target value calculation unit 530 is configured to determine the return of the combination, specifically including: performing a dimensionality reduction process on the third gradient sum value and the second gradient sum value; performing an inner product operation based on the result of the dimensionality reduction process to obtain the return.
[0135] In one embodiment, each sample in the multiple data sets and the support set includes a task instruction and a corresponding expected output.
[0136] It should be understood that for the convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing this specification, the functions of each module can be implemented in the same or multiple software and / or hardware. In addition, for the description of the device, reference can also be made to the foregoing description of the method, which will not be elaborated herein.
[0137] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute Figure 2 or Figure 3 the described method.
[0138] According to an embodiment of still another aspect, there is also provided a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, it implements Figure 2 or Figure 3 the described method.
[0139] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium, or transmitted as one or more instructions or codes on a computer-readable medium.
[0140] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for evaluating data value, comprising: Receiving an evaluation request for the value of an evaluation target data set in a number of fields; Determining first gradient data of a pre-trained large language model on the target data set by parsing the evaluation request; For any target field among the number of fields, obtaining second gradient data of the large language model on the support set of the target field; Determining the target value of the target data set in the target field based on the inner product between the first gradient data and the second gradient data, and incorporating it into the processing result of the evaluation request.
2. The method according to claim 1, wherein Obtaining the second gradient data of the large language model on the support set of the target field includes: Reading the pre-computed second gradient data from memory.
3. The method according to claim 1, wherein, The first gradient data includes a first gradient sum value between the gradients of each sample in the target data set; the second gradient data includes a second gradient sum value between the gradients of each sample in the support set; wherein, determining the target value of the target data set in the target field includes: Calculating the inner product between the first gradient sum value and the second gradient sum value as the target value.
4. The method according to claim 3, wherein Calculating the inner product between the first gradient sum value and the second gradient sum value as the target value includes: Performing dimensionality reduction processing on the first gradient sum value and the second gradient sum value respectively; Performing an inner product operation based on the results of the dimensionality reduction processing to obtain the target value.
5. A method for evaluating data value, where multiple data parties are involved in holding multiple data sets, and the method is applied to a computing platform, comprising: Obtaining gradient data of a pre-trained large language model on each of the multiple data sets; Obtaining gradient data of the large language model on the support set of a predetermined field; Taking each of the data sets as a target data set to be evaluated, and calculating its corresponding Shapley value as its target value in the predetermined field; The calculation of the Shapley value includes: for each data set combination involved, determining the gain of the combination based on the inner product between the gradient data corresponding to the combination and the gradient data on the support set.
6. The method according to claim 5, wherein Obtaining gradient data of a pre-trained large language model on each of the multiple data sets includes: Receiving multiple gradient data from the multiple data parties correspondingly, where any one of the gradient data is determined by the corresponding data party based on the large language model deployed locally by it and the data set held by it.
7. The method according to claim 5, wherein, The computing platform is a trusted platform independent of the multiple data parties; wherein, obtaining gradient data of a pre-trained large language model on each of the multiple data sets includes: Receiving the multiple data sets from the multiple data parties correspondingly; Locally calculating the gradient data of the large language model on each of the multiple data sets.
8. The method according to claim 5, wherein Obtaining gradient data of the large language model on the support set of a predetermined field includes: Locally calculating the gradient data of the large language model on the support set; or, Reading the pre-computed gradient data of the support set from memory.
9. The method according to claim 5, wherein For each combination of datasets involved, determine the benefit of the combination based on the inner product between the gradient data corresponding to the combination and the gradient data on the support set, including: Determine the third gradient sum value between the gradient data corresponding to the combination; Based on the inner product between the third gradient sum value and the second gradient sum value corresponding to the gradient data on the support set, determine the benefit of the combination.
10. The method according to claim 9, wherein, The gradient data on each dataset includes the gradient sum value between the gradients corresponding to each sample in the dataset; wherein, determining the third gradient sum value between the gradient data corresponding to the combination includes: Sum the gradient sum values corresponding to each dataset in the combination to obtain the third gradient sum value.
11. The method according to claim 9, wherein The gradient data on each dataset includes the gradients of each sample in the dataset; wherein, determining the third gradient sum value between the gradient data corresponding to the combination includes: Remove duplicates from the gradients of all samples corresponding to the combination; Sum the remaining gradients after the duplicate removal process to obtain the third gradient sum value.
12. The method according to claim 9, wherein, The gradient data on the support set includes the gradient sum value between the gradients corresponding to each sample in the support set, and this gradient sum value is used as the second gradient sum value.
13. The method according to claim 9, wherein Determining the benefit of the combination includes: Perform dimensionality reduction processing on the third gradient sum value and the second gradient sum value; Based on the result of the dimensionality reduction processing, perform an inner product operation to obtain the benefit.
14. The method according to claim 5, wherein Each sample in the multiple datasets and the support set includes a task instruction and a corresponding expected output.
15. An apparatus for evaluating data value, comprising: An evaluation request receiving unit configured to receive an evaluation request for the value of the target dataset in several fields; An evaluation request parsing unit configured to determine the first gradient data of the pre-trained large language model on the target dataset by parsing the evaluation request; A gradient data determining unit configured to obtain, for any target field in the several fields, the second gradient data of the large language model on the support set of the target field; A target value determining unit configured to determine the target value of the target dataset in the target field based on the inner product between the first gradient data and the second gradient data, and include it in the processing result of the evaluation request.
16. An apparatus for evaluating data value, where multiple data parties hold multiple datasets correspondingly, and the apparatus is integrated into a computing platform, comprising: A first gradient obtaining unit configured to obtain the gradient data of the pre-trained large language model on each of the multiple datasets; A second gradient obtaining unit configured to obtain the gradient data of the large language model on the support set of a predetermined field; A target value calculating unit configured to use each of the datasets as the target dataset to be evaluated, and calculate its corresponding Shapley value as its target value in the predetermined field; The calculation of the Shapley value includes: for each combination of datasets involved, determining the benefit of the combination based on the inner product between the gradient data corresponding to the combination and the gradient data on the support set.
17. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed on a computer, the computer is caused to execute the method according to any one of claims 1-14.
18. A computing device, comprising a memory and a processor, wherein, Executable code is stored in the memory, and when the processor executes the executable code, the method according to any one of claims 1-14 is implemented.