Data attribution method and device based on dimensionality reduction projection

Through the data attribution method based on dimensionality reduction projection, the problem of training data affecting recognition and quantification in the pre-training process of large language models is solved, and the model transparency is enhanced and the training data is optimized, which improves model performance and efficiency.

CN120448807APending Publication Date: 2025-08-08SHANGHAI HEHE INFORMATION TECH DEV +3
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510513166.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

How to effectively identify and quantify the impact of training data on model performance during pre-training of large language models to enhance model transparency and data management.

Method used

The data attribution method based on dimensionality reduction projection is adopted, and the loss gradient characteristics of the training data and evaluation data on different benchmark language models are calculated, and the data attribution scores are calculated to clarify the contribution size of the training data.

Benefits of technology

The contribution relationship between different training data and evaluation data is clarified, the transparency of the model is enhanced, the selection of training data is optimized, the performance and training efficiency of the model is improved, and the cost of computing resources and storage is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448807A_ABST
    Figure CN120448807A_ABST
Patent Text Reader

Abstract

The invention discloses a data attribution method based on dimensionality reduction projection. The method comprises the following steps: S1, preparing a training data set, an evaluation data set and a plurality of benchmark language models; s2, calculating loss gradient characteristics of each training data and each evaluation data on different reference models; and S3, correcting loss gradient features of each training data and each evaluation data on different reference models. And S4, carrying out multiple times of dimensionality reduction projection on the corrected loss gradient features of each training data and each evaluation data on different reference models. And S5, calculating a data attribution score between each piece of training data and each piece of evaluation data. According to the method, the contribution of different training data and different benchmark language models on different evaluation data is defined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a data attribution technology, and in particular to a data attribution method in large language model (LLM) pre-training, which is used to identify and quantify the specific impact of different training data on the performance of large language models. Background Art

[0002] Attribution analysis is a systematic evaluation method used to identify and quantify the contribution of different factors to the final results.

[0003] In the field of artificial intelligence, large language models have attracted widespread attention due to their outstanding ability to handle complex language tasks. A pre-training model (PTM) refers to training a model using a large amount of unlabeled data to equip it with certain prior knowledge and common sense, thereby improving its performance on various tasks. Large language models are able to capture and learn deep-level features of language by pre-training on large-scale datasets. However, as large language models continue to expand in size, their training process becomes increasingly opaque, which limits researchers' understanding of changes in model performance. Researchers hope to attribute data to large language models during pre-training, trace the relationship between model performance and training data, and determine which training data contributes to specific prediction results. Data attribution helps understand the relationship between training data and model performance, thereby helping researchers better select training data. Summary of the Invention

[0004] The technical problem to be solved by this application is: how to effectively identify and quantify the impact of training data (i.e., training samples) on the performance of the model on the evaluation dataset during the pre-training process of LLM, so as to enhance model transparency and data management.

[0005] In order to solve the above technical problems, the present application proposes a data attribution method based on dimensionality reduction projection, comprising the following steps. Step S1: Prepare a training data set, an evaluation data set, and multiple benchmark language models. Step S2: Calculate the loss gradient features of each training data and each evaluation data on different benchmark models. Step S3: Correct the loss gradient features of each training data and each evaluation data on different benchmark models to obtain the corrected loss gradient features of each training data and each evaluation data on different benchmark models. Step S4: Perform multiple dimensionality reduction projections on the corrected loss gradient features of each training data and each evaluation data on different benchmark models to obtain the normalized loss gradient feature information of each training data and each evaluation data on different benchmark language models. Step S5: Calculate the data attribution score between each training data and each evaluation data.

[0006] Furthermore, in step S1, the training data set X is composed of multiple training data x; each training data x is represented by multiple phrases a1, a2, ..., a n A text sequence composed of (a1, a2, ..., a n ), n represents the number of phrases that the training data x is represented by. The evaluation data set Y consists of multiple evaluation data y, and each evaluation data y is also represented as a text sequence consisting of multiple phrases. Multiple benchmark language models m1, m2, ..., m r The set of all benchmark language models M = {m1,m2,…,m r}, r represents the number of benchmark language models; the benchmark language models are all large language models using a "decoder-only" structure.

[0007] Furthermore, in step S2, first, in any benchmark language model m k Calculate the loss Loss(x,θ) of any training data x k ); Among them, θ k is the baseline language model m k Parameters of p(x,θ k ) represents the baseline language model m k The probability of generating the training data x; the baseline language model m k The training data x is generated by the text sequence of the training data x (a1, a2, ..., a n ) is generated phrase by phrase in the order of j |a<j,θ k ) means given the previous j phrases (a1, a2, ..., a j-1 ) under the condition that model m k Generate the next phrase a j The same calculation method is used to calculate the loss of any evaluation data on any benchmark language model. k Calculate the loss Loss(x,θ) of any training data x k ) about the parameters θ of the baseline language model k The loss gradient feature Is a vector whose each component is the loss Loss(x,θ) of the training data x k ) The partial derivative of each model parameter is calculated as follows: Among them, θ k =(θ k,1 ,θ k,2 ,...,θ k,q) is the baseline language model m k The parameter vector of q represents the baseline language model m k The number of parameters included, θ k,i Represents the baseline language model m k The loss gradient characteristics of the parameters of any benchmark language model are calculated using the same method.

[0008] Furthermore, step S3 specifically includes the following sub-steps S31-S32. Step S31: Calculate the second-order moment of the loss gradient feature of each training data and each evaluation data on different baseline models with respect to each parameter of the baseline language model. Step S32: Correct the loss gradient feature of each training data and each evaluation data on different baseline models.

[0009] Furthermore, in step S31, for any training data x in any benchmark language model m k The loss gradient feature on Calculate the loss gradient feature Relative to the baseline language model m k A parameter θ k,i The second moment of The process of calculating the second-order moment is as follows: Among them, mean_i is the mean of all training data x in the training dataset X about the baseline language model m k A parameter θ k,i The loss gradient feature The mean of It refers to the loss Loss(x,θ) of any training data x k ) for the baseline language model m k The i-th parameter θ k,i The partial derivative of N is the number of training data x contained in the training dataset X; any evaluation data y in any benchmark language model m k The loss gradient feature on the baseline language model m k Any parameter θ k,i The second moment of is calculated in the same way.

[0010] Furthermore, in step S32, any training data x is used in any benchmark language model m k The loss gradient feature on Relative to the baseline language model m k A parameter θ k,i The gradient of the training data x is divided by the baseline language model m kThe loss gradient feature on Relative to the baseline language model m k A parameter θ k,i The second moment of The square root of the training data x in the benchmark language model m k The corrected loss gradient feature on Any evaluation data y on any benchmark language model m k The same correction method is used for the loss gradient features on .

[0011] Furthermore, the step S4 specifically includes the following sub-steps S41-S45. Step S41: Generate multiple random projection matrices (P1, P2, ..., P B ); B represents the number of random projection matrices, that is, the number of dimensionality reduction projections. Step S42: For any training data x in any benchmark language model m k The corrected loss gradient feature on Perform multiple dimensionality reduction projections, and perform this operation B times to obtain the loss gradient feature information after B dimensionality reduction projections; then concatenate the loss gradient feature information after B dimensionality reduction projections in the order of dimensionality reduction projections to obtain the training data x in the benchmark language model m k The loss gradient feature information G after multiple dimensionality reduction projections θk (x). For any evaluation data y in any benchmark language model m k Perform multiple dimensionality reduction projections on the corrected loss gradient features to obtain the loss gradient feature information after the multiple dimensionality reduction projections of the evaluation data on the benchmark language model. Step S43: Concatenate the loss gradient feature information after the b-th dimensionality reduction projection of all training data to obtain the total gradient information matrix after the b-th dimensionality reduction projection of all training data, denoted as g θk (X) b The same method is used to obtain the total gradient information matrix after any dimensionality reduction projection of all evaluation data. Step S44: For any training data x in any benchmark language model m k The loss gradient feature information G after multiple dimensionality reduction projections on θk (x) Perform Hessian approximation calculation. Perform Hessian approximation calculation on the loss gradient feature information after multiple dimensionality reduction projections of any evaluation data on any benchmark language model. Step S45: Perform Hessian approximation calculation on any training data x after Hessian approximation calculation on any benchmark language model m k The loss gradient feature information after multiple dimensionality reduction projections on Normalize the loss gradient feature information obtained by concatenating multiple dimensionality reduction projections of any evaluation data on any benchmark language model after Hessian approximation.

[0012] Furthermore, in step S42, the training data x is k The loss gradient feature information after the b-th dimensionality reduction projection Among them, ⊙ is matrix multiplication; G θk (x) = concat(g θk (x)1,g θk (x)2,…,g θk (x) B ); where concat(·) represents a concatenation operation, and the value range of b is a positive integer between 1 and B. In step S43, the total gradient information matrix g after the b-th dimension reduction projection of all training data is θk (X) b =concat(g θk (x1) b ,g θk (x2) b ,…); where concat(·) represents the concatenation operation, the data x i ∈X, b represents the b-th dimensionality reduction projection, and the value range of b is a positive integer between 1-B.

[0013] Furthermore, in step S5, any training sample x i For any evaluation sample y j Contribution score in, Represents any training data x i Normalized loss gradient feature information on any baseline language model; Represents any evaluation data y j Normalized loss gradient feature information on any baseline language model; |||| represents the computational norm; · represents vector dot product; * represents real number multiplication.

[0014] The present application also proposes a data attribution device based on dimensionality reduction projection, comprising a data set and model preparation unit, a loss gradient feature calculation unit, a loss gradient feature correction unit, a dimensionality reduction projection calculation unit, and a data attribution calculation unit. The data set and model preparation unit is used to prepare a training data set, an evaluation data set, and multiple benchmark language models. The loss gradient feature calculation unit is used to calculate the loss gradient features of each training data and each evaluation data on different benchmark models. The loss gradient feature correction unit is used to correct the loss gradient features of each training data and each evaluation data on different benchmark models to obtain the corrected loss gradient features of each training data and each evaluation data on different benchmark models. The dimensionality reduction projection calculation unit is used to perform multiple dimensionality reduction projections on the corrected loss gradient features of each training data and each evaluation data on different benchmark models to obtain normalized loss gradient feature information of each training data and each evaluation data on different benchmark language models. The data attribution calculation unit is used to calculate the data attribution score between each training data and each evaluation data.

[0015] The technical effect achieved by this application is: clarifying the contribution of different training data and different benchmark language models on different evaluation data. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flowchart of the data attribution method based on dimensionality reduction projection proposed in this application.

[0017] Figure 2 yes Figure 1 Specific flow diagram of step S3 in .

[0018] Figure 3 yes Figure 1 Specific flow diagram of step S4 in .

[0019] Figure 4 It is a structural diagram of the data attribution device based on dimensionality reduction projection proposed in this application.

[0020] Explanation of the reference numerals in the figure: data set and model preparation unit 1, loss gradient feature calculation unit 2, loss gradient feature correction unit 3, dimensionality reduction projection calculation unit 4, data attribution calculation unit 5. DETAILED DESCRIPTION

[0021] See also Figure 1 ,The data attribution method based on dimensionality reduction projection proposed in this application includes the following steps.

[0022] Step S1: Prepare the data set and the baseline language model. This step specifically includes: preparing a training data set, denoted as X. The training data set X consists of multiple training data x. A training data x is represented by multiple phrases a1, a2, ..., a n A text sequence composed of (a1, a2, ..., a n ), n represents the number of phrases that the training data x is represented by. Prepare an evaluation dataset, denoted as Y. The evaluation dataset Y consists of multiple evaluation data y, and each evaluation data y is also represented as a text sequence composed of multiple phrases. Prepare multiple benchmark language models m1, m2, ..., m r , recorded as M={m1,m2,…,m r}, where M represents the set of all baseline language models and r represents the number of baseline language models. The baseline language models are all large language models (LLMs). All baseline language models are not required to be identical; for example, they can be any large language model with a decoder-only architecture. A baseline model is a basic model used to evaluate and compare the performance of new models in machine learning or deep learning research and applications.

[0023] Step S2: Calculate the loss gradient characteristics of each training data and each evaluation data on different baseline models. This step specifically includes: calculating the loss of each training data and each evaluation data on each baseline language model, and then using the backpropagation algorithm to calculate the loss gradient characteristics of each training data loss and each evaluation data loss with respect to the parameters of the corresponding baseline language model.

[0024] The following example illustrates how to use any baseline language model m k Up (m k ∈M) calculates the loss Loss(x,θ) of any training data x k ). Among them, θ k is the baseline language model m k The parameter of p(x,θ k ) represents the baseline language model m k The probability of generating the training data x. The baseline language model m k The training data x is generated by the text sequence of the training data x (a1, a2, ..., a n ) is generated one by one in the order of, and the probability of generating a phrase at each step is expressed as p(a j |a<j,θ k ). p(a j |a<j,θ k ) means given the previous j phrases (a1, a2, ..., aj-1 ) under the condition that model m k Generate the next phrase a j Conditional probability refers to the probability that event A occurs given the occurrence of event B. The same calculation method is used to calculate the loss of any evaluation data on any benchmark language model, so I will not repeat it here.

[0025] As mentioned above, a certain baseline language model generates certain training data. This should be understood from the perspective of loss calculation. Suppose there is a large language model used to generate sentences, and there is a training data of the text sequence (I, like, apple). The model will first try to generate "I", then generate "like" based on "I", and finally generate "apple" based on "I like". In this process, the probability of each step will affect the final loss calculation. The loss function of the model Loss(x,θ k ) is calculated by calculating the probability p(x,θ) that the model generates the training data k The higher the probability that the model generates the training data, the better the model fits the training data and the smaller the loss.

[0026] The following example illustrates how to use any baseline language model m k Up (m k ∈M) calculates the loss Loss(x,θ) of any training data x k ) about the parameters θ of the baseline language model k The loss gradient feature Is a vector whose each component is the loss Loss(x,θ) of the training data x k ) The partial derivative of each model parameter is calculated as follows: Among them, θ k =(θ k,1 ,θ k,2 ,...,θ k,q ) is the baseline language model m k The parameter vector of q represents the baseline language model m k The number of parameters included, θ k,i Represents the baseline language model m k The i-th parameter of . Here, R represents the set of real numbers. This means The result of this calculation is a one-dimensional vector of size q, where each element is a real number. The loss gradient characteristics of the parameters of any benchmark language model are calculated using the same method, so this will not be repeated here.

[0027] Step S3: Correct the loss gradient features of each training data and each evaluation data on different benchmark models. This step is used to improve the accuracy and reliability of the loss gradient features, and specifically includes the following sub-steps S31-S32, such as Figure 2 shown.

[0028] Step S31: Calculate the second-order moment of the loss gradient feature of each training data and each evaluation data on different benchmark models relative to each parameter of the benchmark language model. k The loss gradient feature on Calculate the loss gradient feature Relative to the baseline language model m k A parameter θ k,i The second moment of The process of calculating the second-order moment is as follows: Among them, mean_i is the mean of all training data x in the training dataset X about the baseline language model m k A parameter θ k,i The loss gradient feature The mean of , used to normalize the loss gradient features. It refers to the loss Loss(x,θ) of any training data x k ) for the baseline language model m k The i-th parameter θ k,i The partial derivative of . N is the number of training data x contained in the training dataset X. The second-order moment reflects the variance of the loss gradient feature and is used to adjust the size of the loss gradient feature to reduce the impact of outliers. k The loss gradient feature on the baseline language model m k Any parameter θ k,i The second-order moment of is calculated using the same method and will not be described in detail.

[0029] Step S32: Correct the loss gradient characteristics of each training data and each evaluation data on different benchmark models. k The loss gradient feature on Relative to the baseline language model m k A parameter θ k,i The gradient of the training data x is divided by the baseline language model m k The loss gradient feature on Relative to the baseline language model m k A parameter θ k,i The second moment of The square root of the training data x in the benchmark language model m k The corrected loss gradient feature on The specific calculation method is as follows: Any evaluation data y on any benchmark language model m k The loss gradient features on adopt the same correction method and will not be described in detail.

[0030] Step S4: Perform multiple dimensionality reduction projections on the corrected loss gradient features of each training data and each evaluation data on different benchmark models. This step uses the random projection algorithm that conforms to the Johnson–Lindenstrauss theorem (also known as the Johnson–Lindenstrauss lemma) to perform dimensionality reduction projections on the corrected loss gradient features of each training data and each evaluation data on different benchmark models. The purpose of dimensionality reduction projection is to reduce the dimensionality of the data and improve computational efficiency while retaining the important information of the loss gradient features. Figure 3 This step specifically includes the following sub-steps S41-S45.

[0031] Step S41: Generate multiple random projection matrices (P1, P2, ..., P B ). Each projection matrix is randomly generated, and each randomly generated projection matrix is different. Each random projection matrix generated P b (The value range of b is a positive integer between 1-B) The elements in are all from the standard normal distribution The independent sampling in P all obeys the standard normal distribution, where d represents the dimension after dimensionality reduction. b ∈R q×d Where R represents a set of real numbers. q represents a certain benchmark language model m k The number of parameters included, that is, the dimension before dimensionality reduction. b ∈R q×d Denotes the random projection matrix P b is a q×d two-dimensional matrix, where every element is a real number. B represents the number of random projection matrices, i.e., the number of dimensionality reduction projections.

[0032] Step S42: For any training data x in any benchmark language model m k The corrected loss gradient feature on Perform multiple dimensionality reduction projections and record the training data x in the benchmark language model m k The loss gradient feature information after the b-th dimensionality reduction projection is g θk (x) b , with lowercase x in brackets. The calculation method is Where ⊙ is matrix multiplication. g θk (x) b ∈R d , which means g θk (x) b It is a one-dimensional vector of size d, and each element in the vector is a real number.

[0033] The above operation is performed B times to obtain B loss gradient feature information after dimensionality reduction projection. Then the loss gradient feature information after B dimensionality reduction projections is spliced in the order of dimensionality reduction projection to obtain the training data x in the benchmark language model m k The loss gradient feature information G after multiple dimensionality reduction projections θk (x). It is calculated as G θk (x) = concat(g θk (x)1,g θk (x)2,…,g θk (x) B ). Where concat(·) represents the concatenation operation. g θk (x) b Indicates that the training data x in the benchmark language model m k The loss gradient feature information after the b-th dimensionality reduction projection on G, the value range of b is a positive integer between 1-B. θk (x)∈R (B*d) , which means that G θk (x) is a one-dimensional vector of size B*d, and each element in the vector is a real number.

[0034] For any evaluation data y in any benchmark language model m k The corrected loss gradient features on the benchmark language model are projected multiple times with reduced dimension to obtain the loss gradient feature information after multiple reduced dimension projections of the evaluation data on the benchmark language model. The same method is used for calculation and will not be repeated here.

[0035] Step S43: Concatenate the loss gradient feature information after any (e.g., the bth) dimensionality reduction projection of all training data to obtain the total gradient information matrix after the bth dimensionality reduction projection of all training data, denoted as g θk (X) b , with a capital X in the brackets. The calculation method is g θk (X) b =concat(g θk (x1) b ,g θk (x2) b ,…). Where concat(·) represents the concatenation operation, data xi ∈X, b represents the b-th dimension reduction projection. g θk (X) b ∈R r×d This means that g θk (X) b It is a two-dimensional matrix of size r×d, where r refers to the number of training data and each element in the two-dimensional matrix is a real number.

[0036] The total gradient information matrix after any dimensionality reduction projection of all evaluation data is calculated using the same method and will not be repeated here.

[0037] Step S44: For any training data x in any benchmark language model m k The loss gradient feature information G after multiple dimensionality reduction projections on θk (x) Perform Hessian approximation, which is calculated as follows: in, It is any training data x after Hessian approximation calculation in any benchmark language model m k The loss gradient feature information after multiple dimensionality reduction projections on , ⊙ is the matrix multiplication. Where diag(·) represents the operation of constructing a diagonal matrix. Represents the total gradient information matrix g after the b-th dimensionality reduction projection of all training data θk (X) b The transposed matrix of . The superscript -1 indicates the matrix inversion. Hessian(θ k )∈R (B*d)×(B*d) , which means that the Hessian(θ k ) is a two-dimensional matrix of size (B*d)×(B*d), where each element is a real number. * represents intra-dimensional multiplication, for example, a*b represents a one-dimensional vector of size a*b. × represents inter-dimensional multiplication, for example, a×b represents a two-dimensional matrix with the first dimension of size a and the second dimension of size b. The purpose of this step is to decouple the loss gradient features resulting from the concatenation of multiple dimensionality reduction projections and to calculate the independent contribution of each dimension's loss gradient feature in explaining the change in the model output.

[0038] The same algorithm is used to perform Hessian approximation calculation on the loss gradient feature information after multiple dimensionality reduction projections of any evaluation data on any benchmark language model, which will not be repeated here.

[0039] Step S45: Normalization. Any training data x in any benchmark language model m k The normalized loss gradient feature information on in, Represents any training data x after Hessian approximation calculation in any benchmark language model m k The loss gradient feature information after multiple dimensionality reduction projections on The norm used in this document is the Euclidean norm.

[0040] The same algorithm is used to normalize the loss gradient feature information of any evaluation data after Hessian approximation on any benchmark language model, and will not be repeated here.

[0041] Step S5: Calculate the data attribution score between each training data and each evaluation data. Using steps S1 to S4, calculate the normalized loss gradient feature information of all training data x in the training dataset X on any benchmark language model, and calculate the normalized loss gradient feature information of all evaluation data y in the evaluation dataset Y on any benchmark language model. i The normalized loss gradient feature information on any benchmark language model is recorded as x i ∈X. Any evaluation data y j The normalized loss gradient feature information on any benchmark language model is recorded as y j ∈Y. Calculate any training sample x i For any evaluation sample y j The contribution score I(x i ,y j ), the data attribution fraction, is calculated as follows: Here, r represents the number of baseline language models. Represents the training sample x i In the kth baseline language model m k The normalized loss gradient feature information on . express The norm of . Represents the evaluation sample y j In the kth baseline language model m k The normalized loss gradient feature information on . express The norm of . · represents the vector dot product, which is the sum of the elements at corresponding positions of two vectors after multiplication. * represents ordinary real number multiplication.

[0042] I(x i ,y j )The higher the score, the better the training data x i and evaluation data jThe more relevant it is on the set M of all benchmark language models. This means that using training samples x at the level of the set M of all benchmark language models i After training, the set M of all benchmark language models is j The greater the performance increase.

[0043] I(x i ,y j )The lower the score, the more training data x i and evaluation data j The less relevant it is on the set M of all benchmark language models. This means that using training samples x at the level of the set M of all benchmark language models i After training, the set M of all benchmark language models is j The greater the performance degradation.

[0044] I(x i ,y j ) score can indicate the performance of the training data and the baseline language model on the evaluation data, and then trace the relationship between the model performance and the training data, and determine which training samples contribute to a specific prediction result. Specifically, I(x i ,y j )The higher the training data x i For the evaluation data y i The greater the contribution of I(x i ,y j )The lower the training data x i For the evaluation data y i The smaller the contribution.

[0045] See also Figure 4 The data attribution device based on dimensionality reduction projection proposed in this application includes a data set and model preparation unit 1, a loss gradient feature calculation unit 2, a loss gradient feature correction unit 3, a dimensionality reduction projection calculation unit 4, and a data attribution calculation unit 5. Figure 4 The device shown corresponds to Figure 1 The method shown.

[0046] The dataset and model preparation unit 1 is used to prepare a training dataset, an evaluation dataset, and multiple benchmark language models.

[0047] The loss gradient feature calculation unit 2 is used to calculate the loss gradient features of each training data and each evaluation data on different benchmark models.

[0048] The loss gradient feature correction unit 3 is used to correct the loss gradient features of each training data and each evaluation data on different benchmark models to obtain the corrected loss gradient features of each training data and each evaluation data on different benchmark models.

[0049] The dimensionality reduction projection calculation unit 4 is used to perform multiple dimensionality reduction projections on the corrected loss gradient features of each training data and each evaluation data on different benchmark models to obtain normalized loss gradient feature information of each training data and each evaluation data on different benchmark language models.

[0050] The data attribution calculation unit 5 is used to calculate the data attribution score between each training data and each evaluation data.

[0051] This application clarifies the performance of different training data and different baseline language models on different evaluation data by calculating the data attribution score between each training data and each evaluation data, and then clarifies which training samples contribute the most to a specific prediction result.

[0052] This application brings the following beneficial effects to the pre-training process of large language models.

[0053] First, it enhances the transparency of large language models and reveals the relationship between training data and the performance of large language models. The data attribution score clearly shows the impact of different training data on the performance of the baseline language model on the evaluation data. This allows researchers to intuitively understand which training data has a positive or negative effect on the large language model on specific evaluation tasks, breaking the "black box" of the large language model training process and enhancing the interpretability of the internal mechanisms of the large language model. In the training of large language models, due to the large amount of data and the complex model structure, it is difficult to directly analyze how the training data affects the performance of the large language model. This application provides quantitative indicators that allow researchers to deeply explore the role of different training data in the learning process of large language models. For example, some training data may help large language models learn specific language patterns, knowledge or semantic understanding, while others may interfere with the learning of large language models, thereby better understanding the behavior and decision-making process of large language models.

[0054] Second, optimize the selection of training data. After clarifying the contribution of training data to model performance, researchers can select training data more specifically according to actual needs. In order to improve the performance of large-scale language models on specific evaluation tasks, training data with high correlation with evaluation data and high data attribution scores can be selected to reduce the interference of irrelevant or negatively correlated data, improve the quality and utilization efficiency of training data, and thus optimize the training effect of large-scale language models. In the training of large-scale language models, the collection, preprocessing and storage of training data all require a lot of resources. By screening out key training data through the method of this application, it is possible to avoid the use of a large amount of redundant or inefficient data, reduce the cost of data processing and storage, and at the same time reduce the computing resources and time required for large-scale language model training, thereby improving training efficiency.

[0055] Third, improve the performance of large language models and optimize their training results. Reasonable selection of training data allows large language models to focus more on data that has a positive impact on performance, avoiding learning noise or interference information, leading to better convergence during training and improving the accuracy and stability of large language models on various tasks. By removing training data that negatively impacts the performance of large language models, the risk of overfitting during training is reduced, enabling large language models to better learn common patterns and features in the data, improving their generalization capabilities on unseen data, and enhancing their reliability in practical applications.

[0056] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A data attribution method based on dimensionality reduction projection, characterized in that: The method includes the following steps: Step S1: Prepare training datasets, evaluation datasets, and multiple benchmark language models; Step S2: Calculate the loss gradient characteristics of each training data and each evaluation data on different benchmark models; Step S3: Correcting the loss gradient features of each training data and each evaluation data on different benchmark models to obtain the corrected loss gradient features of each training data and each evaluation data on different benchmark models; Step S4: Perform multiple dimensionality reduction projections on the corrected loss gradient features of each training data and each evaluation data on different benchmark models to obtain normalized loss gradient feature information of each training data and each evaluation data on different benchmark language models; Step S5: Calculate the data attribution score between each training data and each evaluation data.

2. The data attribution method based on dimensionality reduction projection according to claim 1 is characterized in that: In step S1, the training data set X is composed of multiple training data x; each training data x is represented by multiple phrases a1, a2, ..., a n A text sequence composed of (a1, a2, ..., a n ), n represents the number of phrases that the training data x is represented by; The evaluation data set Y consists of multiple evaluation data y, and each evaluation data y is also represented as a text sequence consisting of multiple phrases; Multiple baseline language models m1, m2, ..., m r The set of all benchmark language models M = {m1,m2,…,m r }, r represents the number of benchmark language models; the benchmark language models are all large language models using a "decoder only" structure.

3. The data attribution method based on dimensionality reduction projection according to claim 2 is characterized in that: In step S2, first, in any benchmark language model m k Calculate the loss Loss(x,θ) of any training data x k ); Among them, θ k is the baseline language model m k Parameters of p(x,θ k ) represents the baseline language model m k The probability of generating the training data x; the baseline language model m k The training data x is generated by the text sequence of the training data x (a1, a2, ..., a n ) is generated phrase by phrase in the order of j |a <j ,θ k ) means given the previous j phrases (a1, a2, ..., a j-1 ) under the condition that model m k Generate the next phrase a j The same calculation method is used to calculate the loss of any evaluation data on any benchmark language model; Then, in any baseline language model m k Calculate the loss Loss(x,θ) of any training data x k ) about the parameters θ of the baseline language model k The loss gradient feature Is a vector whose each component is the loss Loss(x,θ) of the training data x k ) The partial derivative of each model parameter is calculated as follows: Among them, θ k =(θ k,1 ,θ k,2 …,θ k,q ) is the baseline language model m k The parameter vector of q represents the baseline language model m k The number of parameters included, θ k,i Represents the baseline language model m k The loss gradient characteristics of the parameters of any benchmark language model are calculated using the same method.

4. The data attribution method based on dimensionality reduction projection according to claim 3 is characterized in that: The step S3 specifically includes the following sub-steps S31-S32; Step S31: Calculate the second-order moment of the loss gradient feature of each training data and each evaluation data on different baseline models relative to each parameter of the baseline language model; Step S32: Correct the loss gradient characteristics of each training data and each evaluation data on different benchmark models.

5. The data attribution method based on dimensionality reduction projection according to claim 4 is characterized in that: In step S31, for any training data x in any benchmark language model m k The loss gradient feature on Calculate the loss gradient feature Relative to the baseline language model m k A parameter θ k,i The second moment of The process of calculating the second-order moment is as follows: Among them, mean_i is the mean of all training data x in the training dataset X about the baseline language model m k A parameter θ k,i The loss gradient feature The mean of It refers to the loss Loss(x,θ) of any training data x k ) for the baseline language model m k The i-th parameter θ k,i The partial derivative of N is the number of training data x contained in the training dataset X; any evaluation data y in any benchmark language model m k The loss gradient feature on the baseline language model m k Any parameter θ k,i The second moment of is calculated in the same way.

6. The data attribution method based on dimensionality reduction projection according to claim 5 is characterized in that: In step S32, any training data x is placed in any benchmark language model m k The loss gradient feature on Relative to the baseline language model m k A parameter θ k,i The gradient of the training data x is divided by the baseline language model m k The loss gradient feature on Relative to the baseline language model m k A parameter θ k,i The second moment of The square root of the training data x in the benchmark language model m k The corrected loss gradient feature on Any evaluation data y on any benchmark language model m k The same correction method is used for the loss gradient features on .

7. The data attribution method based on dimensionality reduction projection according to claim 6 is characterized in that: The step S4 specifically includes the following sub-steps S41-S45; Step S41: Generate multiple random projection matrices (P1, P2, ..., P B ); B represents the number of random projection matrices, that is, the number of dimensionality reduction projections; Step S42: For any training data x in any benchmark language model m k The corrected loss gradient feature on Perform multiple dimensionality reduction projections, and perform this operation B times to obtain the loss gradient feature information after B dimensionality reduction projections; then concatenate the loss gradient feature information after B dimensionality reduction projections in the order of dimensionality reduction projections to obtain the training data x in the benchmark language model m k The loss gradient feature information G after multiple dimensionality reduction projections θk (x); For any evaluation data y in any benchmark language model m k Perform multiple dimensionality reduction projections on the corrected loss gradient features on the benchmark language model to obtain the loss gradient feature information after the multiple dimensionality reduction projections of the evaluation data on the benchmark language model; Step S43: Concatenate the loss gradient feature information after the b-th dimensionality reduction projection of all training data to obtain the total gradient information matrix after the b-th dimensionality reduction projection of all training data, denoted as g θk (X) b ; The same method is used to obtain the total gradient information matrix after any dimensionality reduction projection of all evaluation data; Step S44: For any training data x in any benchmark language model m k The loss gradient feature information G after multiple dimensionality reduction projections on θk (x) Perform Hessian approximation calculation; Perform Hessian approximation on the loss gradient feature information of any evaluation data after multiple dimensionality reduction projections on any benchmark language model; Step S45: Apply Hesse approximation to any training data x in any benchmark language model m. k The loss gradient feature information after multiple dimensionality reduction projections on Perform normalization processing; The loss gradient feature information of any evaluation data after Hessian approximation calculation on any benchmark language model is normalized after multiple dimensionality reduction projections.

8. The data attribution method based on dimensionality reduction projection according to claim 7 is characterized in that: In step S42, the training data x is k The loss gradient feature information after the b-th dimensionality reduction projection Among them, ⊙ is matrix multiplication; G θk (x) = concat(g θk (x)1,g θk (x)2,…,g θk (x) B ); where concat(·) represents a concatenation operation, and the value of b is a positive integer between 1 and B; In step S43, the total gradient information matrix g after the b-th dimensionality reduction projection of all training data is θk (X) b =concat(g θk (x1) b ,g θk (x2) b ,…); where concat(·) represents the concatenation operation, the data x i ∈X, b represents the b-th dimensionality reduction projection, and the value range of b is a positive integer between 1-B.

9. The data attribution method based on dimensionality reduction projection according to claim 7 is characterized in that: In step S5, any training sample x i For any evaluation sample y j Contribution score in, Represents any training data x i Normalized loss gradient feature information on any baseline language model; Represents any evaluation data y j Normalized loss gradient feature information on any baseline language model; |||| represents the computational norm; · represents vector dot product; * represents real number multiplication.

10. A data attribution device based on dimensionality reduction projection, characterized in that: It includes data set and model preparation unit, loss gradient feature calculation unit, loss gradient feature correction unit, dimensionality reduction projection calculation unit, and data attribution calculation unit; The dataset and model preparation unit is used to prepare a training dataset, an evaluation dataset, and multiple benchmark language models; The loss gradient feature calculation unit is used to calculate the loss gradient features of each training data and each evaluation data on different benchmark models; The loss gradient feature correction unit is used to correct the loss gradient features of each training data and each evaluation data on different benchmark models to obtain the corrected loss gradient features of each training data and each evaluation data on different benchmark models; The dimensionality reduction projection calculation unit is used to perform multiple dimensionality reduction projections on the corrected loss gradient features of each training data and each evaluation data on different benchmark models to obtain normalized loss gradient feature information of each training data and each evaluation data on different benchmark language models; The data attribution calculation unit is used to calculate the data attribution score between each training data and each evaluation data.

Citation Information

Cited By

  • Large model privacy knowledge decoupling forgetting method and system based on influence function decomposition

    CN121615182A

  • Influence function decomposition-based large model privacy knowledge decoupling forgetting method and system

    CN121615182B