Large Model Training Method, Medium and System Based on Differential Privacy Mechanism

By performing data preprocessing, grouping annotation, gradient noise addition and privacy budget adjustment in large model training, the unfairness problem caused by the differential privacy mechanism is solved, and the fairness of user privacy protection and model training is achieved.

CN119494408BActive Publication Date: 2025-07-18XIAMEN UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510066183.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-07-18
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The unfairness caused by the differential privacy mechanism in large-scale training, especially on groups with small cardinality or sparse data distribution, resulting in unfairness in decision-making and statistical analysis.

Method used

By obtaining historical data for preprocessing and grouping annotation, initializing large language model parameters, calculating the loss function gradient and performing gradient noise addition, calculating the value of the comprehensive unfairness indicator, adjusting the privacy budget to control the noise intensity, ensuring that the unfairness indicator is within the preset range, and forming the final model.

Benefits of technology

While protecting user privacy, it avoids unfairness caused by differential privacy mechanisms and ensures fairness of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119494408B_ABST
    Figure CN119494408B_ABST
Patent Text Reader

Abstract

The present invention discloses a large model training method, medium and system based on differential privacy mechanism. The method includes: S101, obtaining historical data, performing preprocessing, and grouping and annotating the preprocessed historical data to form a training data set; S102, initializing the parameters of the large language model; S103, training based on the training data set and calculating gradients; S104, adding noise to the gradients to obtain the noisy gradients, and calculating the corresponding comprehensive unfairness index value based on the noisy gradients; S105, determining whether the comprehensive unfairness index value is within a preset value range; S106, if the comprehensive unfairness index value is within the preset value range, determining whether the current large language model meets the training requirements; if so, using the current large language model as the final model; if not, returning to step S103. It can effectively protect user privacy, and at the same time, avoid the generation of unfair phenomena caused by using the differential privacy mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of language processing technologies, and particularly relates to a large model training method, medium, and system based on a differential privacy mechanism. Background Art

[0002] With the increasing demand for data privacy protection, differential privacy (DP), as an important privacy protection technology, has gradually been widely applied in data analysis and machine learning. Differential privacy protects data privacy by adding random noise to the calculation results to ensure that it is difficult for attackers to infer whether the data of a single individual exists in the dataset through the output.

[0003] In related technologies, when using the differential privacy mechanism for user privacy protection, the consideration of fairness is lacking. That is to say, the differential privacy mechanism protects privacy by adding random noise to the large model training parameters, but the randomness of this noise will have inconsistent effects on different data groups in the original training set. Specifically, groups with a smaller cardinality or sparse data distribution are more vulnerable to noise, resulting in these groups being systematically marginalized in decision-making and statistical analysis, and thus generating unfair phenomena. Summary of the Invention

[0004] The present invention aims to at least partly solve one of the technical problems in related technologies. For this purpose, an object of the present invention is to propose a large model training method based on a differential privacy mechanism, which can effectively protect user privacy and at the same time avoid the generation of unfair phenomena caused by using the differential privacy mechanism.

[0005] In a first aspect, an embodiment of the present invention proposes a large model training method based on a differential privacy mechanism, including the following steps: S101, obtaining historical data, preprocessing the historical data, and grouping and annotating the preprocessed historical data to form a training dataset; S102, initializing the parameters of the large language model; S103, training the large language model based on the training dataset and calculating the gradient of the loss function with respect to the model parameters; S104, adding noise to the gradient to obtain a noisy gradient, and calculating the corresponding comprehensive unfairness index value based on the noisy gradient; S105, determining whether the comprehensive unfairness index value is within a preset value range; S106, if the comprehensive unfairness index value is within the preset value range, determining whether the current large language model meets the training requirements; if so, taking the current large language model as the final model; if not, returning to step S103.

[0006] According to the large model training method based on the differential privacy mechanism in the embodiments of the present invention, first, in S101, historical data is obtained, and the historical data is preprocessed, and the preprocessed historical data is grouped and labeled to form a training data set; then, in S102, the parameters of the large language model are initialized; then, in S103, the large language model is trained based on the training data set, and the gradient of the loss function with respect to the model parameters is calculated; then, in S104, noise is added to the gradient to obtain a noisy gradient, and the corresponding comprehensive unfairness index value is calculated based on the noisy gradient; then, in S105, it is determined whether the comprehensive unfairness index value is within a preset value range; then, in S106, if the comprehensive unfairness index value is within the preset value range, it is determined whether the current large language model meets the training requirements; if so, the current large language model is used as the final model; if not, return to step S103. Thus, effective protection of user privacy is achieved, and at the same time, the generation of unfair phenomena caused by the use of the differential privacy mechanism is avoided.

[0007] In some embodiments, preprocessing the historical data includes: cleaning the historical data, and labeling the cleaned historical data; performing word segmentation and word vector mapping on the labeled historical data, and performing sequence length processing and normalization processing on the mapped word vectors.

[0008] In some embodiments, training the large language model based on the training data set and calculating the gradient of the loss function with respect to the model parameters includes: for each batch of data in the training data set, performing forward propagation to obtain the corresponding prediction result; calculating the error between the prediction result and the true label based on the cross-entropy loss function, and calculating the gradient of the loss function with respect to the model parameters according to the error using the chain rule.

[0009] In some embodiments, the cross-entropy loss function is expressed by the following formula:

[0010] ;

[0011] where represents the cross-entropy loss function, represents the batch size, represents the number of classes, represents the true label of the th sample in the th class, represents the predicted probability;

[0012] The gradient is calculated by the following formula:

[0013] ;

[0014] Among them, represents the gradient, represents the set of model parameters, represents the th loss of the sample.

[0015] In some embodiments, gradient noise addition is performed to obtain a noisy gradient, and a corresponding unfairness index value is calculated based on the noisy gradient, including: clipping the gradient, generating initial noise, and calculating the noisy gradient according to the initial noise and the clipped gradient; calculating an initial unfairness index value based on the noisy gradient, and calculating a corresponding average unfairness index value according to the initial unfairness index values obtained from multiple trainings; introducing a time decay factor to decay the average unfairness index value, and performing a weighted average calculation on the decayed average unfairness index value to obtain a comprehensive unfairness index value.

[0016] In some embodiments, the initial unfairness index value is calculated by the following formula:

[0017] ;

[0018] Among them, represents the initial unfairness index value, represents the penalty coefficient, represents the data population of the noise, represents the population of the noise, represents the sign function, which is used to represent the direction of the noise, represents the data population corresponding noisy gradient, represents the data population corresponding noisy gradient;

[0019] The average unfairness index is calculated by the following formula:

[0020] ;

[0021] Among them, represents the average unfairness index, represents the total number of iterative trainings, represents the th fairness calculation;

[0022] The average unfairness index value is decayed by the following formula:

[0023] ;

[0024] Among them, Represents the average unfairness metric value after attenuation, represents the time interval between the current time and the th training time;

[0025] The weighted average of the average unfairness metric value after attenuation is calculated through the following formula:

[0026] ;

[0027] ;

[0028] where, represents the unfairness metric value after weighted average, represents the proportion of the degree of unfairness between group and group in the entire data set, represents the comprehensive unfairness metric value.

[0029] In some embodiments, if the comprehensive unfairness metric value is not within the preset value range, the differential privacy model is modified to adjust the privacy budget of each group; the noise intensity is updated according to the adjusted privacy budget, and step S104 is returned.

[0030] In some embodiments, the privacy budget is adjusted through the following formula:

[0031] ;

[0032] ;

[0033] where, represents the adjusted privacy budget, represents the initial privacy budget, represents the comprehensive unfairness metric value, represents the adjustment function, represents the adjustment coefficient;

[0034] ;

[0035] where, represents the updated noise intensity, represents the threshold for gradient norm clipping.

[0036] In a second aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a large model training program based on the differential privacy mechanism is stored. When the large model training program based on the differential privacy mechanism is executed by a processor, the above-mentioned large model training method based on the differential privacy mechanism is implemented.

[0037] In a third aspect, an embodiment of the present invention provides a large model training system based on differential privacy mechanism, including: an acquisition module for acquiring historical data, preprocessing the historical data, and grouping and annotating the preprocessed historical data to form a training data set; an initialization module for initializing the parameters of a large language model; a training module for training the large language model based on the training data set and calculating the gradient of the loss function with respect to the model parameters; a fairness calculation module for adding noise to the gradient to obtain a noisy gradient and calculating the corresponding comprehensive unfairness index value based on the noisy gradient; a judgment module for judging whether the comprehensive unfairness index value is within a preset value range; the judgment module is further configured to judge whether the current large language model meets the training requirements when the comprehensive unfairness index value is within the preset value range, and when the judgment result is yes, take the current large language model as the final model, and when the judgment result is no, return to the step of training the large language model based on the training data set.

[0038] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a schematic flowchart of a large model training method based on differential privacy mechanism according to an embodiment of the present invention;

[0040] Figure 2 is a schematic flowchart of a large model training method based on differential privacy mechanism according to another embodiment of the present invention;

[0041] Figure 3 is a schematic block diagram of a large model training system based on differential privacy mechanism according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0043] A large model training method based on differential privacy mechanism according to an embodiment of the present invention will be described below with reference to the accompanying drawings.

[0044] Please refer to Figure 1 , Figure 1 is a schematic flowchart of a large model training method based on differential privacy mechanism according to an embodiment of the present invention, as shown in Figure 1As shown in the figure, the large model training method based on the differential privacy mechanism includes the following steps:

[0045] S101. Obtain historical data, preprocess the historical data, and group and label the preprocessed historical data to form a training data set.

[0046] In some embodiments, preprocessing the historical data includes: cleaning the historical data, and performing data annotation on the cleaned historical data; performing word segmentation and word vector mapping on the annotated historical data, and performing sequence length processing and normalization processing on the mapped word vectors.

[0047] As an example, first, obtain a large amount of text data (i.e., historical data) from multiple trusted data sources to ensure the diversity and representativeness of the data. Among them, the data sources may include but are not limited to: public corpora (such as Wikipedia, news articles, books, etc.); domain data sets (such as text data in professional fields such as medical, financial, legal, etc.); user-generated content (such as, with user authorization, collecting text data on platforms such as social media and forums, etc.).

[0048] Next, perform data cleaning on the obtained historical data; specifically, the process of data cleaning may include:

[0049] 1. Remove noise and special symbols, that is, use regular expression matching to delete meaningless content such as garbled characters, emojis, HTML tags, etc.;

[0050] 2. Process duplicate data, that is, detect and delete duplicate sentences or paragraphs to prevent the model from overfitting to duplicate content;

[0051] 3. Correct spelling mistakes, that is, use a spelling check tool to correct spelling mistakes in the text data;

[0052] 4. Filter sensitive information, identify and delete sensitive information that may contain personal privacy, such as names, addresses, ID numbers, etc.

[0053] Then, perform data annotation on the cleaned historical data. Specifically, data annotation may include adding labels and sensitive attribute annotation; first, for tasks that require supervised learning, add corresponding labels; for example, sentiment polarity, topic category, etc.; then, according to a predefined set of sensitive attributes A (such as gender, race, age, region, etc.), label the group to which each data sample belongs for subsequent fairness evaluation.

[0054] Next, preprocess the labeled historical data. First, use a tokenization tool to split the text data into a sequence of words or sub-words; then, map the words to word vectors using pre-trained word embeddings or through an embedding layer; then, unify the sequence length to L. For sequences shorter than L, pad them with special symbols, and for sequences longer than L, truncate them; then, standardize the word vectors so that they follow a distribution with a mean of 0 and a variance of 1 to improve the stability of model training.

[0055] In some embodiments, after completing the data preprocessing, the data is further grouped and labeled; specifically, the training dataset can be divided into multiple groups according to the sensitive attributes labeled in the preprocessing. ; where:

[0056] ;

[0057] Among them, represents the value of the th sensitive attribute, such as male or female in gender.

[0058] Next, count the data volume of each group , the total data volume is , then the weight of group is:

[0059] ;

[0060] It should be noted that in subsequent fairness evaluation and optimization, the weights can be used for weighted averaging to reflect the proportion of different groups in the dataset.

[0061] S102, Initialize the parameters of the large language model.

[0062] As an example, based on the Transformer architecture, build a large language model and determine the following hyperparameters: the number of layers L, that is, the number of stacked Transformer encoders; the number of hidden units: that is, the dimension of the hidden state vector at each position; the dimension of the feed-forward network: that is, the dimension of the feed-forward fully connected network; the number of attention heads: that is, the number of heads in the multi-head attention mechanism, and the dimension of each head is . Then, use the Xavier initialization method to initialize the model parameters. For the weight matrix , its elements are randomly initialized according to the following distribution:

[0063] ;

[0064] Among them, and respectively represent the input dimension and the output dimension.

[0065] S103, training the large language model based on the training dataset, and calculating the gradient of the loss function with respect to the model parameters.

[0066] In some embodiments, training the large language model based on the training dataset and calculating the gradient of the loss function with respect to the model parameters includes: for each batch of data in the training dataset, performing forward propagation to obtain the corresponding prediction results; calculating the error between the prediction results and the true labels based on the cross-entropy loss function, and calculating the gradient of the loss function with respect to the model parameters according to the error using the chain rule.

[0067] In some embodiments, the cross-entropy loss function is expressed by the following formula:

[0068]

[0069] where, represents the cross-entropy loss function, represents the batch size, represents the number of classes, represents the th true label of the th sample in the

[0070] The gradient is calculated by the following formula:

[0071] ;

[0072] where, represents the gradient, represents the set of model parameters, represents the th sample's loss.

[0073] As an example, first, the word vectors of the input sequence are passed through an embedding layer and positional encoding to obtain an initial input representation :

[0074] ;

[0075] where, represents the initial input representation, represents the word vector, represents the positional encoding (Positional Encoding), which is used to encode the position information of the word in the sequence.

[0076] Next, is sequentially passed through L Transformer encoder layers to obtain the final representation :

[0077] ;

[0078] Then, input into the fully connected layer to output the corresponding prediction result through the fully connected layer :

[0079] ;

[0080] Among them, represents the prediction result, represents the weight matrix of the fully connected layer, represents the bias vector of the fully connected layer.

[0081] Next, use the cross-entropy loss function to calculate the error value between the prediction result and the true label:

[0082] ;

[0083] Among them, represents the cross-entropy loss function, represents the batch size, represents the number of classes, represents the th sample's true label in the th class, represents the predicted probability;

[0084] Then, use the chain rule to calculate the gradient of the loss function with respect to the model parameters:

[0085] ;

[0086] Among them, represents the gradient, represents the set of model parameters, represents the loss of the th sample.

[0087] S104. Perform gradient noise addition to obtain the noisy gradient, and calculate the corresponding comprehensive unfairness index value based on the noisy gradient.

[0088] In some embodiments, gradient noise addition is performed to obtain a noisy gradient, and a corresponding unfairness metric value is calculated based on the noisy gradient, including: clipping the gradient, generating initial noise, and calculating the noisy gradient according to the initial noise and the clipped gradient; calculating an initial unfairness metric value based on the noisy gradient, and calculating a corresponding average unfairness metric value according to the initial unfairness metric values obtained from multiple trainings; introducing a time decay factor to decay the average unfairness metric value, and performing a weighted average calculation on the decayed average unfairness metric value to obtain a comprehensive unfairness metric value.

[0089] In some embodiments, the initial unfairness metric value is calculated by the following formula:

[0090] ;

[0091] where, represents the initial unfairness metric value, represents the penalty coefficient, represents the data population of the noise, represents the population of the noise, represents the sign function, which is used to represent the direction of the noise, represents the data population corresponding noisy gradient, represents the data population corresponding noisy gradient;

[0092] The average unfairness metric is calculated by the following formula:

[0093] ;

[0094] where, represents the average unfairness metric, represents the total number of iterative trainings, represents the th fairness calculation;

[0095] The average unfairness metric value is decayed by the following formula:

[0096] ;

[0097] where, represents the decayed average unfairness metric value, represents the time interval between the current time and the th training time;

[0098] The decayed average unfairness metric value is weighted averaged by the following formula:

[0099] ;

[0100] ;

[0101] Among them, represents the unfairness index value after weighted average, represents the group and the group The proportion of the degree of unfairness between them in the entire data set, represents the comprehensive unfairness index value.

[0102] As an example, first, to prevent abnormally large gradients from affecting the training of the model, the gradients of each sample are clipped first:

[0103] ;

[0104] Among them, represents the clipped gradient, represents the original gradient, represents the preset gradient norm threshold.

[0105] Then, according to the definition of differential privacy Define, add noise that conforms to the Gaussian noise mechanism:

[0106] ;

[0107] Among them, represents the intensity of the noise, and the calculation formula is:

[0108] ;

[0109] Among them, represents the failure probability, usually taking a relatively small value, for example, .

[0110] Then, add noise to the clipped gradient to obtain the noisy gradient:

[0111] ;

[0112] Next, for any two data groups and , define the unfairness in one training, that is, the initial unfairness index value is:

[0113] ;

[0114] Among them, represents the initial unfairness index value, denotes a penalty coefficient, greater than 1, used to amplify the impact of reverse noise; preferably, it can take , denotes the noise of the data population . denotes the noise of the population . denotes the sign function, used to represent the direction of the noise, denotes the noisy gradient corresponding to the data population . denotes the noisy gradient corresponding to the data population .

[0115] Then, calculate the average unfairness index:

[0116] ;

[0117] wherein, denotes the average unfairness index, denotes the total number of iterative trainings, denotes the th fairness calculation.

[0118] Next, introduce the time decay factor:

[0119] ;

[0120] wherein, denotes the value of the average unfairness index after decay, denotes the time interval between the current time and the time of the th training.

[0121] Then, considering the population weights, calculate the weighted average unfairness:

[0122] ;

[0123] Finally, obtain the comprehensive unfairness index:

[0124] ;

[0125] wherein, denotes the value of the unfairness index after weighted averaging, denotes the population and the population the proportion of the degree of unfairness between them in the entire dataset, denotes the value of the comprehensive unfairness index.

[0126] S105, determine whether the value of the comprehensive unfairness index is within the preset value range.

[0127] S106. If the comprehensive unfairness index value is within the preset value range, determine whether the current large language model meets the training requirements; if so, use the current large language model as the final model; if not, return to step S103.

[0128] In some embodiments, if the comprehensive unfairness index value is not within the preset value range, modify the differential privacy model to adjust the privacy budget for each group; update the noise intensity according to the adjusted privacy budget, and return to step S104.

[0129] In some embodiments, the privacy budget is adjusted by the following formula:

[0130] ;

[0131] ;

[0132] where, represents the adjusted privacy budget, represents the initial privacy budget, represents the comprehensive unfairness index value, represents the adjustment function, represents the adjustment coefficient;

[0133] ;

[0134] where, represents the updated noise intensity, represents the threshold for gradient norm clipping.

[0135] As an example, first, dynamically adjust the privacy budget for each group according to the fairness index :

[0136] ;

[0137] ;

[0138] where, represents the adjusted privacy budget, represents the initial privacy budget, represents the comprehensive unfairness index value, represents the adjustment function, represents the adjustment coefficient;

[0139] Then, update the noise intensity according to the new privacy budget:

[0140] ;

[0141] Then, construct the unfairness matrix:

[0142] ;

[0143] Next, perform eigenvalue decomposition on the matrix:

[0144] ;

[0145] Among them, represents the eigenvector matrix, represents the focusing matrix, which contains eigenvalues.

[0146] Then, construct the fairness loss function:

[0147] ;

[0148] Among them, 1 is a matrix with all elements being 1, and are weight coefficients.

[0149] Next, perform gradient descent to optimize the privacy budget:

[0150] ;

[0151] Among them, represents the learning rate, is the gradient of the loss function with respect to the privacy budget, which can be calculated by the chain rule.

[0152] Then, according to the optimized privacy budget , update the noise intensity , and regenerate the noise .

[0153] Next, use the new noise to update the noisy gradient:

[0154] ;

[0155] Then, average the noisy gradients within the batch to obtain the global gradient:

[0156] ;

[0157] Next, use the optimizer to update the model parameters:

[0158] ;

[0159] Among them, represents the learning rate.

[0160] In some embodiments, there can be multiple ways to determine whether the current large language model meets the training requirements. For example: determining whether the current number of training rounds has reached the preset number of rounds; or, determining whether the loss function of the model has converged, that is, whether the loss change between adjacent iterations is less than the preset threshold; or, whether the performance metrics of the model meet the expected requirements; the determination methods for whether the training requirements are met are not limited herein.

[0161] In some embodiments, the training method further includes validating and evaluating the model.

[0162] As an example, first, evaluate the model performance, use the validation set, and evaluate the performance metrics of the model:

[0163] ;

[0164] where represents the accuracy rate.

[0165] In addition, it should be noted that metrics such as precision, recall, and F1-score can be selected according to the task requirements.

[0166] Next, for different data groups, calculate the performance metrics of the model on each group and check whether there are significant differences.

[0167] First, calculate the mean difference:

[0168] ;

[0169] Equal opportunity:

[0170] ;

[0171] If and are close to 0, it indicates that the model has better fairness among different groups.

[0172] As a specific embodiment of the present invention, as Figure 2 shown, the large model training method based on the differential privacy mechanism disclosed in the present invention specifically includes the following steps:

[0173] S201, obtain historical data, preprocess the historical data, and group and label the preprocessed historical data to form a training data set.

[0174] S202, initialize the parameters of the large language model.

[0175] S203, train the large language model based on the training data set and calculate the gradient of the loss function with respect to the model parameters.

[0176] S204, clip the gradient and generate initial noise, and use the initial noise as the current noise.

[0177] S205, calculate the noise-added gradient based on the current noise and the clipped gradient.

[0178] S206, calculate the initial unfairness index value based on the noise-added gradient.

[0179] S207, calculate the corresponding average unfairness index value according to the initial unfairness index values obtained from multiple trainings.

[0180] S208, introduce a time decay factor to decay the average unfairness index value.

[0181] S209, perform a weighted average calculation on the decayed average unfairness index value to obtain a comprehensive unfairness index value.

[0182] S210, determine whether the comprehensive unfairness index value is within a preset value range; if so, execute step S211; if not, execute step S212.

[0183] S211, determine whether the current large language model meets the training requirements; if so, execute step S214; if not, return to step S203.

[0184] S212, modify the differential privacy model to adjust the privacy budget of each group.

[0185] S213, update the current noise according to the adjusted privacy budget and return to step S205.

[0186] S214, use the current large language model as the final model.

[0187] As an example, first, obtain historical data, preprocess the historical data, and group and label the preprocessed historical data to form a training dataset; then, initialize the parameters of the large language model; then, train the large language model based on the training dataset and calculate the gradient of the loss function with respect to the model parameters; then, for each sample, the calculated gradient , perform a gradient clipping operation to limit the two-norm of the gradient not to exceed a preset threshold:

[0188] ;

[0189] Then, generate initial noise: According to the definition of differential privacy, add noise that conforms to the Gaussian noise mechanism:

[0190] ;

[0191] Among them, represents the intensity of noise, and the calculation formula is:

[0192] ;

[0193] Among them, represents the failure probability, usually taking a relatively small value. For example, .

[0194] Next, calculate the noise-added gradient:

[0195] ;

[0196] Then, for any two data groups and , define the unfairness in one training, that is, the initial unfairness index value is:

[0197] ;

[0198] Among them, represents the initial unfairness index value, represents the penalty coefficient, which is greater than 1 and is used to amplify the influence of the reverse noise; preferably, it can be taken as , represents the noise of the data group , represents the noise of the group , represents the sign function, which is used to represent the direction of the noise, represents the data group corresponding noise-added gradient, represents the data group corresponding noise-added gradient.

[0199] Then, calculate the average unfairness index:

[0200] ;

[0201] Among them, represents the average unfairness index, represents the total number of iterative trainings, represents the th fairness calculation.

[0202] Next, introduce the time decay factor:

[0203] ;

[0204] Among them, represents the value of the average unfairness index after attenuation, represents the current time and the The time interval between consecutive training times.

[0205] Then, considering the group weights, calculate the weighted average unfairness:

[0206] ;

[0207] Finally, obtain the comprehensive unfairness index:

[0208] ;

[0209] Among them, represents the value of the unfairness index after weighted average, represents the group and the group The proportion of the degree of unfairness between them within the entire dataset, represents the value of the comprehensive unfairness index.

[0210] Purpose: To quantitatively evaluate the degree of unfairness between different groups during model training and provide a basis for subsequent adjustments.

[0211] Next, determine whether the comprehensive unfairness index is within an acceptable range, that is, whether it is within the preset value range;

[0212] If so, continue with model training without making adjustments;

[0213] If not, further adjust the differential privacy model.

[0214] Specifically, according to the comprehensive unfairness index, dynamically adjust the privacy budget:

[0215] ;

[0216] ;

[0217] Among them, represents the adjusted privacy budget, represents the initial privacy budget, represents the value of the comprehensive unfairness index, represents the adjustment function, represents the adjustment coefficient;

[0218] Next, update the noise intensity:

[0219] ;

[0220] Then, using the updated noise intensity, regenerate the noise:

[0221] ;

[0222] Next, using the new noise, update the noisy gradient:

[0223] ;

[0224] Then, using the updated noisy gradient, recalculate the comprehensive unfairness index value, and determine whether the adjusted model meets the fairness requirements; if not, continue to adjust the privacy budget and noise until the requirements are met.

[0225] In summary, according to the large model training method based on the differential privacy mechanism of the embodiments of the present invention, first, in S101, obtain historical data, preprocess the historical data, and group and label the preprocessed historical data to form a training data set; then, in S102, initialize the large language model parameters; then, in S103, train the large language model based on the training data set, and calculate the gradient of the loss function with respect to the model parameters; then, in S104, add noise to the gradient to obtain a noisy gradient, and calculate the corresponding comprehensive unfairness index value based on the noisy gradient; then, in S105, determine whether the comprehensive unfairness index value is within a preset value range; then, in S106, if the comprehensive unfairness index value is within the preset value range, determine whether the current large language model meets the training requirements; if so, use the current large language model as the final model; if not, return to step S103. Thus, effective protection of user privacy is achieved, and at the same time, the generation of unfair phenomena caused by using the differential privacy mechanism is avoided.

[0226] In a second aspect, an embodiment of the present invention proposes a computer-readable storage medium, on which a large model training program based on the differential privacy mechanism is stored. When the large model training program based on the differential privacy mechanism is executed by a processor, the large model training method based on the differential privacy mechanism as described above is implemented.

[0227] In a third aspect, as Figure 3 shown, an embodiment of the present invention proposes a large model training system based on the differential privacy mechanism. The large model training system based on the differential privacy mechanism includes: an acquisition module 10, an initialization module 20, a training module 30, a fairness calculation module 40, and a judgment module 50.

[0228] Among them, the acquisition module 10 is used to obtain historical data, preprocess the historical data, and group and label the preprocessed historical data to form a training data set;

[0229] The initialization module 20 is used to initialize the large language model parameters;

[0230] The training module 30 is used to train the large language model based on the training dataset and calculate the gradient of the loss function with respect to the model parameters;

[0231] The fairness calculation module 40 is used to add noise to the gradient to obtain the noisy gradient and calculate the corresponding comprehensive unfairness index value based on the noisy gradient;

[0232] The judgment module 50 is used to judge whether the comprehensive unfairness index value is within a preset value range;

[0233] The judgment module 50 is also used to, when the comprehensive unfairness index value is within the preset value range, judge whether the current large language model meets the training requirements, and when the judgment result is yes, use the current large language model as the final model, and when the judgment result is no, return to the step of training the large language model based on the training dataset.

[0234] It should be noted that the above description of the large model training method based on the differential privacy mechanism also applies to this large model training system based on the differential privacy mechanism, and will not be elaborated here.

[0235] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or otherwise processing as appropriate, and then storing it in a computer memory.

[0236] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0237] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0238] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0239] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0240] In the present invention, unless otherwise clearly specified or limited, terms such as "installed", "connected", "coupled", "fixed", etc. shall be construed broadly. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two components or the interaction relationship between two components, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0241] In the present invention, unless otherwise clearly specified or limited, a first feature being "on" or "under" a second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, a first feature being "above", "over" and "on top of" a second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the horizontal height of the first feature is higher than that of the second feature. A first feature being "under", "below" and "beneath" a second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the horizontal height of the first feature is less than that of the second feature.

[0242] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A large model training method based on differential privacy mechanism, characterized in that, It includes the following steps: S101, Obtain text historical data, preprocess the text historical data, and group and label the preprocessed text historical data to form a training dataset; S102, Initialize the parameters of the large language model; S103, Train the large language model based on the training dataset and calculate the gradient of the loss function with respect to the model parameters; S104, Add noise to the gradient to obtain a noisy gradient, and calculate the corresponding comprehensive unfairness index value based on the noisy gradient; S105, Determine whether the comprehensive unfairness index value is within a preset value range; S106, If the comprehensive unfairness index value is within the preset value range, determine whether the current large language model meets the training requirements; If yes, use the current large language model as the final model; if no, return to step S103; Among them, adding noise to the gradient to obtain a noisy gradient and calculating the corresponding unfairness index value based on the noisy gradient includes: Clip the gradient, generate initial noise, and calculate the noisy gradient according to the initial noise and the clipped gradient; Calculate the initial unfairness index value based on the noisy gradient, and calculate the corresponding average unfairness index value according to the initial unfairness index values obtained from multiple trainings; Introduce a time decay factor to decay the average unfairness index value, and perform weighted average calculation on the decayed average unfairness index value to obtain the comprehensive unfairness index value; The initial unfairness index value is calculated by the following formula: ; Among them, represents the initial unfairness index value, represents the penalty coefficient, represents the data population noise, represents the population noise, represents the sign function, which is used to represent the direction of the noise, represents the data population corresponding noise-added gradient, represents the data population corresponding noise-added gradient; The average unfairness index is calculated by the following formula: ; Among them, represents the average unfairness index, represents the total number of iterative trainings, represents the th fairness calculation, and the random error is reduced by averaging multiple calculations; The average unfairness index value is decayed by the following formula: ; Among them, represents the average unfairness index value after attenuation, represents the time interval between the current time and the th training time; The weighted average of the decayed average unfairness index value is calculated by the following formula: ; ; Among them, represents the unfairness index value after weighted average, represents the group and the group the proportion of the degree of unfairness between them in the entire dataset, represents the comprehensive unfairness index value.

2. The large model training method based on the differential privacy mechanism according to claim 1, characterized in that Preprocessing the text historical data includes: Clean the text historical data and perform data annotation on the cleaned text historical data; Perform word segmentation and word vector mapping on the annotated text historical data, and perform sequence length processing and normalization processing on the mapped word vectors.

3. The large model training method based on the differential privacy mechanism according to claim 1, characterized in that Training the large language model based on the training dataset and calculating the gradient of the loss function with respect to the model parameters includes: For each batch of data in the training dataset, perform forward propagation to obtain the corresponding prediction result; Calculate the error between the prediction result and the true label based on the cross-entropy loss function, and calculate the gradient of the loss function with respect to the model parameters according to the error using the chain rule.

4. The large model training method based on the differential privacy mechanism according to claim 3, wherein, The cross-entropy loss function is expressed by the following formula: ; Among them, represents the cross-entropy loss function, represents the batch size, represents the number of classes, represents the th sample's true label in the th class, represents the predicted probability; The gradient is calculated by the following formula: ; Among them, denotes the gradient, denotes the set of model parameters, denotes the loss of the 5. The large model training method based on the differential privacy mechanism according to claim 1, wherein If the comprehensive unfairness index value is not within the preset value range, modify the differential privacy model to adjust the privacy budget of each group; Update the noise intensity according to the adjusted privacy budget and return to step S104.

6. The large model training method based on the differential privacy mechanism according to claim 5, characterized in that, The privacy budget is adjusted by the following formula: ; ; Among them, represents the adjusted privacy budget, represents the initial privacy budget, represents the comprehensive unfairness index value, represents the adjustment function, represents the adjustment coefficient; ; Among them, represents the updated noise intensity, represents the threshold for gradient norm clipping.

7. A computer-readable storage medium, characterized in that, It stores a large model training program based on the differential privacy mechanism. When the large model training program based on the differential privacy mechanism is executed by a processor, it implements the large model training method based on the differential privacy mechanism described in any one of claims 1-6.

8. A large model training system based on the differential privacy mechanism, characterized in that, It includes: An acquisition module, which is used to acquire text historical data, preprocess the text historical data, and group and label the preprocessed text historical data to form a training data set; An initialization module, which is used to initialize the parameters of the large language model; A training module, which is used to train the large language model based on the training data set and calculate the gradient of the loss function with respect to the model parameters; A fairness calculation module, which is used to add noise to the gradient to obtain a noisy gradient, and calculate the corresponding comprehensive unfairness index value based on the noisy gradient; A judgment module, which is used to judge whether the comprehensive unfairness index value is within a preset value range; The judgment module is further configured to, when the comprehensive unfairness index value is within the preset value range, judge whether the current large language model meets the training requirements, and when the judgment result is yes, use the current large language model as the final model, and when the judgment result is no, return to the step of training the large language model based on the training data set; Among them, adding noise to the gradient to obtain a noisy gradient, and calculating the corresponding unfairness index value based on the noisy gradient includes: Clipping the gradient, generating initial noise, and calculating the noisy gradient according to the initial noise and the clipped gradient; Calculating an initial unfairness index value based on the noisy gradient, and calculating the corresponding average unfairness index value according to the initial unfairness index values obtained from multiple trainings; Introducing a time decay factor to decay the average unfairness index value, and performing a weighted average calculation on the decayed average unfairness index value to obtain a comprehensive unfairness index value; The initial unfairness index value is calculated by the following formula: ; Among them, represents the initial unfairness index value, represents the penalty coefficient, represents the data population noise, represents the population noise, represents the sign function, which is used to represent the direction of the noise, represents the data population corresponding noisy gradient, represents the data population corresponding noisy gradient; The average unfairness index is calculated by the following formula: ; Among them, represents the average unfairness index, represents the total number of iterative trainings, represents the th fairness calculation, and the random error is reduced by averaging multiple calculations; The average unfairness index value is decayed by the following formula: ; Among them, represents the average unfairness index value after attenuation, represents the time interval between the current time and the th training time; The weighted average of the decayed average unfairness index value is calculated by the following formula: ; ; Among them, represents the unfairness index value after weighted average, represents the group and the group The proportion of the degree of unfairness between them in the entire dataset, represents the comprehensive unfairness index value.

Citation Information

Patent Citations

  • Gradient disturbance machine learning fairness method and system

    CN116150619A