Transfer learning method and device based on privacy sample importance evaluation

By jointly evaluating the multi-teacher model and the student model, high-value samples are screened and processed, solving the problem of unreasonable selection of privacy samples and improving the model training efficiency and fitting ability.

CN120910558AActive Publication Date: 2025-11-07CHONGQING UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511003806.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-07
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

In existing technologies, unreasonable selection of privacy samples leads to wasted privacy budgets and low model training efficiency.

Method used

Multiple teacher models are used to predict samples from a common training set. The voting frequency difference and the uncertainty score of the student model are calculated. The two are combined for joint evaluation to select high-value samples and perform privacy processing to train the student model.

Benefits of technology

It achieves accurate characterization of high-value samples, excludes simple consensus samples, saves privacy budget, extends training cycle, and improves the training efficiency and fitting ability of student models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910558A_ABST
    Figure CN120910558A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and provides a privacy sample importance evaluation-based transfer learning method, which comprises the following steps of: predicting samples in a public training set in sequence by utilizing a plurality of teacher models to obtain a plurality of voting prediction results of each sample; calculating the voting frequency difference of each sample according to the prediction result; predicting an uncertainty score of each sample by using a to-be-trained student model; and screening from the public sample training set according to the voting frequency difference and the uncertainty score to obtain a target sample training set composed of high-value samples, and training a to-be-trained student model by using the target sample training set to obtain a trained student model. According to the method, the samples with high training values can be screened out through double screening of two key dimensions, the adaptability of the student model to complex boundary samples is improved, privacy budget is saved, the fitting ability and generalization ability of the student model are enhanced, and the training efficiency of the student model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a transfer learning based on privacy sample importance evaluation. BACKGROUND

[0002] In the rapid development of artificial intelligence, the performance of deep learning and machine learning models largely depends on high-quality large-scale training data, but many high-value data are in privacy-sensitive fields, and the data in such fields need to be kept secret. In the prior art, differential privacy protection technology is often used to process such privacy data to achieve data privacy protection during the training process. However, in the prior art, when selecting a noisy sample, a random sampling method or a single indicator evaluation is often used to select a sample for noise addition, which cannot distinguish between invalid samples and valid samples, so that the invalid samples are given a privacy budget when the differential privacy method is used to add noise, resulting in waste of privacy budget and reduction of training efficiency.

[0003] It can be seen that the existing technology has the problem of unreasonable selection of privacy training samples, resulting in waste of privacy budget and low model training efficiency. SUMMARY

[0004] Therefore, the present application provides a transfer learning method and device based on privacy sample importance evaluation to solve the problem of unreasonable selection of privacy samples in the prior art, resulting in waste of privacy budget and low model training efficiency.

[0005] In a first aspect, the present application provides a transfer learning method based on privacy sample importance evaluation, which comprises:

[0006] A plurality of teacher models are used to predict samples in a public training set in turn to obtain a plurality of voting prediction results corresponding to each sample, wherein each teacher model is trained by a different privacy training set; a voting frequency difference of each sample is calculated according to the plurality of voting prediction results; an uncertainty score of each sample is predicted by a student model to be trained; the importance of each sample is jointly evaluated according to the voting frequency difference and the uncertainty score to obtain a joint evaluation result of each sample; target sample training sets are obtained by sample screening from the public sample training set according to the joint evaluation result, and the target sample training sets are privacy-processed, and the student model to be trained is trained by using the privacy-processed target sample training sets to obtain a trained student model.

[0007] In a second aspect, the present application provides a transfer learning device based on privacy sample importance evaluation, which comprises:

[0008] The first prediction module is configured to sequentially predict samples in the public training set by using a plurality of teacher models to obtain a plurality of voting prediction results corresponding to each sample, wherein each teacher model is trained by a different private training set; the calculation module is configured to calculate a voting frequency difference of each sample according to the plurality of voting prediction results; the second prediction module is configured to predict an uncertainty score of each sample by using a student model to be trained; the evaluation module is configured to jointly evaluate the importance of each sample according to the voting frequency difference and the uncertainty score to obtain a joint evaluation result of each sample; the training module is configured to perform sample screening from the public sample training set to obtain a target sample training set according to the joint evaluation result, perform privacy processing on the target sample training set, train the student model to be trained by using the target sample training set after the privacy processing, and obtain a trained student model.

[0009] In a third aspect, an electronic device is provided, and the electronic device includes:

[0010] at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the steps of the above method.

[0011] In a fourth aspect, a computer storage medium is provided, and the computer storage medium stores a computer program, and the computer program is executed by a processor to enable the processor to perform the steps of the above method.

[0012] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0013] According to the technical solution provided in this application, multiple teacher models are used to predict samples in a common training set sequentially, obtaining multiple voting prediction results for each sample. Each teacher model is trained on a different privacy training set. Based on the multiple voting prediction results, the voting frequency difference for each sample is calculated. The uncertainty score for each sample is predicted using the student model to be trained. The importance of each sample is jointly evaluated based on the voting frequency difference and uncertainty score, resulting in a joint evaluation result for each sample. Based on the joint evaluation result, a target sample training set is obtained by filtering samples from the common sample training set. The target sample training set is then subjected to privacy processing. The student model to be trained is then trained using the privacy-processed target sample training set to obtain a trained student model. On the one hand, by using two key dimensions—the teacher model and the student model—to evaluate samples, not only is a precise characterization of high-value samples achieved, but simple samples with already reached simple consensus are also excluded, saving privacy budget. This ensures that the effective training period of the student model can be extended under the same privacy budget, thereby improving the training efficiency of the student model. On the other hand, through dual screening across two key dimensions, not only can high-value training samples be selected, but the adaptability of the student model to complex boundary samples can also be improved, thereby enhancing the model's fitting and generalization abilities. This avoids the problems of unreasonable selection of privacy training samples, leading to wasted privacy budgets and low training efficiency, which exist in existing technologies. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating a transfer learning method based on privacy sample importance assessment provided in an embodiment of this application.

[0015] Figure 2 This is a schematic diagram of the structure of a transfer learning device based on privacy sample importance assessment provided in an embodiment of this application;

[0016] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Detailed Implementation

[0017] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0018] In the description of the present application, it should be understood that the terms "longitudinal", "transverse", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship shown in the drawings, which are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.

[0019] In the description of the present application, unless otherwise specified and limited, it should be noted that the terms "mounting", "connection", "connection" should be understood broadly, for example, it can be mechanical connection or electrical connection, it can be the communication between two elements, it can be direct connection or indirect connection through intermediate medium, and the specific meaning of the above terms can be understood by those skilled in the art according to the specific circumstances.

[0020] The execution subject of the transfer learning method based on privacy sample importance evaluation in the present application includes but is not limited to at least one of the electronic devices capable of being configured to execute the method provided by the embodiments of the present application, such as a server and a terminal. In other words, the method can be executed by software or hardware installed in a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be a stand-alone server, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms.

[0021] A transfer learning method and device based on privacy sample importance evaluation according to the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0022] Figure 1 is a flowchart of a transfer learning method based on privacy sample importance evaluation provided by the embodiments of the present application, as Figure 1 shown, the method comprises:

[0023] S101, using a plurality of teacher models to sequentially predict the samples in the public training set to obtain a plurality of voting prediction results corresponding to each sample, wherein each teacher model is trained by a different privacy training set;

[0024] S102, calculating the voting frequency difference of each sample according to the plurality of voting prediction results;

[0025] S103, using a student model to be trained to predict the uncertainty score of each sample;

[0026] S104, jointly evaluate the importance of each sample according to the voting frequency difference and the uncertainty score, to obtain a joint evaluation result of each sample;

[0027] S105, according to the joint evaluation result, sample screening is performed from the public sample training set to obtain a target sample training set, and the target sample training set is subjected to privacy processing, and the student model to be trained is trained by using the target sample training set subjected to privacy processing, to obtain a trained student model.

[0028] Specifically, the plurality of teacher models are teacher models that have been trained, and each teacher model is trained by a different privacy training set, that is, each privacy training set is independent of each other, and the training samples in the privacy training set are non-overlapping, avoiding a single model contacting all privacy samples and reducing the risk of privacy leakage.

[0029] It can be understood that the plurality of teacher models are used to sequentially predict the samples in the public training set, to obtain a plurality of voting prediction results corresponding to each sample, that is, each sample will be processed by a plurality of teacher models to obtain respective prediction results, for example, if there are 10 teacher models and 100 samples, each sample will receive 10 independent prediction results (such as class judgment in a classification task), and these results collectively constitute a plurality of voting prediction results of the sample.

[0030] It should be noted that in the prior art, the selection of samples usually depends on random sampling or a single uncertainty measure. Random sampling means that the voting prediction results of the teacher models are processed without distinction. In the voting stage, only the highest vote label is concerned, and noise is uniformly added on all samples, without considering the difference between the voting of the teacher models. For example, on some samples, the voting among the teacher models is highly dispersed, indicating that the sample is difficult to judge; while on some other samples, the voting opinions of the teacher models are highly consistent, and should be preferentially used for student training. However, the current mechanism cannot distinguish and process this, resulting in excessive allocation of privacy noise budget and low efficiency. Secondly, a single uncertainty measure cannot comprehensively reflect the value of the sample to the learning effect of the student model. For example, taking the softmax entropy as an example, although it can reflect the subjective uncertainty of the student model in predicting a sample, it cannot reflect the voting consistency among the teacher models. Some high-entropy samples may have formed a highly consistent voting result among the teachers, and if these samples continue to be aggregated with differential privacy, not only the valuable privacy budget is wasted, but also the privacy consumption and student benefit are seriously unequal.

[0031] To avoid the above problems, the embodiment calculates the voting frequency difference of each sample according to multiple voting prediction results; meanwhile, the uncertainty score of each sample is predicted by using the student model to be trained; the importance of each sample is jointly evaluated in combination with the two dimensions of the voting frequency difference and the uncertainty score, to obtain the joint evaluation result of each sample, so as to jointly evaluate from the prediction dimension of the teacher model and the prediction dimension of the student model, and realize fine evaluation of the importance of the sample. It can be understood that the use of multiple voting prediction results to calculate the voting frequency difference of each sample, and the joint evaluation of the importance of each sample in combination with the two dimensions of the voting frequency difference and the uncertainty score will be described in detail later, and will not be described here.

[0032] Further, after obtaining the joint evaluation result of each sample, the sample set with higher importance can be screened from the public sample training set according to the joint evaluation result to form a target sample training set, at this time, the privacy processing of the target sample training set can effectively improve the use efficiency of the privacy budget, and finally the student model to be trained is trained by using the target sample training set after differential processing, to obtain the trained student model.

[0033] According to the technical scheme provided by the embodiment of the application, multiple teacher models are used to sequentially predict the samples in the public training set to obtain multiple voting prediction results corresponding to each sample, wherein each teacher model is trained by a different privacy training set; the voting frequency difference of each sample is calculated according to the multiple voting prediction results; the uncertainty score of each sample is predicted by using the student model to be trained; the importance of each sample is jointly evaluated according to the voting frequency difference and the uncertainty score to obtain the joint evaluation result of each sample; the target sample training set is obtained by sample screening from the public sample training set according to the joint evaluation result, and the target sample training set is privacy-processed, and the student model to be trained is trained by using the target sample training set after privacy processing, to obtain the trained student model. On the one hand, by evaluating the samples from the two key dimensions of the teacher model and the student model, not only the accurate description of the high-value samples is realized, but also the simple samples that have reached a simple consensus are excluded, and the privacy budget is saved, so that the effective training period of the student model can be prolonged and the training efficiency of the student model can be improved under the same privacy budget. On the other hand, through the double screening of the two key dimensions, not only the samples with high training value can be screened, but also the adaptability of the student model to complex boundary samples can be improved, so that the fitting ability and the generalization ability of the student model are enhanced. The problem of unreasonable selection of privacy training samples in the prior art, resulting in waste of privacy budget and low training efficiency, is avoided.

[0034] In some embodiments, calculating the voting frequency difference of each sample according to multiple voting prediction results comprises:

[0035] The plurality of voting prediction results corresponding to each sample are counted to obtain a voting frequency vector of each sample, wherein the voting frequency vector comprises the number of votes of each sample in each prediction category; and a difference between the number of votes of the two prediction categories with the highest number of votes in the voting frequency vector is calculated as a voting frequency difference of each sample.

[0036] Specifically, the voting frequency vector represents the number of votes of each sample in each prediction category; after obtaining the plurality of voting prediction results of each sample, the number of votes of each sample in each prediction category can be obtained according to the plurality of voting prediction results, so as to obtain the number distribution of the prediction categories of each sample, and further obtain the voting frequency vector of each sample.

[0037] As an example, in a classification task, a specific voting frequency vector can be represented as: v(x)=[n1(x),n2(x),...,n C (x)];wherein v(x) is the voting frequency vector of sample x, c is the total number of categories of the classification task, n1(x) is the number of votes of a certain category, and further, if the number of votes of A category is 5, the number of votes of B category is 3, and the number of votes of C category is 2 in the 10 voting prediction results of sample x. At this time, n A (x)=5, n B (x)=3, and n C (x)=2.

[0038] Further, a difference between the number of votes of the two prediction categories with the highest number of votes in the voting frequency vector is calculated as a voting frequency difference of each sample.

[0039] It should be noted that the voting frequency difference is used to quantify the prediction consistency of the teacher model. By counting the difference in the frequency of different voting prediction results (such as classification labels) in voting, the difference between the number of votes of different options is quantified to measure the consistency of the teacher model in predicting samples. If the voting frequency difference is large, the label of the sample can be reliably determined, and the teacher model does not need to be called again to determine the label, which can be directly used to train the student model and reduce repeated queries of private data. If the voting frequency difference is small, the teacher model is inconsistent in predicting the label, and cannot accurately determine which result the prediction result is biased towards. It is more difficult to learn the accurate features of such samples than other samples, and these samples are more important because using these samples can speed up the convergence speed and accuracy of the model, so these samples need to be further processed to ensure that the student model learns the key information later. In this way, some simple samples that reach a voting consensus can be preliminarily excluded, and the privacy budget can be preliminarily saved.

[0040] Taking the previous example, at this time, the two prediction categories with the highest number of votes in the voting frequency vector are category A with 5 votes and category B with 3 votes, that is, categories A and B will be used as the basis for calculating the voting frequency difference of the sample, and the voting frequency difference of each sample is calculated, which can be calculated using the following formula:

[0041] Gap(x) = n max (x) - n second (x) ;

[0042] wherein Gap(x) is the voting frequency difference of sample x; n max (x) is the prediction category with the highest number of votes; n second (x) is the second prediction category with the highest number of votes. It can be understood that the frequency difference between 5 votes of category A and 3 votes of category B is 2 votes, which is small, indicating that the sample is relatively complex and needs to be further processed to ensure that the student model learns the key information.

[0043] According to the technical scheme provided in the embodiments of the present application, the voting frequency vector of each sample is obtained by counting the plurality of voting prediction results corresponding to each sample, wherein the voting frequency vector includes the number of votes of each sample in each prediction category. The difference between the number of votes of the two prediction categories with the highest number of votes in the voting frequency vector is calculated, and the difference is used as the voting frequency difference of each sample. The calculated voting frequency difference can preliminarily measure the importance of the sample in the teacher dimension, and at the same time quantify the prediction consistency of the teacher model, providing a basis for the teacher model for subsequent importance evaluation, so as to filter out important samples through the subsequent joint screening mechanism.

[0044] In some embodiments, the uncertainty score of each sample is predicted by using a student model to be trained, including: inputting the samples in the public training set into the student model to be trained in sequence for inference to obtain student prediction results corresponding to each sample; determining the uncertainty score of each sample according to the student prediction results corresponding to each sample.

[0045] Specifically, the samples in the public training set are input into the student model to be trained in sequence for inference, and the student model to be trained outputs student prediction results for the samples. For example, the student prediction results can be softmax output, at this time if a picture of a cat is input, the model may output cat: 0.6, dog: 0.3, bird: 0.1, which is the original output without probability conversion; the original output of the model usually does not directly have the meaning of probability, so it needs to be converted into a probability distribution through softmax output to convert the original output into a distribution with the sum of all category probabilities being 1.

[0046] Further, the uncertainty score of each sample is determined according to the student prediction result corresponding to each sample. Specifically, the entropy value corresponding to each sample is calculated according to the student prediction result (softmax output). The entropy value can be used to measure the prediction uncertainty of the student model for the sample. The higher the entropy value, the more dispersed the distribution (the stronger the uncertainty), indicating that the student model has low confidence in the category prediction of the sample, and the sample is a difficult sample for the student model. The lower the entropy value, the more concentrated the distribution (the stronger the certainty). In this example, the entropy value of each sample is determined as the uncertainty score of the sample, that is, the entropy value and the uncertainty score are equal in value.

[0047] According to the technical scheme provided in the embodiment, the samples in the public training set are input into the student model to be trained for inference to obtain the student prediction result corresponding to each sample; and the uncertainty score of each sample is determined according to the student prediction result corresponding to each sample, so that the samples difficult for the student model to learn (samples with high entropy values) can be screened out, and the basis in terms of the student model is provided for subsequent sample importance evaluation, so that important samples can be screened out through a subsequent joint screening mechanism.

[0048] In some embodiments, the importance of each sample is jointly evaluated according to the voting frequency difference and the uncertainty score to obtain a joint evaluation result of each sample, including: determining a frequency difference weight factor of each sample according to the voting frequency difference; and determining the joint evaluation result of each sample according to the frequency difference weight factor and the uncertainty score.

[0049] Specifically, the frequency difference weight factor can represent the importance degree of the sample. If the voting frequency difference of a sample is larger, the frequency difference weight factor at this time is lower, indicating that the importance degree of the sample is lower; and if the voting frequency difference of a sample is smaller, the frequency difference weight factor at this time is higher, indicating that the importance degree of the sample is higher.

[0050] As an example, when the voting frequency difference is large (the voting prediction results of the teacher models tend to be consistent), the frequency difference weight factor tends to 0, indicating that the sample is a simple sample and does not need to waste the privacy resource zone for labeling. When the voting frequency difference is small (the voting prediction results of the teacher models differ greatly), the frequency difference weight factor tends to 1, indicating that the sample is a difficult sample and is also a difficult problem for the teacher model, and the privacy resource zone needs to be spent for labeling.

[0051] Further, the frequency difference weight factor of each sample is determined according to the voting frequency difference, including: determining an attenuation control coefficient corresponding to each sample; and performing exponential attenuation processing on the voting frequency difference by using the attenuation control coefficient of each sample to obtain the frequency difference weight factor of each sample.

[0052] Specifically, the attenuation control coefficient is a linear coefficient in the exponential function, which can determine the degree of "stretching" or "compression" of the voting frequency difference on the entire exponential function.

[0053] As an example, for example, when the attenuation control coefficient is very small (tending to 0), no matter how large the voting frequency difference is, the exponential term is basically equal to 1, at this time almost no attenuation occurs, the screening effect of the voting frequency difference is closed, only the uncertainty score is in effect. When the attenuation control coefficient is very large (such as 1 or higher), the voting frequency difference is slightly increased, and the exponential term is sharply reduced, and the attenuation speed is very fast.

[0054] It can be understood that the attenuation control coefficient can change the step size of the voting frequency difference in the exponential function. If the suppression of the voting frequency difference is required to be strong, the attenuation control coefficient is adjusted to be large, at this time the system will be more strict in selecting high-divergence samples (the voting frequency difference is larger). If the suppression of the voting frequency difference is required to be weak, the attenuation control coefficient is adjusted to be small, at this time the influence of the uncertainty score is greater. The specific value of the attenuation control coefficient can be set according to actual needs, and the specific value of the attenuation control coefficient is not specifically limited in the embodiment.

[0055] Further, according to the frequency difference weight factor and the uncertainty score, the joint evaluation result of each sample is determined. The joint evaluation result of each sample is calculated by using the following formula:

[0056] Score(x)=Entropy(x)·exp(-α·Gap(x));

[0057] Wherein, Score(x) is the joint evaluation result of sample x, Entropy(x) is the uncertainty score of sample x, Gap(x) is the voting frequency difference of sample x, exp(-α·) is the exponential attenuation processing, and a is the attenuation control coefficient; exp(-α·Gap(x)) is the frequency difference weight factor of sample x.

[0058] At this time, only when the uncertainty score Entropy(x) predicted by the student model is higher and the frequency difference weight factor exp(-a·Gap(x)) predicted by the teacher model is larger, the joint evaluation result Score(x) will take a larger value, indicating that the importance of the sample is higher and it is worth being selected. It can be understood that the Score(x) joint evaluation result obtained at this time combines the voting frequency difference of the teacher model and the uncertainty score of the student model, and realizes fine evaluation of the sample by using uncertainty in two dimensions, which facilitates subsequent selection of appropriate samples for training the student model, effectively improves the use efficiency of the privacy budget, improves the training efficiency of the model, and at the same time, since the sample has higher uncertainty, it can more directly improve the learning effect of the student model on the boundary sample. Compared with the random sampling and single uncertainty screening method, the training value of the sample is improved in the embodiment, so that after the sample is used for training, the fitting ability and generalization ability of the student model can be enhanced.

[0059] In some embodiments, the target sample training set is subjected to privacy processing, and the target sample training set subjected to privacy processing is used to train the student model to be trained, comprising:

[0060] The voting frequency vector of each sample in the target sample training set is subjected to noise processing by using the differential privacy method; a pseudo label is constructed for each sample subjected to noise processing to obtain a pseudo label sample training set, wherein the pseudo label represents the predicted category of the maximum value in the voting frequency vector of each sample; and the pseudo label sample training set is used to train the student model to be trained.

[0061] Specifically, the voting frequency vector of each sample in the target sample training set is added with noise by using a differential privacy algorithm (such as Laplace mechanism or Gaussian mechanism). It can be avoided to deduce the specific voting behavior of an individual from the voting frequency vector. In the above example, taking sample x as an example, other samples can refer to sample x. For the voting frequency vector (n A (x)=5, n B (x)=3, n C (x)=2), it can be changed to (n A (x)=5.8, n B (x)=2.5, n C (x)=1.1) after adding noise. The noise can be added by using the following formula:

[0062]

[0063] v j (x) is the value without adding noise, is the value after adding noise, represents noise obeying normal distribution with mean 0 and variance σ2.

[0064] Further, a pseudo label is constructed for each sample after adding noise, where the pseudo label represents the predicted class with the highest number of votes in the vote frequency vector of each sample, for example, the vote frequency vector (n A (x) = 5.8, n B (x) = 2.5, n C (x) = 1.1), where A is the predicted class with the maximum value in the vote frequency vector, and the pseudo label of the sample is A. After constructing the pseudo label for each sample, a pseudo label sample training set is obtained. Finally, the pseudo label sample training set is used to train the student model to be trained, so that the model learns the mapping from the sample features to the pseudo label. In this way, the model learns the data pattern after privacy protection, avoiding direct contact with sensitive original data.

[0065] In some embodiments, before the plurality of teacher models are used to sequentially predict the samples in the public training set, the method further comprises: obtaining a sample training set corresponding to a target privacy field; splitting the sample training set corresponding to the target privacy field to obtain a plurality of privacy training sets, wherein the privacy samples in each privacy training set are disjoint; for each privacy training set, constructing and training a teacher model corresponding thereto to obtain the plurality of teacher models.

[0066] Specifically, the target privacy field is a field that requires privacy measures for sample data, for example, the sample training set corresponding to the target privacy field can be a sample training set in the medical field, the financial field, the government field, etc. The sample training set can also be sample data that needs to be kept secret in other fields. In the target privacy field, both high-value privacy data and direct contact with all privacy data by a single model are required to train the model (to reduce the risk of privacy leakage), so the sample training set needs to be split. Specifically, the sample training set can be divided into a plurality of privacy training sets by using splitting methods such as random allocation, grouping by data source, etc., as long as there is no overlapping sample (disjoint) between the privacy training sets. In this way, each privacy training set only contains part of the data, and a single teacher model can only access one privacy training set, avoiding any teacher model from mastering all the privacy information.

[0067] It should be noted that the structure of the teacher model can be a fixed deep neural network, or can be modularized according to the actual task (such as convolutional neural network, residual network, attention mechanism, etc.). The training process adopts a standard supervised learning paradigm, and the model parameters are updated based only on the statistical information of the sub-data set, without sharing intermediate variables or model gradients, ensuring that the training process complies with the local privacy principle.

[0068] Further, for each privacy training set, a corresponding teacher model is constructed and trained, specifically, in the medical field, the privacy training set can be a medical image privacy training set or an electronic health record training set, taking the medical image privacy training set as an example, the sample can be an X-ray film, a CT scan picture or an MRI image. Further, the embodiment selects a cnn convolutional neural network to construct the teacher model, and the teacher model specifically includes a feature extraction module, a classification module and an output module, wherein the feature extraction module is composed of a plurality of stackable convolution units, each convolution unit includes a convolution layer, a nonlinear function activation layer, a normalization layer and a down-sampling layer, in addition, in order to enhance the translation invariance and the feature compactness, a max-pooling layer can be arranged after the plurality of stackable convolution units. Further, the classification module receives the high-dimensional tensor output by the feature extraction module and converts it into a distinguishable class space distribution. The classification module includes a flattening operation and a plurality of fully connected layers, and maps the features through a nonlinear activation function.

[0069] Further, the training target of each independent teacher model can be to minimize the Kullback-Leibler divergence (KL Divergence) loss function:

[0070]

[0071] Where Nm represents the number of samples in the mthsub dataset, K is the total number of classification categories, q ik is the soft label distribution of sample i on category k (for example, derived from the prior distribution of the teacher ensemble or known soft label), p ik is the prediction probability of the teacher model that the sample belongs to the kthcategory:

[0072]

[0073] Where s ik is the logit value output by the model.

[0074] Alternatively, the training target of each independent teacher model can be a standard multi-class cross-entropy loss function:

[0075]

[0076] Where n i represents the number of samples in the ithsubset; C is the total number of categories of the classification task; 1(y j =k) is an indicator function that takes 1 when the true label of the sample x j is the kthcategory, and 0 otherwise; represents the prediction probability of the teacher model T i that the sample x j belongs to the kthcategory:

[0077]

[0078] wherein, denotes the model T i the logit output value for class k.

[0079] It should be noted that after the completion of the construction and training of the teacher model, each teacher model only serves as a black box predictor in subsequent integrated voting and no longer accepts any form of retraining or parameter optimization. This design avoids the re-exposure of the original training data in the joint aggregation process, thereby improving the overall system's resistance to privacy attacks.

[0080] In this way, since the privacy training set of each teacher model is independent, the knowledge learned by each teacher model is characteristic of the privacy training set, and the multiple teacher models eventually obtained each master part of the knowledge of the target privacy field, and the distributed protection of the privacy data is realized due to the non-overlapping data.

[0081] According to the technical scheme provided in the embodiments of the present application, a sample training set corresponding to a target privacy field is obtained; the sample training set corresponding to the target privacy field is split to obtain multiple privacy training sets, wherein the privacy samples in each privacy training set are mutually exclusive; for each privacy training set, a corresponding teacher model is constructed and trained to obtain multiple teacher models. Through data splitting, a single model cannot access all privacy data, and even if a certain teacher model has a leakage risk, only part of the data is leaked, greatly reducing the overall privacy leakage hazard. At the same time, multiple teacher models learn from different data subsets, and their prediction results can cover more scenarios, providing a more comprehensive knowledge source for subsequent training of student models.

[0082] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the processes of the embodiments of the present application.

[0083] All the optional technical solutions described above can be combined to form optional embodiments of the present application, which will not be described one by one here.

[0084] The following is an apparatus embodiment of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the apparatus embodiments of the present application, please refer to the method embodiments of the present application.

[0085] Figure 2 is a structural schematic diagram of the transfer learning based on the importance evaluation of the privacy sample provided by the embodiments of the present application. As Figure 2 shown, the apparatus comprises:

[0086] The first prediction module 201 is configured to sequentially predict samples in the public training set by using a plurality of teacher models to obtain a plurality of voting prediction results corresponding to each sample, wherein each teacher model is trained by a different private training set;

[0087] The calculation module 202 is configured to calculate a voting frequency difference of each sample according to the plurality of voting prediction results;

[0088] The second prediction module 203 is configured to predict an uncertainty score of each sample by using a student model to be trained;

[0089] The evaluation module 204 is configured to jointly evaluate the importance of each sample according to the voting frequency difference and the uncertainty score to obtain a joint evaluation result of each sample;

[0090] The training module 205 is configured to perform sample screening from the public sample training set to obtain a target sample training set according to the joint evaluation result, perform privacy processing on the target sample training set, train the student model to be trained by using the target sample training set after the privacy processing, and obtain a trained student model.

[0091] In some embodiments, the calculation module 202 is further configured to count the plurality of voting prediction results corresponding to each sample to obtain a voting frequency vector of each sample, wherein the voting frequency vector includes the number of votes of each sample in each prediction category. The difference between the number of votes of the two prediction categories with the highest number of votes in the voting frequency vector is calculated, and the difference is taken as the voting frequency difference of each sample.

[0092] In some embodiments, the second prediction module 203 is further configured to sequentially input the samples in the public training set into the student model to be trained for inference to obtain a student prediction result corresponding to each sample; and determine the uncertainty score of each sample according to the student prediction result corresponding to each sample.

[0093] In some embodiments, the evaluation module 204 is further configured to determine a frequency difference weight factor of each sample according to the voting frequency difference; and determine the joint evaluation result of each sample according to the frequency difference weight factor and the uncertainty score.

[0094] In some embodiments, the evaluation module 204 is further configured to determine a decay control coefficient corresponding to each sample; and perform exponential decay processing on the voting frequency difference by using the decay control coefficient of each sample to obtain a frequency difference weight factor of each sample.

[0095] In some embodiments, the evaluation module 204 is further configured to calculate the joint evaluation result of each sample by using the following formula: Score(x) = Entropy(x) exp(-a Gap(x)); wherein Score(x) is the joint evaluation result of the sample x; Entropy(x) is the uncertainty score of the sample x; Gap(x) is the voting frequency difference of the sample x; exp(-a ) is an exponential decay process; a is an attenuation control coefficient; and exp(-a Gap(x)) is a frequency difference weight factor of the sample x.

[0096] In some embodiments, the training module 205 is further configured to add noise to the voting frequency vector of each sample in the target sample training set by using differential privacy method; construct a pseudo label for each sample after the noise adding process to obtain a pseudo label sample training set, wherein the pseudo label represents the predicted class that is the maximum value in the voting frequency vector of each sample; and train the student model to be trained by using the pseudo label sample training set.

[0097] In some embodiments, the first prediction module 201 is further configured to obtain a sample training set corresponding to a target privacy field; split the sample training set corresponding to the target privacy field to obtain a plurality of privacy training sets, wherein the privacy samples in each privacy training set are disjointed; and construct and train a teacher model corresponding to each privacy training set to obtain a plurality of teacher models.

[0098] According to the technical scheme provided in the embodiments of the present application, a plurality of teacher models are used to sequentially predict samples in the public training set, and a plurality of voting prediction results corresponding to each sample are obtained, wherein each teacher model is trained by a different private training set; according to the plurality of voting prediction results, a voting frequency difference of each sample is calculated; a student model to be trained is used to predict an uncertainty score of each sample; the importance of each sample is jointly evaluated according to the voting frequency difference and the uncertainty score, and a joint evaluation result of each sample is obtained; according to the joint evaluation result, sample screening is performed on the public sample training set to obtain a target sample training set, the target sample training set is subjected to privacy processing, the student model to be trained is trained by using the target sample training set subjected to privacy processing, and a trained student model is obtained. On the one hand, by evaluating the samples by referring to two key dimensions of teacher models and student models, not only the accurate description of high-value samples is realized, but also simple samples that have reached a simple consensus are excluded, and the privacy budget is saved, so that the effective training period of the student model can be prolonged and the training efficiency of the student model can be improved under the same privacy budget. On the other hand, through the double screening of the two key dimensions, not only the samples with high training value can be screened out, but also the adaptability of the student model to complex boundary samples can be improved, so that the fitting ability and the generalization ability of the student model are enhanced. The problem of unreasonable selection of private training samples in the prior art, which leads to waste of privacy budget and low training efficiency, is avoided.

[0099] Figure 3 is a schematic diagram of an electronic device provided by an embodiment of the present application. As shown in Figure 3 is a structural schematic diagram of an electronic device based on the privacy sample importance evaluation-based transfer learning method provided by an embodiment of the present application. The electronic device can include a processor 30, a memory 31, a communication bus 32, and a communication interface 33, and can further include a computer program stored in the memory 31 and executable on the processor 30, such as a privacy sample importance evaluation-based transfer learning method program.

[0100] In some embodiments, the processor 30 can be composed of integrated circuits, for example, can be composed of a single packaged integrated circuit, or can be composed of a plurality of packaged integrated circuits with the same function or different functions, including one or more combinations of central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 30 is the control core (Control Unit) of the electronic device, which connects all components of the electronic device through various interfaces and lines, executes or runs programs or modules stored in the memory 31 (such as executing the privacy sample importance evaluation-based transfer learning method, etc.), and calls data stored in the memory 31 to perform various functions of the electronic device and process data.

[0101] The memory 31 includes at least one type of readable storage medium, including a flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., an SD or a DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. The memory 31 can be an internal storage unit of the electronic device in some embodiments, such as a mobile hard disk of the electronic device. The memory 31 can also be an external storage device of the electronic device in other embodiments, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 31 can include both the internal storage unit and the external storage device of the electronic device. The memory 31 can be used not only to store application software and various data installed in the electronic device, such as the code of the migration learning method program based on the privacy sample importance evaluation, but also to temporarily store data that has been output or will be output.

[0102] The communication bus 32 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The bus is provided to enable connection and communication between the memory 31 and the at least one processor 30, etc.

[0103] The communication interface 33 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface can include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the electronic device and other electronic devices. The user interface can be a display (Display), an input unit (such as a keyboard (Keyboard)), and optionally, the user interface can also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) touch, etc. Among them, the display can also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device and to display a visualized user interface.

[0104] Figure 3 Only the electronic device with components is shown, and those skilled in the art can understand that, Figure 3The illustrated structure does not constitute a limitation on the electronic device, which can include fewer or more components than those shown, or a combination of some components, or different arrangement of the components.

[0105] For example, although not shown, the electronic device can further include a power supply (such as a battery) to supply power to each component. Preferably, the power supply can be logically connected to the at least one processor 30 through a power management device, so that the power management device implements functions such as charge management, discharge management, and power consumption management. The power supply can also include one or more direct current or alternating current power supplies, a recharging device, a power supply failure detection circuit, a power supply converter or inverter, a power supply status indicator, and the like. The electronic device can also include various sensors, a Bluetooth module, a Wi-Fi module, and the like, which will not be described here.

[0106] It should be understood that the embodiments are for illustration only and are not limited in scope by the structure described.

[0107] Further, the modules / units integrated in the electronic device, if implemented in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. The computer readable storage medium can be volatile or non-volatile. For example, the computer readable medium can include any entity or device capable of carrying computer program codes, recording media, U disks, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROM).

[0108] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "example", "specific example", "one implementation", "one preferred implementation", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0109] Although embodiments of the present application have been shown and described, those skilled in the art can understand that various changes, modifications, substitutions and variations can be made to the embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the claims and their equivalents.

Claims

1. A privacy sample importance evaluation based transfer learning method, characterized in that, The method comprises the following steps: using a plurality of teacher models to sequentially predict samples in a public training set, to obtain a plurality of voting prediction results corresponding to each sample, wherein each teacher model is trained by a different private training set; calculating a voting frequency difference of each sample according to the plurality of voting prediction results; predicting an uncertainty score of each sample by using a student model to be trained; jointly evaluating the importance of each sample according to the voting frequency difference and the uncertainty score, to obtain a joint evaluation result of each sample; performing sample screening from the public sample training set according to the joint evaluation result to obtain a target sample training set, performing privacy processing on the target sample training set, and training the student model to be trained by using the privacy-processed target sample training set to obtain a trained student model.

2. The method of claim 1, wherein, The method further comprises the following steps: counting the plurality of voting prediction results corresponding to each sample to obtain a voting frequency vector of each sample, wherein the voting frequency vector comprises the number of votes of each sample in each prediction category; calculating the difference between the number of votes of the two prediction categories with the highest number of votes in the voting frequency vector, and taking the difference as the voting frequency difference of each sample.

3. The method of claim 1, wherein, The method further comprises the following steps: inputting the samples in the public training set into the student model to be trained for inference to obtain a student prediction result corresponding to each sample; determining the uncertainty score of each sample according to the student prediction result corresponding to each sample.

4. The method of claim 1, wherein, The method further comprises the following steps: determining a frequency difference weight factor of each sample according to the voting frequency difference; determining the joint evaluation result of each sample according to the frequency difference weight factor and the uncertainty score.

5. The method of claim 4, wherein, The method further comprises the following steps: determining an attenuation control coefficient corresponding to each sample; performing exponential attenuation processing on the voting frequency difference by using the attenuation control coefficient of each sample to obtain the frequency difference weight factor of each sample.

6. The method of claim 5, wherein, The method further comprises the following steps: calculating the joint evaluation result of each sample by using the following formula: Score(x)=Entropy(x)·exp(-α·Gap(x)); wherein Score(x) is the joint evaluation result of sample x; Entropy(x) is the uncertainty score of sample x; Gap(x) is the voting frequency difference of sample x; exp(-α·) is exponential attenuation processing; a is an attenuation control coefficient; and exp(-α·Gap(x)) is the frequency difference weight factor of sample x.

7. The method of claim 2, wherein, The privacy processing is performed on the target sample training set, the target sample training set after the privacy processing is used for training the student model to be trained, and the method comprises the following steps: The voting frequency vector of each sample in the target sample training set is processed by adding noise by using a differential privacy method; A pseudo label is constructed for each sample after the noise processing, and a pseudo label sample training set is obtained, wherein the pseudo label represents a predicted category with the maximum value in the voting frequency vector of each sample; The student model to be trained is trained by using the pseudo label sample training set.

8. The method of claim 1, wherein, Before the plurality of teacher models are used to sequentially predict the samples in the public training set, the method further comprises the following steps: A sample training set corresponding to a target privacy field is obtained; The sample training set corresponding to the target privacy field is split to obtain a plurality of privacy training sets, wherein the privacy samples in each privacy training set are mutually exclusive; For each privacy training set, a corresponding teacher model is constructed and trained to obtain the plurality of teacher models.

9. A privacy sample importance evaluation based transfer learning device, comprising: The method comprises the following steps: A first prediction module is configured to sequentially predict the samples in the public training set by using a plurality of teacher models to obtain a plurality of voting prediction results corresponding to each sample, wherein each teacher model is trained by a different privacy training set; A calculation module is configured to calculate a voting frequency difference of each sample according to the plurality of voting prediction results; A second prediction module is configured to predict an uncertainty score of each sample by using a student model to be trained; An evaluation module is configured to jointly evaluate the importance of each sample according to the voting frequency difference and the uncertainty score to obtain a joint evaluation result of each sample; A training module is configured to perform sample screening from the public sample training set to obtain a target sample training set according to the joint evaluation result, perform privacy processing on the target sample training set, and train the student model to be trained by using the target sample training set after the privacy processing to obtain a trained student model.

10. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the method for privacy sample importance evaluation based migration learning according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • CoVID-19 chest X-ray image classification learning-oriented training data privacy protection method

    CN115482435A

  • Semi-supervised 3D target detection method, system and equipment based on digital twinning

    CN117649515A

  • Differential privacy federal learning method and system based on double knowledge transfer

    CN119622824A

  • Image classification method and system based on F1-score weighted voting

    CN119992217A

  • Seabed sediment classification method based on multilevel comparative learning and uncertainty measurement

    CN120105205A