Transfer learning method and device based on privacy sample importance evaluation

By jointly evaluating the voting frequency difference of the multi-teacher model and the uncertainty score of the student model, high-value samples are selected, which solves the problem of unreasonable selection of privacy samples and improves the model training efficiency and fitting ability.

CN120910558BActive Publication Date: 2026-01-27CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511003806.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2026-01-27
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

In existing technologies, unreasonable selection of privacy samples leads to wasted privacy budgets and low model training efficiency.

Method used

Multiple teacher models are used to predict the samples by voting. The voting frequency difference and the uncertainty score of the student model are calculated. The two are combined for joint evaluation to screen out high-value samples and perform privacy processing. The target sample training set is then constructed for training the student model.

Benefits of technology

It achieves accurate characterization of high-value samples, excludes simple consensus samples, saves privacy budget, extends the effective training cycle of student models, and improves training efficiency and fitting ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910558B_ABST
    Figure CN120910558B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of artificial intelligence, and provides a transfer learning method based on privacy sample importance evaluation, which comprises the following steps: using multiple teacher models to predict samples in a public training set in turn to obtain multiple voting prediction results of each sample; calculating the voting frequency difference of each sample according to the prediction results; using a student model to be trained to predict the uncertainty score of each sample; selecting a target sample training set composed of high-value samples from the public sample training set according to the voting frequency difference and the uncertainty score, and using the target sample training set to train the student model to be trained to obtain a trained student model. The application can filter out samples with high training value through double screening in two key dimensions, improve the adaptability of the student model to complex boundary samples, save privacy budget, enhance the fitting ability and generalization ability of the student model, and improve the training efficiency of the student model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a transfer learning method based on the assessment of the importance of privacy samples. Background Technology

[0002] In the rapid development of artificial intelligence, the performance of deep learning and machine learning models largely depends on high-quality, large-scale training data. However, much of this valuable data resides in privacy-sensitive areas, requiring confidentiality. Existing technologies often employ differential privacy protection techniques to handle this sensitive data and protect data privacy during the training process. However, current differential privacy protection methods often use random sampling or single-metric evaluation to select samples for noise addition. This fails to distinguish between invalid and valid samples, resulting in the allocation of privacy budgets to invalid samples during subsequent noise addition using differential privacy methods. This leads to wasted privacy budgets and reduced training efficiency.

[0003] It is evident that existing technologies suffer from problems such as unreasonable selection of privacy training samples, resulting in wasted privacy budgets and low model training efficiency. Summary of the Invention

[0004] In view of this, this application provides a transfer learning method and apparatus based on the assessment of the importance of privacy samples, in order to solve the problems in the prior art where unreasonable selection of privacy samples leads to wasted privacy budget and low model training efficiency.

[0005] A first aspect of this application provides a transfer learning method based on privacy sample importance assessment, the method comprising:

[0006] Multiple teacher models are used to predict samples in a common training set sequentially, resulting in multiple voting predictions for each sample. Each teacher model is trained on a different privacy training set. Based on the multiple voting predictions, the voting frequency difference for each sample is calculated. The uncertainty score for each sample is predicted using the student model to be trained. The importance of each sample is jointly evaluated based on the voting frequency difference and uncertainty score, resulting in a joint evaluation result for each sample. Based on the joint evaluation result, a target sample training set is obtained by filtering samples from the common sample training set. The target sample training set is then privacy-processed, and the student model to be trained is trained using the privacy-processed target sample training set to obtain the trained student model.

[0007] A second aspect of this application provides a transfer learning apparatus based on privacy sample importance assessment, the apparatus comprising:

[0008] The first prediction module is configured to use multiple teacher models to sequentially predict samples in a common training set, obtaining multiple voting prediction results for each sample, wherein each teacher model is trained on a different privacy training set; the calculation module is configured to calculate the voting frequency difference for each sample based on the multiple voting prediction results; the second prediction module is configured to use the student model to be trained to predict the uncertainty score for each sample; the evaluation module is configured to jointly evaluate the importance of each sample based on the voting frequency difference and the uncertainty score, obtaining a joint evaluation result for each sample; the training module is configured to select a target sample training set from the common sample training set based on the joint evaluation result, perform privacy processing on the target sample training set, and use the privacy-processed target sample training set to train the student model to be trained, obtaining a trained student model.

[0009] A third aspect of this application provides an electronic device, the electronic device comprising:

[0010] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the steps of the above method.

[0011] A fourth aspect of this application provides a computer storage medium storing a computer program that, when executed by a processor, can perform the steps of the above-described method.

[0012] The beneficial effects of the embodiments in this application compared with the prior art are:

[0013] According to the technical solution provided in this application, multiple teacher models are used to predict samples in a common training set sequentially, obtaining multiple voting prediction results for each sample. Each teacher model is trained on a different privacy training set. Based on the multiple voting prediction results, the voting frequency difference for each sample is calculated. The uncertainty score for each sample is predicted using the student model to be trained. The importance of each sample is jointly evaluated based on the voting frequency difference and uncertainty score, resulting in a joint evaluation result for each sample. Based on the joint evaluation result, a target sample training set is obtained by filtering samples from the common sample training set. The target sample training set is then subjected to privacy processing. The student model to be trained is then trained using the privacy-processed target sample training set to obtain a trained student model. On the one hand, by using two key dimensions—the teacher model and the student model—to evaluate samples, not only is a precise characterization of high-value samples achieved, but simple samples with already reached simple consensus are also excluded, saving privacy budget. This ensures that the effective training period of the student model can be extended under the same privacy budget, thereby improving the training efficiency of the student model. On the other hand, through dual screening across two key dimensions, not only can high-value training samples be selected, but the adaptability of the student model to complex boundary samples can also be improved, thereby enhancing the model's fitting and generalization abilities. This avoids the problems of unreasonable selection of privacy training samples, leading to wasted privacy budgets and low training efficiency, which exist in existing technologies. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating a transfer learning method based on privacy sample importance assessment provided in an embodiment of this application.

[0015] Figure 2 This is a schematic diagram of the structure of a transfer learning device based on privacy sample importance assessment provided in an embodiment of this application;

[0016] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Detailed Implementation

[0017] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0018] In the description of this invention, it should be understood that the terms "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0019] In the description of this invention, unless otherwise specified and limited, it should be noted that the terms "installation", "connection" and "linking" should be interpreted broadly. For example, they can refer to mechanical or electrical connections, or internal connections between two components. They can be direct connections or indirect connections through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms according to the specific circumstances.

[0020] The execution entity of the transfer learning method based on privacy sample importance assessment in this application includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in the embodiments of this application: a server, a terminal, etc. In other words, this method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0021] The following describes in detail, with reference to the accompanying drawings, a transfer learning method and apparatus based on privacy sample importance assessment according to embodiments of this application.

[0022] Figure 1 This is a flowchart illustrating a transfer learning method based on privacy sample importance assessment provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes:

[0023] S101, using multiple teacher models to predict samples in a common training set in sequence, to obtain multiple voting prediction results for each sample, where each teacher model is trained on a different privacy training set;

[0024] S102, Calculate the voting frequency difference for each sample based on multiple voting prediction results;

[0025] S103, using the student model to be trained to predict the uncertainty score for each sample;

[0026] S104, The importance of each sample is jointly evaluated based on the voting frequency difference and uncertainty score to obtain the joint evaluation result of each sample;

[0027] S105. Based on the joint evaluation results, the target sample training set is obtained by screening samples from the public sample training set, and privacy processing is performed on the target sample training set. The student model to be trained is trained using the privacy-processed target sample training set to obtain the trained student model.

[0028] Specifically, the multiple teacher models are pre-trained teacher models, and each teacher model is trained from a different privacy training set. That is, each privacy training set is independent of each other, and the training samples in the privacy training sets do not overlap, thus avoiding a single model from accessing all privacy samples and reducing the risk of privacy leakage.

[0029] It is understandable that multiple teacher models are used to predict samples in a common training set in turn, resulting in multiple voting prediction results for each sample. That is, each sample will be processed separately by multiple teacher models to obtain their own prediction results. For example, if there are 10 teacher models and 100 samples, each sample will receive 10 independent prediction results (such as category judgment in a classification task). These results together constitute multiple voting prediction results for that sample.

[0030] It's important to note that in existing technologies, sample selection typically relies on random sampling or a single uncertainty measure. Random sampling treats the voting predictions of teacher models indiscriminately. During the voting phase, only the highest-voted label is considered, and noise is uniformly added to all samples without considering the differences in voting among teacher models. For example, on some samples, teacher model voting is highly dispersed, indicating that the sample is difficult to judge; while on other samples, teacher model voting is highly consistent, and these should be prioritized for student training. However, current mechanisms cannot differentiate between these, resulting in an overly amortized privacy noise budget and inefficiency. Secondly, a single uncertainty measure cannot fully reflect the value of a sample to the student model's learning effect. For example, while softmax entropy can reflect the subjective uncertainty of a student model's prediction for a particular sample, it does not reflect the consistency of voting among teacher models. Some high-entropy samples may have already formed highly consistent voting results among teachers. Continuing to perform differential privacy aggregation on these samples not only wastes valuable privacy budgets but may also lead to a severe mismatch between privacy costs and student benefits.

[0031] To avoid the aforementioned problems, this embodiment calculates the voting frequency deviation for each sample based on multiple voting prediction results; simultaneously, it uses the student model to be trained to predict the uncertainty score for each sample; and combines the voting frequency deviation and uncertainty score to jointly evaluate the importance of each sample, obtaining a joint evaluation result for each sample. This allows for a joint evaluation from both the teacher model's and student model's prediction dimensions, achieving a refined assessment of sample importance. It is understood that the calculation of the voting frequency deviation for each sample based on multiple voting prediction results, and the joint evaluation of the importance of each sample by combining the voting frequency deviation and uncertainty score, will be explained in detail later; therefore, it will not be elaborated upon here.

[0032] Furthermore, after obtaining the joint evaluation results for each sample, samples with higher importance can be selected from the common sample training set based on the joint evaluation results to form the target sample training set. At this time, privacy processing of the target sample training set can effectively improve the efficiency of privacy budget utilization. Finally, the target sample training set after differential processing is used to train the student model to be trained, and the trained student model is obtained.

[0033] According to the technical solution provided in this application, multiple teacher models are used to predict samples in a common training set sequentially, obtaining multiple voting prediction results for each sample, wherein each teacher model is trained from a different privacy training set; the voting frequency difference of each sample is calculated based on the multiple voting prediction results; the uncertainty score of each sample is predicted using the student model to be trained; the importance of each sample is jointly evaluated based on the voting frequency difference and uncertainty score, obtaining a joint evaluation result for each sample; based on the joint evaluation result, a target sample training set is obtained by filtering samples from the common sample training set, and the target sample training set is subjected to privacy processing; the student model to be trained is trained using the privacy-processed target sample training set to obtain a trained student model. On the one hand, by using two key dimensions, teacher model and student model, to evaluate samples, not only is a precise characterization of high-value samples achieved, but also simple samples that have reached simple consensus are excluded, saving privacy budget, thereby ensuring that the effective training cycle of the student model can be extended under the same privacy budget, and improving the training efficiency of the student model. On the other hand, through dual screening across two key dimensions, not only can high-value training samples be selected, but the adaptability of the student model to complex boundary samples can also be improved, thereby enhancing the model's fitting and generalization abilities. This avoids the problems of unreasonable selection of privacy training samples, leading to wasted privacy budgets and low training efficiency, which exist in existing technologies.

[0034] In some embodiments, calculating the voting frequency difference for each sample based on multiple voting prediction results includes:

[0035] The multiple voting prediction results corresponding to each sample are statistically analyzed to obtain the voting frequency vector of each sample, wherein the voting frequency vector includes the number of votes for each sample in each prediction category; the difference between the number of votes in the two prediction categories with the highest number of votes in the voting frequency vector is calculated, and the difference is used as the voting frequency difference for each sample.

[0036] Specifically, the voting frequency vector represents the number of votes for each sample in each predicted category. After obtaining multiple voting prediction results for each sample, the number of votes for each sample in each predicted category can be obtained based on the multiple voting prediction results, thereby obtaining the distribution of the number of predicted categories for each sample, and then obtaining the voting frequency vector for each sample.

[0037] As an example, in a classification task, the specific vote frequency vector can be represented as: v(x)=[n1(x),n2(x),...,n C [(x)]; where v(x) is the voting frequency vector of sample x, c is the total number of categories in the classification task, and n1(x) is the number of votes for a certain category. For example, if in the 10 predicted voting results for sample x, category A has 5 votes, category B has 3 votes, and category C has 2 votes, then n A (x) = 5, n B (x) = 3, n C (x) = 2.

[0038] Further, the difference in the number of votes for the two predicted categories with the highest number of votes in the voting frequency vector is calculated, and the difference is used as the voting frequency difference for each sample.

[0039] It's important to note that voting frequency is used to quantify the predictive consistency of the teacher model. By statistically analyzing the differences in the frequency of different voting predictions (such as classification labels) in the voting, it quantifies the gap in the number of votes for different options, thus measuring the consistency of the teacher model's predictions for the sample. A large voting frequency indicates that the label of the sample can be reliably determined, eliminating the need for additional calls to the teacher model for re-judgment, and allowing direct use to train the student model, reducing repeated queries of privacy data. A small voting frequency indicates that the sample is more complex, and the teacher model's predictions for the label are inconsistent, failing to accurately determine which outcome the prediction leans towards. Learning the accurate features of such samples is more difficult than other samples, and their importance is higher because effectively utilizing these samples can accelerate the model's convergence speed and accuracy. Therefore, these samples require further processing to ensure that the student model subsequently learns key information. This process can initially exclude some simple samples that reach voting consensus, thus initially saving on privacy budget.

[0040] Continuing the previous example, the two predicted categories with the highest number of votes in the vote frequency vector are now Class A (5 votes) and Class B (3 votes). Class A and Class B will serve as the basis for calculating the vote frequency difference for this sample. The vote frequency difference for each sample can be calculated using the following formula:

[0041] Gap(x) = n max (x)-n second (x);

[0042] Where Gap(x) is the voting frequency difference of sample x; n max (x) represents the prediction category with the most votes; n second (x) represents the prediction category with the second-highest number of votes. Understandably, the frequency difference between 5 votes for category A and 3 votes for category B is 2 votes. This small frequency difference indicates that the sample is relatively complex and requires further processing to ensure the student model learns the key information.

[0043] According to the technical solution provided in the embodiments of this application, the multiple voting prediction results corresponding to each sample are statistically analyzed to obtain a voting frequency vector for each sample. The voting frequency vector includes the number of votes for each sample in each prediction category. The difference in the number of votes for the two prediction categories with the highest number of votes in the voting frequency vector is calculated, and this difference is used as the voting frequency difference for each sample. The calculated voting frequency difference can initially measure the importance of a sample from a teacher's perspective, while also quantifying the predictive consistency of the teacher model. This provides a foundation for subsequent importance assessment from a teacher model perspective, enabling the selection of important samples through a subsequent joint screening mechanism.

[0044] In some embodiments, predicting the uncertainty score of each sample using a student model to be trained includes: sequentially inputting samples from a common training set into the student model to be trained for inference to obtain a student prediction result corresponding to each sample; and determining the uncertainty score of each sample based on the student prediction result corresponding to each sample.

[0045] Specifically, samples from the common training set are sequentially input into the student model to be trained for inference. The student model will then output the student's prediction results for the samples. For example, the student's prediction result can be a softmax output. If a picture of a cat is input, the model might output: cat: 0.6, dog: 0.3, bird: 0.1. This is the raw output without probability transformation. The raw output of the model usually does not directly have a probability meaning, so it needs to be transformed into a probability distribution through softmax, converting the raw output into a distribution where the sum of the probabilities of all categories is 1.

[0046] Furthermore, based on the student prediction results for each sample, the uncertainty score for each sample is determined. Specifically, based on the student prediction results (softmax output), the entropy value for each sample is calculated. The entropy value can be used to measure the uncertainty of the student model's prediction of the sample. The higher the entropy value, the more dispersed the distribution (stronger the uncertainty), indicating that the student model has low confidence in predicting the category of this sample, and that the sample is a hard example for the student model. The lower the entropy value, the more concentrated the distribution (stronger the certainty). In this example, the entropy value of each sample is determined as the uncertainty score of that sample, that is, the entropy value and the uncertainty score are numerically equal.

[0047] According to the technical solution provided in this embodiment, the samples in the common training set are input into the student model to be trained for inference to obtain the student prediction result corresponding to each sample; based on the student prediction result corresponding to each sample, the uncertainty score of each sample is determined, which can screen out samples that are difficult for the student model to learn (high entropy samples), providing a foundation for the student model in terms of subsequent sample importance assessment, so as to screen out important samples through subsequent joint screening mechanisms.

[0048] In some embodiments, the importance of each sample is jointly evaluated based on the voting frequency difference and the uncertainty score to obtain a joint evaluation result for each sample, including: determining a frequency difference weighting factor for each sample based on the voting frequency difference; and determining the joint evaluation result for each sample based on the frequency difference weighting factor and the uncertainty score.

[0049] Specifically, the frequency weighting factor can characterize the importance of a sample. If the voting frequency of a sample is larger, then the frequency weighting factor is lower, indicating that the sample is less important. Conversely, if the voting frequency of a sample is smaller, then the frequency weighting factor is higher, indicating that the sample is more important.

[0050] As an example, when the voting frequency difference is large (the teacher model's voting predictions tend to be consistent), the frequency difference weighting factor will tend to 0, indicating that the sample is a simple sample and does not need to be labeled with privacy resource areas. When the voting frequency difference is small (the teacher model's voting predictions vary greatly), the frequency difference weighting factor will tend to 1, indicating that the sample is a difficult sample, which is also a difficult problem for the teacher model and requires labeling with privacy resource areas.

[0051] Furthermore, based on the voting frequency difference, the frequency difference weighting factor for each sample is determined, including: determining the attenuation control coefficient corresponding to each sample; and using the attenuation control coefficient of each sample to perform exponential attenuation processing on the voting frequency difference to obtain the frequency difference weighting factor for each sample.

[0052] Specifically, the attenuation control coefficient is the linear coefficient in the exponential function, which determines the degree to which the voting frequency difference "stretches" or "compresses" the entire exponential function.

[0053] As an example, when the decay control coefficient is very small (approaching 0), the exponent term is essentially equal to 1 regardless of the voting frequency difference. In this case, decay hardly occurs, the filtering effect of the voting frequency difference is turned off, and only the uncertainty score is at play. When the decay control coefficient is very large (e.g., 1 or higher), a slight increase in the voting frequency difference will cause the exponent term to decrease drastically, and the decay rate is very fast.

[0054] It is understandable that the attenuation control coefficient can change the step size of the voting frequency difference in the exponential function. If a strong suppression of the voting frequency difference is needed, the attenuation control coefficient is increased, and the system will more strictly select high-dissent samples (those with larger voting frequencies). If a weaker suppression of the voting frequency difference is desired, the attenuation control coefficient is decreased, and the impact of uncertainty scoring is greater. The specific value of the attenuation control coefficient can be set according to actual needs; this embodiment does not specify a particular value for the attenuation control coefficient.

[0055] Further, based on the frequency difference weighting factor and the uncertainty score, the joint evaluation result for each sample is determined. Specifically, the joint evaluation result for each sample is calculated using the following formula:

[0056] Score(x)=Entropy(x)·exp(-α·Gap(x));

[0057] Wherein, Score(x) is the joint evaluation result of sample x, Entropy(x) is the uncertainty score of sample x, Gap(x) is the voting frequency difference of sample x, exp(-α·) is the exponential decay treatment, α is the decay control coefficient, and exp(-α·Gap(x)) is the frequency difference weighting factor of sample x.

[0058] At this point, after multiplying the uncertain score Entropy(x) by the frequency difference weighting factor exp(-α·Gap(x)), the joint evaluation result Score(x) will only take a large value if the uncertain score Entropy(x) predicted by the student model is high and the frequency difference weighting factor exp(-α·Gap(x)) predicted by the teacher model is large. This indicates that the sample is of high importance and worthy of being selected. It is understandable that the joint evaluation result Score(x) obtained at this point integrates the voting frequency difference of the teacher model and the uncertain score of the student model. It utilizes the uncertainty of two dimensions to achieve a refined evaluation of the sample, facilitating the subsequent selection of suitable samples for training the student model based on the evaluation results. This effectively improves the efficiency of privacy budget utilization and model training efficiency. Furthermore, due to the high uncertainty of the sample, it can more directly improve the learning performance of the student model on boundary samples. Compared to random sampling and single uncertainty screening methods, this embodiment improves the training value of the sample, thus enhancing the fitting and generalization abilities of the student model after training with the sample.

[0059] In some embodiments, privacy processing is performed on the target sample training set, and the privacy-processed target sample training set is used to train the student model to be trained, including:

[0060] The differential privacy method is used to add noise to the voting frequency vector of each sample in the target sample training set; a pseudo-label is constructed for each sample after noise addition, and a pseudo-label sample training set is obtained, where the pseudo-label represents the predicted category with the maximum value in the voting frequency vector of each sample; the pseudo-label sample training set is used to train the student model to be trained.

[0061] Specifically, differential privacy algorithms (such as Laplace's algorithm and Gaussian algorithm) are used to add noise to the voting frequency vector of each sample in the target sample training set. This avoids inferring an individual's specific voting behavior from the voting frequency vector. Continuing the previous example, taking sample x as an example, other samples can be referenced from sample x. For the voting frequency vector (n...) of sample x... A (x) = 5, n B (x) = 3, n C (x)=2), after adding noise, it may become (n) A (x) = 5.8, n B (x) = 2.5, n C (x) = 1.1). Specifically, noise can be added using the following formula:

[0062]

[0063] v j (x) represents the value without added noise. The value after adding noise. This represents noise that follows a normal distribution with a mean of 0 and a variance of σ².

[0064] Furthermore, a pseudo-label is constructed for each sample after noise addition, where the pseudo-label represents the predicted category with the highest number of votes in the voting frequency vector of each sample. For example, the voting frequency vector (n) after noise addition... A (x) = 5.8, n B (x) = 2.5, n C (x) = 1.1), where class A is the predicted class with the maximum value in the voting frequency vector. The pseudo-label for this sample is then class A. After constructing pseudo-labels for each sample, a pseudo-labeled sample training set is obtained. Finally, the pseudo-labeled sample training set is used to train the student model, enabling the model to learn the mapping from sample features to pseudo-labels. In this way, the model learns a privacy-preserving data pattern, avoiding direct access to sensitive raw data.

[0065] In some embodiments, before using multiple teacher models to predict samples in a common training set sequentially, the method further includes: obtaining a sample training set corresponding to the target privacy domain; splitting the sample training set corresponding to the target privacy domain to obtain multiple privacy training sets, wherein the privacy samples in each privacy training set are mutually exclusive; and constructing and training a corresponding teacher model for each privacy training set to obtain multiple teacher models.

[0066] Specifically, the target privacy domain refers to the domain in which sample data needs to be kept confidential. For example, the sample training set corresponding to the target privacy domain can be a sample training set from fields such as medicine, finance, and government affairs. At the same time, the sample training set can also be sample data from other domains that need to be kept confidential. In the target privacy domain, it is necessary to use high-value privacy data to train the model, while avoiding a single model directly accessing all privacy data (reducing the risk of privacy leakage). Therefore, the sample training set needs to be split. Specifically, splitting methods such as random allocation or grouping by data source can be used to divide the sample training set into multiple privacy training sets. It is only necessary to ensure that there are no overlapping samples between the privacy training sets (they are mutually exclusive). In this way, each privacy training set contains only a portion of the data, and a single teacher model can only access one privacy training set, preventing any teacher model from possessing all privacy information.

[0067] It should be noted that the teacher model can have a fixed deep neural network structure, or it can be modularly designed according to the actual task (such as convolutional neural networks, residual networks, attention mechanisms, etc.). The training process adopts a standard supervised learning paradigm, and the model parameters are updated only based on the statistical information of a subset of the dataset. Intermediate variables or model gradients are not shared, ensuring that the training process complies with the principle of local privacy.

[0068] Furthermore, for each privacy training set, a corresponding teacher model is constructed and trained. Specifically, in the medical field, the privacy training set can be a medical image privacy training set or an electronic health record training set. Taking the medical image privacy training set as an example, the samples can be X-ray films, CT scan images, or MRI images. Further, this embodiment selects a CNN (Convolutional Neural Network) to construct the teacher model. This teacher model specifically includes a feature extraction module, a classification module, and an output module. The feature extraction module consists of several stackable convolutional units. Each convolutional unit includes a convolutional layer, a nonlinear activation layer, a normalization layer, and a downsampling layer. In addition, to enhance translation invariance and feature compactness, a max-pooling layer can be set after the stackable convolutional units. Further, the classification module receives the high-dimensional tensor output by the feature extraction module and converts it into a distinguishable class space distribution. This classification module includes a flattening operation and multiple fully connected layers, and maps the features through a nonlinear activation function.

[0069] Furthermore, the training objective for each individual teacher model can be to minimize the Kullback-Leibler divergence (KL Divergence) loss function:

[0070]

[0071] Where Nm represents the number of samples in the m-th subset, K is the total number of classification categories, and q ik Let p be the soft label distribution of sample i in class k (e.g., a prior distribution derived from teacher ensembles or known soft labels). ik The teacher model predicts the probability that a sample belongs to the k-th class:

[0072]

[0073] Among them, s ik This is the logit value output by the model.

[0074] Alternatively, each individual teacher model can be trained using the standard multi-class cross-entropy loss function:

[0075]

[0076] Where n i Let represent the number of samples in the i-th subset; C is the total number of categories in the classification task; 1(y j =k) ​​is an indicator function, when sample x j The value is 1 if the actual label is of class k, and 0 otherwise. Teacher model T i For sample x j Predicted probability of belonging to class k:

[0077]

[0078] in, Model T i The logit output value for category k.

[0079] It should be noted that after the teacher models are built and trained, each teacher model only participates in subsequent ensemble voting as a black-box predictor and is no longer subject to any form of retraining or parameter tuning. This design avoids the re-exposure of the original training data during the joint aggregation process, thereby improving the overall system's resistance to privacy attacks.

[0080] In this way, since each teacher model has an independent privacy training set, the knowledge it learns will carry the characteristics of that privacy training set. Ultimately, multiple teacher models each master a portion of the knowledge in the target privacy domain, and because the data does not overlap, distributed protection of privacy data is achieved.

[0081] According to the technical solution provided in this application, a sample training set corresponding to the target privacy domain is obtained; the sample training set corresponding to the target privacy domain is split to obtain multiple privacy training sets, wherein the privacy samples in each privacy training set are mutually exclusive; for each privacy training set, a corresponding teacher model is constructed and trained to obtain multiple teacher models. Through data splitting, a single model cannot access all privacy data, and even if a certain teacher model has a risk of leakage, only a portion of the data is leaked, significantly reducing the overall harm of privacy leakage. Simultaneously, multiple teacher models learn from different data subsets, and their prediction results can cover more scenarios, providing a more comprehensive knowledge source for subsequent training of student models.

[0082] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the process of the embodiments of this application.

[0083] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0084] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0085] Figure 2 This is a schematic diagram of a transfer learning structure based on privacy sample importance assessment provided in an embodiment of this application. Figure 2 As shown, the device includes:

[0086] The first prediction module 201 is configured to use multiple teacher models to predict samples in a common training set in turn, and obtain multiple voting prediction results for each sample, wherein each teacher model is trained on a different privacy training set.

[0087] The calculation module 202 is configured to calculate the voting frequency difference for each sample based on multiple voting prediction results;

[0088] The second prediction module 203 is configured to predict the uncertainty score for each sample using the student model to be trained.

[0089] Evaluation module 204 is configured to jointly evaluate the importance of each sample based on voting frequency difference and uncertainty score, and obtain a joint evaluation result for each sample.

[0090] Training module 205 is configured to select a target sample training set from the public sample training set based on the joint evaluation results, perform privacy processing on the target sample training set, and use the privacy-processed target sample training set to train the student model to be trained, thereby obtaining a trained student model.

[0091] In some embodiments, the calculation module 202 is further configured to statistically analyze the plurality of voting prediction results corresponding to each sample to obtain a voting frequency vector for each sample, wherein the voting frequency vector includes the number of votes for each sample in each prediction category. The difference between the number of votes in the two prediction categories with the highest number of votes in the voting frequency vector is calculated, and this difference is used as the voting frequency difference for each sample.

[0092] In some embodiments, the second prediction module 203 is further configured to sequentially input samples from the common training set into the student model to be trained for inference, to obtain the student prediction result corresponding to each sample; and to determine the uncertainty score of each sample based on the student prediction result corresponding to each sample.

[0093] In some embodiments, the evaluation module 204 is further configured to determine a frequency difference weighting factor for each of the samples based on the voting frequency difference; and to determine the joint evaluation result for each of the samples based on the frequency difference weighting factor and the uncertainty score.

[0094] In some embodiments, the evaluation module 204 is further configured to determine the attenuation control coefficient corresponding to each sample; and to perform exponential attenuation processing on the voting frequency difference using the attenuation control coefficient of each sample to obtain the frequency difference weighting factor of each sample.

[0095] In some embodiments, the evaluation module 204 is further configured to calculate the joint evaluation result for each sample using the following formula: Score(x) = Entropy(x)·exp(-α·Gap(x)); where Score(x) is the joint evaluation result for sample x; Entropy(x) is the uncertainty score for sample x; Gap(x) is the voting frequency difference for sample x; exp(-α·) is the exponential decay treatment; α is the decay control coefficient; and exp(-α·Gap(x)) is the frequency difference weighting factor for sample x.

[0096] In some embodiments, the training module 205 is further configured to add noise to the voting frequency vector of each sample in the target sample training set using differential privacy; construct pseudo-labels for each sample after noise addition to obtain a pseudo-labeled sample training set, wherein the pseudo-labels represent the predicted category with the maximum value in the voting frequency vector of each sample; and train the student model to be trained using the pseudo-labeled sample training set.

[0097] In some embodiments, the first prediction module 201 is further configured to acquire a sample training set corresponding to the target privacy domain; split the sample training set corresponding to the target privacy domain to obtain multiple privacy training sets, wherein the privacy samples in each privacy training set are mutually exclusive; and construct and train a corresponding teacher model for each privacy training set to obtain multiple teacher models.

[0098] According to the technical solution provided in this application, multiple teacher models are used to predict samples in a common training set sequentially, obtaining multiple voting prediction results for each sample. Each teacher model is trained on a different privacy training set. Based on the multiple voting prediction results, the voting frequency difference for each sample is calculated. The uncertainty score for each sample is predicted using the student model to be trained. The importance of each sample is jointly evaluated based on the voting frequency difference and uncertainty score, resulting in a joint evaluation result for each sample. Based on the joint evaluation result, a target sample training set is obtained by filtering samples from the common sample training set. The target sample training set is then subjected to privacy processing. The student model to be trained is then trained using the privacy-processed target sample training set to obtain a trained student model. On the one hand, by using two key dimensions—the teacher model and the student model—to evaluate samples, not only is a precise characterization of high-value samples achieved, but simple samples with already reached simple consensus are also excluded, saving privacy budget. This ensures that the effective training period of the student model can be extended under the same privacy budget, thereby improving the training efficiency of the student model. On the other hand, through dual screening across two key dimensions, not only can high-value training samples be selected, but the adaptability of the student model to complex boundary samples can also be improved, thereby enhancing the model's fitting and generalization abilities. This avoids the problems of unreasonable selection of privacy training samples, leading to wasted privacy budgets and low training efficiency, which exist in existing technologies.

[0099] Figure 3 This is a schematic diagram of an electronic device provided in an embodiment of this application. Figure 3 The diagram shown is a schematic representation of an electronic device based on a transfer learning method for privacy sample importance assessment, according to an embodiment of the present invention. The electronic device may include a processor 30, a memory 31, a communication bus 32, and a communication interface 33. It may also include a computer program stored in the memory 31 and executable on the processor 30, such as a program for a transfer learning method based on privacy sample importance assessment.

[0100] In some embodiments, the processor 30 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 30 is the control unit of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. It executes programs or modules stored in the memory 31 (e.g., executing transfer learning methods based on privacy sample importance assessment) and calls data stored in the memory 31 to perform various functions of the electronic device and process data.

[0101] The memory 31 includes at least one type of readable storage medium, including flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 31 can be an internal storage unit of an electronic device, such as a portable hard drive. In other embodiments, the memory 31 can be an external storage device of the electronic device, such as a plug-in portable hard drive, SmartMediaCard (SMC), SecureDigital (SD) card, FlashCard, etc. Furthermore, the memory 31 can include both internal and external storage units of the electronic device. The memory 31 can be used not only to store application software and various types of data installed on the electronic device, such as the code of a transfer learning method program based on privacy sample importance assessment, but also to temporarily store data that has been output or will be output.

[0102] The communication bus 32 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable communication between the memory 31 and at least one processor 30, etc.

[0103] Communication interface 33 is used for communication between the aforementioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, Bluetooth interface, etc.), typically used to establish communication connections between the electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), or optionally, a standard wired or wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the electronic device and to display a visual user interface.

[0104] Figure 3 Only electronic devices with components are shown; those skilled in the art will understand that... Figure 3The structure shown does not constitute a limitation on the electronic device and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0105] For example, although not shown, the electronic device may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to at least one processor 30 via a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be elaborated further here.

[0106] It should be understood that the embodiments are for illustrative purposes only and are not limited to this structure in the scope of the patent application.

[0107] Furthermore, if the modules / units integrated into an electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, a computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, and read-only memory (ROM).

[0108] In the description of this specification, the references to terms such as "an embodiment," "some embodiments," "example," "specific example," "a preferred embodiment," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0109] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A transfer learning method based on privacy sample importance assessment, characterized in that, include: Multiple teacher models are used to predict samples in a common training set in turn to obtain multiple voting prediction results for each sample. Each teacher model is trained on a different privacy training set, and the common training set is images. Based on the multiple voting prediction results, calculate the voting frequency difference for each sample; The uncertainty score for each sample is predicted using the student model to be trained; The importance of each sample is jointly evaluated based on the voting frequency difference and the uncertainty score to obtain a joint evaluation result for each sample. This joint evaluation includes: Based on the voting frequency difference, a frequency difference weighting factor is determined for each sample, wherein determining the frequency difference weighting factor for each sample based on the voting frequency difference includes: determining the attenuation control coefficient corresponding to each sample; and using the attenuation control coefficient of each sample to perform exponential attenuation processing on the voting frequency difference to obtain the frequency difference weighting factor for each sample. Based on the frequency difference weighting factor and the uncertainty score, the joint evaluation result of each sample is determined, wherein determining the joint evaluation result of each sample based on the frequency difference weighting factor and the uncertainty score includes: calculating the joint evaluation result of each sample using the following formula: Score(x) = Entropy(x) · exp(-α · Gap(x)); where Score(x) is the joint evaluation result of sample x; Entropy(x) is the uncertainty score of sample x; Gap(x) is the voting frequency difference of sample x; exp(-α · ) is the exponential decay treatment; α is the decay control coefficient; exp(-α ⋅ Gap(x)) is the frequency difference weighting factor of sample x; Based on the joint evaluation results, a target sample training set is obtained by screening samples from the public training set, and privacy processing is performed on the target sample training set. The student model to be trained is then trained using the privacy-processed target sample training set to obtain the trained student model.

2. The method according to claim 1, characterized in that, The step of calculating the voting frequency difference for each sample based on the multiple voting prediction results includes: The voting frequency vector of each sample is obtained by statistically analyzing the multiple voting prediction results corresponding to each sample, wherein the voting frequency vector includes the number of votes for each sample in each prediction category; Calculate the difference in the number of votes for the two predicted categories with the highest number of votes in the voting frequency vector, and use the difference as the voting frequency difference for each sample.

3. The method according to claim 1, characterized in that, The step of predicting the uncertainty score for each sample using the student model to be trained includes: The samples in the common training set are input into the student model to be trained for inference, and the student prediction result corresponding to each sample is obtained. The uncertainty score for each sample is determined based on the student prediction results corresponding to each sample.

4. The method according to claim 2, characterized in that, The step of performing privacy processing on the target sample training set and training the student model to be trained using the privacy-processed target sample training set includes: The voting frequency vector of each sample in the target sample training set is noise-added using the differential privacy method; For each of the samples after noise processing, a pseudo-label is constructed to obtain a pseudo-label sample training set, wherein the pseudo-label represents the predicted category with the maximum value in the voting frequency vector of each sample; The student model to be trained is trained using the pseudo-labeled sample training set.

5. The method according to claim 1, characterized in that, Before using multiple teacher models to sequentially predict samples in a common training set, the process also includes: Obtain the sample training set corresponding to the target privacy domain; The sample training set corresponding to the target privacy domain is split to obtain multiple privacy training sets, wherein the privacy samples in each privacy training set are mutually exclusive. For each privacy training set, a corresponding teacher model is constructed and trained to obtain the multiple teacher models.

6. A transfer learning device based on privacy sample importance assessment, characterized in that, include: The first prediction module is configured to use multiple teacher models to predict samples in a common training set in turn, and obtain multiple voting prediction results for each sample, wherein each teacher model is trained on a different privacy training set, and the common training set is an image. The calculation module is configured to calculate the voting frequency difference for each of the multiple voting prediction results; The second prediction module is configured to predict the uncertainty score for each of the samples using the student model to be trained. The evaluation module is configured to jointly evaluate the importance of each sample based on the voting frequency difference and the uncertainty score, to obtain a joint evaluation result for each sample. The joint evaluation of the importance of each sample based on the voting frequency difference and the uncertainty score, to obtain a joint evaluation result for each sample, includes: Based on the voting frequency difference, a frequency difference weighting factor is determined for each sample, wherein determining the frequency difference weighting factor for each sample based on the voting frequency difference includes: determining the attenuation control coefficient corresponding to each sample; and using the attenuation control coefficient of each sample to perform exponential attenuation processing on the voting frequency difference to obtain the frequency difference weighting factor for each sample. Based on the frequency difference weighting factor and the uncertainty score, the joint evaluation result of each sample is determined, wherein determining the joint evaluation result of each sample based on the frequency difference weighting factor and the uncertainty score includes: calculating the joint evaluation result of each sample using the following formula: Score(x) = Entropy(x) · exp(-α · Gap(x)); where Score(x) is the joint evaluation result of sample x; Entropy(x) is the uncertainty score of sample x; Gap(x) is the voting frequency difference of sample x; exp(-α · ) is the exponential decay treatment; α is the decay control coefficient; exp(-α ⋅ Gap(x)) is the frequency difference weighting factor of sample x; The training module is configured to select a target sample training set from the public training set based on the joint evaluation results, perform privacy processing on the target sample training set, and use the privacy-processed target sample training set to train the student model to be trained, thereby obtaining a trained student model.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the transfer learning method based on privacy sample importance assessment as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Differential privacy federal learning method and system based on double knowledge transfer

    CN119622824A

  • Seabed sediment classification method based on multilevel comparative learning and uncertainty measurement

    CN120105205A