A differential privacy knowledge transfer method and system for medical data
By adopting differential privacy knowledge transfer methods of multi-model voting and Gaussian noise processing in medical data, the problem that differential privacy machine learning model is difficult to maintain efficient availability in medical data, and the model availability improvement under strong privacy protection is achieved.
Patent Information
- Application Number
- CN202211296875.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-10-21
AI Technical Summary
In medical data, differential privacy machine learning models are difficult to maintain efficient model availability while maintaining privacy protection, especially under the high privacy and sensitivity requirements of electronic medical records.
A differential privacy knowledge migration method for medical data is proposed. By dividing private medical data into multiple copies and training multiple classification models, voting mechanisms and Gaussian noise are used for data prediction and label aggregation, and tag optimization and cleaning are combined with k-NN algorithm to ensure that the usability of the model is improved without reducing privacy protection.
Through this method, the correctness of the label after noise is improved, the usability of the model is significantly improved, and the problem of low model accuracy under the requirements of strong privacy is solved.
Smart Images

Figure CN115985433B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of differential privacy machine learning technology, and in particular to a differential privacy knowledge transfer method and system for medical data. Background Art
[0002] In recent years, with the development of medical informatization, many paper medical records have been converted into electronic medical records, resulting in a large amount of electronic medical record data. For each medical center, it is now required to put their own electronic medical record data into use. Nowadays, various machine learning and deep learning algorithms and models have been widely used, but the model itself will implicitly remember the relevant information of the original data. This is unacceptable for sensitive medical data.
[0003] In order to ensure the security of data, an important measure taken by existing technologies is to protect the parameters generated by machine learning or deep learning models through differential privacy. Differential privacy technology adds noise to privacy-sensitive data, making it impossible for attackers to analyze whether any individual exists in the data set from the obtained data, thereby protecting the privacy of individuals.
[0004] Electronic medical records have the characteristics of large data volume, multiple attributes and fast updates, so the demand for mining the value of them is becoming more and more urgent, and machine learning naturally meets this demand. A large number of studies have used ML technology to conduct brain-related research, such as applying high-dimensional nonlinear pattern classification methods to functional magnetic resonance imaging images to distinguish spatial patterns of brain activity associated with lies and truth; a computer-assisted classification method combining conventional and perfusion magnetic resonance imaging for differential diagnosis of brain tumor types and grades; using SVM to detect epileptic seizures by analyzing scalp EEG and building patient-specific classifiers; the added value of various machine learning algorithms (such as SVM, NN and random forest (RF)) in predicting the prognosis of moderate to severe traumatic brain injury (TBI); using improved CSP and transfer learning algorithms to improve the accuracy of EEG signal classification and speed up training time, etc. Patterns of medical images can be identified by ML technology, allowing radiologists to make informed decisions based on radiological information, such as basic radiography, computed tomography (CT), MRI, positron emission tomography (PET) images and radiology reports. Researchers have proposed a sequential reinforcement learning technique for improving performance when using SVM to detect microcalcification (MC) clusters in mammograms, among others. ML and pattern recognition algorithms have a significant impact on brain imaging, and in the long run, technological developments in the field of ML and radiology can benefit each other. Deep learning (DL) is a branch of ML that deals with algorithms (i.e., ANNs) inspired by the biology and function of the brain. DL has quickly become the preferred method for evaluating medical images in the field of medical imaging, which has led to an increasing number of related studies covering neuropathology, abdomen, lungs, heart, retina, musculoskeletal, and breast, among others. In order to provide privacy protection, differential privacy must be adopted for the entire model. However, since the noise of differential privacy will affect the machine learning model itself, differential privacy machine learning will inevitably have a trade-off between model utility and privacy protection. Although reducing the noise scale can achieve some improvement in model performance, it will inevitably reduce protection in the future.
[0005] Existing machine learning methods based on differential privacy will choose to sacrifice a certain degree of privacy protection in order to ensure the usability of the model. Due to the high privacy and sensitivity of electronic medical records, this approach cannot be adopted. Therefore, how to maintain the usability of the model under the influence of strong noise without sacrificing privacy protection has become a problem that needs to be solved. Summary of the invention
[0006] In order to maintain the usability of the model under the influence of strong noise without sacrificing privacy protection, the present invention proposes a differential privacy knowledge transfer method and system for medical data, the method comprising the following steps:
[0007] Data users have public unlabeled medical data, and data owners have private medical data stored locally;
[0008] The data owner divides the private medical data into n parts, and uses each part of the data to train a medical diagnosis classification model using logistic regression. The n models together form a medical diagnosis teacher model.
[0009] The data user sends the unlabeled medical data to the data owner, who uses the trained teacher model to predict all the unlabeled medical data and gives the classification results for each sample. The classification results are expressed as a series of label probabilities.
[0010] Each model votes for the classification structure it has derived and votes for the label with the highest probability in the voting results. After the voting results of n models are aggregated, Gaussian noise perturbation is added to the aggregated voting results, and the label with the most votes after the perturbation is sent to the data user;
[0011] The data user labels the unlabeled data according to the received prediction results, and uses the k-NN algorithm to cluster the labels and optimize the labels of the unlabeled data;
[0012] The data user uses the obtained labeled data to perform local training, uses logistic regression to train its own classification model, obtains the student model, and uses the student model to classify the medical data on the data user side.
[0013] Furthermore, the data owner adds Gaussian noise to the obtained prediction results to aggregate the prediction results, that is, to set noise parameters, including the privacy budget ε and the privacy scale δ to train the voting results. The training process includes the following steps:
[0014] Randomly select a random number from the Gaussian distribution composed of the privacy budget ε and the privacy scale δ, and add the random number as noise to the voting result;
[0015] Calculate the privacy loss of the voting results after adding noise, and determine whether ε-c ≤ 0. If so, end the training, otherwise continue to add noise until ε-c ≤ 0.
[0016] Furthermore, the privacy loss after adding noise is expressed as:
[0017]
[0018] Among them, c(o; M, aux, d, d') represents the privacy loss after adding noise, o represents the voting result after inputting noise, M represents the logistic regression training algorithm, aux represents the training parameters of the logistic regression training algorithm, d and d' represent two data sets and the two data sets differ by only one data; Pr[M(aux, d) = o] represents the probability that the voting result of data set d is o calculated by the logistic regression training algorithm M with the network parameters aux, and Pr[M(aux, d') = o] represents the probability that the voting result of data set d' is o calculated by the logistic regression training algorithm M with the network parameters aux.
[0019] Furthermore, the noise added to the prediction result has a mean of 0 and a variance of σ 2 A random number from a Gaussian distribution with a variance σ that satisfies:
[0020]
[0021] Here, s represents sensitivity.
[0022] Furthermore, the privacy loss is controlled by adjusting the noise parameters (ε, δ). The relationship between the noise parameters and the privacy loss can be expressed as:
[0023]
[0024] Among them, α M (λ; aux, d, d') represents the privacy loss of two data sets d and d' that differ by only one data point when the logistic regression training algorithm M is used with the system parameter set aux; λ represents.
[0025] Furthermore, the k-NN algorithm is used to clean the labels, and obtaining the labels of the unlabeled data includes the following steps: calculating the distance between the unlabeled data, for a sample in the unlabeled data, aggregating the labels of the k samples closest to the sample, setting a threshold t, and when the proportion of the label of the current sample in the k samples is greater than the set threshold t, the label of the current sample is relabeled as the union of the labels of its nearest k samples; if the proportion of the label of the current sample in the k samples is not greater than the set threshold t, the label of the current sample is not relabeled.
[0026] Furthermore, the data user has trusted data stored locally. After the data user completes the training of the local model using the obtained labeled data, the trusted data is input into the trained model for testing. If the loss of the model is less than the set threshold, the knowledge transfer is completed, otherwise the knowledge transfer is performed again.
[0027] The present invention also provides a differential privacy knowledge transfer system for medical data, which is used to implement a differential privacy knowledge transfer method for medical data, including a data user server and a data owner server. The data owner server divides local privacy data into n parts, and each part obtains a classification model based on logistic regression training; when the data user server requests service from the data user server, the data request server sends its local unlabeled data to the data owner, and the data owner uses the local n models to predict the labels of each unlabeled data respectively, and each model votes for the label with the highest probability in the prediction result, and aggregates the voting results of the n classification models, adds Gaussian noise to the aggregated result, and uses the label with the highest votes after adding noise as the label of the corresponding data. The data owner sends the predicted labels of all unlabeled data to the data user server; the data user server cleans the received labels, labels the local unlabeled data with the cleaned labels, and trains the local classification model using the labeled data.
[0028] Compared with the solutions in the prior art, the present invention improves the correctness of the noisy labels, greatly improves the usability of the model without reducing the privacy protection of differential privacy, and effectively solves the problem of low model accuracy in the prior art under the strong privacy requirements of medical centers. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is a flow chart of a differential privacy knowledge transfer method for medical data of the present invention;
[0030] Figure 2 The flowchart of the student who is the data user in the present invention re-labeling the data set through the k-NN algorithm. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0032] This paper proposes a differential privacy knowledge transfer method for medical data. Figure 1 , specifically including the following steps:
[0033] Data users have public unlabeled medical data, and data owners have private medical data stored locally;
[0034] The data owner divides the private medical data into n parts, and uses each part of the data to train a medical diagnosis classification model using logistic regression. The n models together form a medical diagnosis teacher model.
[0035] The data user sends the unlabeled medical data to the data owner, who uses the trained teacher model to predict all the unlabeled medical data and gives the classification results for each sample. The classification results are expressed as a series of label probabilities.
[0036] Each model votes for the classification structure it has derived and votes for the label with the highest probability in the voting results. After the voting results of n models are aggregated, Gaussian noise perturbation is added to the aggregated voting results, and the label with the most votes after the perturbation is sent to the data user;
[0037] The data user labels the unlabeled data according to the received prediction results, and uses the k-NN algorithm to cluster the labels and optimize the labels of the unlabeled data;
[0038] The data user uses the obtained labeled data to perform local training, uses logistic regression to train its own classification model, obtains the student model, and uses the student model to classify the medical data on the data user side.
[0039] Furthermore, the data owner adds Gaussian noise to the obtained prediction results to aggregate the prediction results, that is, to set noise parameters, including the privacy budget ε and the privacy scale δ to train the voting results. The training process includes the following steps:
[0040] Randomly select a random number from the Gaussian distribution composed of the privacy budget ε and the privacy scale δ, and add the random number as noise to the voting result;
[0041] Calculate the privacy loss of the voting results after adding noise, and determine whether ε-c ≤ 0. If so, end the training, otherwise continue to add noise until ε-c ≤ 0.
[0042] Furthermore, the privacy loss after adding noise is expressed as:
[0043]
[0044] Among them, c(o; M, aux, d, d') represents the privacy loss after adding noise, o represents the voting result after inputting noise, M represents the logistic regression training algorithm, aux represents the training parameters of the logistic regression training algorithm, d and d' represent two data sets and the two data sets differ by only one data; Pr[M(aux, d) = o] represents the probability that the voting result of data set d is o calculated by the logistic regression training algorithm M with the network parameters aux, and Pr[M(aux, d') = o] represents the probability that the voting result of data set d' is o calculated by the logistic regression training algorithm M with the network parameters aux.
[0045] Furthermore, the noise added to the prediction result has a mean of 0 and a variance of σ 2 A random number from a Gaussian distribution with a variance σ that satisfies:
[0046]
[0047] Here, s represents sensitivity.
[0048] Furthermore, the privacy loss is controlled by adjusting the noise parameters (ε, δ). The relationship between the noise parameters and the privacy loss can be expressed as:
[0049] Furthermore, the k-NN algorithm is used to clean the labels, and obtaining the labels of the unlabeled data includes the following steps: calculating the distance between the unlabeled data, for a sample in the unlabeled data, aggregating the labels of the k samples closest to the sample, setting a threshold t, and when the proportion of the label of the current sample in the k samples is greater than the set threshold t, the label of the current sample is relabeled as the union of the labels of its nearest k samples; if the proportion of the label of the current sample in the k samples is not greater than the set threshold t, the label of the current sample is not relabeled.
[0050] Furthermore, the data user has trusted data stored locally. After the data user completes the training of the local model using the obtained labeled data, the trusted data is input into the trained model for testing. If the loss of the model is less than the set threshold, the knowledge transfer is completed, otherwise the knowledge transfer is performed again.
[0051] The present invention also proposes a differential privacy knowledge transfer system for medical data, which is used to implement a differential privacy knowledge transfer method for medical data, including a data user server and a data owner server. The data owner server divides local privacy data into n parts, and each part obtains a classification model based on logistic regression training; when the data user server requests service from the data user server, the data request server sends its local unlabeled data to the data owner, and the data owner uses the local n models to predict the labels of each unlabeled data respectively, and each model votes for the label with the highest probability in the prediction result, and aggregates the voting results of the n classification models, adds Gaussian noise to the aggregated result, and uses the label with the highest votes after adding noise as the label of the corresponding data. The data owner sends the predicted labels of all unlabeled data to the data user server; the data user server cleans the received labels, labels the local unlabeled data with the cleaned labels, and uses the labeled data to train the local classification model.
[0052] The present invention adopts a differential privacy framework based on k-NN, which aims to solve the problem that in medical scenarios, when the data of the medical center is not stored in the warehouse, it needs to provide services to other medical centers or third parties. In this embodiment, the differential privacy framework based on k-NN generally includes three parts, namely, a teacher model belonging to the data owner, an aggregation mechanism belonging to the data owner, and a student model of the data user. The data owner is an institution such as a medical center that owns the private data, and the data user is a medical center or other third-party service agency that does not own the private data and wants to use the private data to train a machine learning model.
[0053] The private data owned by the data owner is the data that needs to be protected in this application. This data includes but is not limited to the patient's identity information, medical records and other data. This application refers to this type of data as electronic medical data. The medical center as the data owner represents the electronic medical data as D = {D1, D2, ..., D N}, D represents the collection of electronic medical data, D N represents the Nth electronic medical data, N is the total number of electronic medical data, these electronic medical data are divided into n parts, and the same model is trained for each part, and n models are obtained by training, which is expressed as M = {M1, M2, ..., M n}, M represents a set of n models, M n Represents the model trained based on the nth piece of data. Each model is used to make predictions based on the data user's query. These n models are collectively referred to as teacher models.
[0054] The data owner's aggregation mechanism is to aggregate the voting results from the teacher model, add Gaussian noise for perturbation, and provide the corresponding query results to the data user based on the query.
[0055] Data users need to disclose a portion of non-private unlabeled data sets to the medical center as the data owner, allowing the data owner to make predictions for each piece of data and aggregate the prediction results and feed them back to the data user. After the data user receives the prediction results, he or she will label the unlabeled data according to the prediction results and train the local model based on the labeled data.
[0056] In view of the above scheme, this embodiment provides a specific implementation method, including the following steps:
[0057] 1. The data user has a public unlabeled dataset, and the medical center has private local data. The data user has a small validation set, which is data that the data user considers to be credible.
[0058] For the teacher model:
[0059] 2. The medical center manually divides its data into n parts. In this embodiment, n∈{3,5,10,15,20,25,30,50}. For each piece of data, a local model is trained. For convenience, the same training algorithm can be selected for all. For example, in this embodiment, a logistic regression algorithm is selected for training. For each unlabeled sample disclosed by the data user, each model of the medical center gives its own prediction result.
[0060] 3. Aggregation Mechanism Collect the prediction results of all medical centers and aggregate them, which is called voting result in this embodiment. Gaussian noise is added to the voting result of each piece of data.
[0061] The medical center calculates the privacy loss this time according to the RDP method. The specific process is as follows:
[0062] The medical center collects all the prediction results, aggregates them, generates a histogram, and adds Gaussian noise to the statistical results of the histogram.
[0063] For the noise parameters (ε, δ), different δ is used according to the actual data set, generally set to 1 / n, where n is the size of the data set; for the ε parameter, set ε∈{0.1, 0.2, 0.3, 0.4, 0.5, 1}. The medical center calculates the corresponding privacy loss based on different ε parameters. The larger the privacy loss, the less secure the model is, but the less disturbance the voting results are subject to, the more realistic the voting results received by the data users are, and the stronger the usability is; the smaller the privacy loss, the more secure the model is, but the greater the disturbance the voting results are subject to, and the weaker the usability is. The medical center selects different ε based on the acceptable privacy loss.
[0064] According to the current noise scale, the privacy loss is calculated according to the RDP method.
[0065] According to the actual situation, we add noise of different scales to the voting results and send the voting results to the data users.
[0066] For the student model:
[0067] 4. After the data user gets the voting results, he / she will label the original public unlabeled dataset according to the voting results, and the data user will get the complete dataset.
[0068] 5. Since the labels are disturbed by differential privacy noise, which leads to a certain degree of deviation, the data user uses the k-NN algorithm to process the labeled data set to eliminate the label flipping caused by differential privacy noise. According to the post-processing of differential privacy, this process will not reduce the protection of differential privacy. Figure 2 , the specific steps are as follows:
[0069] For each sample in the re-labeled data set, the k samples closest to it in the data set are calculated. In this embodiment, the Euclidean distance calculation method is selected to calculate the distance between two samples (other distance calculation methods such as cosine distance and Hamming distance can also be used). In addition, the k value can be adjusted according to actual conditions. This embodiment tests the effect of k values between 1-15, and those skilled in the art can choose the best one.
[0070] For these k samples, a threshold t is set. When the label ratio of the majority samples among these k samples is greater than the threshold, the original label is relabeled as the majority label at this time; otherwise, no operation is performed.
[0071] After relabeling all samples, data users obtain a label-sanitized dataset. According to the post-processing nature of differential privacy, this process will not reduce the privacy protection of differential privacy.
[0072] 6. The data user trains on the disinfected data set to obtain his own model and verifies the effect of his model on his own validation set. At this point, the data user actually obtains the value of the medical center's data and completes the knowledge transfer. At the same time, the medical center's data never leaves the local area during this process and is protected by differential privacy.
[0073] In this embodiment, the user adjusts the privacy loss of the model by adjusting the noise parameter according to the relationship between the noise parameter and the privacy loss. When the user sets the noise parameter (ε, δ), the smaller the value of (ε, δ), the larger the current noise scale, the stronger the noise, the greater the degree of disturbance to the original result, and the better the privacy protection effect. According to the value of (ε, δ), a value is taken from the Gaussian distribution and added to the original result. Excessive noise will cause the noisy result to deviate too much from the original result and make the data completely unusable. Smaller noise will make the privacy protection insufficient and it is easy for attackers to infer personal privacy information. Usually, the value of ε does not exceed 10. The smaller the better while ensuring the usability of the model; the value of δ is usually set to 1 / n, where n is the amount of data.
[0074] In actual situations, it is necessary to continuously adjust (ε, δ) to strike a balance between model availability and privacy protection. Generally speaking, before training a differential privacy model, a total ε is set, called the privacy budget, which is the degree of privacy leakage that users can accept. As the model is trained, noise is continuously added during the process. Each time noise is added, a privacy loss c is caused. The privacy loss c at this time is subtracted from the privacy budget. When the privacy budget ε is reduced to 0, the model stops training; the larger the ε setting, the lower the model security, but the better the model training effect; the smaller the ε setting, the higher the model security, but the model training effect will deteriorate.
[0075] The availability of the model can be reflected by the accuracy of the classification model, and the degree of privacy protection needs to be reflected by the calculation of privacy loss. In this embodiment, the RDP method is used to calculate the privacy loss, and the specific process is as follows:
[0076] Let M represent the logistic regression training algorithm, d, d' represent two adjacent medical data sets (there is only one sample difference between the two medical data sets). The difference is to prevent the privacy detection program from distinguishing whether the sample of the two data sets is involved in the model training during the training process. For example, for any sample x, if the d data set includes sample x, d' = d\{x}, that is, the difference between d' and d is that it does not include sample x;
[0077] After differential privacy processing, the probability that the model is built on d or d' is similar in the eyes of the adversary, so it is impossible to distinguish whether it is d or d', which can protect the privacy of the individual; aux represents all other parameters in the training algorithm. For any output result o after adding noise, the privacy loss is defined as:
[0078]
[0079] Let λ∈{1, 2, ..., 32}, define α M (λ; aux, d, d') = logE o~M(aux,d) [exp(λc(o;M,aux,d,d'))];
[0080] According to the combination theorem and the tail limit theorem:
[0081]
[0082] For the set noise parameters (ε, δ), for the prediction results of each sample, we add Gaussian noise. Gaussian noise is a random number selected from a Gaussian distribution. The variance σ of the Gaussian distribution satisfies the following relationship:
[0083]
[0084] Among them, s is the sensitivity, which is generally set to 1, and (ε, δ) are pre-set parameters.
[0085] This embodiment also proposes a differential privacy knowledge transfer system for medical data. The system is used to implement a differential privacy knowledge transfer method for medical data, including a data user server and a data owner server. The data owner server divides local private data into n parts, each of which obtains a classification model based on logistic regression training; when the data user server requests service from the data user server, the data request server sends its local unlabeled data to the data owner, and the data owner uses the local n models to predict the labels of each unlabeled data respectively, and each model votes for the label with the highest probability in the prediction result, and aggregates the voting results of the n classification models, adds Gaussian noise to the aggregated result, and uses the label with the highest votes after adding noise as the label of the corresponding data. The data owner sends the predicted labels of all unlabeled data to the data user server; the data user server cleans the received labels, labels the local unlabeled data with the cleaned labels, and trains the local classification model with the labeled data.
[0086] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A differential privacy knowledge transfer method for medical data, characterized in that: The specific steps include: Data users have public unlabeled medical data, and data owners have private medical data stored locally; The data owner divides the private medical data into n parts, and uses each part of the data to train a medical diagnosis classification model using logistic regression. The n models together form a medical diagnosis teacher model. The data user sends the unlabeled medical data to the data owner, who uses the trained teacher model to predict all the unlabeled medical data and gives the classification results for each sample. The classification results are expressed as a series of label probabilities. Each model votes for the classification structure it has derived and votes for the label with the highest probability in the voting results. After the voting results of n models are aggregated, Gaussian noise perturbation is added to the aggregated voting results, and the label with the most votes after the perturbation is sent to the data user; The data user labels the unlabeled data according to the received prediction results, and uses the k-NN algorithm to cluster the labels and optimize the labels of the unlabeled data; The data user uses the obtained labeled data to perform local training, uses logistic regression to train its own classification model, obtains the student model, and uses the student model to classify the medical data on the data user side.
2. According to claim 1, a differential privacy knowledge transfer method for medical data is characterized in that: The data owner adds Gaussian noise to the obtained prediction results to aggregate the prediction results, that is, to set the noise parameters, including the privacy budget ε and the privacy scale δ to train the voting results. The training process includes the following steps: Randomly select a random number from the Gaussian distribution composed of the privacy budget ε and the privacy scale δ, and add the random number as noise to the voting result; Calculate the privacy loss of the voting results after adding noise, and determine whether ε-c ≤ 0. If so, end the training, otherwise continue to add noise until ε-c ≤ 0.
3. A differential privacy knowledge transfer method for medical data according to claim 2, characterized in that: The privacy loss after adding noise is expressed as: Among them, c(o; M, aux, d, d') represents the privacy loss after adding noise, o represents the voting result after inputting noise, M represents the logistic regression training algorithm, aux represents the training parameters of the logistic regression training algorithm, d and d' represent two data sets and the two data sets differ by only one data; Pr[M(aux, d) = o] represents the probability that the voting result of data set d is o calculated by the logistic regression training algorithm M with the network parameters aux, and Pr[M(aux, d') = o] represents the probability that the voting result of data set d' is o calculated by the logistic regression training algorithm M with the network parameters aux.
4. A differential privacy knowledge transfer method for medical data according to claim 2, characterized in that: The noise added to the prediction result has a mean of 0 and a variance of σ 2 A random number from a Gaussian distribution with a variance σ that satisfies: Here, s represents sensitivity.
5. The differential privacy knowledge transfer method for medical data according to claim 2, characterized in that: The privacy loss is controlled by adjusting the noise parameters (ε, δ). The relationship between the noise parameters and the privacy loss can be expressed as: Among them, α M (λ; aux, d, d') represents the privacy loss of two data sets d and d' that differ by only one data point when the logistic regression training algorithm M is used with the system parameter set aux; λ represents.
6. The differential privacy knowledge transfer method for medical data according to claim 1, characterized in that: The k-NN algorithm is used to clean the labels, and obtaining the labels of the unlabeled data includes the following steps: calculating the distance between the unlabeled data, for a sample in the unlabeled data, aggregating the labels of the k samples closest to the sample, setting a threshold t, and when the proportion of the label of the current sample in the k samples is greater than the set threshold t, the label of the current sample is relabeled as the union of the labels of its nearest k samples; if the proportion of the label of the current sample in the k samples is not greater than the set threshold t, the label of the current sample is not relabeled.
7. The differential privacy knowledge transfer method for medical data according to claim 1, characterized in that: The data user stores trusted data locally. After the data user completes the training of the local model using the obtained labeled data, the trusted data is input into the trained model for testing. If the loss of the model is less than the set threshold, the knowledge transfer is completed, otherwise the knowledge transfer is performed again.
8. A differential privacy knowledge transfer system for medical data, characterized in that: The system is used to implement a differential privacy knowledge transfer method for medical data as described in claim 1, comprising a data user server and a data owner server, the data owner server divides local privacy data into n parts, each of which obtains a classification model based on logistic regression training; when the data user server requests service from the data user server, the data request server sends its local unlabeled data to the data owner, the data owner uses the local n models to predict the label of each unlabeled data respectively, each model votes for the label with the highest probability in the prediction result, and aggregates the voting results of the n classification models, adds Gaussian noise to the aggregated result, and uses the label with the highest votes after adding the noise as the label of the corresponding data, and the data owner sends the predicted labels of all unlabeled data to the data user server; The data user server cleans the received labels, labels the local unlabeled data with the cleaned labels, and uses the labeled data to train the local classification model.
9. A differential privacy knowledge transfer system for medical data according to claim 8, characterized in that: Aggregate the voting results of n classification models and add Gaussian noise to the aggregated results. That is, the data owner adds Gaussian noise to the obtained prediction results to aggregate the prediction results, that is, set noise parameters, including privacy budget ε and privacy scale δ to train the voting results. The training process includes the following steps: Randomly select a random number from the Gaussian distribution composed of the privacy budget ε and the privacy scale δ, and add the random number as noise to the voting result; Calculate the privacy loss of the voting results after adding noise, and determine whether ε-c ≤ 0. If so, end the training, otherwise continue to add noise until ε-c ≤ 0.
10. A differential privacy knowledge transfer system for medical data according to claim 8, characterized in that: The data user server uses the k-NN algorithm to clean the received labels, including: calculating the distance between unlabeled data, for a sample in the unlabeled data, aggregating the labels of the k samples closest to the sample, setting a threshold t, and when the proportion of the label of the current sample in the k samples is greater than the set threshold t, the label of the current sample is relabeled as the union of the labels of its most recent k samples; if the proportion of the label of the current sample in the k samples is not greater than the set threshold t, the label of the current sample is not relabeled.
Citation Information
Patent Citations
Adversarial domain adaptive differential privacy protection method in multi-source domain migration
CN112800471A
Distributed privacy protection method and system based on semi-supervised-transfer learning
CN114927190A