A defense method against attribute inference attacks in machine learning

By constructing a disguised dataset and iterative training methods, voting models are generated, and the problem of attribute inference attacks in machine learning models is solved, which significantly improves the privacy, security and utility of the model.

CN115329984BActive Publication Date: 2025-06-06NANJING YIZHI CYBERSPACE TECH INNOVATION INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211078605.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-05
Publication Date
2025-06-06
Estimated Expiration
2042-09-05

AI Technical Summary

Technical Problem

The existing technology cannot effectively defend against attribute reasoning attacks in machine learning models, resulting in the risk of model privacy leakage.

Method used

By constructing a disguised dataset, the feature space is the same as the original dataset but the statistical distribution is different. Combined with iterative training methods, a new voting model is generated to fuzz the global attribute characteristics of the model and reduce the success rate of attribute inference attacks.

Benefits of technology

It significantly reduces the threat of machine learning models to attribute inference attacks, ensures the privacy and security of the model, and improves the security of the model while ensuring the effectiveness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115329984B_ABST
    Figure CN115329984B_ABST
Patent Text Reader

Abstract

The present invention discloses a defense method for attribute inference attacks in machine learning, including: constructing a disguised data set based on an original data set, wherein the disguised data set has the same feature space as the original data set but has a different statistical distribution; using the original data set and the disguised data set to train a machine learning model to obtain a voting model; reconstructing a new disguised data set based on the original data set, using the voting model to screen the new disguised data set, using the output of the voting model as its new label for each sample in the new disguised data set, completing the reconstruction of the new disguised data set, training the reconstructed disguised data set and the original data set together to generate a new voting model; repeating iterations until the number of iterations reaches a maximum number of iterations. The present invention improves the security of the model while ensuring the utility of the machine learning model, and ultimately realizes a training method for a machine learning model with high utility and high security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning, and aims at the security problem of machine learning models, and proposes a defense method against attribute inference attacks in machine learning. Background Art

[0002] In recent years, with the rapid development of artificial intelligence, machine learning capabilities have increased rapidly, and the great success of machine learning has become the most important driving force of artificial intelligence. The sharp increase in data volume and powerful computing infrastructure, as well as the progress in training complex machine learning models, have greatly increased the application of machine learning in software systems, such as image recognition, speech recognition, natural language processing, malware identification, etc. The widespread application of machine learning has brought great convenience to people's lives. However, developing an excellent machine learning model requires a lot of investment in computing time and manpower. This has promoted the establishment of the online market - Machine Learning as a Service (MLaaS) platform, and shared machine learning models have become more and more popular. On the MLaaS platform, platform providers use their strong professional capabilities in artificial intelligence and deep learning to provide users with powerful computing resources and rich machine learning APIs. Users provide their own private data sets and train them on the platform to obtain the corresponding machine learning models. Users use the obtained models for sharing and trading.

[0003] However, the trained machine learning model itself may have privacy leakage security issues. Due to the complexity and difficulty of interpreting the machine learning model itself, the machine learning model may remember some additional feature information about the training set during the training process. Usually, this part of information is sensitive. If it is leaked, it will inevitably lead to serious privacy security issues, which has attracted widespread attention from scholars at home and abroad. For machine learning models, attackers use specific attack strategies to extract private information hidden in the model under the white box conditions of obtaining the model itself or the black box conditions of obtaining the model prediction API, or infer sensitive information contained in the model training process. Existing research divides attacks into four categories: model extraction attack, model inversion attack, poisoning attack, and adversarial attack.

[0004] Reference 1 (Tramèr F, Zhang F, Juels A, et al. Stealing machine learning models via prediction apis [C] / / 25th {USENIX} Security Symposium ({USENIX} Security 16). 2016: 601-618) discloses that model extraction attacks are attacks on machine learning models themselves. This attack uses the model prediction API to copy the model without knowing the training data and algorithm, thereby obtaining sensitive information such as the structure, parameters, and hyperparameters of the model itself. Reference 2 (Xiao H, Biggio B, Brown G, et al. Is feature selection secure against training data poisoning [C] / / International Conference on Machine Learning. 2015: 1689-1698) discloses that poisoning attacks are attacks on model utility, causing the model to produce misclassification. This attack pollutes the model training set, deviates from the model decision boundary, or injects a backdoor into the model, causing the model to generate incorrect prediction labels for certain specific samples, reducing the prediction accuracy of the model and destroying the usability of the model. Reference 3 (Goodfellow IJ, Shlens J, Szegedy C. Explaining and harnessing adversarial examples [J]. arXiv preprint arXiv: 1412.6572, 2014) discloses that adversarial attacks are also attacks on model utility, causing the model to misclassify. However, unlike poisoning attacks, this attack does not affect the training process of the target model itself, but exploits the model's own vulnerabilities to attack. This attack uses a specific attack strategy to add perturbations to the original input samples to generate malicious samples, causing the model to generate incorrect prediction labels for this sample, while the human eye cannot distinguish between the original sample and the malicious sample. Model inversion attacks attack the model training set and steal information related to the training set. This attack infers the attributes of the training set by using the model itself or the model prediction API. The attributes obtained are those that the model owner does not want to disclose, that is, the attribute information is sensitive. This attack uses an additional attack model to extract sensitive information about the training set hidden in the model structure, parameters, or predictions. Model inference attacks are specifically divided into two types: membership inference attack and property inference attack.Reference 4 (Shokri R, Stronati M, Song C, et al. Membership inference attacks against machine learning models [C] / / 2017 IEEE Symposium on Security and Privacy (SP). IEEE, 2017: 3-18) discloses that membership inference attacks can infer whether a specific sample is included in the model training set under the white-box condition of obtaining the model structure and parameters or under the black-box condition of using the model prediction API. Reference 5 (Ateniese G, Mancini LV, Spognardi A, et al. Hacking smartmachines with smarter ones: How to extract meaningful data from machine learning classifiers [J]. International Journal of Security and Networks, 2015, 10 (3): 137-150) discloses that attribute inference attacks can infer the global attributes of the model training set under the white-box condition of obtaining the model structure and parameters.

[0005] Attribute inference attacks are attacks on model training sets. Attackers use specific attack strategies to train an attack model to infer the global attributes of the target model training set. For example, the ratio of male to female in the face recognition model training set is 1:1, and the proportion of a certain type of patient in the entire data set in the medical diagnosis model training set. This type of attack is widely used in various machine learning scenarios, such as support vector machine models, hidden Markov models, transfer learning, federated learning, etc. There is currently no defense strategy for the above attacks.

[0006] The invention with publication number CN113259369A discloses a data set authentication method and system based on machine learning member inference attack. The invention can only make probabilistic judgments on member fingerprint data based on the authentication model, thereby determining whether the suspicious model is trained by the Internet of Things data set, and cannot solve the problem of attribute inference attacks on machine learning models.

[0007] The invention with publication number CN111310819A discloses a data screening method, device, equipment and readable storage medium. The invention screens the data set based on the error range configured by the coordinator, so that the training data of the participants are similar but different, and can make full use of the diversity of the training data owned by the participants, maximize the advantages of federated learning, and train better models. The invention requires reasonable configuration of the error range, and at the same time cannot solve the problem of attribute reasoning attacks on machine learning models. Summary of the invention

[0008] Technical problem to be solved: The present invention studies the defense method of attribute reasoning attacks in machine learning and proposes a defense method of attribute reasoning attacks for the first time. On the premise of ensuring the effectiveness of the machine learning model, the security of the model is improved, and finally a training method of a high-utility and high-security machine learning model is realized.

[0009] Technical solution:

[0010] A method for defending against attribute inference attacks in machine learning, the method for building a machine learning model comprising the following steps:

[0011] S1, selecting or constructing a machine learning model according to the application scenario, generating an original data set corresponding to the machine learning model; constructing a disguised data set based on the original data set, wherein the disguised data set has the same feature space as the original data set but has a different statistical distribution;

[0012] S2, preprocessing the original data set and disguised data set to eliminate the noise in them;

[0013] S3, using the preprocessed original data set and disguised data set to train the machine learning model, and the trained machine learning model is named the voting model;

[0014] S4, reconstruct a new disguised dataset based on the original dataset, use the voting model to screen the new disguised dataset, use the output of the voting model as its new label for each sample in the new disguised dataset, complete the reconstruction of the new disguised dataset, train the reconstructed disguised dataset and the original dataset together, generate a new voting model, and increase the number of iterations by one;

[0015] S5, repeat step S4 until the number of iterations reaches the maximum number of iterations , is a positive integer greater than 1.

[0016] Furthermore, in step S1, the process of constructing a disguised data set based on the original data set includes the following sub-steps:

[0017] S11, assuming that the given original data set is ,in Indicates the original data set No. i Sample features, Indicates i Sample features correspond to labels, original data set Total N Such feature-label pairs are extracted to obtain the features, labels and their corresponding value ranges of the original data set;

[0018] S12, find out whether there is a publicly available dataset with the same feature space as the original dataset but with a different statistical distribution. If so, extract the number of n If the sample is , a disguised data set is generated, otherwise, the process goes to step S13;

[0019] S13, using the obtained features and labels to construct the full range feature space of the original data set, the number of random sampling in the full range feature space is n The sample construction generates a disguised dataset , Represents the camouflage dataset Middle j Sample features, Indicates j The sample features correspond to labels.

[0020] Furthermore, in step S2, data cleaning, data reduction and data transformation are performed on the samples in the original data set and the disguised data set in turn, so as to complete the preprocessing of the original data set and the disguised data set.

[0021] Furthermore, in step S3, the process of training the machine learning model using the preprocessed original data set and the disguised data set and naming the trained machine learning model as a voting model includes the following sub-steps:

[0022] Initializing the machine learning model , indicating the model The characteristics Mapping to labels ;

[0023] Improve the loss function of the machine learning model and get ;in is a constant, and 1, Represents the calculation for the original data set Features in Model decision label and the true label The loss between Represents the calculation for the disguised dataset Features in Model decision label and the true label Losses between

[0024] By minimizing the improved loss function , and use the stochastic gradient descent method to gradually complete the model learning process until the model accuracy reaches the preset threshold and then the training is terminated to obtain the machine learning model ; During the model training process, the original data set is used and camouflage dataset Two data sets are involved in the training of the model; during training, by constantly adjusting size, so that the machine learning model can be used in the original dataset and camouflage dataset The balance between utility and feature representation is achieved, so that the model not only guarantees the model utility on the original data set, but the model's feature representation is more inclined to show that the model is disguising the data set. Obtained through training.

[0025] Furthermore, in step S4, one iteration process includes the following steps:

[0026] Assume that the voting model obtained in the previous iteration is ;

[0027] Re-obtain a sample size of n The camouflage dataset ,in Indicates that in the first iteration, the disguised data set No. j Sample features, Indicates that the dataset is disguised at the first iteration Middle j tags.

[0028] For the disguised dataset Each sample in is input into the voting model In the example above, we get the output corresponding to the sample; if the voting model If the output of is the same as the original label of the sample, it remains unchanged. If it is different, the voting model is used. The output of is used as the new label of the sample;

[0029] Generating a new camouflage dataset ,in Indicates j Sample characteristics Voting Model Corrected label:

[0030] Using the original dataset And the newly arrived disguised dataset Participate in machine learning model training to obtain new voting models .

[0031] Furthermore, the machine learning model construction method also includes the following steps:

[0032] S6, repeat the iteration and get After the voting model, this The voting models are screened in the full range dataset to obtain the final disguised dataset.

[0033] Further, in step S6, using this The process of screening the full range data set by a voting model to obtain the final disguised data set includes the following steps:

[0034] S61, randomly select a data sample in the training set input space of the target model, Each voting model makes a prediction for this sample. The prediction process is as follows: vote on the sample label, and finally use the prediction result with the highest number of votes as the label of the sample;

[0035] S62, repeat step S61 until a batch of samples are obtained to form a disguised data set.

[0036] Furthermore, the machine learning model includes a support vector machine model, a hidden Markov model, a transfer learning model and a federated learning model.

[0037] Furthermore, the application scenarios include target recognition, task classification, probability prediction and recommendation optimization.

[0038] Beneficial effects:

[0039] First, the defense method against attribute reasoning attacks in machine learning of the present invention proposes a new way of constructing a data set, so that the machine learning model trained using the data set can not only guarantee the utility on the original data set, but also blur the global attribute characteristics of the original data set, so that the attribute reasoning attack will incorrectly judge the global attribute information of the original data set.

[0040] Second, the defense method against attribute inference attacks in machine learning of the present invention improves the loss function during training of traditional machine learning models, so that two different but similar data sets can participate in the training of the machine learning model at the same time.

[0041] Third, the defense method against attribute reasoning attacks in machine learning of the present invention is the first to propose a defense method against attribute reasoning attacks. The machine learning model trained by this method can significantly reduce the threat of attribute reasoning attacks, avoid the privacy of the model from being leaked, and ensure the privacy security of the machine learning model.

[0042] Fourth, the defense method against attribute inference attacks in machine learning of the present invention can be widely applied to various typical machine learning models in the industrial field, including support vector machine models, hidden Markov models, transfer learning models and federated learning models, to complete learning tasks in different application scenarios, such as target recognition, task classification, probability prediction and recommendation optimization, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 A flowchart of a method for defending against attribute inference attacks in machine learning according to an embodiment of the present invention;

[0044] Figure 2 Schematic diagram of the iterative process. DETAILED DESCRIPTION

[0045] The following examples will enable those skilled in the art to more fully understand the present invention, but are not intended to limit the present invention in any way.

[0046] This embodiment intuitively proposes to use the "original data set + disguised data set iterative training" method during the model training process, so that the final model can guarantee the utility on the original data set, and make the characteristics of the model itself biased towards the disguised data set, resulting in the failure of attribute reasoning attacks. Based on this point of view, this embodiment conducts research on attribute reasoning attacks and defenses, and proposes an attribute reasoning attack defense method.

[0047] See also Figure 1 This embodiment proposes a defense method for attribute inference attacks in machine learning. The machine learning model construction method includes the following steps:

[0048] S1, select or construct a machine learning model according to the application scenario, and generate an original data set corresponding to the machine learning model; construct a disguised data set based on the original data set, and the disguised data set has the same feature space as the original data set but a different statistical distribution.

[0049] S2, preprocess the original data set and the disguised data set to eliminate the noise therein.

[0050] S3, uses the preprocessed original data set and disguised data set to train the machine learning model, and names the trained machine learning model as the voting model.

[0051] S4, reconstruct a new disguised dataset based on the original dataset, use the voting model to screen the new disguised dataset, use the output of the voting model as its new label for each sample in the new disguised dataset, complete the reconstruction of the new disguised dataset, train the reconstructed disguised dataset and the original dataset together, generate a new voting model, and increase the number of iterations by one.

[0052] S5, repeat step S4 until the number of iterations reaches the maximum number of iterations m.

[0053] The specific implementation steps of the defense method of this embodiment are divided into three parts, namely:

[0054] (1) Construction of the camouflage dataset

[0055] In the attribute inference attack defense method proposed in the invention, it is necessary to use an additional disguised data set whose statistical distribution is different from that of the original data set. Therefore, this step constructs a disguised data set that meets the conditions. First, if there is a data set that is similar to the original data set and is public, it can be used directly as a disguised data set. If such a data set does not exist, it is necessary to randomly construct the value of each sample feature in the full range input space of the original data set to artificially construct a batch of samples and form a data set with an average distribution as the statistical distribution to serve as a disguised data set.

[0056] The specific steps are as follows:

[0057] (11) Assume that the given original data set is ,The features of the data set are obtained through feature extraction, labels and their corresponding value ranges.

[0058] (12) Use the obtained features and labels to construct a full-range feature space of the data set. Then, randomly select n samples from the feature space to construct a new data set - the disguised data set. .

[0059] (13) If there is a public dataset that has a different statistical distribution from the original dataset but has the same feature space, we can directly extract n samples from it as the disguised dataset used for training in step (2).

[0060] (2) Model training for camouflage dataset defense strategies

[0061] First, the original data set and the disguised data set obtained in step (1) are preprocessed by data cleaning, data reduction, data transformation, etc. to eliminate the large amount of noise in the data set and obtain a standard, clean, and continuous data set to facilitate the training of the machine learning model. Secondly, the loss function of the traditional machine learning model training is changed, and the "original data set + disguised data set" is used to participate in the training to complete the model training, so as to achieve the two goals of ensuring the model utility and blurring the model features.

[0062] The specific steps are as follows:

[0063] (21) For the original data set And the disguised data set obtained in step (1) Perform preprocessing operations such as data cleaning, data reduction, and data transformation to remove a large amount of noise in the data set and obtain a data set that is conducive to model training.

[0064] (22) Initialize the machine learning model During model training, the original dataset is used and camouflage dataset Two data sets are involved in the training of the model. By minimizing the improved loss function , and use the stochastic gradient descent method to gradually complete the model learning process until the model accuracy reaches the preset threshold and then the training is terminated to obtain the machine learning model .

[0065] (23) Among them, is a constant, and 1. During training, by constantly adjusting size, so that the machine learning model can be used in the original dataset and camouflage dataset The balance between utility and feature representation is achieved, so that the model guarantees the model utility on the original data set, and the model's feature representation is more inclined to show that the model is disguising the data set. This causes the attack model to misjudge the attributes of the target model training set, rendering the attack ineffective and ensuring the privacy of the target model.

[0066] (3) Attribute inference attack defense method based on iterative voting to construct random disguised datasets

[0067] Through the training of step (2), a trained machine learning model is obtained, which is referred to as a voting model in the present invention. Then, the method of step (1) is used again to obtain a disguised data set, and the voting model is used to screen the disguised data set, that is, the output of the voting model is used as the new label for each sample in the disguised data set. After screening all samples once, a new disguised data set is obtained. We call the whole process an iteration. After n iterations, the model obtained in the last iteration is the machine learning model we need to defend against attribute inference attacks.

[0068] After the machine learning model uses the training method of step (1), the features of the model can gradually shift away from the feature representation on the original data set. However, the shift after one round of training is not enough, so the present invention further uses an iterative training defense strategy to expand its shift.

[0069] The specific steps are as follows:

[0070] (31) The defender uses the original dataset , and step (1) to obtain the disguised data set , after step (2), we get a trained machine learning model In this embodiment, the model is referred to as a voting model.

[0071] (32) Repeat step (1) to obtain a disguised data set with a sample size of n. For each sample in the disguised dataset, input it into the voting model , and get the output corresponding to the sample. If the output of the voting model is the same as the original label of the sample, it remains unchanged. If different, the output of the model is used as the new label of the sample. Repeat this operation for each sample to get a new disguised dataset Then, the disguised dataset is used as a new disguised dataset for model training.

[0072] (33) Use the original data set and the newly obtained disguised data set to go through step (2) again to obtain a new voting model, and use the voting model to screen the data again. Repeat this process n times to obtain A voting model. The voting model is used to filter the full range data set to obtain the final disguised data set. For example, a data sample is randomly selected from the training set input space of the target model. Each voting model makes a prediction for this sample, that is, votes on the sample label, and finally uses the prediction result with the highest number of votes as the label of the sample. This process is repeated to obtain a batch of samples to form a disguised data set.

[0073] In this embodiment, the machine learning model is not limited. This embodiment is not optimized for the machine learning model, but for the training process of the machine learning model, so that the security performance of the final machine learning model is higher. In theory, the defense method of this embodiment can be applied to the training process of typical machine learning models commonly used in the current industrial field, such as face recognition models, disease diagnosis models, path optimization models, small target detection models, action recognition models, etc., and the corresponding data sets can select face image sample libraries, disease diagnosis sample libraries, etc. Take the face recognition model as an example. During the training process, in addition to the original face image sample library as the original data set, several new face image sample libraries with the same feature space but different statistical distributions will be generated as disguised data sets; at the same time, several voting models are generated, and the final voting model is used as the face recognition model after training. Similarly, when the disease diagnosis sample library is used as the original data set, several new disease diagnosis sample libraries with the same feature space but different statistical distributions will be generated as disguised data sets, as well as voting models trained according to different disguised data sets, and the final voting model will be used as the disease diagnosis model. The training process of the path optimization model, the small target detection model, and the action recognition model is similar. As long as there is a complete machine learning model structure and a relatively complete original data set, the method of this embodiment can be used for training, so that the final model can not only guarantee the effectiveness on the original data set, but also make the characteristics of the model itself tend to disguise the data set, resulting in the failure of attribute reasoning attacks.

[0074] The above are only preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should be regarded as the protection scope of the present invention.

Claims

1. A defense method against attribute inference attacks in machine learning, It is characterized in that The method comprises the following steps: S1, selecting or constructing a machine learning model according to an application scenario, and generating an original data set corresponding to the machine learning model; the application scenario includes at least target recognition, task classification, probability prediction and recommendation optimization; the machine learning model includes at least a face recognition model, a disease diagnosis model, a path optimization model, a small target detection model, and an action recognition model, and the corresponding original data set includes at least a face image sample library and a disease diagnosis sample library; the process of constructing a disguised data set based on the original data set includes the following sub-steps: S11, assuming that the given original data set is in Indicates the original data set The i-th sample feature, Indicates the label corresponding to the i-th sample feature, the original data set There are N such feature-label pairs in total. The features, labels and their corresponding value ranges of the original data set are obtained through feature extraction. S12, searching whether there is a publicly available data set with the same feature space as the original data set but with a different statistical distribution, if so, extracting n samples from the found data set to generate a disguised data set, otherwise, proceeding to step S13; S13, using the obtained features and labels to construct a full-range feature space of the original data set, randomly extracting n samples from the full-range feature space to construct a disguised data set Where n is a positive integer greater than 1, Represents the camouflage dataset The jth sample feature in Indicates the label corresponding to the jth sample feature; S2, preprocessing the original data set and disguised data set to eliminate the noise in them; S3, using the preprocessed original data set and disguised data set to train the machine learning model, and the trained machine learning model is named the voting model; S4, reconstruct a new disguised dataset based on the original dataset, use the voting model to screen the new disguised dataset, use the output of the voting model as its new label for each sample in the new disguised dataset, complete the reconstruction of the new disguised dataset, train the reconstructed disguised dataset and the original dataset together, generate a new voting model, and increase the number of iterations by one; S5, repeat step S4 until the number of iterations reaches a maximum number of iterations m, where m is a positive integer greater than 1.

2. The method for defending against attribute inference attacks in machine learning according to claim 1, It is characterized in that In step S2, data cleaning, data reduction and data transformation are performed on the samples in the original data set and the disguised data set in turn, so as to complete the preprocessing of the original data set and the disguised data set.

3. The method for defending against attribute inference attacks in machine learning according to claim 1, It is characterized in that In step S3, the process of training the machine learning model using the preprocessed original data set and the disguised data set and naming the trained machine learning model as a voting model includes the following sub-steps: Initialize the machine learning model f θ : Represents the model f θ The characteristics Mapping to labels Improve the loss function of the machine learning model to obtain the improved loss function Where α is a constant, and 0<α<1, Represents the calculation for the original data set Features in Model decision label and the true label The loss between Represents the calculation for the disguised dataset Features in Model decision label and the true label Losses between By minimizing the improved loss function θ * , and use the stochastic gradient descent method to gradually complete the model learning process until the model accuracy reaches the preset threshold and then the training is terminated to obtain the machine learning model During the model training process, the original data set is used and camouflage dataset Two data sets are involved in the training of the model. During training, by continuously adjusting the size of α, the machine learning model is allowed to and camouflage dataset The balance between utility and feature representation is achieved, so that the model not only guarantees the model utility on the original data set, but the model's feature representation is more inclined to show that the model is disguising the data set. Obtained through training.

4. The method for defending against attribute inference attacks in machine learning according to claim 1, It is characterized in that In step S4, one iteration process includes the following steps: Assume that the voting model obtained in the previous iteration is Re-obtain a disguised dataset with n samples in Indicates that in the first iteration, the disguised data set The jth sample feature, Indicates that the dataset is disguised at the first iteration The jth label in For the disguised dataset Each sample in is input into the voting model In the example above, we get the output corresponding to the sample; if the voting model If the output of is the same as the original label of the sample, it remains unchanged. If it is different, the voting model is used. The output of is used as the new label of the sample; Generating a new camouflage dataset in Represents the jth sample feature Voting Model Corrected label: Using the original dataset And the newly arrived disguised dataset Participate in machine learning model training to obtain new voting models 5. The method for defending against attribute inference attacks in machine learning according to claim 1, It is characterized in that The machine learning model construction method also The following steps are involved: S6, after repeated iterations and m voting models are obtained, these m voting models are used to screen the full range data set to obtain the final disguised data set.

6. The method for defending against attribute inference attacks in machine learning according to claim 1, It is characterized in that In step S6, the process of using the m voting models to screen the full range data set to obtain the final disguised data set includes the following steps: S61, randomly select a data sample in the training set input space of the target model, and m voting models make predictions for this sample respectively. The prediction process is: vote on the sample label, and finally use the prediction result with the highest number of votes as the label of the sample; S62, repeat step S61 until a batch of samples are obtained to form a disguised data set.

7. The method for defending against attribute inference attacks in machine learning according to claim 1, It is characterized in that The machine learning models include support vector machine models, hidden Markov models, transfer learning models and federated learning models.

8. The method for defending against attribute inference attacks in machine learning according to claim 1, It is characterized in that The application scenarios include target recognition, task classification, probability prediction and recommendation optimization.

Citation Information

Patent Citations

  • Data screening method, device and equipment and readable storage medium

    CN111310819A

  • Data set authentication method and system based on machine learning member inference attack

    CN113259369A