Poisoned member inference attack method oriented to black box federal learning
By implementing poisoning attacks in the black box federated learning scenario and using the K-mean clustering unsupervised binary classifier, the problem of member reasoning attacks in the existing technology that rely on prior knowledge is solved, and the effective attack effect in the black box environment is achieved.
Patent Information
- Application Number
- CN202510115140.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing member reasoning attack methods rely on prior knowledge of federated learning models and training samples, and cannot be effectively implemented in black box federated learning scenarios, limiting the applicability of the attack.
Through the variations in poisoning attacks and predicting confidence, the attacker implements member reasoning attacks in the black box federated learning scenario, and uses K-mean clustering unsupervised binary classifiers to classify to achieve the attack effect.
Without any prior knowledge of the model and training samples, member reasoning attacks can be successfully implemented in federated learning, ensuring good attack performance.
Smart Images

Figure CN120046142A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and in particular, to a poisoning member inference attack method for black-box federated learning. Background Art
[0002] As a novel machine learning framework, federated learning highlights the protection of data privacy. In this paradigm, multiple clients can independently carry out model training on their respective local datasets without disclosing or transmitting the original private training data. The clients exchange update information about model training, such as model parameters or gradients, to jointly build a global model shared by all participants in this way. Federated learning alleviates the storage and computational burdens caused by distributed data collection. Importantly, since it avoids direct data sharing, federated learning effectively resolves the problems of data privacy and user confidentiality that often occur in traditional centralized machine learning methods. Given the strong privacy protection ability of federated learning, it has been widely applied in many practical scenarios, including keyboard prediction, smart cities, and personalized recommendation systems. These practical applications not only demonstrate the practicality of federated learning but also further confirm its excellent performance in data privacy protection. However, recent related research has exposed many security challenges faced by federated learning models, such as source inference attacks, model inversion attacks, and attribute inference attacks. These attack methods attempt to extract sensitive information contained in the training dataset, thus posing a serious threat to data privacy. To protect the data privacy of federated clients, some research work suggests setting the federated learning model as a black box using a secure execution environment to defend against privacy leakage. The model is saved in the secure execution environment for training and inference, enhancing the privacy protection ability of federated learning.
[0003] The present invention focuses on member inference attacks. The most critical goal of this attack is to infer whether a given sample (i.e., the target sample) has ever been part of the training set of a federated learning model. Once the attack is successful, it is very likely to cause extremely serious privacy risks. For example, when multiple medical institutions jointly train a federated learning model, if an attacker can successfully confirm whether a sample participated in the training, it is very likely to indirectly disclose sensitive information such as the patient's medical history and treatment records. This may not only violate personal privacy but also likely have an adverse impact on the reputation and daily business operations of medical institutions. Therefore, deeply understanding and effectively defending against member inference attacks is crucial and indispensable for ensuring the security of federated learning models and data privacy protection.
[0004] Although membership inference attacks have received extensive research attention in recent years, most of the current research work is based on the same assumption: that is, the attacker has mastered the internal information of the target model (model structure, parameters or training loss function, etc.), or has mastered the information of its training data (data samples or data distribution, etc.). However, in real scenarios, developing a federated learning model often requires huge costs and professional knowledge, which involves many links, such as data collection, data set annotation, model selection and parameter tuning. This shows that it is usually difficult for attackers to easily obtain prior knowledge of samples. At the same time, in actual applications, in order to further enhance the security of the model and the protection of privacy, the internal parameters of the federated learning model and the update of the client are strictly protected in the two stages of training and reasoning using relevant technologies such as secure executable environments. Under such strict protection measures, malicious clients can hardly peek into the internal information of the federated learning model, and they can only access the model through black box permissions. Therefore, the applicability of existing membership reasoning attacks under federated learning is largely limited, which poses a considerable challenge to membership reasoning attacks in federated learning. Summary of the invention
[0005] The purpose of the present invention is to provide a poisoned member inference attack method for black-box federated learning. This method aims to solve the problem that the existing member inference attack methods rely on prior knowledge of models and training samples in the current black-box federated learning scenario. By poisoning the target sample and quantifying the poisoning impact, the member information of a single sample implied in the model is extracted to successfully implement the attack. The present invention can achieve better attack effects without any prior knowledge of models and training samples.
[0006] The technical solution of the present invention is as follows: a poisoned member reasoning attack method for black-box federated learning, which implements an attack in a black-box federated learning scenario by using poisoning attacks and changes in prediction confidence; a malicious federated learning client launches a certain degree of poisoning attack, and infringes on the member privacy of other benign clients during the continuous training and iteration of the global model, and sets the attacker without any prior knowledge of the model and training samples, and can launch an attack only by using the black-box authority of the federated learning model.
[0007] Furthermore, the specific steps include:
[0008] Step 1: The attacker and the assistant generate poisoned samples for the target sample respectively;
[0009] Step 2: The assistant uses its own poisoned samples to participate in the federated learning model training, affecting the decision boundary of the target samples in the global model to assist the attacker in launching an attack; the attacker uses its own poisoned samples to participate in the federated learning model training and continuously queries the global model for the prediction results of the target samples;
[0010] Step 3: The attacker calculates a series of metric values affected by the poisoning attack based on the prediction results of the target samples;
[0011] Step 4: Construct an unsupervised binary classifier of K-means clustering based on the series of metric values of each target sample in Step 3 for implementing the membership inference attack.
[0012] Furthermore, Step 1 is specifically as follows: propose a poisoning strategy for target samples under federated learning, and divide the poisoning process into two parts, namely the attacker part and the assistant part;
[0013] The attacker part plays an inferring role, aiming to launch a membership inference attack against the target sample (x ii , y ii ) to pry into the membership privacy of other clients; the attacker modifies the label of the target sample as the poisoned sample, that is, (x ii , y pp ), where the poisoned label y pp ≠y ii , forming a poisoned dataset;
[0014] The assistant part plays a catalytic role, aiming to offset the weakening caused by federated learning aggregation as much as possible; the poisoned sample of the assistant is (x * , y pp ), where Adam{·} represents the Adam optimizer, and F θ (x) represents the prediction result of the global model θ for the sample, and x is initialized as a randomly generated noise sample.
[0015] Furthermore, Step 2 is specifically as follows: after the assistant and the attacker respectively prepare their own poisoned datasets, they normally participate in the federated training and aggregation processes; the assistant is used to alleviate the problem of poor poisoning attack effect caused by federated learning model aggregation; under the synergistic effect of the assistant and the attacker, a poisoning attack that meets the metric requirements is carried out; at the same time, the attacker continuously queries the global model for the prediction results of the target samples and records them.
[0016] Furthermore, Step 3 is specifically as follows: after the attacker obtains the prediction results of the target samples in each round, use the following metric value calculation method Mentr PPPP =-(1 - Fθ (x) cc ) log(F θ (x) cc ) - F θ (x) pp log(1 - F θ (x) pp ) Measure the impact of poisoning, where F θ (x) cc represents the predicted value of the global model output on the true label, and F θ (x) pp represents the predicted value of the global model output on the poisoned label; Arrange and combine this series of measurement values in the order of the number of rounds of federated learning iterations and save them.
[0017] Furthermore, the specific content of step 4 is to construct a binary classifier according to the series of measurement values obtained in step 3. A series of measurement values of a target sample form a vector as follows:
[0018]
[0019] where s ii represents the trajectory of the measurement value of the i-th target sample, and n is the number of iterations of the global model; Modify the formula for calculating the distance of the K-means clustering algorithm to better leak member information. The modified distance formula is as follows: where r kk is the k-th cluster center, s iiii and r kkii are the values of s ii and r kk in the j-th dimension respectively; deriv(·) is a formula for calculating the slope of a vector, and the slope reflects the change of Mentr PPPP during the process of federated learning iteration, improving the classification ability of the K-means clustering algorithm; The vector slope is calculated using the central difference method: where h is the offset distance forward and backward for calculating the slope of the current dimension; The K-means clustering algorithm classifies s PPPP with a larger Mentr ii as non-member sample cases and smaller ones as member sample cases, realizing a membership inference attack.
[0020] Advantages of the present invention: The present invention uses poisoning against the target samples to cause the global model to misjudge the target samples to a certain extent. Then, during the training iteration process of federated learning, the prediction results of the target samples are continuously queried. Next, the measurement method designed by the present invention is used to quantify the membership privacy of the target samples of other benign participants to the greatest extent. Finally, the unsupervised binary classifier K-means clustering algorithm is used to classify these quantified values, thereby realizing a membership inference attack without prior knowledge of the model and training samples and ensuring good attack performance. Description of the Drawings
[0021] Figure 1 It is the attack flowchart during the federated learning training stage of the present invention.
[0022] Figure 2 It is the flowchart of the method of the present invention.
[0023] Figure 3 It is the binary classification flowchart using the K-means clustering unsupervised model of the present invention.
[0024] Figure 4 It is the visualization diagram of the change difference of the membership inference attack measurement value of the present invention; (a) is a series of measurement values of the member samples along with the federated learning iteration process; (b) is a series of measurement values of the non-member samples along with the federated learning iteration process.
[0025] Figure 5 It is the visualization diagram of some samples of the poisoned samples of the helper of the present invention; (a) is the schematic diagram of the target sample; (b) is the schematic diagram of the poisoned sample generated for the target sample. Detailed Embodiment
[0026] In order to eliminate the attacker's dependence on the prior knowledge of the model and training samples and at the same time implement a membership inference attack with better performance, the present invention proposes a poisoning membership inference attack method for black-box federated learning, which violates the membership privacy of the federated learning model without using the prior knowledge of the model and training samples. The present invention will be further described below with reference to the drawings and embodiments.
[0027] The present invention provides a poisoning membership inference attack method for black-box federated learning, including the following steps:
[0028] Step 1: The attacker and the helper respectively generate poisoned samples for the target samples;
[0029] Step 2: The attacker and the helper use the poisoned samples and the federated learning global model for training and normally participate in the aggregation of the federated learning model. The attacker continuously queries the prediction results of the target samples from the global model;
[0030] Step 3: The attacker calculates a series of metric values affected by the poisoning attack based on the prediction results of the target samples;
[0031] Step 4: Construct an unsupervised binary classifier of K-means clustering based on the series of metric values of each target sample in Step 3.
[0032] The more specific steps are as follows:
[0033] Specifically, Step 1 is that the present invention proposes a poisoning strategy for target samples under federated learning, and a poisoning mechanism that interferes with the decision boundary to counteract the influence of federated learning aggregation. We divide the poisoning process into two parts, namely the attacker part and the helper part. The attacker part plays a role in inference, aiming to launch a membership inference attack against the target sample (x ii ,y ii ) to pry into the membership privacy of other clients. The attacker modifies the label of the target sample as the poisoned sample, that is, (x ii ,y pp ), where the poisoned label y pp ≠y ii , forming a poisoned dataset; the helper part plays a catalytic role, aiming to counteract as much as possible the weakening brought by federated learning aggregation. The poisoned sample of the helper is (x * ,y pp ), where is optimized by the Adam optimizer, and F θ (x) represents the prediction confidence vector of the model for the sample, and x is initialized as a randomly generated noise sample;
[0034] Specifically, Step 2 is that after the helper and the attacker respectively prepare their poisoned datasets, they normally participate in the federated training and aggregation process. It should be particularly noted that the helper makes up for the deficiency of the weakening of the poisoning attack effect brought by some other benign clients. Under the synergistic effect of the helper and the attacker, a poisoning attack that meets the metric requirements is achieved. At the same time, the attacker continuously queries the prediction results of the target samples from the global model and records them;
[0035] Specifically, Step 3 is that after the attacker obtains the prediction results of the target samples in each round, using the metric method designed by the present invention
[0036] Mentr PPPP =-(1 - F θ (x) cc )log(F θ (x) cc ) - F θ (x) pp log(1 - F θ (x)pp )
[0037] Measure the impact of poisoning, where F θ (x) cc represents the predicted value of the model output on the true label, and F θ (x) pp represents the predicted value on the poisoned label. Arrange and save this series of measurement values in the order of the number of rounds of federated learning iterations;
[0038] Specifically, in step 4, use the series of measurement values obtained in step 3 to construct a binary classifier. A series of measurement values of a target sample form a vector as follows:
[0039]
[0040] where s ii represents the trajectory of the measurement value of the i-th target sample, and n is the number of iterations of the global model. In order to make the model classification accuracy higher, the present invention modifies the formula for calculating the distance of the K-means clustering algorithm to better leak member information. The modified distance formula is as follows: dist(s iiii , r kkii ) = (s iiii - r kkii ) 2 + (deriv(s iiii ) - deriv(r kkii )) 2 , where r kk is the k-th cluster center, and s iiii and r kkii are the values of s ii and r kk on the j-th dimension respectively. deriv(·) is a formula for calculating the slope of a vector. The present invention uses the central difference method to estimate the derivative of the vector, such as: where h is the offset distance forward and backward for calculating the slope of the current dimension. The present invention uses as the basis for clustering by the K-means clustering algorithm to implement a membership inference attack.
[0041] See Figure 1 , the present invention is a membership inference attack initiated by a federated learning client. The target of the attack is to violate the member privacy of benign clients. Among them, malicious clients include two roles, namely, an assistant and an attacker. The main function of the assistant is to assist the attacker in further attacking the target sample, and the attacker is the main body that initiates the membership inference attack.
[0042] See Figure 2, A poisoning member inference attack method for black-box federated learning, which implements a member inference attack by quantifying the impact of target samples on poisoning attacks during the federated learning training process; First, malicious participants generate a part of poisoned samples for target samples and construct a poisoned dataset; Second, the malicious client uses the poisoned samples to train the global model and normally participates in the federated learning aggregation process; Then, the attacker continuously queries the prediction results of the target samples and saves them; Finally, the attacker uses a series of prediction results to calculate the measure of the poisoning impact and constructs a K-means clustering unsupervised binary classifier. More detailed steps include:
[0043] Step 1, the attacker and the assistant respectively generate poisoned samples for the target samples;
[0044] Step 2, the attacker and the assistant use the poisoned samples and the federated learning global model for training and normally participate in the federated learning model aggregation. The attacker continuously queries the prediction results of the target samples from the global model;
[0045] Step 3, the attacker uses the prediction results of the target samples to calculate a series of metric values of the impact of the poisoning attack;
[0046] Step 4, use a series of metric values of each target sample in Step 3 to construct a K-means clustering unsupervised binary classifier.
[0047]
[0048]
[0049] Specifically, Step 3 is as shown in Algorithm 1. In Figure 4 , the inventor demonstrated the change differences of member samples and non-member samples in Mentr PPPP during the federated learning training process. It can be clearly found that the Mentr PPPP of member samples tends to be smaller and does not increase significantly, but for non-member samples, Mentr PPPP gradually becomes larger as the model training progresses.
[0050] The existing mainstream datasets CIFAR-100, CIFAR-10, MNIST, Purchase-100, and Fashion-MNIST are used to conduct attacks on the network models CNN and MLP to test the performance of membership inference attacks. CNN and MLP are used as the global models of federated learning, and the attack effects are tested using five mainstream datasets CIFAR-100, CIFAR-10, MNIST, Purchase-100, and Fashion-MNIST. Table 1 shows the four performance metrics of the poisoning membership inference attack proposed in the present invention on different datasets and models. The poisoning membership inference attack method in Table 1 shows the results when the malicious client is the attacker. It can be observed that this method has achieved significant attack effects on multiple models and datasets. The poisoning membership inference attack method for black-box federated learning can achieve good results in attack accuracy, recall rate, precision rate, and F1 score without any prior knowledge of the model and training samples, which indicates that the present invention is effective. Table 2 shows the four performance metrics of the membership inference attack White-Box proposed in "Comprehensive privacy analysis of deep learning: Passive and active white-box inference attacks against centralized and federated learning" on different datasets and models. Under the same experimental settings, even though Nasr et al.'s membership inference attack has prior knowledge of the internal parameters of the model, their attack performance is weaker than that of the present invention, which further proves the excellent attack ability of the present invention. In addition, Table 3 shows the influence of the poisoning strategy of the present invention on the test accuracy of different datasets and models. Due to the aggregation property of federated learning, poisoning has little impact on the accuracy of the model. Therefore, even when the dataset is partially poisoned, a certain degree of model test accuracy can be guaranteed, but there is a risk of leaking member privacy. These results demonstrate the effectiveness of the attack method of the present invention while minimizing the impact on model performance as much as possible.
[0051] Table 1
[0052] Dataset + Model Accuracy Precision Recall F1 Score CIFAR-10 + CNN 0.851 0.818 0.989 0.872 CIFAR-100 + CNN 0.854 0.709 0.975 0.823 MNIST + CNN 0.763 0.731 0.761 0.745 Fashion-MNIST + CNN 0.752 0.725 0.890 0.799 Purchase-100 + MLP 0.731 0.672 0.807 0.733
[0053] Table 2
[0054]
[0055]
[0056] Table 3
[0057]
[0058] In summary, the present invention is a poisoning member inference attack method for black-box federated learning. The poisoning member inference attack is implemented in the following manner: quantifying the impact of target samples on poisoning attacks during the federated learning training process. First, a malicious participant generates a certain number of poisoned samples for the target samples, and then constructs a poisoned dataset. Secondly, the malicious client uses the poisoned samples to train the global model and normally participates in the aggregation process of federated learning. Then, the malicious client continuously queries and saves the prediction results of the target samples. Finally, the malicious client calculates the poisoning impact using a series of prediction results and constructs a K-means clustering unsupervised binary classifier to achieve a better attack effect.
[0059] The above content is only a preferred embodiment of the present invention and does not impose any formal restrictions on the present invention. As long as the content does not deviate from the technical solution of the present invention, any simple modification, equivalent change, and modification made to the above embodiments in accordance with the technical essence of the present invention still fall within the scope covered by the technical solution of the present invention.
Claims
1. A poisoned member reasoning attack method for black-box federated learning, characterized in that: The attack is carried out in the black-box federated learning scenario by using poisoning attacks and differences in changes in prediction confidence. The malicious federated learning client launches a certain degree of poisoning attack and violates the privacy of members of other benign clients during the continuous training and iteration of the global model. The attacker does not need to have any prior knowledge of the model and training samples, and can launch an attack by only using the black-box permissions of the federated learning model.
2. The poisoned member reasoning attack method for black-box federated learning according to claim 1 is characterized in that: The specific steps are as follows: Step 1: The attacker and the assistant generate poisoned samples for the target sample respectively; Step 2: The assistant uses his own poisoned samples to participate in the training of the federated learning model, affecting the decision boundary of the target sample in the global model and assisting the attacker in launching the attack; the attacker uses his own poisoned samples to participate in the training of the federated learning model and continuously queries the global model for the prediction results of the target sample; Step 3: The attacker calculates a series of metrics of the impact of the poisoning attack based on the prediction results of the target sample; Step 4: Construct a K-means clustering unsupervised binary classifier based on a series of metric values of each target sample in step 3 to implement membership inference attack.
3. The poisoned member reasoning attack method for black-box federated learning according to claim 2 is characterized in that: Specifically, step 1 proposes a poisoning strategy for target samples under federated learning, and divides the poisoning process into two parts, namely the attacker part and the assistant part; The attacker part plays the role of reasoning, and the purpose is to launch an attack on the target sample (x i ,y i ) member reasoning attack to spy on the privacy of members of other clients; the attacker modifies the label of the target sample as a poisoned sample, that is, (x i ,y p ), where the poisoned label y p ≠y i , which constitutes the poisoning dataset; The assistant part plays a catalytic role in order to offset the weakening brought by federated learning aggregation as much as possible; the poisoned sample of the assistant is (x * ,y p ),in Adam{·} represents the Adam optimizer, F θ (x) represents the prediction result of the global model θ for the sample, and x is initialized to a randomly generated noise sample.
4. The poisoned member reasoning attack method for black-box federated learning according to claim 2 is characterized in that: Specifically, step 2 includes: after the assistant and the attacker prepare their own poisoned data sets, they participate in the federated training and aggregation process normally; the assistant is used to alleviate the problem of poor poisoning attack effect caused by federated learning model aggregation; under the synergy of the assistant and the attacker, a poisoning attack that meets the measurement requirements is carried out; at the same time, the attacker continuously queries the global model for the prediction results of the target sample and records them.
5. The poisoned member reasoning attack method for black-box federated learning according to claim 2 is characterized in that: Specifically, step 3 is as follows: after the attacker obtains the prediction results of each round of the target sample, he uses the following metric calculation method Mentr PE =-(1-F θ (x) c )log(F θ (x) c )-F θ (x) p log(1-F θ (x) p ) measures the poisoning effect, where F θ (x) c represents the predicted value of the global model output on the true label, F θ (x) p Represents the predicted value of the global model output on the poisoned label; this series of measurement values are arranged and saved in the order of federated learning iteration rounds.
6. The poisoned member reasoning attack method for black-box federated learning according to claim 5 is characterized in that: Specifically, step 4 is to construct a binary classifier according to a series of metric values obtained in step 3. A series of metric values of a target sample form a vector as follows: Among them, s i represents the trajectory of the metric value of the i-th target sample, and n is the number of iterations of the global model; the formula for calculating the distance of the K-means clustering algorithm is modified to better leak member information. The modified distance formula is as follows: Among them, r k is the kth cluster center, s ij and r kj They are i and r k The value in the jth dimension; deriv(·) is the formula used to calculate the slope of the vector, which reflects Mentr PE As the federated learning iterative process changes, the classification ability of the K-means clustering algorithm is improved; the vector slope is calculated using the central difference method: Where h is used to calculate the forward and backward offset distance of the slope of the current dimension; the K-means clustering algorithm converts Mentr PE Larger i Classified as a non-member sample, the smaller s i Classify into member sample cases and implement member reasoning attack.