Machine learning poisoning attack defense method and system based on GAN, and equipment medium

Through the GAN-based machine learning poison attack defense method, the global model is used to generate virtual test data to identify and eliminate malicious model updates, solving the problems of degraded defense effects and privacy leakage in the existing technology, and achieving efficient poison attack defense.

CN120582833APending Publication Date: 2025-09-02XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510674231.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing machine learning poisoning attack defense methods decrease when the number of malicious parties increases or data heterogeneity increases, and there is a risk of high computing overhead and privacy leakage, making it difficult to effectively identify malicious behavior.

Method used

Using the GAN-based machine learning poisoning attack defense method, by replacing the model parameters of the discriminator with global model parameters, the generator generates virtual test data, comparing the differences in output results, identifying and eliminating the model updates of malicious participants, and combining the federated averaging algorithm for global model aggregation.

Benefits of technology

It improves the security and privacy of the system, enhances robustness and expansion, and can maintain good defense effects under high proportion of malicious participants and slight data heterogeneity, reducing computing overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120582833A_ABST
    Figure CN120582833A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of information security, and particularly relates to a machine learning poisoning attack defense method and system based on a GAN, and an equipment medium. According to the method, firstly, a participant downloads a global model from an aggregation server and then carries out training, a last local model is updated, then the aggregation server replaces model parameters of an internal discriminator with model parameters of the global model, adversarial training is carried out through the discriminator and a generator, and the generator with the updated parameters is obtained; the aggregation server processes the virtual test data generated by the generator after the parameters are updated, malicious local models are detected to be updated, and then the malicious local models are rejected to form a benign model set; the system comprises a local model training and uploading module, an antagonism training module, an anomaly detection module and a global model aggregation module. According to the method, the security and privacy of the system can be improved under the condition of a large number of malicious participants or slight data heterogeneity, and the robustness and expansibility of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information security technology, and specifically relates to a GAN-based machine learning poisoning attack defense method, system, and device medium. Background Art

[0002] Machine learning is a distributed machine learning technology that allows multiple devices or servers to share models for collaborative training while ensuring data privacy. In this model, only model updates, not raw data, are shared, effectively mitigating the risk of data leakage. Machine learning is widely used in smartphones, IoT devices, and other fields, improving model generalization and privacy. However, due to the local nature of data storage, machine learning systems are vulnerable to poisoning attacks. Attackers can manipulate participating devices to upload malicious updates or inject malicious data, disrupting the global model training process, resulting in performance degradation or the insertion of backdoors. Poisoning attacks can be categorized into data poisoning and model poisoning based on their attack methods. Data poisoning attacks disrupt model training by inserting malicious data into local data, while model poisoning attacks directly tamper with locally trained model updates, affecting the global model through abnormal updates. Because each device only shares local model updates, traditional security mechanisms struggle to effectively detect these malicious activities, making attacks more stealthy. Therefore, preventing malicious actors from participating in machine learning and improving machine learning's defenses against poisoning attacks are critical issues that need to be addressed.

[0003] Existing defense methods for machine learning poisoning attacks, whether robust aggregation algorithms or anomaly detection methods based on the statistical characteristics of participant model parameters, become less effective as the number of malicious participants increases or data heterogeneity intensifies. They also require certain computational overhead or access to local information, increasing deployment complexity and the risk of privacy leakage. Under different aggregation strategies, the stability and effectiveness of the defense mechanism often show unstable performance.

[0004] In the prior art, the invention with patent publication number CN118246009A and titled “A Method for Defending Against Poisoning Attacks in Federated Learning” discloses a gradient detection scheme based on the local outlier factor (LOF). The local outlier factor (LOF) algorithm is used to detect gradient anomalies and remove outliers to achieve defense against poisoning attacks. This method maintains model accuracy at a malicious client ratio of 10%-30%, and has the ability to identify isolated malicious gradients. However, it has the following problems: it relies on the assumption that malicious gradients are sparsely distributed in space, and cannot cope with dense gradients forged by a large number of coordinated attacks; at the same time, it needs to access the original gradient, which may leak local data features.

[0005] In the prior art, the invention with the patent publication number CN116155611A and the name of “A method for defending against poisoning attacks in a federated learning system based on RDP” discloses an anomaly detection scheme based on random distance prediction (RDP). The server detects anomalies by analyzing the output of the client's fully connected layer and constructing an RDP network composed of a twin neural network and a mapping neural network to achieve poisoning defense. This method detects poisoning attacks through the output of the fully connected layer and has the ability to identify non-cooperative attacks. However, it has the following problems: RDP is sensitive to the distribution of the output of the fully connected layer, and has a high false positive rate under Non-IID data; at the same time, it is necessary to train dual neural networks (SiameseNet and MappingNet), and the time complexity is O(n 2 ); The output of the fully connected layer may expose the characteristics of the client's local data. Summary of the Invention

[0006] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to propose a GAN-based machine learning poisoning attack defense method, system, and device medium. This method trains the generator by replacing the model parameters of the discriminator with a global model, and eliminates malicious participants by comparing the differences in the output results of virtual test data generated by the generator before and after the model update. This solves the problem that the effectiveness of existing federated learning poisoning attack defense methods decreases when the number of malicious participants increases or the heterogeneity of data intensifies, and the problem that existing defense methods require a certain amount of computing overhead or access to local information, which increases deployment complexity and privacy leakage.

[0007] To achieve the above object, the technical solutions adopted by the present invention are as follows:

[0008] First, a GAN-based machine learning poisoning attack defense method includes the following steps:

[0009] S1. The participant downloads the i-th (i=1, 2, ..., n) round global model from the aggregation server and uses the gradient descent algorithm to train the local model on the participant's local data and obtain the local model update;

[0010] S2. The participant uploads the updated local model described in step S1 to the aggregation server;

[0011] S3. After the local model is updated and uploaded in step S2, the aggregation server replaces the model parameters of the internal discriminator with the model parameters of the i-th (i=1, 2, ..., n)-th round global model, and the discriminator after the model parameter replacement is subjected to adversarial training with the generator to obtain a generator with updated parameters;

[0012] S4. The aggregation server updates the global model in the i-th (i=1, 2, ..., n) round and the local model uploaded to the aggregation server in step S2, and processes the virtual test data generated by the generator after the parameter update in step S3, respectively, to obtain the output results before the update and the output results after the update, and performs anomaly detection on the difference between the output results before the update and the output results after the update to confirm malicious local model updates in the local model updates, and removes the malicious local model updates from the local model updates in this step to form a benign model set;

[0013] S5. The aggregation server uses the federated averaging algorithm (FedAvg) to perform global model aggregation on the benign model set described in step S4, completes the i-th (i=1, 2, ..., n)-th round of global model training, and generates the i+1-th (i=1, 2, ..., n)-th round of global model training;

[0014] S6. Repeat steps S1 to S5 until the i+1th (i=1, 2, ..., n)th round of global model converges in step S5.

[0015] Furthermore, the participants in step S1 include normal participants and malicious participants. Normal participants perform benign model training and obtain benign local model updates. Normal participants all follow the same training rules set by the system.

[0016] Furthermore, after the malicious party described in step S1 downloads the global model of round i (i=1, 2, ..., n) from the aggregation server, the malicious party performs malicious model training. The malicious model training can be performed through collusion or individually, where the independent malicious model training includes modifying the data set, modifying the training algorithm, or modifying the uploaded model update, and ultimately generating a malicious local model update.

[0017] Furthermore, the loss function of the discriminator in step S3 is expressed as follows:

[0018] L Dis =-logD(x)-log(1-D(G(noise)))

[0019] Among them, L Dis represents the loss of the discriminator, G represents the generator, noise represents random noise, and D represents the discriminator after model parameter replacement.

[0020] Furthermore, the loss function of the generator in step S3 is expressed as follows:

[0021] L Gen = -logD(G(noise))

[0022] Among them, L Genrepresents the loss of the generator, G represents the generator, noise represents random noise, and D represents the discriminator after model parameter replacement.

[0023] Furthermore, the number of rounds of adversarial training in step S3 is not less than 3 times.

[0024] Furthermore, the anomaly detection in step S4 determines that a malicious local model update is being updated after determining that the difference exceeds a set threshold.

[0025] Furthermore, the aggregated global model obtained by the federated averaging algorithm in step S5 is expressed as follows: n

[0026]

[0027] Among them, M G is the global model, n is the set of benign participant models I begin The number of, α is the corresponding weight, m i Represents a collection of participant models.

[0028] In the second aspect, a GAN-based machine learning poisoning attack defense system is provided, which applies the poisoning attack defense method. The defense system includes a local model training and uploading module, an adversarial training module, an anomaly detection module, and a global model aggregation module:

[0029] Local model training and upload module: The participant downloads the i-th (i=1, 2, ..., n) round global model from the aggregation server, and uses the gradient descent algorithm to perform local model training on the participant's local data to obtain a local model update, and then uploads the local model update to the aggregation server;

[0030] Adversarial training module: After the local model is updated and uploaded, the aggregation server replaces the model parameters of the internal discriminator with the model parameters of the global model in round i (i = 1, 2, ..., n). The discriminator after model parameter replacement performs adversarial training with the generator to obtain the generator with updated parameters;

[0031] Anomaly detection module: The aggregation server updates the global model in round i (i=1, 2, ..., n) and the local model uploaded to the aggregation server, and processes the virtual test data generated by the generator after the updated parameters, respectively, to obtain the output results before the update and the output results after the update, and performs anomaly detection on the differences between the output results before the update and the output results after the update to confirm malicious local model updates in the local model updates, and eliminates the malicious local model updates from the local model updates in this step to form a benign model set;

[0032] Global model aggregation module: The aggregation server uses the federated averaging algorithm (FedAvg) to perform global model aggregation on the benign model set, completes the i-th (i=1, 2,…, n) round of global model training, and generates the i+1-th (i=1, 2,…, n) round of global model training. The four module functions are repeated until the i+1-th (i=1, 2,…, n) round of global model converges.

[0033] In a third aspect, an electronic device includes a memory and a processor:

[0034] Memory: used for storing a computer program for implementing the poisoning attack defense method;

[0035] Processor: used to implement the poisoning attack defense method when executing the computer program.

[0036] In a fourth aspect, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the poisoning attack defense method is implemented.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] 1. This method replaces the discriminator's model parameters with global model parameters in step S3, and further determines whether a malicious model exists by generating virtual test data based on the global model, rather than relying on model information between participants. This improves the security and privacy of the system, avoids the leakage of participants' model information, and has low computational overhead.

[0039] 2. Through the processing of steps S1 to S4, this method can maintain a good defense effect under conditions of a high proportion of malicious participants and slight data heterogeneity, thereby increasing the robustness of the system. It can also be combined with different aggregation robustness aggregation algorithms to improve the scalability of the system. Simulation experiments have verified the effectiveness of this method in defending against poisoning attacks in different scenarios.

[0040] In summary, the present invention replaces the model parameters of the discriminator with the global model parameters, and further generates virtual test data based on the global model to determine whether there is a malicious model. This can ensure that the security and privacy of the system are improved in the case of a large number of malicious participants or slight data heterogeneity, and increase the robustness and scalability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is the architecture diagram of the GAN-based machine learning poisoning attack defense algorithm proposed in this invention.

[0042] Figure 2This is a flow chart of the GAN-based machine learning poisoning attack defense method proposed in this invention.

[0043] Figures 3a to 3d Schematic diagrams of the defense effects of this method against BBA, LIE, LFA, and DBA in scenario one.

[0044] Figures 4a to 4d Schematic diagrams of the defense effects of this method against BBA, LIE, LFA, and DBA in scenario 2.

[0045] Figures 5a to 5d Schematic diagrams of the defense effects of this method against BBA, LIE, LFA, and DBA in scenario three.

[0046] Figures 6a to 6d Schematic diagrams of the defense effects of this method against BBA, LIE, LFA, and DBA in scenario four.

[0047] Figures 7a to 7d Schematic diagrams of the Krum defense method's defense effectiveness against BBA, LIE, LFA, and DBA in scenarios one and two, respectively.

[0048] Figures 8a to 8d Schematic diagrams of the defense effects of the foolsgold defense method against BBA, LIE, LFA, and DBA in scenarios one and two, respectively. DETAILED DESCRIPTION

[0049] The following is combined with Figure 1 To the attached Figure 8d The present invention is described in further detail:

[0050] like Figure 1 and Figure 2 As shown in the figure, the machine learning system consists of an aggregation server and multiple participants, where the aggregation server is trustworthy. The tasks of the aggregation server and participants in the machine learning training process in the architecture diagram are:

[0051] Aggregation Server: The aggregation server's primary responsibility is to issue learning tasks and coordinate the training and update aggregation of local models among all participants. Throughout the training process, the aggregation server not only performs routine model aggregation tasks but also employs the defense algorithm described in this invention to screen received local model updates. Specifically, the aggregation server ensures the security and stability of the global model by detecting and rejecting malicious model updates, preventing malicious attackers from interfering with the global model.

[0052] Participants: Participants are divided into benign and malicious participants. Benign participants perform local model training according to system requirements and generate local model updates based on their training results. Model updates by benign participants can generally improve the performance of the global model and promote the normal development of the system. Malicious participants, through collusion or individual action, deliberately conduct malicious model training with the aim of undermining the learning objectives of the global model or reducing the overall performance of the global model. Malicious participants may affect model performance by manipulating training data or maliciously adjusting model parameters, or even implanting backdoors in the model, causing the model to produce erroneous outputs for specific categories of samples, thereby affecting the stability and reliability of the entire system. Whether benign or malicious, after completing local model training, the local model update will be submitted to the aggregation server.

[0053] The generative adversarial network algorithm is a deep learning architecture consisting of two neural networks, a generator and a discriminator, which compete with each other through adversarial training, ultimately enabling the generator to generate fake data that is very close to real data. The purpose of the generator is to generate data that is as real as possible from random noise. The task of the generator is to generate a fake sample from a random noise vector (usually high-dimensional). The generator continuously adjusts the parameters by learning the distribution of training data so that the samples it generates are closer and closer to the distribution of real data. The task of the discriminator is to determine whether the input sample is real data or generated data. It outputs a probability value, which indicates the probability that the sample belongs to real data. The goal of the generator is to make the discriminator think that the sample it generates is real as much as possible. That is, to maximize the probability of the discriminator's wrong judgment. Its loss function is expressed as follows:

[0054]

[0055] where z is the value from the noise distribution p z (z), G(z) is the fake sample generated by the generator, and D(G(z)) is the judgment probability of the discriminator on the generated sample.

[0056] The goal of the discriminator is to distinguish between real data and generated data as accurately as possible. Its goal is to maximize the probability of judging real samples and the probability of judging generated samples as false. Its loss function is expressed as follows:

[0057]

[0058] Among them, p data (x) is the true data distribution, D(x) is the judgment probability of the discriminator on the true sample, and D(G(z)) is the judgment probability of the discriminator on the generated sample.

[0059] By alternately training the generator and discriminator, the game eventually reaches a Nash equilibrium. Ideally, the generator generates highly realistic data, while the discriminator is unable to distinguish between real and fake data. At this point, the generator's generated data is indistinguishable from real data, and the training process has reached convergence.

[0060] The anomaly detection solution proposed in this paper makes the following assumptions about the machine learning system:

[0061] (1) The aggregation server is considered trustworthy and has a relatively small normal dataset for training the generator model. At the same time, the aggregation server does not directly review the model parameters of the participants.

[0062] (2) Malicious actors can train malicious models to launch attacks aimed at reducing the accuracy of the global model or implanting backdoors in the global model. Such attacks can be carried out by malicious actors alone or through collaborative efforts.

[0063] The symbols and their meanings used in the scheme proposed in the present invention are shown in Table 1.

[0064] Table 1 Symbols

[0065]

[0066]

[0067] The system framework involved in this method consists of an aggregation server and n participants. The set of honest participants is U = {u1,u2,u3,...,u n}, there are m malicious participants in the system, then the set of malicious participants is The local dataset is d i , the amount of data in each local data set is the same, and the local models of the participants are consistent with the training algorithm. The aggregation server has a normal data set d server .

[0068] The specific defense scheme of the present invention is as follows:

[0069] Based on the different tasks of each member, the system architecture can be divided into an aggregation server and participants, each of which has specific responsibilities and functions. Before training begins, the aggregation server is responsible for configuring model parameters and algorithm flow, and selecting several participants for training. Participants download the global model from the aggregation server. After completing local training, they upload the updated model parameters to the aggregation server. The aggregation server performs anomaly detection on the received model, confirms its validity, and then aggregates the global model.

[0070] First, a GAN-based machine learning poisoning attack defense method includes the following steps:

[0071] S1. Local model training: The participant downloads the i-th (i=1, 2, ..., n)-th round global model from the aggregation server and uses the gradient descent algorithm to perform local model training on the participant's local data and obtain a local model update; the gradient descent algorithm can be a stochastic gradient descent algorithm or an adam algorithm. The participants include normal participants and malicious participants. Participant selection criteria: The aggregation server will first select a number of qualified participants from all participants to participate in this round of training. The selection criteria can be random selection or selection based on conditions such as the participant's computing power and communication capabilities;

[0072] Specifically, in step S1, the normal participants conduct benign model training and obtain benign local model updates. The normal participants all follow the same training rules set by the system, such as learning rate, number of iterations, model structure, etc., to ensure the effectiveness of the training;

[0073] Specifically, in step S1, after the malicious participant downloads the global model of round i (i=1, 2, ..., n) from the aggregation server, the malicious participant conducts malicious model training. The malicious model training can be conducted through collusion or individually. Independent malicious model training includes modifying the dataset, modifying the training algorithm, and modifying the uploaded model update, ultimately generating a malicious local model update.

[0074] S2. Local model upload: The participants upload the local model updates described in step S1 to the aggregation server. The uploaded local model updates should include the weight changes or parameter adjustments calculated during the training process. After the local model training described in step S1 is completed, the normal participants and the malicious participants will each generate a benign local model update and a malicious local model update and upload them to the aggregation server.

[0075] S3. Adversarial training: After the local model update in step S2 is uploaded, the aggregation server replaces the model parameters of the internal discriminator with the model parameters of the global model of round i (i=1, 2, ..., n). The discriminator after the model parameter replacement is subjected to adversarial training with the generator. Specifically, adversarial training is to input the real data in the aggregation server and the virtual training data generated by the generator into the discriminator, and calculate the corresponding cross entropy loss to obtain the generator with updated parameters.

[0076] When the model accuracy of the global model in round i+1 (i=1, 2, ..., n) is ≥ 20%, after the fifth round of global model training, the discriminator in the aggregation server in step S3 may choose not to replace the model parameters. In this case, the virtual test data can distinguish between benign local model updates and malicious local model updates.

[0077] Specifically, the loss function of the discriminator in step S3 is expressed by formula (3):

[0078] L Dis =-logD(x)-log(1-D(G(noise))) (3)

[0079] Among them, L Dis represents the loss of the discriminator, G represents the generator, noise represents random noise, and D represents the discriminator after model parameter replacement;

[0080] Specifically, the loss function of the generator in step S3 is expressed as follows:

[0081] L Gen = -logD(G(noise)) (4)

[0082] Among them, L Gen Represents the loss of the generator, G represents the generator, noise represents random noise, and D represents the discriminator after the model parameters are replaced;

[0083] Specifically, the number of rounds of adversarial training in step S3 is not less than 3;

[0084] S4, anomaly detection: The aggregation server processes the virtual test data generated by the generator after the parameter update in step S3 through the i-th (i=1, 2, ..., n)-th round global model and the local model update uploaded to the aggregation server in step S2 (including benign local model updates and malicious local model updates), obtains the output results before the update and the output results after the update, and performs anomaly detection on the difference between the output results before the update and the output results after the update to confirm the malicious local model update in the local model update, and eliminates the malicious local model update from the local model update in step S4 to form a benign model set to ensure the security of the system;

[0085] Specifically, the judgment criteria in anomaly detection are: when the difference exceeds the set threshold, the model will be identified as a malicious model;

[0086] S5. Global model aggregation: The aggregation server uses the federated averaging algorithm (FedAvg) to perform global model aggregation on the set of benign models described in step S4, completes the i-th (i=1, 2, ..., n) round of global model training, and generates the i+1-th (i=1, 2, ..., n) round of global model training. The federated averaging algorithm generates a new round of global models by weighted averaging the local model update parameters uploaded by the participants. The weight of each participant is usually determined based on the amount of local data or computing power to ensure that the new round of global models can maintain a balance between different data distributions.

[0087] Specifically, the aggregated global model obtained by the federated averaging algorithm in step S5 is expressed by formula (5):

[0088]

[0089] Among them, M G is the global model, n is the set of benign participant models I begin The number of, α is the corresponding weight, m i Represents a collection of participant models;

[0090] S6. Repeat steps S1 to S5 until the i+1th (i=1, 2, ..., n)th round of global model converges in step S5, or the i+1th (i=1, 2, ..., n)th round of global model reaches the set number of iterations p, where p is between 50 and 150.

[0091] In the second aspect, a GAN-based machine learning poisoning attack defense system is provided, which applies the poisoning attack defense method. The defense system includes a local model training and uploading module, an adversarial training module, an anomaly detection module, and a global model aggregation module:

[0092] Local model training and upload module: The participant downloads the i-th (i=1, 2, ..., n) round global model from the aggregation server, and uses the gradient descent algorithm to perform local model training on the participant's local data to obtain a local model update, and then uploads the local model update to the aggregation server;

[0093] Adversarial training module: After the local model is updated and uploaded, the aggregation server replaces the model parameters of the internal discriminator with the model parameters of the global model in round i (i = 1, 2, ..., n). The discriminator after model parameter replacement performs adversarial training with the generator to obtain the generator with updated parameters;

[0094] Anomaly detection module: The aggregation server updates the global model in round i (i=1, 2, ..., n) and the local model uploaded to the aggregation server, and processes the virtual test data generated by the generator after the updated parameters, respectively, to obtain the output results before the update and the output results after the update, and performs anomaly detection on the differences between the output results before the update and the output results after the update to confirm malicious local model updates in the local model updates, and eliminates the malicious local model updates from the local model updates in this step to form a benign model set;

[0095] Global model aggregation module: The aggregation server uses the federated averaging algorithm (FedAvg) to perform global model aggregation on the benign model set, completes the i-th (i=1, 2,…, n) round of global model training, and generates the i+1-th (i=1, 2,…, n) round of global model training. The four module functions are repeated until the i+1-th (i=1, 2,…, n) round of global model converges.

[0096] In a third aspect, an electronic device includes a memory and a processor:

[0097] Memory: used for storing a computer program for implementing the poisoning attack defense method;

[0098] Processor: used to implement the poisoning attack defense method when executing the computer program.

[0099] In a fourth aspect, a computer-readable storage medium stores a computer program, which implements the poisoning attack defense method when executed by a processor; the computer-readable storage medium includes: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk, etc., which can store program codes.

[0100] Example

[0101] This method is verified by setting specific parameters and procedures through Example 1 and using simulation. The specific process is as follows:

[0102] 1. Participant Selection

[0103] Aggregator: From the n participants, it selects participants with equal probability (randomly selecting 10%-40%) to participate in this round of training. At the same time, it sets parameters such as the global model learning rate and the number of local training rounds.

[0104] 2. Local model training and uploading

[0105] Normal participant u i (u i ∈UU q ): Download the current round global model from the aggregation server, and then use local data to communicate with normal participants u i The global model is trained locally to get the local model update Normal participant u i Update the benign local model Sent to the aggregation server.

[0106] Malicious party q i (q i∈U q ): After downloading the global model of the current round from the aggregation server, the malicious participant conducts malicious model training (possibly using its local malicious data to collude with the global model to conduct malicious model training, or directly conduct independent malicious model training), and finally generates a malicious local model update The malicious party updates the local model with a malicious Upload to the aggregation server.

[0107] 3. Anomaly Detection

[0108] Aggregation server: After receiving the update of the participant model, the aggregation server adds the global model to the participant model update to obtain the participant model set m, and then performs anomaly detection according to the following algorithm 1.

[0109]

[0110]

[0111] In this method, the discriminator model parameters in the aggregation server are first replaced with those of the global model. The discriminator, after model parameter replacement, undergoes adversarial training with the generator to produce an updated generator. Next, the updated generator generates virtual test data. The output of the virtual test data before and after the update of the local model of each participant is compared. Models with a difference exceeding a certain threshold are identified as malicious. Finally, a set of benign models, after removing the malicious models, is returned.

[0112] Since the discriminator is replaced by the global model, the generator's goal becomes to make it impossible for the global model to distinguish between real data and generated data. The generator generates virtual test data based on feedback from the global model, so the generated virtual test data is closely related to the parameters of the global model. If the client's updated local model is benign, the difference between the updated local model and the global model will be small, and the output of the benign local model on the virtual test data will have a high similarity with the output of the global model on the virtual test data. However, malicious local models are often trained with malicious data or deliberately designed to deviate from benign local models. Therefore, there will be a significant difference between the output of the malicious local model on the generated virtual test data and the output of the global model. This is because the generator does not generate data specifically to make it impossible for the malicious local model to distinguish between real data and virtual test data.

[0113] The discriminator's goal is to distinguish between real data and the virtual test data generated by the generator. This component is primarily intended to maintain the quality of the virtual test data generated by the generator, ensuring a uniform distribution of its categories. This ensures that when an attacker launches a targeted attack, the difference in the output of the virtual test data generated by the generator before and after the model update is sufficiently significant. This prevents the instability of defense effectiveness caused by poor generated data quality, making it impossible to distinguish between the attack and data quality issues.

[0114] For benign model updates, the output of virtual test data should ensure that labels are evenly distributed across categories and that the differences between the output results are small. If a label reversal attack occurs, the number of inconsistencies between the output results of the model before and after the update will exceed 80%. In the case of a backdoor attack, the probability of a certain category in the output results after the model update will exceed 90%. For precision poisoning attacks, the output categories of the updated data will be reduced to 1 to 3 categories. The robustness of the machine learning system is improved by eliminating malicious model updates (where "malicious model update elimination" is reflected).

[0115] 4. Simulation verification

[0116] The following section verifies the effectiveness, robustness, and efficiency of the present invention through simulation experiments. The experimental environment of Example 1 is shown in Table 1, and the dataset uses the Cifar10 dataset. The CIFAR-10 dataset is an image classification dataset widely used in machine learning and computer vision research. It contains 60,000 32x32 pixel color images divided into 10 categories, with 6,000 images in each category. The dataset is divided into 50,000 training images and 10,000 test images. Due to its moderate size and diversity, CIFAR-10 is often used to test the performance of image classification algorithms and is one of the classic benchmark datasets for deep learning models.

[0117] Table 1 Experimental environment

[0118]

[0119] (1) Experimental setup

[0120] Table 2 Experimental parameter settings

[0121]

[0122] The experiment in this embodiment simulates 100 participants participating in the training, as shown in Table 2, and a total of 95 rounds of training. The data distribution of the participants can be IID or NON-IID. 10 participants are selected in each round, and malicious participants can account for 20% / 80% of the total participants.

[0123] (2) Experimental scene setting

[0124] Table 3 Experimental scenarios

[0125] Scene Number Data distribution Proportion of malicious participants Attack Methods Scenario 1 IID 20% LFA, BBA, DBA, LIE Scenario 2 IID 80% LFA, BBA, DBA, LIE Scenario 3 NON-IID 20% LFA, BBA, DBA, LIE Scene 4 NON-IID 80% LFA, BBA, DBA, LIE

[0126] Table 4 Thresholds

[0127] Attack Methods Number of abnormalities LFA 80% of each category was replaced by another category BBA 90% output is the same DBA 90% output is the same LIE The number of categories is less than 4 or 90% of the output is the same

[0128] Table 3 shows the different experimental scenarios set up by this solution, and Table 4 shows the threshold settings for anomaly detection. These tests respectively test the defense effectiveness against poisoning attacks when a large number of malicious participants launch catastrophic attacks and when the data distribution of the participants varies. Poisoning attacks are categorized as label inversion attacks (LFA), blind backdoor attacks (BBA), distributed backdoor attacks (DBA), and little bit attacks (LIE).

[0129] (3) Evaluate the effectiveness of the plan

[0130] like Figures 3a to 3d In Scenario 1, the impact of four different malicious client attack strategies on global model accuracy after applying this method to the machine learning system is shown. Without defenses deployed, BBA and DBA attacks improve backdoor task accuracy, reaching a 99% success rate within a few rounds. However, after deploying our solution, the backdoor success rate drops to approximately 10%, while the main task accuracy remains unaffected. For LIE and LFA attacks, global model accuracy drops significantly. LIE attacks can also embed a backdoor, achieving 100% accuracy. After deploying our solution, the main task accuracy remains largely unaffected, with the backdoor success rate dropping to approximately 10%.

[0131] like Figures 4a to 4d In Scenario 2, after applying this method to the machine learning system, the impact of four different attack strategies by a malicious client on global model accuracy is shown. Without defenses deployed, attacks can damage global model accuracy or embed backdoors. A large number of attackers can slow aggregation and reduce model accuracy. Under BBA and DBA attacks, the backdoor task accuracy quickly reaches 100%, while the main task accuracy drops by approximately 10%. If the attacker does not attack in the first five rounds, the virtual test dataset is heavily contaminated, making it difficult to detect malicious parameters. After deploying our proposed solution, backdoor accuracy drops to 15%. For LIE and LFA attacks, model accuracy drops to around 50%, and LIE attacks are unable to embed backdoors. After deploying the proposed solution, the main task accuracy is barely affected, and the backdoor success rate drops to 15%.

[0132] like Figures 5a to 5dIn Scenario 3, the impact of four different malicious client attack strategies on global model accuracy after applying this method to the machine learning system is shown. Under non-IID conditions, model accuracy slightly decreases, and the attack is less effective. For example, the backdoor is difficult to embed, resulting in low model accuracy. Under BBA and DBA attacks, backdoor accuracy fluctuates dramatically, ultimately reaching 100%. After deploying our solution, backdoor accuracy drops below 20%. Under LIE and LFA attacks, model accuracy decreases rapidly; however, after deploying our solution, model accuracy remains nearly constant at the normal training level.

[0133] like Figures 6a to 6d In Scenario 4, the machine learning system, after applying this method, shows the impact of four malicious client attack strategies on global model accuracy. When malicious clients predominate, the attack is highly effective, contaminating the model from the outset. Under BBA and DBA attacks, the backdoor is quickly embedded, and the global model backdoor accuracy reaches 100%. Under LIE and LFA attacks, the model quickly collapses, with accuracy dropping to approximately 10%. Backdoor accuracy fluctuates significantly. After deploying this solution, the model quickly recovers to normal levels.

[0134] Comparative experiment

[0135] Figures 7a to 7d The Krum algorithm defense method demonstrates its effectiveness in scenarios 1 and 2. Figures 8a to 8d The defensive effect of the foolsgold defense method in scenarios one and two is demonstrated; the krum solution is easily bypassed by carefully designed lie attacks when there are a small number of clients, it is difficult to detect backdoor attacks, and it is easily affected by label reversal attacks, and the defense effect is poor. In the case of a large number of attackers, the attack effect is irresistible within a few rounds, resulting in a decrease in the accuracy of the global model and the embedding of backdoors. The FoolsGold solution performs better when there are a small number of attackers, the backdoor can be eliminated within a certain number of rounds, and the accuracy remains high; but in the case of a large number of attackers, irresistible attacks will occur within a few rounds, and the backdoor can still be embedded. In comparison, the solution of the present invention has more advantages in defense effect.

[0136] The working principle of the present invention is:

[0137] The aggregation server of the present invention first selects some participants for local training. Normal participants use benign data to update the model, while malicious participants upload tampered parameters. The global model is used as a discriminator for adversarial training with the generator. After generating virtual test data, malicious updates are identified and eliminated by comparing the output differences of each participant's model on the virtual test data. Finally, only the benign models are aggregated to generate a new global model. Specifically, the generator generates virtual test data that is highly adapted to the current global model parameters through adversarial training, so that the output results of the benign client model for these data are highly similar to the global model, while the malicious client model will produce a significantly deviated output difference. This difference is determined by a preset quantitative threshold to identify potential malicious model updates. Since the generator is only optimized for the global model, the malicious client cannot predict the generation logic of the virtual test data, making it difficult to bypass detection through targeted attacks.

Claims

1. A GAN-based machine learning poisoning attack defense method, characterized in that: The following steps are involved: S1. The participant downloads the i-th (i=1, 2, ..., n) round global model from the aggregation server and uses the gradient descent algorithm to train the local model on the participant's local data and obtain the local model update; S2. The participant uploads the updated local model described in step S1 to the aggregation server; S3. After the local model is updated and uploaded in step S2, the aggregation server replaces the model parameters of the internal discriminator with the model parameters of the i-th (i=1, 2, ..., n)-th round global model, and the discriminator after the model parameter replacement is subjected to adversarial training with the generator to obtain a generator with updated parameters; S4. The aggregation server updates the global model in the i-th (i=1, 2, ..., n) round and the local model uploaded to the aggregation server in step S2, and processes the virtual test data generated by the generator after the parameter update in step S3, respectively, to obtain the output results before the update and the output results after the update, and performs anomaly detection on the difference between the output results before the update and the output results after the update to confirm malicious local model updates in the local model updates, and removes the malicious local model updates from the local model updates in this step to form a benign model set; S5. The aggregation server uses the federated averaging algorithm (FedAvg) to perform global model aggregation on the benign model set described in step S4, completes the i-th (i=1, 2, ..., n)-th round of global model training, and generates the i+1-th (i=1, 2, ..., n)-th round of global model training; S6. Repeat steps S1 to S5 until the i+1th (i=1, 2, ..., n)th round of global model converges in step S5.

2. The poisoning attack defense method according to claim 1, wherein: The participants in step S1 include normal participants and malicious participants. The normal participants perform benign model training and obtain benign local model updates. The normal participants all follow the same training rules set by the system.

3. The poisoning attack defense method according to claim 2, wherein: After the malicious participant described in step S1 downloads the global model of round i (i=1, 2, ..., n) from the aggregation server, the malicious participant performs malicious model training and generates malicious local model updates. The malicious model training is performed through collusion or individually, where independent malicious model training includes modifying the data set, modifying the training algorithm, or modifying the uploaded model update.

4. The poisoning attack defense method according to claim 1, wherein: The loss function of the discriminator in step S3 is expressed as follows: L Dis =-logD(x)-log(1-D(G(noise))) Among them, L Dis represents the loss of the discriminator, G represents the generator, noise represents random noise, and D represents the discriminator after model parameter replacement.

5. The poisoning attack defense method according to claim 1, wherein: The loss function of the generator in step S3 is expressed as follows: L Gen =-logD(G(noise)) Among them, L Gen represents the loss of the generator, G represents the generator, noise represents random noise, and D represents the discriminator after the model parameters are replaced.

6. The poisoning attack defense method according to claim 1, wherein: The anomaly detection in step S4 determines that a malicious local model update is being updated after determining that the difference exceeds a set threshold.

7. The poisoning attack defense method according to claim 1, wherein: Step S5 Among them, M G is the global model, n is the set of benign participant models I begin The number of, α is the corresponding weight, m i Represents a collection of participant models.

8. A GAN-based machine learning poisoning attack defense system, applying the poisoning attack defense method according to claims 1 to 7, characterized in that: The defense system includes a local model training and uploading module, an adversarial training module, an anomaly detection module, and a global model aggregation module: Local model training and upload module: The participant downloads the i-th (i=1, 2, ..., n) round global model from the aggregation server, and uses the gradient descent algorithm to perform local model training on the participant's local data to obtain a local model update, and then uploads the local model update to the aggregation server; Adversarial training module: After the local model is updated and uploaded, the aggregation server replaces the model parameters of the internal discriminator with the model parameters of the global model in round i (i = 1, 2, ..., n). The discriminator after model parameter replacement performs adversarial training with the generator to obtain the generator with updated parameters; Anomaly detection module: The aggregation server updates the global model in round i (i=1, 2, ..., n) and the local model uploaded to the aggregation server, and processes the virtual test data generated by the generator after the updated parameters, respectively, to obtain the output results before the update and the output results after the update, and performs anomaly detection on the differences between the output results before the update and the output results after the update to confirm malicious local model updates in the local model updates, and eliminates the malicious local model updates from the local model updates in this step to form a benign model set; Global model aggregation module: The aggregation server uses the federated averaging algorithm (FedAvg) to perform global model aggregation on the benign model set, completes the i-th (i=1, 2,…, n) round of global model training, and generates the i+1-th (i=1, 2,…, n) round of global model training. The four module functions are repeated until the i+1-th (i=1, 2,…, n) round of global model converges.

9. An electronic device comprising a memory and a processor, characterized in that: Memory: used for storing a computer program for implementing the poisoning attack defense method; Processor: used to implement the poisoning attack defense method when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the poisoning attack defense method is implemented.

Citation Information

Patent Citations

  • Federal learning system poisoning attack defense method based on RDP

    CN116155611A

  • Federal learning poisoning attack defense method

    CN118246009A