Federal backdoor detection method and device based on virtual sample analysis

By generating virtual samples and optimizing using gradient inverse method, combining logical function values and multi-category behavior analysis, the detection difficulties of advanced backdoor attacks in federated learning are solved, the accuracy and stability of detection are improved, and data heterogeneous scenarios are adapted to data heterogeneous scenarios, and model and data security are protected.

CN120296735APending Publication Date: 2025-07-11HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510380555.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to accurately detect advanced backdoor attacks that adjust the adaptability of parameters, especially in data heterogeneous scenarios, which affects the reliability and security of federated learning models.

Method used

By generating virtual samples and optimizing using gradient inverse method, we identify clients with abnormal logical function values, combine multi-category behavior analysis, dynamically adjust the detection threshold, and distinguish between unbalanced benign model and backdoor model.

Benefits of technology

Effectively detect advanced backdoor attacks, reduce false positive rates, improve detection stability and accuracy, adapt to data heterogeneous scenarios, and protect model reliability and data privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296735A_ABST
    Figure CN120296735A_ABST
Patent Text Reader

Abstract

The invention discloses a federated backdoor detection method and device based on virtual sample analysis, and the method comprises the following steps: in a round of federated training process, a central server distributes a global model to clients, and the clients train local models and upload the local models to the central server; the central server initializes a random virtual sample, and updates the initial virtual sample through a gradient reverse method to generate a global virtual sample; based on the output distribution of the local model, further optimizing the global virtual sample by using a gradient reverse method to obtain a local virtual sample; calculating a logic function value of each local model on a local virtual sample, identifying clients with abnormal and prominent logic function values, and preliminarily screening the clients as potential backdoor clients; and further analyzing logic function values of the potential backdoor clients on other types of samples, screening out the potential backdoor clients with unbalanced data, and identifying a final backdoor attack. The backdoor attack detection rate can be improved, and the false positive rate in a data exception scene can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of backdoor attack detection in federated learning, and particularly to a federated backdoor detection method and device based on virtual sample analysis. Background Art

[0002] In recent years, deep learning has achieved remarkable results in many fields, but its development faces the dual challenges of data requirements and privacy protection. On the one hand, training a reliable and effective deep learning model usually requires a large amount of data. On the other hand, with the awakening of people's privacy awareness and more and more institutions regarding data as a strategic resource, the acquisition and sharing of data have become increasingly difficult, and it is difficult for trainers to gather a sufficient amount of reliable data for model training in the same place, which seriously restricts the efficiency of model training and deployment.

[0003] In this context, federated learning, as a privacy-preserving distributed learning paradigm, has emerged and been widely promoted. Federated learning takes the core idea of "data does not move, the model moves", allowing multiple parties to participate jointly and train collaboratively. Each party uses its local dataset to train the model and uploads the model update to the central server. The central server aggregates the model updates from all parties to form a global model, which not only realizes the protection of data privacy but also promotes the circulation of data value. Federated learning has been applied in many industries such as healthcare, finance, and intelligent devices. For example, in the healthcare field, different medical institutions can jointly train a disease diagnosis model through federated learning without sharing sensitive patient data; in the financial industry, banks can use federated learning to improve the accuracy of anti-fraud models while protecting customer privacy.

[0004] However, the distributed architecture of federated learning and the opacity of the local model training process make it vulnerable to backdoor attacks. Backdoor attacks are a serious threat to model security. Attackers inject malicious samples containing specific triggers into the training data, causing the model to output according to the attacker's intention when encountering the trigger, while the model performs normally on normal input data. This seriously undermines the reliability of the model and may lead to serious privacy and security problems. Given that federated learning is often applied to privacy-sensitive high-value scenarios such as healthcare and financial services, backdoor attacks against federated learning should be strongly prevented.

[0005] The existing detection schemes are generally divided into two categories: model parameter detection and model behavior detection. Detection schemes based on model parameters, such as dimensionality reduction clustering of model updates through principal component analysis technology, or anomaly detection using statistical features of model updates, such as cosine similarity and Euclidean distance, have low computational overhead and are easy to implement. However, they are insufficient in resisting advanced backdoor attacks that use constraint conditions such as gradient projection constraints and model update directions during the backdoor training process. Detection methods based on model behavior focus more on the behavior performance of some neurons or the overall model. Most of them can provide better detection effects, but are more vulnerable to data heterogeneity. For example, the detection scheme proposed by Clement et al. reduces the impact of backdoor updates by assigning lower aggregation weights to updates with high pairwise cosine similarity. Li et al. implanted an outlier watermark in the global model based on the outlier characteristics of backdoor samples, and detected the existence of backdoors through the sensitivity of local models to the watermark.

[0006] The limitations of existing solutions are mainly manifested in: most existing works are difficult to accurately detect advanced backdoor attacks with parameter adaptation adjustment, including gradient projection backdoor attacks and chameleon backdoor attacks, etc. That is, when users have a benign dataset, attackers can master the parameter distribution of the benign model by training the benign model. During the training process of the backdoor model, they can specifically adjust the distribution of model updates, thus bypassing the federated backdoor defense scheme based on parameters. Most existing works do not consider the heterogeneity problem of benign local models in the data heterogeneity scenario. When there are significant differences in the data distributions of the local datasets held by benign clients, there will also be obvious differences between benign local models. This change in heterogeneity will directly lead to a significant increase in the false positive rate during the backdoor detection process, resulting in a slowdown in the convergence of the global model. Summary of the Invention

[0007] In view of this, it is necessary to provide a federated backdoor detection method and device based on virtual sample analysis to effectively solve the technical problems that existing technologies are difficult to accurately detect advanced backdoor attacks with parameter adaptation adjustment and false positive rates in the data heterogeneity scenario.

[0008] The present invention provides a federated backdoor detection method based on virtual sample analysis, including the following steps:

[0009] Step S1, during a round of federated training, the central server distributes the global model to each client, and the client trains the local model and uploads it to the central server;

[0010] Step S2. In each training round, the central server initializes random virtual samples and updates the initial virtual samples through the gradient reversal method. After multiple training iterations, global virtual samples aligned with the target categories are generated. Based on the output distribution of the local models, the global virtual samples are further optimized using the gradient reversal method to obtain local virtual samples.

[0011] Step S3. Calculate the logical function values of each local model on the local virtual samples, identify the clients with abnormally prominent logical function values, and preliminarily screen them as potential backdoor clients.

[0012] Step S4. Further analyze the logical function values of the potential backdoor clients on other category samples, screen out the potential backdoor clients with data imbalance, and identify the final backdoor attack.

[0013] Preferably, step S1 is specifically:

[0014] In a round of federated training, the central server randomly selects a batch of clients from all the clients as the participants in this round of federated learning and distributes the global model of the current round.

[0015] The clients participating this time initialize the local models with the parameters of the global model and train on the local datasets.

[0016] The clients upload the trained local models to the central server, and the central server receives the local models uploaded by the selected clients.

[0017] Preferably, in step S2, updating the initial virtual samples through the gradient reversal method and generating global virtual samples aligned with the target categories after multiple training iterations is specifically:

[0018] The central server performs gradient clipping on the local models.

[0019] For each category, the central server randomly initializes a batch of initial virtual samples with the same input shape.

[0020] Based on the initial virtual samples, the central server uses the gradient reversal method to train the global virtual samples for the output space of the global model.

[0021] After several rounds of training iterations, the global virtual samples of the current round are obtained.

[0022] Preferably, the central server performing gradient clipping on the local models is specifically:

[0023] The central server performs gradient clipping on the updated parameters of each local model according to the set threshold:

[0024]

[0025] where θ i ' is the parameter after clipping, and θ i represents the parameter before clipping, ⊙ represents the Hadamard product, min() represents taking the minimum value, and ||θ i ||₂ represents the L2 norm of the model parameters, and τ is the set threshold;

[0026] By gradient clipping, the parameter range of the local model is restricted to avoid model amplification attacks.

[0027] Preferably, the gradient reversal method is used to train the global virtual samples, and the optimization objective of the training is:

[0028]

[0029] where V global,c represents the global virtual sample, l() is the classification accuracy constraint, f θ () is the global model, v is the virtual sample, c is the target class, λ is the diversity coefficient, and D() is the diversity constraint.

[0030] Preferably, in step S2, based on the output distribution of the local model, the gradient reversal method is further used to optimize the global virtual sample to obtain a local virtual sample, specifically:

[0031] After all local models are updated, for each client, based on the output distribution of its local model and the difference between the local model output and the global model output, the gradient reversal method is further used to update the global virtual sample to generate a local virtual sample that better adapts to the local data characteristics, and the optimization objective is:

[0032]

[0033] where V lobal,i is the local virtual sample, and f θ (v) i represents the confidence that the input virtual sample v is classified as class i by the model f θ (), and f θ (v) k represents the confidence of the model on the k-th class when the input is v, and C represents the set of target classes.

[0034] Preferably, step S3 is specifically:

[0035] The central server calculates the logical function values of each local model on the local virtual samples to obtain the prediction confidence of each local model for each class;

[0036] Calculate the difference between the maximum value and the second maximum value of the logical function values for each category, and obtain the confidence difference of the local virtual samples of the local model for each category.

[0037] Mark the categories in the local model with a confidence difference greater than the difference threshold as potential backdoor target classes, and mark the corresponding local model as a potential backdoor model.

[0038] If there is no category in the local model with a confidence difference greater than the difference threshold, mark the corresponding local model as a benign model and save the corresponding local model in the list to be aggregated.

[0039] Preferably, step S4 is specifically as follows:

[0040] For potential backdoor target classes, analyze whether there is a stacking effect in the distribution of the logical function values of the corresponding client on samples of other categories.

[0041] For each potential backdoor client, count the number of categories with a stacking effect. If the number of categories exceeds the set value, determine that the corresponding potential backdoor model is a backdoor model; otherwise, it is a benign model with data imbalance.

[0042] The central server filters the model updates of the backdoor models, and uses an aggregation algorithm to aggregate the model updates of the benign models to obtain a new round of global models.

[0043] Preferably, for potential backdoor target classes, analyzing whether there is a stacking effect in the distribution of the logical function values of the corresponding client on samples of other categories is specifically as follows:

[0044] Calculate the sum of the logical function values of each potential backdoor target class on the local virtual samples of other categories.

[0045] Use the absolute median difference algorithm to calculate the anomaly score for each category. If the anomaly score is less than the set value, there is a stacking effect in the distribution of the logical function values of the corresponding potential backdoor client on the samples of the corresponding category.

[0046] The present invention also provides a federated backdoor detection device based on virtual sample analysis, including a memory and a processor. A computer program is stored on the memory, and when the computer program is executed by the processor, the federated backdoor detection method based on virtual sample analysis is implemented.

[0047] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention generates virtual samples through gradient reverse technology, without using real data for detection, thus avoiding the risk of privacy leakage and meeting the requirements of data privacy protection in federated learning. At the same time, by analyzing the model behavior on virtual samples, the characteristics of backdoor attacks can be accurately identified. Even in the face of advanced backdoor attacks such as gradient projection attacks and chameleon attacks, abnormal behaviors can be effectively detected. On the other hand, after initially screening backdoor clients, the present invention further analyzes their behavior on other category samples, dynamically adjusts the detection threshold, and effectively distinguishes unbalanced benign models from real backdoor models through abnormal detection of logical function values and multi-category behavior analysis. This mechanism can dynamically adjust the detection strategy according to changes in the federated learning environment, reduce the false positive rate, and improve the stability and reliability of detection. The present invention does not depend on the distribution characteristics of data, can effectively cope with data heterogeneous scenarios, and reduce the false positive rate caused by data distribution differences. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation of the present invention. In the drawings:

[0049] Figure 1 is a flowchart of an embodiment of a federated backdoor detection method based on virtual sample analysis provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The following will specifically describe the preferred embodiments of the present invention with reference to the drawings. The drawings constitute a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.

[0051] Embodiment 1

[0052] Please refer to Figure 1 , a federated backdoor detection method based on virtual sample analysis in this embodiment specifically includes the following steps:

[0053] Step S1: During a round of federated training, the central server distributes the global model to each client, and the client trains the local model and uploads it to the central server;

[0054] Step S2: In each training round, the central server initializes random virtual samples, updates the initial virtual samples through the gradient reverse method, and generates global virtual samples aligned with the target category after multiple training iterations; based on the output distribution of the local model, the global virtual samples are further optimized using the gradient reverse method to obtain local virtual samples;

[0055] Step S3: Calculate the logical function values of each local model on the local virtual samples, identify the clients with abnormally prominent logical function values, and preliminarily screen them as potential backdoor clients;

[0056] Step S4: Further analyze the logical function values of the potential backdoor clients on other category samples, filter out the potential backdoor clients with data imbalance, and identify the final backdoor attack.

[0057] The object of the present invention is to design a federated backdoor detection method based on virtual sample behavior analysis, which can effectively detect advanced backdoor attacks with parameter adaptation adjustment in the federated learning environment, protect the reliability of the model and the security of data, and at the same time meet the detection requirements in the data heterogeneous scenario.

[0058] The present invention generates virtual samples through gradient inversion, without relying on real data as a reference, thus meeting the requirements of the federated learning environment and being able to reveal potential backdoor attack characteristics. It can effectively resist backdoor attacks with parameter adaptation adjustment, such as gradient projection backdoor attacks and chameleon backdoor attacks. Even if the attacker uses the parameter distribution of the benign model for targeted adjustment during the training of the backdoor model, the present invention can accurately detect potential backdoor attacks.

[0059] At the same time, in the data heterogeneous scenario, there are significant differences in the data distributions of the local datasets held by different clients. The local datasets of some clients are extremely imbalanced, resulting in imbalanced benign models, that is, there are also obvious differences between benign local models. Regarding the false positive rate in the data heterogeneous scenario, the present invention reduces false alarms caused by the distribution differences of the local datasets of benign clients by distinguishing imbalanced benign models from backdoor models, improves the accuracy of detection, thus overcoming the common data heterogeneity limitations in federated learning and reducing the interference of heterogeneous benign clients on detection.

[0060] The backdoor attack refers to the attacker injecting malicious samples containing specific triggers into the training data, causing the model to output according to the attacker's intention when encountering the trigger, while the model performs normally on normal input data, seriously damaging the reliability of the model and possibly causing serious privacy and security problems. The present invention aims to timely discover and prevent the damage of advanced backdoor attacks to the federated learning model through an innovative detection mechanism, ensuring that the model can provide accurate and reliable output results in various application scenarios, such as disease diagnosis in healthcare, anti-fraud in financial services, etc., and safeguarding the data security and privacy rights and interests of users.

[0061] The overall technical idea of the present invention is to detect backdoor attacks in federated learning through behavior analysis rather than simply relying on model parameters. This method achieves high-precision backdoor detection through three core steps: using gradient reversal technology to generate virtual samples aligned with target classes for subsequent detection; based on the behavior of virtual samples, identifying clients with abnormally prominent logical function values to preliminarily screen potential backdoor clients; further analyzing the behavior of potential backdoor clients on other class samples to distinguish data imbalance from true backdoor attacks, thereby reducing the false positive rate and improving detection accuracy.

[0062] The first core step is to generate local virtual samples by gradient reversal, which specifically includes two parts: global virtual sample generation and local virtual sample optimization.

[0063] Global virtual sample generation: In each training round of federated learning, when the central server distributes the global model to local clients, for each target class of the classification task, the following operations are performed: ① Initial virtual sample generation: Generate initial random virtual samples using a Gaussian distribution as the starting point of the virtual samples. ② Global virtual sample alignment: Update the initial virtual samples through the gradient reversal method to align them with the target class in the output space of the global model. Specifically, optimize the virtual samples to maximize the score of the target class in the output logical function value of the global model on these samples. After several iterations, the global virtual samples for the current round are obtained.

[0064] Local virtual sample optimization: After collecting all local model updates, for each local client, based on the output distribution of its local model, further update the virtual samples using the gradient reversal method to make them adapt to the characteristics of the local model. The finally obtained local virtual samples will be used for subsequent detection.

[0065] The second core step is to preliminarily screen backdoor clients based on logical function values. Based on the generated local virtual samples, conduct behavior analysis on each local model. This specifically includes two parts: logical function value calculation and anomaly detection.

[0066] Logical function value calculation: Calculate the logical function values of each local model on the local virtual samples to obtain the output scores of each class.

[0067] Anomaly detection: Identify clients with abnormally prominent logical function values for a certain class, that is, clients whose logical function values for this class are significantly higher than those of other classes. Specifically, calculate the difference between the maximum value of the logical function value and other values. If this difference exceeds a preset threshold, mark this class as a potential backdoor target class and mark this client as a potential backdoor client.

[0068] The third core step is to distinguish between unbalanced classes and backdoor models. For the potentially backdoored clients preliminarily screened out, further analyze their behavior on samples of other classes to distinguish between unbalanced benign models and genuine backdoor attacks. This specifically includes two parts: multi-class behavior analysis and a discrimination mechanism.

[0069] Multi-class behavior analysis: For the potentially backdoored target class, analyze the distribution of the logical function values of this client on samples of other classes, and observe whether there is a stacking effect, that is, the potentially target class shows relatively high logical function values on samples of multiple classes.

[0070] Discrimination mechanism: By comparing the distribution of the logical function values of the potentially backdoored target class with those of other classes, determine whether this anomaly is caused by data imbalance. If the anomaly only appears in a single target class and is significantly different from other classes, it is determined as a backdoor attack; if similar anomalies appear in multiple classes, it may be a natural deviation caused by data imbalance, and thus it is excluded.

[0071] The key technical points of the present invention lie in gradient reverse generation of virtual samples, preliminary screening of backdoor clients, and distinguishing between unbalanced classes and backdoor models.

[0072] Gradient reverse generation of virtual samples: The present invention generates virtual samples for each client through the gradient reverse method, ensuring that among the logical function values of the model on these virtual samples, the largest value is significantly greater than the second largest value. This process can reveal potential backdoor attack characteristics and provide a basis for subsequent detection. At the same time, this process does not rely on real data as a reference and can meet the requirements of the federated learning environment.

[0073] Preliminary screening of backdoor clients: Based on the generated virtual samples, the present invention preliminarily screens out those clients with abnormally prominent logical function values on a certain class and marks them as potentially backdoored clients. This screening mechanism utilizes the common characteristics in backdoor attacks, that is, the logical function values of specific classes are abnormally high, thereby effectively narrowing the detection scope. The present invention detects the essential characteristics of the backdoor model rather than the cumulative effect on parameters, so it can overcome the federated backdoor attack with parameter adaptive adjustment.

[0074] Distinguishing between unbalanced classes and backdoor models: For the potentially backdoored clients preliminarily screened out, the present invention further analyzes the logical function values of these clients on samples of other classes. By comparing the logical function values of the target class with those of other classes, it is distinguished whether it is a natural deviation caused by data imbalance or an anomaly caused by a backdoor attack. This discrimination mechanism can effectively reduce false positives and improve the accuracy of detection. Through further screening in this process, the present invention can overcome the limitations of common data heterogeneity in federated learning and reduce the interference of heterogeneous benign clients on detection.

[0075] To illustrate the present invention more specifically, this embodiment takes federated learning in medical image diagnosis as an applicable scenario for illustration. In this scenario, multiple medical institution clients and a central server participate in federated learning to jointly train a deep learning model for disease diagnosis. This deep learning model can correctly indicate the corresponding disease type through the recognition and analysis of medical images. Each medical institution has a local medical image dataset, and these datasets are non-independent and identically distributed, that is, there are significant differences in the class distribution of data from different institutions, and the data volume and quality are also different. In each round of federated learning process, the local dataset is used to train and update the local model and upload it. The central server is responsible for the initialization of the global model and performs the following tasks in different training rounds: distributing the global model to each local client, collecting the updates of all local models, detecting backdoor attacks and discarding the attacked model updates, and aggregating the remaining benign model updates to form a new round of global model.

[0076] The specific steps of step S1 are as follows: In the t-th round of federated training process, the central server distributes the global model, the client trains the local model and uploads it, and the central server collects the local model.

[0077] Step 11: Randomly select local clients and distribute the global model. The central server randomly selects 20 from all participating medical institution clients as the participants in this round of federated learning and distributes the global model of the current round.

[0078] Step 12: The local client trains the local model and uploads it. Each local client initializes the local model with the parameters of the received global model, and uses SGD optimizer with a batch size of 128, a learning rate set to 0.1, and trains for 2 epochs on the local dataset. The optimization objective is set as:

[0079]

[0080] Among them, It means that the optimization objective is to find θ that minimizes the value of the loss function i , L() is the loss function, which is the cross-entropy function in this scenario, and θ i is the local model parameter, and D i is the local dataset.

[0081] For potential backdoor attackers, the optimization objective is set as:

[0082]

[0083] Among them, It means that the optimization objective is to find θ that minimizes the value of L adv () function, and L poison , L adv() represents the optimization function of the attacker, L() is the loss function, and θ i is the local model parameter, and θ poison represents the backdoor model parameter, and θ benign represents the benign model parameter used by the attacker for reference, represents the backdoor dataset, D i is the local dataset, SIM() represents the parameter distribution similarity, and λ represents the balance parameter;

[0084] That is to say, advanced backdoor attackers may add parameter constraints during the training process to reduce the similarity with the benign model, thereby bypassing parameter-based backdoor detection.

[0085] After completing the model training, the local client uploads the local model to the central server.

[0086] Step 13: The central server receives the local model uploaded by the selected client.

[0087] Specifically, the step S2 is specifically as follows:

[0088] The central server performs gradient clipping on the local model;

[0089] For each category, the central server randomly initializes a batch of initial virtual samples with the same input shape;

[0090] Based on the initial virtual samples, the central server uses the gradient reversal method to train the global virtual samples for the output space of the global model;

[0091] After several rounds of training iterations, the global virtual samples of the current round are obtained;

[0092] After all local models are updated, for each client, based on the output distribution of its local model and the difference between the local model output and the global model output, the gradient reversal method is further used to update the global virtual samples to generate local virtual samples that are more adapted to the local data characteristics.

[0093] Step 21: The central server performs gradient clipping on the local model to resist model amplification attacks. The central server corrects the overly large update parameters of each local model according to the set threshold τ to avoid the global model parameters being replaced by a single local model due to a large update amount of that local model. The formula for gradient clipping is as follows:

[0094]

[0095] Among them, θ i ' is the clipped parameter, and θ iDenotes the parameters before cropping, represents the Hadamard product, min() represents taking the minimum value, ||θ i ||2 represents the L2 norm of the model parameters, and τ is the set threshold.

[0096] By gradient clipping, the parameter range of the local model is restricted, avoiding model amplification attacks and facilitating subsequent detection processes.

[0097] Step 22: The central server initializes random virtual samples. For each category, the central server randomly initializes a batch of initial virtual samples with the same shape as the input using a Gaussian distribution.

[0098] Step 23: Based on the initial virtual samples, the central server uses the gradient reversal method to generate global virtual samples aligned with each category in the output space of the global model. By optimizing and updating using the gradient reversal method, these samples can align with the performance of the global model in terms of categories, and finally global virtual samples are obtained. These samples simulate the distribution of the global model in this category, and the optimization objective during the training process is:

[0099]

[0100] Among them, V global,c represents the global virtual sample, l() is the classification accuracy constraint, f θ () is the global model, v is the virtual sample, c is the target category, λ is the diversity coefficient, and D() is the diversity constraint, making the samples more diverse and covering a larger output space of the global model categories.

[0101] Specifically, the virtual samples are iteratively updated using the gradient reversal method, and the iterative process is as follows:

[0102] v t+1 = v t - ηΔ v F(v t , c, θ)

[0103] Among them, v t+1 represents the virtual sample after the (t + 1)-th iteration, v t represents the virtual sample after the t-th iteration, η is the learning rate, Δ v F(v t , c, θ) is the gradient of the optimization term with respect to the current virtual sample v t .

[0104] Step 24: Optimize the global virtual samples based on the output distribution of the local models. In this step, after all local models are updated, each local client optimizes the global virtual samples received from the central server according to the output distribution of its local model to generate local virtual samples that are more adapted to the characteristics of local data, taking into account the differences between the local model output and the global model output, and adjusting the virtual samples through the gradient reversal method to make them better reflect the distribution characteristics of local data. The optimization objective is as follows:

[0105]

[0106] where V lobal,i is the local virtual sample, and f θ (v) i represents the confidence that the input virtual sample v is classified as category i by the model f θ (), and f θ (v) k represents the confidence of the model on the k-th category when the input is v. C represents the set of target categories.

[0107] Specifically, the gradient reversal method is used to iteratively update the virtual samples. The iterative process is as follows:

[0108] v t+1 = v t - ηΔ v G(v t , c, θ)

[0109] where v t+1 represents the virtual sample after the (t + 1)-th iteration, v t represents the virtual sample after the t-th iteration, η is the learning rate, and Δ v G(v t , c, θ) is the gradient of the optimization term with respect to the current virtual sample v t .

[0110] Specifically, step S3 is as follows: The central server uses the local virtual samples collected from each client to analyze and detect potential backdoor attacks. This step involves a detailed analysis of the model output of each client to identify possible backdoor target categories.

[0111] Step 31: Calculate the logical function values. The central server first calculates the logical function values of each local model on the local virtual samples. The logical function values reflect the prediction confidence of the model for each category.

[0112] Step 32: Calculate the difference between the maximum and the second maximum of the logical function values for each category. The server identifies the maximum and the second maximum in each set of logical function values and calculates the difference between them to determine the confidence difference of the model for virtual samples of each category:

[0113]

[0114] where diff i,c is the confidence difference, logit i,c is the predicted confidence, C represents all label types, logit i,c [C] is the maximum of the logical function values, and

[0115] is the second maximum of the logical function values.

[0116] Step 33: Calculate the upper quartile Q3 and the interquartile range IQR of the confidence differences of each local model for each category on virtual samples.

[0117] Step 34: If there is a category in the local model such that diff i,c > Q3 + 1.5IQR, then mark the corresponding category as a potential backdoor target category and the corresponding local model as a potential backdoor model.

[0118] Step 35: If there is no category in the local model such that diff i,c < Q3 + 1.5IQR, then mark the corresponding local model as a benign model and save the parameters of this local model in the list to be aggregated.

[0119] Specifically, step S4 is as follows: Screening of suspicious backdoor categories based on logical function values. In this step, the central server will further screen and confirm potential backdoor categories based on logical function values.

[0120] Step 41: Analyze the distribution of logical function values of potential backdoor clients, and calculate the sum of logical function values of each category on virtual samples of other categories.

[0121] Step 42: Use the absolute median difference algorithm to calculate the anomaly score of each category. If the anomaly score of the potential backdoor category is less than 3, that is, it performs normally, then the detection result is a backdoor model; otherwise, the detection result is a benign model with data imbalance.

[0122] Example 2

[0123] This embodiment provides a federated backdoor detection device based on virtual sample analysis, including a memory and a processor. A computer program is stored on the memory, and when the computer program is executed by the processor, it implements the federated backdoor detection method according to Embodiment 1.

[0124] The federated backdoor detection device based on virtual sample analysis provided in this embodiment is used to implement the federated backdoor detection method based on virtual sample analysis. Therefore, the technical effects possessed by the federated backdoor detection method based on virtual sample analysis are also possessed by the federated backdoor detection device based on virtual sample analysis, and will not be elaborated here.

[0125] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the present invention.

Claims

1. A federated backdoor detection method based on virtual sample analysis, characterized in that, It includes the following steps: Step S1: In a round of federated training, the central server distributes the global model to each client, and the client trains the local model and uploads it to the central server. Step S2: In each training round, the central server initializes random virtual samples and updates the initial virtual samples by the gradient reversal method. After multiple training iterations, global virtual samples aligned with the target category are generated. Based on the output distribution of the local model, the global virtual samples are further optimized using the gradient reversal method to obtain local virtual samples. Step S3: Calculate the logical function values of each local model on the local virtual samples, identify the clients with abnormally prominent logical function values, and initially screen them as potential backdoor clients. Step S4: Further analyze the logical function values of the potential backdoor clients on other category samples, filter out the potential backdoor clients with data imbalance, and identify the final backdoor attack.

2. The method for detecting a federal backdoor based on virtual sample analysis according to claim 1, wherein The specific content of step S1 is as follows: In a round of federated training, the central server randomly selects a batch of clients from all the clients as the participants in this round of federated learning and distributes the global model of the current round. The clients participating this time initialize the local model with the parameters of the global model and train on the local dataset. The client uploads the trained local model to the central server, and the central server receives the local models uploaded by the selected clients.

3. The federal backdoor detection method based on virtual sample analysis according to claim 1, wherein, In step S2, the initial virtual samples are updated by the gradient reversal method, and global virtual samples aligned with the target category are generated after multiple training iterations. The specific content is as follows: The central server performs gradient clipping on the local model. For each category, the central server randomly initializes a batch of initial virtual samples with the same input shape. Based on the initial virtual samples, the central server uses the gradient reversal method to train the global virtual samples for the output space of the global model. After several rounds of training iterations, the global virtual samples of the current round are obtained.

4. The federal backdoor detection method based on virtual sample analysis according to claim 3, characterized in that The specific content of the central server performing gradient clipping on the local model is as follows: The central server performs gradient clipping on the updated parameters of each local model according to the set threshold: Among them, θ i ' is the parameter after clipping, and θ i represents the parameter before clipping, ⊙ represents the Hadamard product, min() represents taking the minimum value, ||θ i ||2 represents the L2 norm of the model parameter, and τ is the set threshold; Through gradient clipping, the parameter range of the local model is restricted to avoid model amplification attacks.

5. The federated backdoor detection method based on virtual sample analysis according to claim 3, wherein, When using the gradient reversal method to train the global virtual samples, the optimization objective of the training is: Among them, V global,c represents the global virtual sample, l() is the classification accuracy constraint, f θ () is the global model, v is the virtual sample, c is the target category, λ is the diversity coefficient, and D() is the diversity constraint.

6. The method for detecting federal backdoors based on virtual sample analysis according to claim 1, wherein In step S2, based on the output distribution of the local model, the global virtual samples are further optimized using the gradient reversal method to obtain local virtual samples. The specific content is as follows: After all local models are updated, for each client, according to the output distribution of its local model, based on the difference between the local model output and the global model output, the global virtual samples are further updated using the gradient reversal method to generate local virtual samples that better adapt to the local data characteristics. The optimization objective is: Among them, V lobal,i is a local virtual sample, and f θ (v) i represents the confidence that the input virtual sample v is classified as class i by the model f θ (), and f θ (v) k represents the confidence of the model on the k-th class when the input is v, and C represents the set of target classes.

7. The federal backdoor detection method based on virtual sample analysis according to claim 1, characterized in that The specific content of step S3 is as follows: The central server calculates the logical function values of each local model on the local virtual samples to obtain the prediction confidence of each local model for each category. Calculate the difference between the maximum value and the second maximum value of the logical function values on each category to obtain the confidence difference of the local virtual samples of each category by the local model; Mark the categories with confidence differences greater than the difference threshold in the local model as potential backdoor target categories, and mark the corresponding local models as potential backdoor models; If there are no categories with confidence differences greater than the difference threshold in the local model, mark the corresponding local model as a benign model and save the corresponding local model in the list to be aggregated.

8. The federal backdoor detection method based on virtual sample analysis according to claim 1, characterized in that The specific content of step S4 is as follows: For potential backdoor target categories, analyze whether there is a stacking effect in the distribution of logical function values of the corresponding client on samples of other categories; For each of the potential backdoor clients, count the number of categories with a stacking effect. If the number of categories exceeds the set value, determine that the corresponding potential backdoor model is a backdoor model; otherwise, it is a benign model with data imbalance; The central server filters the model updates of the backdoor models and uses an aggregation algorithm to aggregate the model updates of the benign models to obtain a new round of global models.

9. The federal backdoor detection method based on virtual sample analysis according to claim 8, characterized in that For potential backdoor target categories, analyze whether there is a stacking effect in the distribution of logical function values of the corresponding client on samples of other categories, specifically: Calculate the sum of the logical function values of each potential backdoor target category on the local virtual samples of other categories; Use the absolute median difference algorithm to calculate the anomaly score of each category. If the anomaly score is less than the set value, there is a stacking effect in the distribution of the logical function values of the corresponding potential backdoor client on the samples of the corresponding category.

10. A federated backdoor detection device based on virtual sample analysis, characterized in that, It includes a memory and a processor. A computer program is stored on the memory, and when the computer program is executed by the processor, it implements the federated backdoor detection method based on virtual sample analysis according to any one of claims 1-9.