Federal learning backdoor attack defense method based on model fingerprints
By using model fingerprint samples to verify client model updates in federated learning, the hidden and destructive problems of backdoor attacks in federated learning are solved, and accurate identification and defense of backdoor attacks are achieved, ensuring the security of model updates and the robustness of the system.
Patent Information
- Application Number
- CN202510068019.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-30
AI Technical Summary
The concealment and destructiveness of backdoor attacks in federated learning make it the main threat to federated learning security. The existing defense methods have limited effectiveness in the face of carefully designed backdoor attacks, making it difficult to accurately identify malicious updates without affecting model performance.
The federated learning backdoor attack defense method based on model fingerprint is adopted. The server generates specific model fingerprint samples and verifies them in the client model update to determine whether the model submitted by the client has been tampered with by the backdoor attack, thereby identifying and eliminating malicious updates.
This method can accurately identify malicious clients and defend against backdoor attacks in multiple complex environments. It has wide adaptability and will not affect the performance of federated learning, significantly improving the accuracy and robustness of backdoor detection.
Smart Images

Figure CN120068082A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and federated learning security, and specifically to a method for detecting and defending against backdoor attacks using model fingerprint technology. Background Art
[0002] As a distributed machine learning method, federated learning allows multiple clients to train models locally and send the updated model parameters to the server for aggregation without sharing the original data, which well protects user data privacy. However, due to the distributed nature of federated learning, malicious participants can manipulate the prediction results by injecting backdoors into the global model. This kind of attack is usually achieved by introducing malicious samples or tampering with model parameters during the local training process, causing the model to produce misclassifications under specific trigger conditions while not affecting the model performance under normal circumstances. The concealment and destructiveness of backdoor attacks make them one of the main threats to the security of federated learning.
[0003] Traditional federated learning defense methods include aggregation rules based on anomaly detection and robustness optimization, etc., but these methods often have limited effects when facing carefully designed backdoor attacks. In addition, it is very difficult for these methods to accurately identify malicious updates without affecting the model performance. Therefore, a new defense mechanism is needed to effectively identify and resist backdoor attacks.
[0004] Based on this, this patent proposes a method for defending against backdoor attacks in federated learning based on model fingerprints. This method judges whether the model submitted by the client has been tampered with by a backdoor attack by generating specific model fingerprint samples on the server side and verifying them in the client model update, so as to identify and eliminate malicious updates. This defense method can accurately identify malicious clients, and at the same time will not affect the performance of federated learning during the defense process, can defend against backdoor attacks in a variety of complex environments, and has broad adaptability.
[0005] In existing related research, methods based on cosine similarity or Euclidean distance detect anomalies by directly calculating the differences between model updates. However, these methods are unstable when facing carefully designed backdoor attacks and are prone to misjudgment or missed judgment. The backdoor defense method based on model watermark uses white-box watermark technology to embed watermarks into the model and extract and verify them later, but this method is likely to affect the normal task performance of the model, and it is difficult to accurately control the watermark embedding capacity. In contrast, the method of the present invention does not need to modify the model parameters, avoiding the impact on the model performance. At the same time, the fingerprint samples have good generalization ability and flexibility, can adapt to different attack scenarios and be updated regularly, significantly improving the accuracy and robustness of backdoor detection. Summary of the Invention
[0006] To effectively resist backdoor attacks in federated learning and improve the security of model updates and the robustness of the system, the present invention proposes a method for defending against backdoor attacks in federated learning based on model fingerprints.
[0007] Specifically, it includes the following steps:
[0008] S1: Generate a set of fingerprint samples with specific characteristics on the server side. These samples can show different responses on normal models and backdoor models for subsequent verification;
[0009] S2: Each client trains a model on local data and submits the trained model update to the server;
[0010] S3: The server uses the fingerprint samples to verify the model update submitted by each client and records the response of the update on the fingerprint samples;
[0011] S4: Compare the consistency between the response of the model update submitted by the client to the fingerprint samples and the expected result to determine whether there is an anomaly;
[0012] S5: Dynamically adjust the aggregation weight of the model update according to the participation times of the client and the historical detection results to adapt to different attack frequencies and the proportion of malicious clients;
[0013] S6: If abnormal client model updates are continuously detected, the server will reject the update and can record the suspicious client, restrict its subsequent participation or conduct more rigorous verification.
[0014] Furthermore, in step S1, when generating fingerprint samples, a set of specific characteristics that can reflect the model performance can be selected, such as random perturbation characteristics generated by slightly perturbing images or non-primary semantic characteristics irrelevant to the main label in image classification tasks. The selection of these characteristics needs to ensure coverage of diverse data types and characteristics to enhance the generalization ability of the fingerprint samples. The generated fingerprint samples and their corresponding standard responses should be securely stored on the server side and updated and expanded regularly according to changes in different attack types and scenarios to ensure that the fingerprint sample library can adapt to dynamic security threats.
[0015] Furthermore, in step S2, the client independently trains a model on the local dataset and follows the federated learning protocol during the training process to ensure that data privacy is not leaked. After training is completed, the client submits the model update to the server, and relevant metadata such as client identifiers, participation rounds, and dataset sizes are attached to the update. In addition, the client can also record training logs, including training rounds, data volume, and model parameter changes, to provide a reference basis for subsequent verification on the server side.
[0016] Further, in step S3, the server can input the generated fingerprint samples into the model updated by the client and evaluate the credibility of the update by observing the model's response to the fingerprint samples. These response results include output categories, confidence levels, and loss values, etc. The server compares these results with the standard fingerprint responses, calculates the consistency score, and uses this to determine whether there are signs of anomalies or backdoor attacks in the model update.
[0017] Further, in step S4, the server identifies anomalies by comparing the consistency between the model's response to the fingerprint samples and the expected standard response. Specifically, the fingerprint samples contain specific features that can produce significant differences between the normal model and the backdoor model. These features can be designed through slight perturbations, specific label mappings, or randomly embedded information. The server compares these actual responses with the expected standard responses and calculates the consistency score. If the consistency score is lower than the preset threshold, it indicates that there may be a backdoor in the model update, and the model's response to the fingerprint samples shows an abnormal deviation. This detection method can effectively distinguish between the normal model and the backdoor model, while ensuring that the detection process is transparent to the client and does not affect the normal training process, helping to timely identify and intercept potential backdoor attacks.
[0018] Further, in step S5, the server can set an initial aggregation weight for each client and dynamically adjust the weight during the training process according to the client's participation times and detection results and other behaviors. For clients with a high participation rate and stable verification results, appropriately increase their aggregation weight; while for clients that are detected as abnormal multiple times or have a low participation rate, reduce their weight, or even set it to zero if necessary. After the weight adjustment, normalization processing is required to ensure that the sum of all client weights is 1. In addition, the server can set an adaptive adjustment period, such as dynamically updating the weight after each round of aggregation or when an anomaly is detected, to improve the system's adaptability and security to the dynamic attack environment.
[0019] Further, in step S6, the server can reject abnormal model updates to prevent them from affecting the security of the global model. Once an abnormal update is detected, the server issues a warning to the corresponding client and records its behavior. If the abnormal behavior persists, the participation of the client can be restricted or it can be added to the blacklist. At the same time, the server can replace the rejected update with the most recent normal update or the backup model to maintain the stability of the training process. In addition, the server can generate a security report, detailing the results of the anomaly detection, including client identification, detection time, and anomaly type, and further verify the suspicious clients to confirm whether there are persistent threats or misjudgment situations.
[0020] The beneficial effects of the present invention are as follows:
[0021] Reducing the false detection rate: The present invention validates the model updates submitted by clients by using fingerprint samples with a specific design, combined with a dynamically adjusted threshold mechanism, effectively reducing the false detection rate. The fingerprint samples have significant differences between the normal model and the backdoor model, enabling the system to more accurately identify the backdoor model while avoiding misjudging normal client updates as abnormal, ensuring high accuracy and reliability of the detection.
[0022] Having little impact on the main task performance: During the verification process, the fingerprint samples used in the present invention are carefully designed to ensure that they do not have an obvious impact on the performance of the main task during normal model training. The fingerprint detection process is independent of the model training process, and the generation and verification of fingerprint samples do not interfere with the performance of the model in actual tasks, thus ensuring the efficient and stable performance of the model under normal conditions.
[0023] Simple operation: The fingerprint generation and verification mechanism of the present invention has a simple process and is easy to integrate into the existing federated learning framework. The client does not need to perform additional complex operations and only needs to submit model updates according to the standard protocol, while the fingerprint verification process on the server side can be automated, effectively reducing the system deployment and maintenance costs.
[0024] Adapting to the actual privacy protection scenario of federated learning: The present invention fully considers the privacy protection requirements of federated learning. By performing fingerprint verification on the server side without accessing the original data and model updates of the client, it ensures data privacy and security. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a schematic flow diagram of the present invention;
[0026] Figure 2 is a schematic flow diagram of model fingerprint generation; DETAILED DESCRIPTION OF THE INVENTION
[0027] S1: Generate a set of fingerprint samples with specific characteristics on the server side. These samples can show different responses on the normal model and the backdoor model for subsequent verification;
[0028] S2: Each client performs model training on local data and submits the trained model updates to the server;
[0029] S3: The server uses the fingerprint samples to verify the model updates submitted by each client and records the responses of the updates on the fingerprint samples;
[0030] S4: Compare the consistency between the responses of the model updates submitted by the client to the fingerprint samples and the expected results to determine whether there is an abnormality;
[0031] S5: Dynamically adjust the aggregation weight of model updates according to the number of times a client participates and its historical detection results to adapt to different attack frequencies and the proportion of malicious clients;
[0032] S6: If abnormal client model updates are continuously detected, the server will reject the update, and may record the suspicious client, restrict its subsequent participation or conduct more stringent verification.
[0033] In step S1, the server generates a set of fingerprint samples with specific characteristics, which are used to detect whether there are abnormalities or backdoors in the model updates submitted by the clients. When generating fingerprint samples, specific characteristics that can reflect model performance are selected. In this instance, randomly perturbed characteristics generated by slightly perturbing images are selected. These characteristics should be diverse to ensure coverage of different data types and scenarios, so as to improve the generalization ability of the fingerprint samples. The generated fingerprint samples should be verified by a standard model to obtain corresponding standard responses and be securely stored on the server side. To adapt to changing security threats, the server also needs to regularly update and expand the fingerprint sample library to ensure that it can handle new attack types and scenarios.
[0034] In step S2, each client independently conducts model training on its local dataset. During the training process, the federated learning protocol must be strictly followed to ensure that data privacy is not leaked. After training is completed, the client submits the model update to the server, along with relevant metadata, such as client identifier, participation round, dataset size, and other information.
[0035] In step S3, after the server receives the model update submitted by the client, it uses the pre-generated fingerprint samples to verify the model. The specific operation is to input the fingerprint samples into the model update submitted by the client and record the response results of the model on these fingerprint samples. These response results include indicators such as the output category, confidence level, and loss value of the model for the fingerprint samples. The purpose of this step is to evaluate whether the model has been contaminated by a backdoor attack through the model's reaction to the fingerprint samples, providing data support for subsequent consistency comparison.
[0036] In step S4, the server compares the actual response of the model update submitted by the client on the fingerprint samples with the expected standard response, and judges whether there are abnormalities in the model update by calculating the consistency score. The specific feature design in the fingerprint samples makes the responses of normal models and backdoor models on these samples significantly different. The server detects backdoor attacks through this difference. If the consistency score is lower than the preset threshold, it indicates that there may be a backdoor in the model update and the response result shows an abnormal deviation. Through this process, the server can effectively identify potential backdoor attacks while ensuring transparency in the client training process and not affecting the normal training process.
[0037] In the step S5, the server dynamically adjusts the aggregation weight of model update according to the participation times of the client and the historical detection results. For clients with a large number of participations and stable verification results, the server will appropriately increase their aggregation weights to encourage these trusted clients to continuously contribute to model updates. For clients that are detected with anomalies multiple times or have a small number of participations, the server will reduce their aggregation weights and even set the weights to zero when necessary to reduce the impact of these clients on the global model. After the weight adjustment, normalization processing is required to ensure that the sum of the weights of all clients is 1. In addition, the server can set an adaptive adjustment period, such as dynamically updating the weights after each round of model aggregation or when an anomaly is detected, to improve the adaptability and security of the system to the dynamic attack environment.
[0038] In the step S6, when the server continuously detects that the model update of a certain client is abnormal, it will reject the update to prevent it from having a negative impact on the global model. If the abnormal behavior of this client persists, the server can restrict its subsequent participation in federated learning or add it to the blacklist to prevent further attacks. To maintain the stability of the training process, the server can use the most recent normal model update or backup model to replace the rejected abnormal update.
Claims
1. A federated learning backdoor attack defense method based on model fingerprint, characterized in that: The method comprises the following steps: S1: Generate a set of fingerprint samples with specific characteristics on the server side. These samples can show different responses on the normal model and the backdoor model for subsequent verification; S2: Each client performs model training on local data and submits the trained model update to the server; S3: The server uses the fingerprint sample to verify the model update submitted by each client and records the response of the update on the fingerprint sample; S4: Compare the consistency between the response of the model update submitted by the client to the fingerprint sample and the expected result to determine whether there is an anomaly; S5: Dynamically adjust the aggregation weight of the model update based on the number of client participation and historical detection results to adapt to different attack frequencies and malicious client ratios; S6: If abnormal client model updates are continuously detected, the server will reject the update and may record suspicious clients, restrict their subsequent participation, or perform stricter verification.
2. The method for defending against backdoor attacks of federated learning based on model fingerprint according to claim 1, characterized in that: The fingerprint sample in step S1 has diverse features, including randomly disturbed features, specific label information, or features generated by slight disturbances, so as to ensure that significantly different responses can be generated on the normal model and the backdoor model.
3. The method for defending against backdoor attacks of federated learning based on model fingerprint according to claim 1, characterized in that: The response recorded by the server in step S3 includes output category, confidence level, loss value or other quantifiable performance indicators.
4. The method for defending against backdoor attacks of federated learning based on model fingerprint according to claim 1, characterized in that: The consistency comparison in step S4 is performed by calculating the similarity score between the fingerprint sample response and the standard response, and comparing it with a preset threshold to determine anomalies.
5. The method for defending against backdoor attacks of federated learning based on model fingerprint according to claim 1, characterized in that: The dynamic weight adjustment in step S5 includes: increasing the aggregate weight for clients with many participation times and normal detection results; reducing or setting the aggregate weight to zero for clients with abnormalities detected or few participation times; normalizing the adjusted weights to ensure that the sum of all client weights is 1, etc.
6. The method for defending against backdoor attacks of federated learning based on model fingerprint according to claim 1, characterized in that: In step S6, after rejecting the abnormal update, the server may use the most recent normal update or backup model to replace the abnormal update to maintain the stability of the federated learning training process.
7. The method for defending against backdoor attacks of federated learning based on model fingerprint according to claim 1, characterized in that: In step S6, the server generates an abnormality detection report, recording the abnormal client's identification, detection time, abnormality type and consistency comparison result.
Citation Information
Cited By
Model fingerprint injection and verification method and device based on bypass network
CN120632840A
Model fingerprinting and verification method and apparatus based on side branch network
CN120632840B