Longitudinal federated learning backdoor defense method and system based on high-confidence deviation

By constructing clean embedded datasets, robust distillation and backdoor detection, the backdoor attack defense problem in vertical federated learning is solved, significantly reducing the success rate of backdoor attacks and improving the security and reliability of the model.

CN120145374APending Publication Date: 2025-06-13WUHAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510301127.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Vertical federated learning (VFL) is vulnerable to backdoor attacks, and due to limited access to data and model, prior art is difficult to effectively defend against backdoor attacks.

Method used

Using a backdoor defense method based on high confidence bias, a model without backdoor is extracted by building a clean embedded dataset, robust distillation and backdoor detection to ensure the accuracy of prediction.

Benefits of technology

It significantly reduces the success rate of backdoor attacks, improves the prediction accuracy of the model for poisoning samples, and ensures the security and reliability of the longitudinal federated learning model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145374A_ABST
    Figure CN120145374A_ABST
Patent Text Reader

Abstract

The invention discloses a longitudinal federated learning backdoor defense method and system based on high-confidence deviation, and aims to solve the problem of backdoor attack defense caused by the fact that data and model access is limited and the embedding authenticity is difficult to verify in the prior art. The method comprises the steps of distillation data set construction, robust distillation and backdoor detection. The distillation dataset building module screens clean embedding by analyzing the difference between the model output clean embedding and poisoning embedding. The robust distillation module then utilizes the outputs and tags of the top layer model to guide training of the backdoor-free model. In the distillation process, synthetic poisoning embedding is used to further enhance the robustness of the distillation model. And finally, the backdoor detection module can identify poisoning embedding, so that the prediction robustness is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence security, and relates to a backdoor defense method and system for vertical federated learning (VFL), and particularly to a vertical federated learning backdoor defense method and system based on high-confidence deviation, which is used to resist backdoor attacks in a multi-party collaboration scenario and ensure the security and reliability of machine learning models. Background Art

[0002] With the increasing global attention to data privacy and security, federated learning (FL) has emerged as a distributed machine learning framework. Federated learning allows multiple data owners (such as different enterprises and organizations) to collaboratively train a shared model without sharing the original data. Federated learning aims to solve the data silo problem, promote the efficient utilization of data and the optimization of models, while complying with strict privacy protection regulations. According to the distribution of data among different data sources, federated learning can be divided into three main types: horizontal federated learning (HFL), vertical federated learning (VFL), and federated transfer learning (FTL). Vertical federated learning (VFL) is applicable to scenarios where data sources share the same sample space but have different feature subsets. For example, a bank and an e-commerce platform may respectively have different features of the same group of users (such as financial transaction records and shopping behaviors), and can collaboratively train a model through VFL without sharing the original data. For a schematic diagram of the VFL scenario, please refer to Figure 6 .

[0003] However, recent studies have shown that VFL is vulnerable to backdoor attacks. Attackers can manipulate its local information, such as data and models, during training to inject backdoors into the top-level model, with the goal of misleading the top-level model to make incorrect predictions using malicious embeddings during inference. However, due to its unique workflow, backdoor defense of VFL faces greater challenges. First, each participant in VFL has partial features, and the defender cannot directly access the raw data or local models of other participants, so it is impossible to pre-process the data or directly monitor the model. Second, the defender relies on the embeddings provided by the passive party for training and inference, but cannot verify the authenticity of these embeddings, which increases the difficulty of detecting and defending against backdoor attacks. At present, although many studies have focused on backdoor defense in centralized machine learning and HFL, due to the particularity of VFL, backdoor defense methods for VFL are still scarce. Therefore, it is of great significance to develop effective backdoor defense strategies specifically for VFL. Summary of the invention

[0004] In order to solve the problem of backdoor attack defense caused by limited data and model access and difficult to verify embedded authenticity in the prior art, the present invention provides a longitudinal federated learning backdoor defense method and system based on high confidence deviation, which extracts a backdoor-free model from the attacked top-level model through distilled data set construction, robust distillation and backdoor detection to ensure the accuracy of prediction.

[0005] The technical solution adopted by the method of the present invention is: a vertical federated learning backdoor defense method based on high confidence deviation, comprising the following steps:

[0006] Step 1: Construct a clean embedding dataset D using embeddings that meet preset criteria c ;

[0007] The embedding that meets the preset criteria is that if each participant-level model C k Embed The confidence score Below a predefined threshold τ 1 , then the embedding satisfies the preset standard; where k∈[K], K represents the passive participant P k quantity;

[0008] Step 2: Using the constructed clean embedding dataset D c , using knowledge distillation technology, while introducing synthetic poisoning embedding D in the distillation process b , knowledge from the top model Moving to a distillation model

[0009] Step 3: Backdoor prediction;

[0010] Compare the prediction results of the distilled model with those of the top-level model . If they are inconsistent, and the participant-level model C k predicts the same class as the top-level model , and the embedding confidence score exceeds the threshold τ 1 , then the embedding is malicious, and the prediction output of the distilled model is selected; otherwise, the prediction output of the top-level model is retained.

[0011] Preferably, during the model training process, for the selection of the threshold τ 1 , first find the highest confidence scores of all participant-level models in each training sample, and then calculate the median of these highest confidence scores in all training samples as the threshold τ 1 .

[0012] Preferably, in step 2, introducing the synthetic poisoned embedding D b during the distillation process, its specific implementation includes the following sub-steps:

[0013] Step 2.1: Determine a set of embeddings E k that generate high confidence in their corresponding participant-level models C k for each passive party P k ; where k ∈ [K], and K represents the number of passive parties P k .

[0014] Step 2.2: Randomly select a clean embedding subset from D c ; for each selected clean embedding randomly select a passive party P k , and replace its embedding with a high-confidence embedding from a different class to synthesize the corresponding poisoned embedding and generate a set of correctly labeled synthetic poisoned embeddings D b .

[0015] Preferably, in step 2.1, define the threshold τ 2 as the Nth percentile of the maximum confidence scores in all training samples, and select the embeddings with confidence scores exceeding this threshold τ 2 as E k ; where N is a preset value.

[0016] Preferably, in step 2.2, where is the model C k for The confidence score of, y i ≠y j Indicates different categories.

[0017] Preferably, during the model training process, the overall loss function adopted is:

[0018]

[0019] where α is a weighting coefficient, e represents the embedded sample, y represents the sample label, S(e) represents the output of the distilled model, T(e) represents the output of the top-level model, and L KL (S(e), T(e)) represents the KL divergence between the top-level model and the distilled model, and L CE (S(e), y) represents the cross-entropy loss function of the distilled model, represents the output of the distilled model for the synthetic poisoned embedding.

[0020] The technical solution adopted by the system of the present invention is: A vertical federated learning backdoor defense system based on high-confidence deviation, including the following modules:

[0021] A distilled dataset construction module for constructing a clean embedded dataset D using the embeddings that meet the preset criteria c ;

[0022] For the embeddings that meet the preset criteria, if each participant-level model C k for the embedding the confidence score is lower than the predefined threshold τ 1 , then the embedding is an embedding that meets the preset criteria; where k ∈ [K], and K represents the number of passive participants P k quantity;

[0023] A robust distillation module for using the constructed clean embedded dataset D c , adopting knowledge distillation technology, and introducing synthetic poisoned embeddings D b during the distillation process, and transferring knowledge from the top-level model to the distilled model

[0024] A backdoor prediction module for comparing the prediction result of the distilled model with the prediction result of the top-level model . If they are inconsistent, and the participant-level model C k and the top-level model predict the same category, and the confidence score of the embedding exceeds the threshold τ 1 , then the embedding is malicious, and the distilled model is selected The predicted output; otherwise, the top-level model will be retained The predicted output.

[0025] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the longitudinal federated learning backdoor defense method based on high-confidence deviation is implemented.

[0026] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the longitudinal federated learning backdoor defense method based on high-confidence deviation is implemented.

[0027] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the longitudinal federated learning backdoor defense method based on high-confidence deviation is implemented.

[0028] Compared with the prior art, the beneficial effects of the present invention are:

[0029] (1) The present invention proposes a new longitudinal federated learning backdoor defense technology, which does not require obtaining the privacy information of other parties, uses confidence to perform data screening and poisoned embedding synthesis, and distills a backdoor-free model. Experiments prove that the defense effect of the present invention is better than the current state-of-the-art defense scheme, significantly reducing the success rate of backdoor attacks and being able to accurately predict anti-backdoors for poisoned embeddings.

[0030] (2) By training a separate model for each passive party, the present invention amplifies the difference in confidence scores. Embeddings that show low confidence in these models are considered clean and used to construct a distilled dataset, solving the problem that the defense party in longitudinal federated learning cannot obtain clean data.

[0031] (3) Relying solely on clean embeddings for distillation makes the distilled model vulnerable to embedding collision attacks. To address this situation, the present invention synthesizes poisoned embeddings against such attacks and introduces them into the distillation process, thus effectively enhancing the robustness of the distilled model.

[0032] (4) The present invention also develops a backdoor detection algorithm, which assigns the identified poisoned samples to the distilled model while retaining the predictions of the top-level model for clean samples, thus ensuring high accuracy of the model for both clean samples and poisoned samples. Description of the Drawings

[0033] The following uses examples and specific implementation manners to further illustrate the technical solution of the present invention. In addition, during the process of illustrating the technical solution, some drawings are also used. For those skilled in the art, without creative efforts, other drawings and the intention of the present invention can also be obtained based on these drawings.

[0034] Figure 1 It is the schematic diagram of the method for the embodiment of the present invention;

[0035] Figure 2 It is the confidence distribution diagram of the top-level model and the participant-level model for the embodiment of the present invention; among them, the left side is the confidence distribution diagram of the top-level model, and the right side is the confidence distribution diagram of the participant-level model;

[0036] Figure 3 It is the data diagram for the comparative analysis of the backdoor defense performance of vertical federated learning for the embodiment of the present invention; among them, ND represents the vertical federated learning framework without taking any defense measures;

[0037] Figure 4 It is the data diagram for the comparative analysis of the backdoor defense performance of vertical federated learning on the Yahoo Answers dataset for the embodiment of the present invention; since the TECB attack method does not support text data, it is not included;

[0038] Figure 5 It is the data diagram of the ablation experiment for synthetic poisoned embeddings in the embodiment of the present invention;

[0039] Figure 6 It is the schematic diagram of the VFL scenario in the prior art of the present invention. Specific implementation manner

[0040] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below in conjunction with the drawings and embodiments. It should be understood that the implementation examples described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0041] Please refer to Figure 1 , a method for backdoor defense of vertical federated learning based on high-confidence deviation provided in this embodiment includes the following steps:

[0042] Step 1: Use the embeddings that meet the preset criteria to construct a clean embedding dataset D c ;

[0043] Defense techniques based on distillation have become an effective strategy to mitigate backdoor attacks. However, a key challenge in applying knowledge distillation in VFL is that the active party cannot obtain a clean dataset. Therefore, the active party must screen out clean embeddings from the potentially poisoned training sample embeddings it receives.

[0044] In one implementation, the confidence difference between the clean embedding and the poisoned embedding is analyzed by establishing a participant hierarchy model. Since each passive party has a different subset of features, the active party can build a model C k for each participant P k . These models are trained simultaneously with the top-level model, using the embeddings E k from all parties as input and the true labels Y as the target. Since these participant hierarchy models only utilize a subset of the features available to the top-level model, they assign a lower confidence to the clean embeddings. However, when the attacker P adv injects poisoned embeddings, the corresponding participant hierarchy model C adv is also attacked. In this case, the influence of the poisoned embeddings in a specific poisoned model C adv is not diluted by the contributions of the honest participants, thus amplifying the backdoor signal.

[0045] As Figure 2 shown, the participant hierarchy model C adv exhibits a significant confidence bias between the poisoned embeddings and the clean embeddings. The confidence scores generated by the poisoned embeddings are much higher than those of most clean embeddings. Utilizing this observation, if the confidence score k of each participant hierarchy model C is lower than a predefined threshold τ 1 , the embedding is considered a clean embedding, as shown in the following equation:

[0046]

[0047] where K represents the number of passive participants, represents the confidence score of the participant hierarchy model C k for the embedding . For the selection of the threshold τ 1 , in this implementation, first, the highest confidence scores of all participant hierarchy models in each training sample are found, and then the median of these highest confidence scores among all training samples is calculated, which is the threshold τ 1 . The embeddings that meet this criterion will be used to construct a distilled dataset D c .

[0048] Step 2: Utilize the constructed clean embedding dataset D c , adopt the knowledge distillation technique, and introduce synthetic poisoned embeddings D b during the distillation process to transfer the knowledge from the top-level model to the distilled model

[0049] In one embodiment, the knowledge distillation technique is adopted, and the constructed clean embedding dataset D is used c , to transfer knowledge from the top-level model to the distilled model However, although the distilled dataset consists only of clean embeddings, the distilled model is still vulnerable to a specific type of attack, which this embodiment refers to as the embedding collision attack. In this case, the attacker injects malicious embeddings that are very similar to the target class, thus creating a collision in the embedding space.

[0050] To enhance the robustness of the distilled model against the embedding collision attack, this embodiment introduces synthetic poisoned embeddings during the distillation process. This process first determines a set of embeddings E that generate high confidence in their respective participant-level models C k for each passive party P k . Specifically, this embodiment defines a threshold τ k , set as the 90th percentile of the maximum confidence scores among all training samples, and selects the embeddings whose confidence scores exceed this threshold. Next, this embodiment randomly selects a subset of clean embeddings from D 2 . For each selected clean embedding c , this embodiment randomly selects a passive party P , and replaces its embedding k with a high-confidence embedding from a different class (i.e., y i ≠y j ), thus synthesizing the corresponding poisoned embedding . This synthesis process can be expressed as: where

[0051]

[0052] is the confidence score of model C for k . Through this process, we generate a set of correctly labeled synthetic poisoned embeddings D . b .

[0053] Step 3: Backdoor prediction;

[0054] In one embodiment, the prediction result of is compared with the prediction result of the top-level model . The basic principle is that and tend to generate similar prediction results for clean embeddings, but diverge for embeddings affected by the backdoor attack. When and When the predictions for a given embedding are different, this embodiment further examines the predictions of the participant-level models. If a participant-level model C k and predict the same class and its confidence score exceeds the threshold τ 1 , this embodiment considers it malicious and chooses to trust the prediction of the distilled model . Otherwise, this embodiment will retain the output of the top-level model .

[0055] In this embodiment, during the model training process, the overall loss function adopted includes the following two parts:

[0056]

[0057] where α is a weighting coefficient, e represents the embedded sample, y represents the sample label, S(e) represents the output of the distilled model, T(e) represents the output of the top-level model, and L KL (S(e), T(e)) represents the KL divergence between the top-level model and the distilled model, and L CE (S(e), y) represents the cross-entropy loss function of the distilled model, represents the output of the distilled model for the synthetic poisoned embedding.

[0058] The first part is applied to the clean embedding and combines the KL divergence and the cross-entropy (CE) loss. The second part trains the distilled model to correctly classify the synthetic poisoned embedding, thereby improving its robustness.

[0059] This embodiment also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the vertical federated learning backdoor defense method based on high-confidence deviation.

[0060] This embodiment also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the vertical federated learning backdoor defense method based on high-confidence deviation.

[0061] This embodiment also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the vertical federated learning backdoor defense method based on high-confidence deviation.

[0062] The following further elaborates on the present invention through specific experiments.

[0063] In the experimental part, five representative vertical federated learning backdoor attack methods were selected, including LR-BA, TECB, GA, SDDP, and VILLAIN, and two existing defense methods, BaDExpert and VFLIP, were used as control groups for comparison. The experimental datasets covered a variety of data modalities, including the image datasets CIFAR-10, CINIC-10, and BHI, the text dataset Yahoo-Answer, and the tabular dataset BM, ensuring the extensiveness and representativeness of the experiments. The evaluation metrics included primary metrics and auxiliary metrics. The primary metrics were the clean accuracy (CA) and the attack success rate (ASR), which were used to evaluate the prediction performance of the model on normal samples and the effectiveness of the backdoor attack, respectively. The auxiliary metrics included the robust accuracy (RA) and the defense effectiveness rating (DER), which were used to evaluate the prediction accuracy of the model on poisoned samples and the balance ability of the defense technology between CA and ASR, respectively. In particular, for the BHI dataset, the F1 score was used as the CA metric to more accurately evaluate the model performance. The experimental results showed that the technical solution ConfiDefense of the present invention significantly reduced the backdoor attack success rate while maintaining a high clean accuracy, demonstrating the effectiveness and practicality of ConfiDefense in the defense against vertical federated learning backdoor attacks.

[0064] This experiment constructed an experimental scenario with four passive parties, where one passive party acted as the attacker. As Figure 3 and Figure 4As shown, the ConfiDefense method proposed by the present invention is significantly superior to the existing comparative methods in terms of defense effect. Specifically, ConfiDefense demonstrates excellent defense performance under various attack strategies, and its two indicators of DER and RA are both superior to the comparative methods. On the CIFAR-10, CINIC, BHI, and BM datasets, the average DER values of ConfiDefense reach 85.75%, 90.49%, 96.97%, and 89.31% respectively, and the CA decreases by only 2.51%, 3.37%, 1.57%, and 2.41%. On the Yahoo Answers dataset, the average DER of ConfiDefense is 7.3% higher than that of VFLIP and 13.96% higher than that of BaDExpert. Compared with VFLIP, ConfiDefense increases the RA by 9.39%, 15.26%, and 7.96% on the CIFAR-10, CINIC, and Yahoo Answers datasets respectively. The performance of BaDExpert is the worst, mainly because its assumption about the clean dataset is unrealistic in the VFL setting. When the model is fine-tuned on a dataset that may contain poisoned embeddings, this will hinder the extraction of backdoors and normal functions, thus reducing the effectiveness of BaDExpert.

[0065] This experiment conducted ablation experiments to evaluate the role of synthetic poisoned embeddings. This experiment compared the performance of ConfiDefense with and without synthetic poisoned embeddings on the CIFAR-10 dataset. As Figure 5 shown, without using synthetic poisoned embeddings, it is difficult for ConfiDefense to defend against LR-BA, TECB, and VILLAIN. In contrast, introducing these synthetic poisoned embeddings for robust distillation can keep the ASR below 3% and increase the DER by an average of 24.44%.

[0066] The present invention verifies the effectiveness and robustness of this method through experiments.

[0067] It should be understood that the above-described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. In addition, the technical features in each embodiment or single embodiment provided by the present invention can be combined with each other arbitrarily to form a feasible technical solution. This combination is not restricted by the order of steps and / or the mode of structural composition, but must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions results in contradictions or cannot be realized, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0068] It should be understood that the above description of the preferred embodiments is relatively detailed, and it should not be considered as a limitation on the protection scope of the present invention. Under the inspiration of the present invention, those of ordinary skill in the art can also make substitutions or deformations without departing from the protection scope defined by the claims of the present invention, and all of them fall within the protection scope of the present invention. The scope of protection claimed by the present invention shall be subject to the appended claims.

Claims

1. A vertical federated learning backdoor defense method based on high confidence deviation, characterized in that: The following steps are involved: Step 1: Construct a clean embedding dataset D using embeddings that meet preset criteria c ; The embedding that meets the preset criteria is that if each participant-level model C k Embed The confidence score If the embedding is lower than the predefined threshold τ1, then the embedding satisfies the preset standard; where k∈[K], K represents the passive participant P k quantity; Step 2: Use the constructed clean embedding dataset D c , using knowledge distillation technology, while introducing synthetic poisoning embedding D in the distillation process b , knowledge from the top model Moving to a distillation model Step 3: Backdoor prediction; The distillation model The prediction results of the top model If there is any inconsistency, and the participant-level model C k With the top model The same category is predicted and embedded If the confidence score exceeds the threshold τ1, the embedding is malicious and the distillation model is selected Otherwise, the top model is retained. The predicted output.

2. The vertical federated learning backdoor defense method based on high confidence deviation according to claim 1 is characterized in that: During the model training process, for the selection of the threshold τ1, we first find the highest confidence score of all participant-level models in each training sample, and then calculate the median of these highest confidence scores in all training samples as the threshold τ1.

3. The vertical federated learning backdoor defense method based on high confidence deviation according to claim 1 is characterized in that: In step 2, the synthetic poisoning insert D is introduced during the distillation process. b , its specific implementation includes the following sub-steps: Step 2.1: For each passive party P k Identify a set of models C at their corresponding actor level k Produces a high confidence embedding E k ; where k∈[K], K represents the passive participant P k quantity; Step 2.2: From D c A subset of clean embeddings is randomly selected from ; for each selected clean embedding Randomly select a passive player P k , and embedding with high confidence from different categories Replace its embed Thus synthesizing the corresponding poisoned embedding Generate a set of correctly labeled synthetic poisoning embeddings D b .

4. The vertical federated learning backdoor defense method based on high confidence deviation according to claim 3 is characterized in that: In step 2.1, define the threshold τ2 as the Nth percentile of the maximum confidence score among all training samples, and select the embedding whose confidence score exceeds the threshold τ2 as E k ; Where N is the preset value.

5. The vertical federated learning backdoor defense method based on high confidence deviation according to claim 3 is characterized in that: In step 2.2, in, It is Model C k right The confidence score of i ≠y j Indicates different categories.

6. The method for backdoor defense of vertical federated learning based on high confidence deviation according to any one of claims 1 to 5, characterized in that: The overall loss function used during model training for: Among them, α is the weighting coefficient, e represents the embedded sample, and y represents the sample label. represents the output of the distillation model, T(e) represents the output of the top model, and L KL (S(e), T(e)) represents the KL divergence between the top model and the distillation model, L CE (S(e),y) represents the cross entropy loss function of the distillation model, represents the output of the distillation model for synthetic poisoned embeddings.

7. A vertical federated learning backdoor defense system based on high confidence deviation, characterized in that: Includes the following modules: The distillation dataset construction module is used to construct a clean embedding dataset D using embeddings that meet preset criteria. c ; The embedding that meets the preset criteria is that if each participant-level model C k Embed The confidence score If the embedding is lower than the predefined threshold τ1, then the embedding satisfies the preset standard; where k∈[K], K represents the passive participant P k quantity; Robust distillation module for leveraging the constructed clean embedding dataset D c , using knowledge distillation technology, while introducing synthetic poisoning embedding D in the distillation process b , knowledge from the top model Moving to a distillation model Backdoor prediction module, used to distill the model The prediction results of the top model If there is any inconsistency, and the participant-level model C k With the top model The same category is predicted and embedded If the confidence score exceeds the threshold τ1, the embedding is malicious and the distillation model is selected Otherwise, the top model is retained. The predicted output.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the vertical federated learning backdoor defense method based on high confidence deviation as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the vertical federated learning backdoor defense method based on high confidence deviation as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the vertical federated learning backdoor defense method based on high confidence deviation as described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Batch normalization mask pruning defense method based on channel contribution degree

    CN121303229A