Method and device for longitudinal federated learning attack defense based on mutual information
By using mutual information defense methods in vertical federated learning, mixed noise reduces the correlation between encoding representation and real tags, solving the security problem of vertical federated learning system, and achieving defense against model improvement, gradient inversion and backdoor attacks.
Patent Information
- Application Number
- CN202211715007.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-12-29
AI Technical Summary
There are poor security problems in vertical federated learning systems, including model improvement attacks, label inference attacks and backdoor attacks, which lead to threats to data privacy and model security.
Through vertical federated learning of attack defense methods based on mutual information, the coded representation of the passive party is obtained and noise is mixed in the regular term module to generate a second coded representation, reducing the mutual information between the coded representation and the real tag to prevent attacks.
It effectively defends against attacks from passive and proactive parties, protects data privacy and model security, and reduces the ability of attackers to recover data and infer tags through model inversion.
Smart Images

Figure CN116074065B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for attacking and defending vertical federated learning based on mutual information. Background Art
[0002] Vertical federated learning can achieve collaborative learning among multiple participants. However, not all parties in vertical federated learning are trustworthy, and there are various security issues in vertical federated learning.
[0003] For example, in the "model improvement" attack, the attacker (the passive party acts as the attacker) has some data with known labels in an equal quantity in each class to be divided. After the vertical federated learning training process ends, the attacker adds a fully connected layer at the end of the current local model and uses the data with known labels to train this fully connected layer, so that the last layer of this fully connected layer as the "improved model" can output predicted labels very close to the true labels, thereby achieving the attack goal of leaking the data labels of the active party.
[0004] For example, in the label inference attack based on gradient inversion, the passive party acts as the attacker and uses the backpropagated gradient information to complete the inference of the label information held by the attacked active party.
[0005] For example, in the backdoor attack launched by the passive party, the passive party changes the final output of the vertical federated learning joint model by changing the output of the local model, affecting the expression effect of the joint model.
[0006] For example, in the data inference attack launched by the active party as the attacker, the active party uses the pre-knowledge of the passive party's model and restores the local data of the passive party through the model inversion method.
[0007] Therefore, maintaining the security of the vertical federated learning system has become an urgent problem to be solved at present. Summary of the Invention
[0008] The present invention provides a method and device for attacking and defending vertical federated learning based on mutual information, which are used to solve the defect of poor security of the vertical federated learning system in the prior art and improve the security of vertical federated learning.
[0009] In a first aspect, the present invention provides a method for attacking and defending vertical federated learning based on mutual information, which is applied to a defense party and includes:
[0010] Obtain a first encoded representation output by the passive party of the vertical federated learning model, where the first encoded representation is obtained by inputting the local data of the passive party into the local model of the passive party;
[0011] Input the first encoded representation into the regularization term module, add noise to the first encoded representation, and obtain the second encoded representation output by the regularization term module. The second encoded representation is used to jointly obtain a prediction result with a third encoded representation, where the third encoded representation is obtained by the active party of the vertical federated learning model inputting the local data of the active party into the local model of the active party;
[0012] Among them, the defender is the active party or the passive party.
[0013] Optionally, the step of inputting the first encoded representation into the regularization term module, adding noise to the first encoded representation, and obtaining the second encoded representation output by the regularization term module includes:
[0014] Obtain the noise mixing parameter based on the first encoded representation;
[0015] Obtain the noise mixing result based on the noise mixing parameter;
[0016] Decode the noise mixing result to obtain the second encoded representation.
[0017] Optionally, the noise mixing parameter includes an expectation parameter and a standard deviation parameter. The expectation parameter and the standard deviation parameter are used to determine the noise normal distribution function, and the noise normal distribution function is used to generate the noise mixing result.
[0018] Optionally, the step of obtaining the noise mixing result based on the noise mixing parameter includes:
[0019] Perform random sampling on the standard normal distribution function to obtain random noise;
[0020] Sample according to the random noise in the noise normal distribution function to obtain the noise mixing result.
[0021] Optionally, the vertical federated learning model is obtained through the following steps:
[0022] Train the initial vertical federated learning model according to the local data samples of the active party, the local data samples of the passive party that are in one-to-one correspondence with the local data samples of the active party, and the sample labels that are in one-to-one correspondence with the local data samples of the active party;
[0023] Update the parameters of the initial vertical federated learning model through a loss function to obtain the vertical federated learning model;
[0024] Among them, the loss function includes a first component and a second component. The first component is used to indicate the difference between the prediction result and the sample label, and the second component is used to indicate the mutual information amount between the noise mixing result and the first encoded representation.
[0025] Optionally, the loss function is:
[0026] L MID = L contra + λI(T, H p ), λ ≥ 0
[0027]
[0028] where L MID is the loss function, L contra is the first component, I(T, H p ) is the second component, Y label is the sample label, is the prediction result output by the vertical federated learning model, T is the noise mixing result, H p is the first encoded representation, CE is the cross-entropy loss function, I is the mutual information function, and λ is a hyperparameter.
[0029] In a second aspect, the present invention further provides a vertical federated learning attack and defense device based on mutual information, which is applied to the defense party and includes:
[0030] An acquisition unit, configured to acquire a first encoded representation output by the passive party of the vertical federated learning model, where the first encoded representation is obtained by inputting the local data of the passive party into the local model of the passive party;
[0031] A defense unit, configured to input the first encoded representation into a regularization term module, mix noise into the first encoded representation, and obtain a second encoded representation output by the regularization term module, where the second encoded representation is used to jointly obtain a prediction result with a third encoded representation, and the third encoded representation is obtained by inputting the local data of the active party of the vertical federated learning model into the local model of the active party;
[0032] where the defense party is the active party or the passive party.
[0033] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the vertical federated learning attack and defense method based on mutual information as described in the first aspect.
[0034] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the vertical federated learning attack and defense method based on mutual information as described in the first aspect.
[0035] Fifth aspect, the present invention further provides a computer program product, including a computer program, which when executed by a processor implements the method for defending against vertical federated learning attacks based on mutual information as described in the first aspect.
[0036] The method and device for defending against vertical federated learning attacks based on mutual information provided by the present invention obtain a second encoded representation by adding noise to the first encoded representation, and reduce the mutual information between the first encoded representation output by the local model of the passive party and the true label held by the active party to defend against related data privacy and model security attacks. The method for defending against vertical federated learning attacks based on mutual information provided by the embodiments of the present invention can not only defend against the attacks where the passive party uses the information of the labels contained in the local model as the attacker and infers the label information held by the active party by using methods such as "model improvement" attacks and gradient inversion, but also defend against the backdoor attacks launched by the passive party; it can also defend against the inference attacks by the active party on the local data of the passive party, realizing the defense against data leakage and backdoor attacks in vertical federated learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] Figure 1 is one of the schematic flowcharts of the method for defending against vertical federated learning attacks based on mutual information provided by the embodiments of the present invention;
[0039] Figure 2 is another schematic flowchart of the method for defending against vertical federated learning attacks based on mutual information provided by the embodiments of the present invention;
[0040] Figure 3 is still another schematic flowchart of the method for defending against vertical federated learning attacks based on mutual information provided by the embodiments of the present invention;
[0041] Figure 4 is one of the comparison diagrams of the defense effects provided by the embodiments of the present invention;
[0042] Figure 5 is another comparison diagram of the defense effects provided by the embodiments of the present invention;
[0043] Figure 6 is still another comparison diagram of the defense effects provided by the embodiments of the present invention;
[0044] Figure 7 is the effect diagram of the data recovery attack provided by the embodiments of the present invention;
[0045] Figure 8 It is a schematic structural diagram of a vertical federated learning attack and defense device based on mutual information provided by an embodiment of the present invention;
[0046] Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific embodiments
[0047] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0048] The following introduces the technical terms and background related to the present invention:
[0049] Loss Function: Also known as the cost function or objective function, it is used to measure the degree of inconsistency between the predicted value f(x) of the model and the true value Y. Generally, it is a non - negative real - valued function expressed as L(Y, f(x)). Generally speaking, the smaller the value of the loss function (i.e., the loss value), the better the model fits, and the stronger the prediction ability for new data. The loss function is the "baton" for training the model in deep learning. It guides the learning of model parameters through the backpropagation of the error generated by the predicted samples and the true sample labels. When the loss value of the loss function gradually decreases (converges), it can be considered that the model training is completed.
[0050] Vertical Federated Learning (VFL): In the case where the users of two data sets overlap more while the user features overlap less, the data sets are split vertically (i.e., along the feature dimension), and the part of the data where the users of both parties are the same but the user features are not completely the same is taken for training.
[0051] Active Party and Passive Party: In VFL, only one of the two users has labeled data, which is called the active party; the other party has unlabeled data, which is called the passive party. Both parties have local models and complete joint modeling through the output of the local models and the corresponding gradient information.
[0052] The following combines Figures 1 - 7 to describe the vertical federated learning attack and defense method based on mutual information provided by the embodiments of the present invention.
[0053] Figure 1It is one of the schematic flowcharts of the vertical federated learning attack defense method based on mutual information provided by the embodiments of the present invention. As Figure 1 shown, the vertical federated learning attack defense method based on mutual information provided by the embodiments of the present invention can be applied to the defense party, including:
[0054] Step 110, obtain a first encoded representation output by the passive party of the vertical federated learning model, where the first encoded representation is obtained by inputting the local data of the passive party into the local model of the passive party;
[0055] Specifically, for the vertical federated learning model, which can also be referred to as the vertical federated learning system, hereinafter simply referred to as VFL, for the introduction of VFL, the active party in VFL, and the passive party in VFL, refer to the above introduction, which will not be elaborated here.
[0056] The defense party is the active party or the passive party, that is, the defense party can be the active party in VFL or the passive party in VFL. Exemplarily, for the case where the active party is the defense party, the passive party is the attacking party, and the active party executes Step 110 and Step 120; correspondingly, for the case where the passive party is the defense party, the active party is the attacking party, and the passive party executes Step 110 and Step 120.
[0057] Step 120, input the first encoded representation into the regularization term module, mix noise into the first encoded representation, and obtain a second encoded representation output by the regularization term module. The second encoded representation is a prediction representation of the prediction result, and the second encoded representation is used to jointly obtain the prediction result with the third encoded representation, where the third encoded representation is obtained by the active party of the vertical federated learning model inputting the local data of the active party into the local model of the active party.
[0058] Specifically, the regularization term module can obtain the second encoded representation by mixing noise into the first encoded representation, and the regularization term module is used to reduce the mutual information between the first encoded representation and the second encoded representation. Mutual Information (MI) represents whether two variables X and Y are related and the strength of the relationship.
[0059] In traditional VFL, the passive party inputs the local data of the passive party into the local model of the passive party, and a first encoded representation can be obtained. The first encoded representation is used to characterize the local data input into the local model of the passive party; the active party inputs the local data of the active party into the local model of the active party, and a third encoded representation can be obtained. The third encoded representation is used to characterize the local data input into the local model of the active party.
[0060] The passive party sends the first encoded representation to the active party, and the active party generates a prediction result through the first encoded representation and the third encoded representation.
[0061] In traditional VFL, when the active party acts as the attacker, since the active party can directly obtain the first encoded representation, and the first encoded representation is related to the local data of the passive party (there is a certain mutual information between the two), the active party uses the prior knowledge of the passive party's model and recovers the local data of the passive party through the method of model inversion.
[0062] Correspondingly, in traditional VFL, when the passive party acts as the attacker, since the passive party holds the first encoded representation and the first encoded representation is used to generate the prediction result, and since the prediction result is the same as or similar to the true label held by the active party, there is also a correlation (a certain mutual information) between the first encoded representation and the true label held by the active party, enabling the passive party to complete the inference of the label information held by the attacked active party using the backpropagated gradient information or the passive party to launch a backdoor attack using the correlation between the first encoded representation and the prediction result.
[0063] It should be understood that in the embodiments of the present invention, when the active party acts as the defender, the active party executes step 120 to obtain the second encoded representation, and obtains the prediction result through the second encoded representation and the third encoded representation; when the passive party acts as the defender, the passive party executes step 120 to obtain the second encoded representation; the passive party sends the second encoded representation to the active party, and the active party obtains the prediction result through the second encoded representation and the third encoded representation. For other steps, reference can be made to the process of vertical federated learning in related technologies, which will not be elaborated here.
[0064] In the embodiments of the present invention, when the passive party acts as the defender, the active party cannot directly obtain the first encoded representation, and the mutual information between the second encoded representation and the local data of the passive party is less. Therefore, when the active party acts as the attacker, it cannot recover the local data of the passive party through model inversion; correspondingly, when the active party acts as the defender, the defender adds noise to the first encoded representation to obtain the second encoded representation, and generates the prediction result from the second encoded representation and the third encoded representation, reducing the degree of correlation between the first encoded representation and the prediction result, and further reducing the degree of correlation between the first encoded representation and the true label held by the active party. The passive party cannot complete the inference of the label information held by the attacked active party using the backpropagated gradient information, nor can it launch a backdoor attack.
[0065] The mutual information regularization defense (MID) method for vertical federated learning provided by the embodiments of the present invention obtains a second encoded representation by adding noise to the first encoded representation, and reduces the mutual information between the first encoded representation output by the local model of the passive party and the true label held by the active party to defend against related data privacy and model security attacks. The mutual information-based vertical federated learning attack defense method provided by the embodiments of the present invention can not only defend against attacks in which the passive party, as an attacker, uses the information of the labels contained in the local model and uses methods such as "model improvement" attacks and gradient inversion to infer the label information held by the active party, but also defend against backdoor attacks launched by the passive party; it can also defend against inference attacks by the active party on the local data of the passive party, realizing the defense of data leakage and backdoor attacks in vertical federated learning.
[0066] Next, a further description will be given of the possible implementation manners of the above steps in specific embodiments.
[0067] Optionally, in step 120, the inputting the first encoded representation into the regularization term module and adding noise to the first encoded representation to obtain the second encoded representation output by the regularization term module includes:
[0068] Step 121, obtaining a noise addition parameter based on the first encoded representation;
[0069] Specifically, the noise addition parameter refers to the parameter of the noise addition result. The noise addition parameter can be the eigenvalue of the noise addition result, and the noise addition result can be obtained through the noise addition parameter.
[0070] Optionally, the noise addition parameter includes an expected parameter μ and a standard deviation parameter σ. The expected parameter μ and the standard deviation parameter σ are used to determine the noise normal distribution function, and the noise normal distribution function is used to generate the noise addition result.
[0071] The noise normal distribution function N(μ,σ 2 ) is a normal distribution function with a mathematical expectation of μ and a variance of σ 2 . By sampling the noise normal distribution function, the noise addition result can be obtained.
[0072] Step 122, obtaining a noise addition result based on the noise addition parameter;
[0073] Specifically, the noise addition result can be referred to as an information bottleneck layer, and the noise addition result controls the amount of information flowing from the first encoded representation H p to the prediction result .
[0074] The noise addition result can be in the form of a matrix, and the matrix can include multiple elements.
[0075] In the case where the noise mixing parameters include an expected parameter μ and a standard deviation parameter σ, the expected parameter μ is numerically equal to the mean of multiple elements in the noise mixing result (matrix), and the standard deviation parameter σ is numerically equal to the square root of the variance of multiple elements in the noise mixing result (matrix).
[0076] Optionally, in step 122, obtaining the noise mixing result based on the noise mixing parameters includes:
[0077] Step 1221, performing random sampling on the standard normal distribution function to obtain random noise;
[0078] Specifically, random noise ε is generated by sampling the standard normal distribution N(0, 1).
[0079] Step 1222, sampling according to the random noise in the noise normal distribution function to obtain the noise mixing result.
[0080] Specifically, the noise mixing result T is obtained by sampling from the noise normal distribution function N(μ, σ 2 ).
[0081] The noise mixing result can be obtained based on the following formula:
[0082] T = ε × σ + μ.
[0083] Step 123, decoding the noise mixing result to obtain the second encoded representation.
[0084] Decoding the information bottleneck layer T to obtain the prediction representation Z (the second encoded representation) of the target to be predicted (i.e., the label data).
[0085] It should be understood that the regularization term module belongs to the vertical federated learning model. The regularization term module will be trained together with the vertical federated learning model. During the training process, the regularization term module can learn how to obtain the noise mixing parameters from the first encoded representation, how to obtain the noise mixing result according to the noise mixing parameters, and how to decode the noise mixing result to obtain the second encoded representation that can generate the correct prediction result.
[0086] In a possible implementation, the regularization term module includes: an encoder model a decoder model and a noise mixing module completed using the reparameterization trick between the two models. In the embodiments of the present invention, the above three parts are combined and denoted as The noise mixing result T is called the information bottleneck layer.
[0087] The encoder model Used to learn the mean μ (expected parameter) and the square root σ (standard deviation parameter) of each element of the information bottleneck layer T from the first encoded representation (the output of the original local model of the passive party as input).
[0088] Noise mixing module: Based on the output of the encoder model (expected parameter and standard deviation parameter), generate random noise ε by sampling from the standard normal distribution N(0, 1), and complete sampling the information bottleneck layer T from N(μ, σ 2 ).
[0089] Decoder model Decode the information bottleneck layer T to obtain the predicted representation of the label data.
[0090] Optionally, the vertical federated learning model is obtained through the following steps of training:
[0091] Train the initial vertical federated learning model according to the local data samples of the active party, the local data samples of the passive party corresponding one-to-one to the local data samples of the active party, and the sample labels corresponding one-to-one to the local data samples of the active party;
[0092] Update the parameters of the initial vertical federated learning model through the loss function to obtain the vertical federated learning model;
[0093] Wherein, the loss function includes a first component and a second component, the first component is used to indicate the difference between the prediction result and the sample label, and the second component is used to indicate the mutual information amount between the noise mixing result and the first encoded representation.
[0094] Specifically, the vertical federated learning model includes the regularization term module. As described above, the regularization term module will be trained together with the vertical federated learning model. Therefore, compared with the traditional vertical federated learning model, the loss function of the vertical federated learning model including the regularization term module will change, and the loss function including the first component and the second component can take into account the training of the regularization term module. The sample label is the correct label corresponding to the local data samples of the active party and the local data samples of the passive party.
[0095] Optionally, the loss function is:
[0096] L MID = L contra + λI(T, H p ), λ ≥ 0
[0097]
[0098] Wherein, L MID is the loss function, L contrais the first component, I(T, H p ) is the second component, Y label is the sample label, is the prediction result output by the vertical federated learning model, T is the noise mixing result, H p is the first coding representation, CE is the cross-entropy loss function, I is the mutual information function, and λ is a hyperparameter.
[0099] Specifically, the larger λ is, the more information in T p is compressed, and theoretically better defense effects can be obtained; but at the same time, there is also a risk of damaging the accuracy of the main task. Therefore, the selection of the hyperparameter λ is relatively crucial.
[0100] In the method for attacking and defending vertical federated learning based on mutual information provided by the embodiments of the present invention, when minimizing L MID , by minimizing the convergence of the VFL joint model is ensured. At the same time, maximizing and minimizing makes I(Y label , H p ) minimized, achieving the defense design goal of the MID method.
[0101] Figure 2 is the second schematic diagram of the process of the method for attacking and defending vertical federated learning based on mutual information provided by the embodiments of the present invention; Figure 2 shows the flowchart of the method for attacking and defending vertical federated learning based on mutual information provided by the embodiments of the present invention in the case where the active party is the defending party, for defending against attacks launched by the passive party. Figure 3 is the third schematic diagram of the process of the method for attacking and defending vertical federated learning based on mutual information provided by the embodiments of the present invention; Figure 3 shows the flowchart of the method for attacking and defending vertical federated learning based on mutual information provided by the embodiments of the present invention in the case where the passive party is the defending party, for defending against attacks launched by the active party.
[0102] To better understand the principle of this defense mechanism, the following combines Figure 2 and Figure 3 to introduce the operating mechanisms of each part. As shown in Figure 2 and Figure 3 , the part outside the box is the forward propagation process of the joint modeling of the basic VFL, and the part inside the box is the regular term module architecture provided by the embodiments of the present invention.
[0103] The part outside the box includes: inputting the local data X p of the passive party into the local model G p of the passive party to obtain the first coding representation H p ; inputting the local data X of the active partya The local model G input to the active party a to obtain the third encoded representation H a ; through the second encoded representation Z and H a to generate the prediction result output by the vertical federated learning model
[0104] The regular term module architecture includes:
[0105] The encoder model The decoder model and the noise mixing module completed using the reparameterization trick between the two models. In the embodiment of the present invention, the above three parts are combined and denoted as The noise mixing result T is called the information bottleneck layer.
[0106] The encoder model is used to take the first encoded representation (the output of the passive party's original local model as input) and learn the mean μ (expected parameter) and the square root of the variance σ (standard deviation parameter) of each element of the information bottleneck layer T.
[0107] The noise mixing module: Based on the output of the encoder model (expected parameter and standard deviation parameter), generate a random noise ε by sampling the standard normal distribution N(0,1) to complete sampling the information bottleneck layer T from N(μ,σ 2 )
[0108] The decoder model decodes the information bottleneck layer T to obtain the representation of the label data.
[0109] It should be understood that the model used by the MID method can be used either on the active party or on the passive party, as Figure 2 and Figure 3 shown respectively. However, overall, no matter which party uses this defense method, the overall model architecture is the same, so the use of the loss function is also the same. During the model training process, it can be regarded as a part of the VFL model for forward and backward propagation, that is, joint modeling and defense can be completed without changing the VFL protocol.
[0110] Figure 4 is one of the defense effect comparison diagrams provided by the embodiment of the present invention, Figure 5 is the second defense effect comparison diagram provided by the embodiment of the present invention, Figure 6 is the third defense effect comparison diagram provided by the embodiment of the present invention, as Figures 4 - 6 shown, Figure 4Shows the comparison results between the MID method and existing methods in defending against model refinement attacks initiated by the passive party on the CIFAR10 dataset, Figure 5 Shows the comparison results between the MID method and existing methods in label inference attacks based on gradient inversion on the MNIST dataset, Figure 6 Shows the comparison results between the MID method and existing methods in backdoor attacks on noisy data on the CIFAR100 dataset. Figures 4 - 6 In, DP-G represents differential privacy with Gaussian noise, DP-L represents differential privacy with Laplace noise, GS represents gradient sparsification, DG represents gradient discretization, CAE represents confusion autoencoder, RVFR represents robust subfeature reconstruction, w / o defense represents no defense, and MID represents the mutual-information-based vertical federated learning attack defense method provided by the present invention. Figures 4 - 6 In, multiple points on each curve are the defense results of the hyperparameters of the corresponding defense method (noise intensity in DP-G and DP-L, sparsification degree in GS, number of discretization intervals in DG, confusion degree in CAE, and λ in MID) at multiple different values. Figures 4 - 6 In, the farther the curve is to the lower right, the better the defense effect and the smaller the impact on the main task accuracy. As Figures 4 - 6 shown, if the attacked active party uses MID for defense, it can obtain good defense effects in the above three attacks while effectively maintaining a high main task accuracy.
[0111] Figure 7 is the effect diagram of the data recovery attack provided by the embodiment of the present invention. As Figure 7 shown, Figure 7 Shows the effect of the MID method in defending against the model inversion-based data recovery attack initiated actively. The original image is held by the passive party, and the inversion results deduced by the active party under different circumstances, such as the results without defense (No Defense) and the results with different values of λ. It can be seen that the passive party using the MID method can effectively prevent the active party from inferring the local data of the passive party.
[0112] The following describes the mutual-information-based vertical federated learning attack defense device provided by the present invention. The mutual-information-based vertical federated learning attack defense device described below can be mutually corresponding and referred to the mutual-information-based vertical federated learning attack defense method described above.
[0113] Figure 8 is the structural schematic diagram of the mutual-information-based vertical federated learning attack defense device provided by the embodiment of the present invention. As Figure 8 shown, the mutual-information-based vertical federated learning attack defense device provided by the embodiment of the present invention includes:
[0114] An acquisition unit 810, configured to acquire a first encoded representation output by a passive party of a vertical federated learning model, where the first encoded representation is obtained by inputting local data of the passive party into a local model of the passive party;
[0115] A defense unit 820, configured to input the first encoded representation into a regularization term module, mix noise into the first encoded representation, and obtain a second encoded representation output by the regularization term module, where the second encoded representation is used to jointly obtain a prediction result with a third encoded representation, and the third encoded representation is obtained by inputting local data of an active party of the vertical federated learning model into a local model of the active party;
[0116] Wherein, the defense party is the active party or the passive party.
[0117] Optionally, the defense unit 820 is configured to obtain a noise mixing parameter based on the first encoded representation;
[0118] The defense unit 820 is configured to obtain a noise mixing result based on the noise mixing parameter;
[0119] The defense unit 820 is configured to decode the noise mixing result to obtain the second encoded representation.
[0120] Optionally, the noise mixing parameter includes an expectation parameter and a standard deviation parameter, and the expectation parameter and the standard deviation parameter are used to determine a noise normal distribution function, and the noise normal distribution function is used to generate the noise mixing result.
[0121] Optionally, the defense unit 820 is configured to perform random sampling on a standard normal distribution function to obtain random noise;
[0122] The defense unit 820 is configured to sample in the noise normal distribution function according to the random noise to obtain the noise mixing result.
[0123] Optionally, the apparatus further includes a training unit;
[0124] The training unit is configured to train an initial vertical federated learning model according to local data samples of the active party, local data samples of the passive party that are in one-to-one correspondence with the local data samples of the active party, and sample labels that are in one-to-one correspondence with the local data samples of the active party;
[0125] The training unit is configured to update parameters of the initial vertical federated learning model through a loss function to obtain the vertical federated learning model;
[0126] Among them, the loss function includes a first component and a second component. The first component is used to indicate the difference between the prediction result and the sample label, and the second component is used to indicate the mutual information amount between the noise-mixed result and the first encoded representation.
[0127] Optionally, the loss function is:
[0128] L MID = L contra + λI(T, H p ), λ ≥ 0
[0129]
[0130] Among them, L MID is the loss function, L contra is the first component, I(T, H p ) is the second component, Y label is the sample label, is the prediction result output by the vertical federated learning model, T is the noise-mixed result, H p is the first encoded representation, CE is the cross-entropy loss function, I is the mutual information function, and λ is a hyperparameter.
[0131] It should be noted here that the above device provided by the embodiments of the present invention can implement all the method steps implemented by the above method embodiments and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiments will not be specifically described in this embodiment.
[0132] Figure 9 Illustrates a schematic physical structure diagram of an electronic device, such as Figure 9As shown, the electronic device may include: a processor 910, a communications interface 920, a memory 930, and a communication bus 940. Among them, the processor 910, the communications interface 920, and the memory 930 complete their mutual communication through the communication bus 940. The processor 910 may call the logical instructions in the memory 930 to execute the vertical federated learning attack defense method based on mutual information, which can be applied to the defense party. The method includes: obtaining a first encoded representation output by the passive party of the vertical federated learning model, where the first encoded representation is obtained by inputting the local data of the passive party into the local model of the passive party; inputting the first encoded representation into a regularization term module to mix noise into the first encoded representation to obtain a second encoded representation output by the regularization term module, and the second encoded representation is used to jointly obtain a prediction result with a third encoded representation, where the third encoded representation is obtained by the active party of the vertical federated learning model inputting the local data of the active party into the local model of the active party; where the defense party is the active party or the passive party.
[0133] In addition, when the logical instructions in the above-mentioned memory 930 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0134] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the mutual-information-based vertical federated learning attack defense method provided by the above-mentioned various methods, which can be applied to the defense party. The method includes: obtaining a first encoded representation output by the passive party of the vertical federated learning model, where the first encoded representation is obtained by inputting the local data of the passive party into the local model of the passive party; inputting the first encoded representation into a regularization term module to mix noise into the first encoded representation to obtain a second encoded representation output by the regularization term module, and the second encoded representation is used to jointly obtain a prediction result with a third encoded representation, where the third encoded representation is obtained by the active party of the vertical federated learning model inputting the local data of the active party into the local model of the active party; where the defense party is the active party or the passive party.
[0135] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the mutual-information-based vertical federated learning attack defense method provided by the above-mentioned various methods, which can be applied to the defense party. The method includes: obtaining a first encoded representation output by the passive party of the vertical federated learning model, where the first encoded representation is obtained by inputting the local data of the passive party into the local model of the passive party; inputting the first encoded representation into a regularization term module to mix noise into the first encoded representation to obtain a second encoded representation output by the regularization term module, and the second encoded representation is used to jointly obtain a prediction result with a third encoded representation, where the third encoded representation is obtained by the active party of the vertical federated learning model inputting the local data of the active party into the local model of the active party; where the defense party is the active party or the passive party.
[0136] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0137] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A vertical federated learning attack and defense method based on mutual information, characterized in that, Applied to the defense side, including: Obtain a first encoded representation output by the passive party of the vertical federated learning model, where the first encoded representation is obtained by inputting the local data of the passive party into the local model of the passive party; Input the first encoded representation into a regularization term module, add noise to the first encoded representation, and obtain a second encoded representation output by the regularization term module. The second encoded representation is used to jointly obtain a prediction result with a third encoded representation, where the third encoded representation is obtained by the active party of the vertical federated learning model inputting the local data of the active party into the local model of the active party; Wherein, the defense side is the active party or the passive party; Wherein, the step of inputting the first encoded representation into a regularization term module, adding noise to the first encoded representation, and obtaining a second encoded representation output by the regularization term module includes: Obtain a noise mixing parameter based on the first encoded representation; Obtain a noise mixing result based on the noise mixing parameter; Decode the noise mixing result to obtain the second encoded representation; Wherein, the noise mixing parameter includes an expectation parameter and a standard deviation parameter, and the expectation parameter and the standard deviation parameter are used to determine a noise normal distribution function, and the noise normal distribution function is used to generate the noise mixing result; Wherein, the step of obtaining a noise mixing result based on the noise mixing parameter includes: Perform random sampling on the standard normal distribution function to obtain random noise; Sample according to the random noise in the noise normal distribution function to obtain the noise mixing result; Wherein, the vertical federated learning model is trained through the following steps: Train an initial vertical federated learning model according to the local data samples of the active party, the local data samples of the passive party corresponding one-to-one to the local data samples of the active party, and the sample labels corresponding one-to-one to the local data samples of the active party; Update the parameters of the initial vertical federated learning model through a loss function to obtain the vertical federated learning model; Wherein, the loss function includes a first component and a second component. The first component is used to indicate the difference between the prediction result and the sample label, and the second component is used to indicate the mutual information amount between the noise mixing result and the first encoded representation.
2. The method for attacking and defending vertical federated learning based on mutual information according to claim 1, wherein, The loss function is: L MID = L contra + λI(T, H p ), λ ≥ 0 Among them, L MID is the loss function, L contra is the first component, I(T, H p ) is the second component, Y label is the sample label, is the prediction result output by the vertical federated learning model, T is the noise-mixed result, H p is the first encoded representation, CE is the cross-entropy loss function, I is the mutual information function, and λ is a hyperparameter.
3. A vertical federated learning attack defense device based on mutual information, characterized in that, Applied to the defense side, including: An acquisition unit for acquiring a first encoded representation output by the passive party of the vertical federated learning model, where the first encoded representation is obtained by inputting the local data of the passive party into the local model of the passive party; A defense unit for inputting the first encoded representation into a regularization term module, adding noise to the first encoded representation, and obtaining a second encoded representation output by the regularization term module. The second encoded representation is used to jointly obtain a prediction result with a third encoded representation, where the third encoded representation is obtained by the active party of the vertical federated learning model inputting the local data of the active party into the local model of the active party; Wherein, the defense side is the active party or the passive party; Wherein, the defense unit is used to obtain a noise mixing parameter based on the first encoded representation; The defense unit is configured to obtain a noise-mixed result based on the noise-mixed parameter; The defense unit is configured to decode the noise-mixed result to obtain the second encoded representation; Wherein, the noise-mixed parameter includes an expected parameter and a standard deviation parameter, the expected parameter and the standard deviation parameter are used to determine a noise normal distribution function, and the noise normal distribution function is used to generate the noise-mixed result; Wherein, the defense unit is configured to perform random sampling on the standard normal distribution function to obtain random noise; The defense unit is configured to sample according to the random noise in the noise normal distribution function to obtain the noise-mixed result; Wherein, the device further includes a training unit; The training unit is configured to train an initial vertical federated learning model according to the local data samples of the active party, the local data samples of the passive party that are in one-to-one correspondence with the local data samples of the active party, and the sample labels that are in one-to-one correspondence with the local data samples of the active party; The training unit is configured to update the parameters of the initial vertical federated learning model through a loss function to obtain the vertical federated learning model; Wherein, the loss function includes a first component and a second component, the first component is used to indicate the difference between the prediction result and the sample label, and the second component is used to indicate the mutual information amount between the noise-mixed result and the first encoded representation.
4. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the mutual-information-based vertical federated learning attack defense method according to any one of claims 1 to 2.
5. A non-transitory computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the mutual-information-based vertical federated learning attack defense method according to any one of claims 1 to 2.
6. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the mutual-information-based vertical federated learning attack defense method according to any one of claims 1 to 2.
Citation Information
Patent Citations
Vertical federated modeling optimization method and device, and readable storage medium
WO2022016964A1